arXiv 2312.06709

AM-RADIO: Agglomerative Vision Foundation Model -- Reduce All Domains Into One

By Mike Ranzinger, Greg Heinrich, et al.

Published 2023-12-10

Wiki summary

Explore the paper's summary, context, and related research on Papiers.

A handful of visual foundation models (VFMs) have recently emerged as the backbones for numerous downstream tasks. VFMs like CLIP, DINOv2, SAM are trained with distinct objectives, exhibiting unique characteristics for various downstream tasks. We find that despite their conceptual differences, these models can be effectively merged into a unified model through multi-teacher distillation. We name this approach AM-RA…

View the original paper on arXiv