Key Concepts
AI Vision, LLMs (Large Language Models), Computer Vision, Edge Computing, Vision Evals (ImageNet, COCO), Big Pre-training, Visual Intelligence, Vision Language Models (VLMs), Convolutional Models, Transformers, Object Detection, Domain Adaptability, Few-Shot Learning, Zero-Shot Learning, Embeddings, Data Sets (Objects 365, RF100-VL), Metrics (mAP - mean Average Precision).
The State of AI Vision: Challenges and Opportunities
Peter Robisho, ML Lead at Rooflow, discusses the current state of AI vision, highlighting its importance and the gap between human and computer vision. He argues that vision is crucial for systems interacting with the real world, but faces unique challenges compared to language models, such as latency requirements and the need for edge computing.
Saturated Vision Evals and Underutilized Pre-training
Robisho argues that current vision evaluations like ImageNet and COCO are saturated, primarily measuring pattern matching rather than true visual intelligence. This leads to vision models not leveraging big pre-training as effectively as language models. He points out that while large language models can be unleashed on the internet to gain significant intelligence, the best vision models are often those specifically trained to solve the eval, limiting their broader applicability.
Vision Models Lack "Smartness"
The core argument is that vision models are not as "smart" as they could be. As an example, even advanced models like Cloud 3.5 and Cloud 4 fail to accurately tell time from an image of a watch, demonstrating a disconnect between conceptual understanding and actual visual perception.
Example: Cloud 3.5 and 4 failing to read the time on a watch image, even when the watch displays the common time of 10:10.
The MMVP Data Set: Measuring the Inability to "See"
The MMVP data set is introduced as a tool to measure the inability of LLMs to "see." It presents seemingly obvious visual questions that models like ChatGPT-4o often get wrong, even hallucinating details to support incorrect answers.
Example: The school bus direction question in the MMVP data set, where the model incorrectly identifies the front or back of the bus and fabricates supporting details.
The data set is constructed using image pairs that are close in CLIP space (vision-language model) but far in DinoV2 space (pure vision model). This highlights the limitations of vision-language pre-training, as CLIP struggles to differentiate images that a pure vision model can easily distinguish. The speaker argues that the captions used in CLIP training often lack the specific details needed to differentiate such images.
Vision-Only Pre-training: A Promising Avenue
DinoV2 is presented as an example of successful vision-only pre-training. Visualizations of its PCA features reveal that it can identify not only obvious features like the mask of a dog but also more nuanced segments and analogous features across different objects (e.g., dog legs and human legs).
The Convolutional vs. Transformer Divide in Object Detection
The speaker addresses the question of why existing pre-trained vision models aren't being leveraged effectively, particularly in object detection. He attributes this to the historical dominance of convolutional models, which don't benefit from big pre-training as much as transformers.
Data: Comparison of YOLOv8 (convolutional) and LWDETR (transformer) on COCO with and without pre-training on Objects 365. YOLOv8 gains only 0.2 mAP, while LWDETR gains 5-7 mAP.
The scale of pre-training in vision is also significantly smaller than in language. Pre-training on Objects 365 (1.6 million images) is considered large in vision but would be a small challenge data set in the LLM world.
RF-DETR: Rooflow's Solution
Rooflow introduces RF-DETR, a new model that leverages the DinoV2 pre-trained backbone for real-time object detection. This is presented as an answer to the lack of effective utilization of big pre-training in visual models.
Data: RF-DETR achieves decent improvements on COCO compared to LWDETR.
R100-VL: A New Benchmark for Visual Intelligence
Rooflow introduces R100-VL, a new data set designed to measure the domain adaptability and visual intelligence of models more comprehensively than COCO. It consists of 100 diverse object detection data sets curated from Rooflow Universe.
Features of R100-VL:
- Diverse camera poses (e.g., aerial views)
- Different visual imaging domains (e.g., microscopes, X-rays)
- Contextualized class names (e.g., "block" in the context of volleyball)
- Visual descriptions and instructions for object identification
The data set is also a vision-language benchmark, allowing for evaluation of models' ability to contextualize class names and generalize across different domains.
Data: YOLOv8 trained on 10 examples per class on R100-VL outperforms state-of-the-art VLMs like Quen V2.5 72B, highlighting the VLMs' weakness in visual generalization.
The Importance of Visual Generalization
The speaker emphasizes that while VLMs are good at generalizing out-of-distribution in the linguistic domain, they struggle with visual generalization. R100-VL aims to drive research in this area and ensure that the visual aspects of VLMs are not neglected.
Q&A Highlights
- RF-DETR can be run and fine-tuned at the edge.
- R100-VL is publicly available on rf100vl.org and Hugging Face.
- Rooflow encourages researchers to contribute labeled data back to the community.
- The data set includes canonical 10-shot splits with class names, annotator instructions, and visual examples.
- Specialist models currently perform best on R100-VL, but the goal is to develop generalist models that can effectively leverage the provided information.
Conclusion
The presentation concludes that while significant progress has been made in AI vision, challenges remain in leveraging big pre-training and achieving true visual intelligence. The R100-VL data set is presented as a valuable tool for driving research in visual generalization and ensuring that VLMs can effectively "see" and understand the visual world. The key takeaway is that current vision models are not as "smart" as they could be, and there is a need for more sophisticated approaches to pre-training and evaluation.
AI summaries can miss context or contain errors. Check important details against the original video.