They Just Built a New Form of AI, and It’s Better Than LLMs

By AI Revolution

Share:

VLJA: A New Paradigm in Vision-Language AI

Key Concepts:

  • VLJA (Vision Language Joint Embedding Predictive Architecture): A novel AI architecture focused on predicting meaning directly through embeddings, rather than generating text tokens.
  • Embeddings: Continuous vector representations of meaning, capturing semantic information without relying on specific wording.
  • Vision Encoder: Component responsible for converting visual input (images/video) into visual embeddings. (VJEPA 2 is used in this implementation)
  • Predictor: Core module that maps visual and text inputs to a semantic embedding representing the answer. (Initialized from Llama 3.21b, without causal masking)
  • Y Encoder/Decoder: Components for encoding/decoding text into/from embeddings, primarily used during training and optional text output.
  • Selective Decoding: A technique enabled by embedding-based prediction, where text is only generated when significant semantic changes are detected.
  • Semantic Space: A structured representation of meaning where similar concepts are clustered together.

1. The Limitations of Traditional Vision-Language Models (VLMs)

Current VLMs, like those used for image captioning and visual question answering, operate by generating text token by token. While effective, this approach has inherent inefficiencies. The models are forced to learn variations in phrasing, word choice, and sentence structure that don’t affect the underlying meaning. This consumes significant training resources. Furthermore, token-by-token generation introduces latency, making these models unsuitable for real-time applications like robotics or live video analysis where continuous understanding is crucial. The semantics are only revealed after the entire text sequence is generated.

2. Introducing VLJA: Predicting Meaning Directly

VLJA, developed by MetaPare, represents a departure from this token-based approach. Instead of predicting words, VLJA predicts embeddings – continuous vectors representing meaning directly. This means the model learns to map visual input and a text query directly to a semantic representation of the answer, bypassing the need for sequential text generation during the core processing stage. As Yan Lun states, this feels like “what comes after the LLM era, not just another upgrade on top of it.”

3. VLJA Architecture and Components

VLJA consists of four key components:

  • Visual Encoder: Uses VJEPA 2 (a self-supervised vision transformer with 304 million parameters, kept frozen during training) to convert images or video frames into visual embeddings. These are continuous vectors, analogous to visual tokens.
  • Predictor: The central module, built using transformer layers initialized from Llama 3.21b without causal masking (allowing full attention between vision and text). It predicts the answer embedding based on visual embeddings and the text query.
  • Y Encoder: Encodes the target text answer into an embedding during training, serving as the learning target. This embedding represents the meaning of the answer, not the exact wording.
  • Y Decoder: Used only at inference time to convert the predicted embedding back into readable text when necessary. It doesn’t participate in training.

4. Training Methodology and Semantic Space Construction

VLJA is trained in a loop: a visual input and query are provided, the Y encoder transforms the correct answer into an embedding, and the predictor attempts to recreate that embedding from the visual input and query. The loss is calculated directly in embedding space. Crucially, the training process encourages the model to build a structured semantic space where similar answers cluster together and different answers remain distinct. This is achieved by pulling predicted meanings towards the correct meaning while maintaining separation between different answers. This contrasts with token space, where multiple valid answers can be vastly different sequences of symbols.

5. Performance Comparison: VLJA vs. Token-Based Models

Researchers conducted a controlled experiment comparing VLJA to a standard token-based model, keeping all parameters constant except the prediction target. The VLJA version, with roughly half the trainable parameters (500 million vs. 1 billion), demonstrated superior performance.

  • Video Captioning: After 5 million samples, VLJA achieved 14.7 CR compared to the token-based model’s 7.1 CR.
  • Classification Accuracy: VLJA reached approximately 35% top-five accuracy, versus roughly 27% for the baseline.
  • Efficiency: The performance gap widened with more training data, indicating a structural advantage for VLJA.

6. Inference and Selective Decoding for Real-Time Applications

VLJA’s embedding-based approach enables selective decoding. Instead of generating text at fixed intervals, the system monitors changes in the semantic embeddings. Text is only decoded when a significant semantic shift is detected. Testing on long procedural videos (average 6 minutes, 143 annotations) showed that VLJA could achieve comparable performance to uniform decoding (1 decode/second) by decoding only once every 2.85 seconds – a 2.85x reduction in decoding operations. This is particularly beneficial for latency-sensitive applications like smart glasses and robotics.

7. Versatility and Task Agnostic Architecture

VLJA’s architecture is versatile, capable of handling generation, classification, retrieval, and discriminative visual question answering without task-specific heads or separate models.

  • Classification: Candidate labels are encoded into embeddings and compared to the predicted embedding.
  • Text-to-Video Retrieval: Text queries are encoded, and videos are ranked by similarity.
  • Discriminative VQA: Candidate answers are embedded, and the nearest one is selected.

VLJA (1.6 billion parameters, trained on 2 billion samples) outperformed CLIP, SIGLIP 2, and Perception Encoder on average across eight video classification and eight text-to-video retrieval datasets, even surpassing models trained on significantly larger datasets (up to 86 billion samples).

8. Advanced Evaluation: World Modeling and Embedding Quality

  • World Modeling: VLJA SFT (Supervised Fine-Tuning) achieved 65.7% accuracy in a world modeling task (identifying the action causing a transition between images), surpassing larger vision-language models and even frontier language models like GPT-4, Claude 3.5, and Gemini 2.
  • Embedding Quality: Analysis using benchmarks like Sugar Crate++ and Vizla demonstrated that VLJA’s Y encoder produces a sharper and more structured semantic space, capable of detecting subtle semantic changes.

9. Ablation Studies and Key Findings

Ablation studies revealed:

  • Removing the large caption-based pre-training stage significantly degraded performance.
  • Strengthening the Y encoder improved alignment.
  • Larger predictors enhanced performance, particularly for VQA.
  • Visually aligned text encoders boosted retrieval and classification.

10. Conclusion: A Shift from Language to Meaning

VLJA represents a significant shift in vision-language AI, moving the focus from generating text to predicting meaning directly. While token-based models remain suitable for tasks requiring deep reasoning and complex planning, VLJA excels in perception-heavy problems, particularly those involving video, real-time input, and continuous understanding. This approach reduces computational cost, improves efficiency, and unlocks new possibilities for real-world applications. As stated in the video, this work feels like “more than just another model iteration” and potentially signals “what comes after the LLM era.”

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video