The Only Embedding Model You Need for RAG

Prompt EngineeringAbout 4 min readJul 2, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Multimodal embeddings, multilingual embeddings, text retrieval, code retrieval, image retrieval, Retrieval Augmented Generation (RAG), dense embeddings, multi-vector representations, LoRA adapters, late chunking, cosine similarity, vision language models.

JA Embedding V4: A Universal Embedding Model

The video focuses on the JA Embedding V4, a new embedding model that is presented as a significant advancement due to its multimodal and multilingual capabilities. It can be used for text, code, and image retrieval tasks using the same model weights, which are available on Hugging Face.

The Role of Embeddings in Retrieval Tasks

Embeddings are crucial for retrieval tasks, especially in RAG systems. Real-world documents often contain images, text, and tables, sometimes in complex layouts. Traditionally, images were converted to text descriptions, leading to information loss.

Alternative Approaches: Call Poly and Colbert-Style Representations

Call Poly converts PDF pages into images and uses an image encoder to project their representation into the same space as text embeddings. This allows using the same model for both text and image queries. Call Poly is inspired by Colbert-style multi-vector representations.

  • Dense Embeddings: Input text is converted into embeddings for each token, then aggregated into a single vector. The output vector size remains constant regardless of the input text size.
  • Multi-Vector Representations: Embeddings are created for each token, resulting in more accurate representations but significantly increasing storage requirements. Call Poly applies a similar approach to vision, dividing pages into patches and computing embeddings for each patch.

Embed Form: A Fixed-Size Multimodal Embedding

CO's Embed Form is a state-of-the-art multimodal embedding approach that provides a fixed-size vector, reducing storage costs compared to Colali.

JA Embedding V4: Combining Multiple Approaches

The JA Embedding V4 combines several ideas into a single model, including LoRA adapters for different tasks.

Nvidia's Multimodal RAG Nemo Retriever

Nvidia released a multimodal RAG Nemo Retriever based on a Llama 3.2 model. It uses a vision encoder to encode images but uses the same embedding space for both text and images. This model is trained on 1 billion parameters.

Architecture of JA Embedding V4

The architecture of JA Embedding V4 is described as fascinating.

  • Input: Processes both text and images.
  • Vision Decoder: Images are fed into a vision decoder based on quen 2.5, a vision language model with 3.8 billion parameters.
  • Language Model Decoder: The same language model decoder is used for both text and images.
  • LoRA Adapters: Includes LoRA adapters for retrieval, code search, and classification. The model selects the appropriate LoRA adapter based on the specified task.
  • Embedding Vector Generation: Can generate either dense embeddings (fixed-size vector) or multi-vector representations (token-level embeddings).
  • Metrushka Embedding Models: Supports generating output embeddings of different sizes, similar to OpenAI's text embedding large model. This allows truncating the output embedding size to save on cost and speed. The model supports dimensions from 128 to 24,048.
  • Token Limit: Supports up to 32,000 tokens, making it suitable for late chunking.

Late Chunking Explained

Late chunking involves generating embeddings for the entire document (using a model that supports long context) and then chunking the document after generating the embeddings. Token-level embeddings are then pulled for each specific chunk, preserving global information.

Practical Considerations and Examples

  • Model Size: The JA Embedding V4 is a relatively large model for an embedding model.
  • Image Support: Supports processing up to 20-megapixel images.
  • Language Support: Supports 29 different languages.
  • Hugging Face: Model weights are available on Hugging Face.
  • Resource Requirements: Some examples may not run on a single T4 GPU due to the model's size.

Example Notebook Demonstration

The example notebook demonstrates using the model for:

  • Multilingual Retrieval: The notebook uses text in English, Spanish, Japanese, Portuguese, German, and Arabic, along with three images inspired by Star Wars. The model correctly identifies the Star Wars-related image as the closest match for all queries.
  • Task Specification: When embedding the query, the task (e.g., retrieval) must be specified to select the appropriate LoRA adapter.
  • Cosine Similarity: Cosine similarity is used to measure the similarity between text and image embeddings.
  • Text Matching: The model can be used for text matching and topic clustering.
  • Code Retrieval: The same model can be used for code retrieval tasks.
  • Multi-Vector Representation: Demonstrates using the model for multi-vector representation, but this may require more VRAM.

Local GPT Vision Repo

The video recommends checking out the local GPT vision repo for an out-of-the-box multimodal RAG system.

Conclusion

The JA Embedding V4 is a versatile and powerful embedding model that combines multiple advanced techniques. Its multimodal, multilingual, and adaptable nature makes it a promising tool for various retrieval tasks. The video emphasizes its potential for RAG systems and highlights its unique features, such as LoRA adapters, variable output embedding sizes, and support for long contexts. The presenter also offers consulting services for companies working on RAG and search systems.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.