Hands-On Vision RAG: Images, Tables & Text

Prompt EngineeringAbout 4 min readApr 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Multimodal Retrieval Augmented Generation (RAG)
  • Vision Language Models (VLMs)
  • Image Embeddings
  • Cohere Embed v4
  • Kall-Pali
  • Embedding Quantization
  • Vector Stores
  • Gemini
  • BLIP2

Multimodal RAG Systems: Processing Images, Text, and Tables

The video addresses the limitations of text-based RAG systems in handling multimodal data, specifically images and tables, within documents. It presents methods for building multimodal RAG systems capable of processing these data types, including both proprietary API-based and local solutions.

Traditional Approach vs. Direct Image Processing

  • Traditional Approach: Parses images and tables, generates captions using a vision language model, and embeds the captions as text. This approach loses contextual information and relies heavily on the VLM's accuracy and prompting.
  • Direct Image Processing (Poly): Encodes images directly using a vision encoder and creates embeddings. However, this results in multi-level embeddings, leading to high memory requirements and limited vector store compatibility.

Cohere Embed v4: Multimodal Embeddings for Improved Retrieval

Cohere's Embed v4 is introduced as a solution for multimodal search, generating fixed-size embeddings for images. Benchmarks show state-of-the-art results on vision-based retrieval, even surpassing models like Kall-Pali.

  • Workflow:
    1. Create embeddings of document images using Embed v4.
    2. Store embeddings in a vector store.
    3. Embed the user query using Embed v4.
    4. Retrieve relevant images based on embedding similarity.
    5. Pass the query and retrieved images to a multimodal model (e.g., Gemini) for answer generation.
  • Cost Considerations: Embed v4 offers dynamic embedding sizes (512 to 1536). Quantization can significantly reduce compute and storage costs while preserving performance. The video references a previous video on embedding quantization and its impact on retrieval performance.

Code Implementation: Cohere API and Local Model (Kall-Pali)

The video demonstrates a code implementation based on an example from Cohere, showcasing both their proprietary API and a local model (Kall-Pali) approach.

Cohere API Implementation

  1. Setup:
    • Install the Cohere Python package.
    • Obtain a Cohere API key.
    • Obtain a Gemini API key.
  2. Image Embedding:
    • Resize images to a maximum resolution.
    • Use embed_v4 from Cohere to generate embeddings for each image.
    • Store the embeddings (currently as a NumPy array, but can be stored in any vector store).
  3. Retrieval:
    • Embed the user query using embed_v4.
    • Calculate the dot product (cosine similarity) between the query embedding and image embeddings.
    • Retrieve the image with the highest similarity score.
  4. Generation:
    • Pass the retrieved image and the original user query to a vision language model (Gemini 2.5 Flash Preview).
    • Generate an answer based on the visual content and the query.
  5. Re-ranking (Optional): Implement a vision-based re-ranker to refine the top-k retrieved images before passing them to the VLM.

Local Model Implementation (Kall-Pali)

  1. Setup:
    • Install the bld package for running Kall-Pali models.
    • Install the pdf2image package for converting PDFs to images.
    • Obtain a Hugging Face API token.
  2. Loading the Model:
    • Use the bld package to load the Kall-Pali model.
    • Point the model to the folder containing the images to create an index.
  3. Retrieval:
    • Use the search function with the user query to compute embeddings and retrieve the top image.
  4. Generation:
    • Feed the retrieved image and the user query to a vision language model (e.g., Gemini).

Examples and Use Cases

The video provides examples using infographics from company websites. It demonstrates how the system can answer questions based on visual data, including:

  • Extracting numerical data (e.g., net profit of Nike).
  • Identifying key information (e.g., largest acquisitions of Google).
  • Performing calculations based on visual data (e.g., net profit of Tesla without interest).

Conclusion

The video highlights the potential of multimodal RAG systems for processing images, text, and tables. By combining vision-based retrieval with vision language models, these systems can answer complex questions based on visual data, overcoming the limitations of traditional text-based approaches. The video provides practical code examples using both proprietary (Cohere) and local (Kall-Pali) solutions, offering a starting point for building multimodal RAG systems. The speaker emphasizes the importance of retrieval in enterprise search and mentions his advanced RAG techniques course and advising services.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.