Key Concepts
- Contextualized Chunk Embeddings
- Retrieval Augmented Generation (RAG)
- Chunking Strategies
- Dense Embeddings
- Late Chunking
- Similarity Computation (No Interaction, Full Interaction, Latent Interaction)
- Cross Encoders
- Bi-Encoders
- Token-Level Embeddings
- Voyage Embedding Models
- Quantization
Standard RAG System and its Limitations
The video begins by addressing the fundamental challenge of chunking data for Retrieval Augmented Generation (RAG) systems. A standard RAG system involves:
- Document Chunking: Dividing a document into smaller, manageable chunks.
- Dense Embedding: Converting each chunk into a dense vector representation using an embedding model.
- Full Text Search (Optional): Indexing the chunks for faster retrieval.
- Retrieval: Using the embeddings to find the most relevant chunks for a given query.
The key problem with this approach is that each chunk is treated in isolation, lacking context from surrounding chunks or the overall document. This can lead to poor retrieval performance.
Approaches to Add Context
Several approaches aim to address the context issue:
1. Contextualized Retrieval Pre-processing (Anthropic's Approach)
- Process: An LLM is used to generate context for each chunk by feeding it the chunk along with the entire document or surrounding chunks. This context is then added to the chunk.
- Embedding & Retrieval: The rest of the embedding process and full-text search index remain the same. Retrieval is performed on these contextualized chunks.
- Drawback: This method is computationally expensive, especially for large documents, due to the numerous LLM calls required.
2. Late Chunking
- Process: Utilizes embedding models with long context windows (e.g., 8,000 or 32,000 tokens). The entire document is fed into the embedding model, generating token-level embeddings. These token embeddings are then combined (pooled) to create chunk-level embeddings.
- Advantage: Aims to preserve both local and global context because the embedding process happens at the document level before chunking.
- Technical Terms: Token-level embeddings, pooling operation.
Similarity Computation Methods
The video then delves into different methods for computing similarity between queries and documents:
1. No Interaction
- Process: Embeddings for documents/chunks and queries are computed independently. Similarity (e.g., cosine similarity) is calculated directly between the resulting embedding vectors.
- Drawback: The embedding model doesn't see the document and query together, leading to potentially poor retrieval.
2. Full Interaction (Cross Encoders)
- Process: Chunks and queries are fed into the same model simultaneously to determine their relatedness.
- Advantage: More accurate than "No Interaction."
- Drawback: Computationally intensive and typically used as a secondary step in a re-ranker.
3. Latent Interaction (ColBERT)
- Process: Token-level embeddings are computed for both documents and queries. Similarity is calculated at the token level rather than the vector level.
- Advantage: More fine-grained than dense embeddings, leading to better retrieval.
- Drawback: Requires more storage and compute due to the multi-vector representation.
- Technical Terms: Multi-vector representation.
Contextualized Chunk Embeddings: A Deep Dive
The core of the video focuses on contextualized chunk embeddings:
- Process:
- Document Chunking: The document is divided into chunks using a chosen strategy.
- Contextualized Embedding: Each chunk, along with the original document, is passed through a contextualized chunk embedding model. This model adds global and local context to each chunk embedding.
- Query Time:
- Chunk-Level Retrieval: The user query is used to retrieve relevant chunks.
- Contextual Information: Each retrieved chunk includes associated contextual information at both local and global levels (e.g., the document it came from and related chunks within that document).
- Difference from Late Chunking: In late chunking, embeddings are computed at the document level before chunking. In contextualized chunk embeddings, the document is chunked first, and then embeddings are computed while preserving global information.
Voyage Embedding Models and Benchmarks
The video highlights Voyage embedding models (now part of MongoDB) as a provider of high-quality embedding models. Benchmarks are presented, showing that contextualized chunk embeddings outperform other embedding models on various datasets at both chunk and document levels.
- Key Observation: While standard dense embedding models tend to improve with larger chunk sizes, the retrieval accuracy of the Voyage Context 3 model decreases as chunk size increases. This suggests that smaller chunk sizes are preferable for this model.
- Quantization: The model supports quantization, allowing for reduced storage costs with minimal impact on retrieval accuracy. Binary quantization (2048 vector size) offers a good balance between storage and accuracy, with only a 2-3% drop in retrieval accuracy compared to the full-size model.
Pricing and Code Example
- Pricing: Voyage Context 3 embeddings cost around $0.18 per million tokens, with the first 200 million tokens free. This is compared to Gemini Embedding 001 at $0.15 per million tokens.
- Code Example: A Google Colab notebook (link provided) demonstrates the implementation of contextualized chunk embeddings. The example involves chunking documents (AWS S3 documentation, Azure documentation, and FinTech compliance requirements) and comparing the retrieval performance of standard dense embeddings versus contextualized chunk embeddings.
- Chunking Strategy: Sentence-level chunking is used in the example. The video suggests adapting the chunking strategy based on the document structure (e.g., paragraph-level or section-by-section chunking).
- Results: The code example shows that contextualized chunk embeddings can retrieve more relevant chunks, especially when the query requires understanding the context of the document.
Conclusion
The video concludes by emphasizing that the choice of RAG system design depends on the specific data and retrieval requirements. Contextualized chunk embeddings are presented as a valuable tool to have in one's arsenal, offering improved retrieval accuracy by preserving both local and global context. The speaker encourages viewers to experiment with the technique and share their experiences and questions in the comments.
AI summaries can miss context or contain errors. Check important details against the original video.





