Forget RAG Pipelines—Build Production Ready Agents in 15 Mins: Nina Lopatina, Rajiv Shah, Contextual

AI EngineerAbout 8 min readJun 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • RAG (Retrieval Augmented Generation): A framework for enhancing LLMs with external knowledge from unstructured data.
  • Managed RAG Service: Treating RAG as a fully managed service, similar to cloud-based vector databases or LLMs.
  • Document Understanding Pipeline: The process of extracting, parsing, and structuring information from various document types (PDFs, images, tables).
  • Grounded Language Model: An LLM specifically trained to prioritize provided context and avoid generating information from its own pre-existing knowledge.
  • Groundness Check: A mechanism to verify that the claims made by the LLM are supported by the provided context documents.
  • Attribution: Providing clear references and bounding boxes to show the source of information used by the LLM.
  • Query Reformulation: Techniques to improve query understanding, including query expansion, decomposition, and multi-turn conversation handling.
  • Re-ranking: Using a model to re-order retrieved documents based on relevance to the query, often with instruction following capabilities.
  • LMUnit (Language Model Unit Testing): A framework for evaluating RAG systems using unit tests and a fine-tuned LLM as a judge.
  • MCP (Model Communication Protocol): A protocol for connecting LLMs and agents, enabling integration with tools like Claude Desktop.

1. Introduction and Team

  • Rajie Sha, Chief Evangelist of Contextual AI, introduces the workshop on simplifying RAG.
  • Nina, from Contextual AI, co-presents, bringing NLP and language modeling expertise.
  • Matthew (Platform Engineer) and John (Solution Architect) are available for technical and integration questions.
  • The core idea is to treat RAG as a managed service, avoiding the need to build and maintain custom pipelines.

2. The Value and Struggles of AI and RAG

  • AI's value is acknowledged, but scaling RAG beyond demos is challenging.
  • Issues include:
    • Difficulties in extracting information from diverse documents.
    • Maintaining accuracy at scale.
    • User dissatisfaction due to query formulation problems.
  • Contextual AI focuses on addressing these enterprise RAG challenges.
  • The founders, Dao and Aman, have a strong background in RAG research (initial RAG paper, Meta).

3. Basics of RAG

  • RAG enables understanding of unstructured enterprise data.
  • A simple RAG pipeline involves:
    • A vector database for storing document embeddings.
    • Cosine similarity for finding relevant documents.
    • An LLM for generating answers based on retrieved context.
  • The workshop aims to demonstrate a more complex and powerful RAG platform.

4. Contextual AI Platform Overview

  • The platform caters to different user levels:
    • No-code: For business users with simple document Q&A needs.
    • Developer: For those who want to orchestrate and customize their RAG pipelines.
    • Modular: For users who want to integrate specific components (e.g., extraction, re-ranking) into existing systems.
  • Contextual AI focuses solely on RAG, unlike companies with broader AI offerings.
  • The platform addresses the messiness and complexity of scaling RAG in production.

5. Building a First Agent (Nina's Demo)

  • The demo uses a notebook and the Contextual AI platform (app.contextual.ai).
  • The process involves:
    1. Loading financial statements from Nvidia and spurious correlation data.
    2. Trying out queries, including quantitative reasoning and data interpretation.
    3. Signing up on the platform and obtaining an API key.
    4. Setting up a workspace (unique name required).
    5. Installing necessary Python packages (pip installs).
    6. Creating a data store to hold the documents.
    7. Uploading the documents to the data store.
    8. Inspecting the parsed documents in the platform UI (raw text, rendered previews).
    9. Creating an agent and querying it through the UI and API.
  • Example queries:
    • "What was Nvidia's annual revenue by fiscal year 2022 to 2025?" (Demonstrates quantitative reasoning across multiple documents).
    • "When did Nvidia's data center revenue overtake gaming revenue?"
    • "What's the correlation between the distance between Neptune and the sun and burglary rates in the US?" (Tests the system's ability to handle spurious correlations).
  • The platform avoids hallucinations and focuses on data from loaded documents.
  • The agent can be customized through the admin panel (system prompt, query understanding, retrieval settings, etc.).

6. Contextual AI Platform Architecture (Rajie's Deep Dive)

  • The platform addresses key RAG challenges: extraction, accuracy, and scaling.
  • It can run in Contextual AI's SAS, a user's VPC, and offers both UI and API access.
  • The architecture includes:
    1. Document Understanding Pipeline: Extracts information from various document types (text, tables, images).
    2. Chunking: Divides documents into smaller pieces for efficient retrieval.
    3. Retrieval: Uses a mixture of retrievers (BM25, embeddings) and a state-of-the-art re-ranker.
    4. Grounded Language Model: Generates answers based on retrieved context, avoiding external knowledge.
  • The platform emphasizes academic benchmarks and customer data for performance evaluation.
  • Use cases extend beyond chatbots, including website Q&A (Qualcomm), automated workflows, and integration with other tools.

7. Preventing Hallucinations

  • Hallucinations can be prevented by:
    • Retrieving clean information through good extraction and retrieval.
    • Using a grounded language model that refers back to the context.
    • Implementing checks for groundness and attribution.

8. Document Understanding Pipeline Details

  • The pipeline involves:
    1. Adding metadata to the document.
    2. Layout analysis (image-only, images with text, tables, charts).
    3. OCR for image-based documents.
    4. Image captioning for multimodal charts and graphs.
    5. Table extraction mode for high-quality table parsing.
    6. Structure analysis (section headers, subsections).
    7. Creating a markdown/JSON version of the document.
    8. Chunking the document into smaller pieces.
    9. Injecting metadata into chunks.
    10. Setting bounding boxes for attribution.
  • The platform offers a playground for testing individual components.

9. Query Path Details

  • The query path involves:
    1. Query translation (if needed).
    2. Query reformulation (expansion, decomposition, multi-turn handling).
    3. Semantic and lexical search (hybrid search).
    4. Data store filter (to focus on relevant documents).
    5. Retrieval (default: 100 chunks).
    6. Re-ranking (using an instruction-following re-ranker, default: top 15).
    7. Filtering (optional).
  • The re-ranker can be instructed to prioritize documents based on specific criteria (e.g., recency).

10. Generation Details

  • The grounded language model is specifically fine-tuned to respect the provided context.
  • It can distinguish between facts and commentary, allowing users to focus on factual information.
  • The groundness check verifies that each claim made by the LLM is supported by the context documents.
  • The UI highlights ungrounded claims in yellow.

11. Evaluation with LMUnit (Nina's Demo)

  • LMUnit is used for natural language unit testing of RAG systems.
  • It involves:
    1. Generating responses from the RAG system.
    2. Creating unit tests that ask specific questions about the responses.
    3. Using the LMUnit model to evaluate the unit tests (scoring on a 1-5 scale).
  • Example unit tests:
    • Does the response accurately extract specific numerical data?
    • Does the agent properly distinguish between correlation and causation?
    • Are multi-document comparisons performed correctly?
    • Are potential limitations or uncertainties acknowledged?
    • Are quantitative claims properly supported?
    • Does the response avoid unnecessary information?
  • The results can be visualized using polar plots to identify areas for improvement.
  • Users can improve the agent by updating the system prompt or other settings based on the unit test results.

12. Connecting to Claude Desktop via MCP (Rajie's Demo)

  • The Contextual AI RAG agent can be integrated with Claude Desktop using MCP.
  • The process involves:
    1. Cloning the Contextual AI MCP server repository.
    2. Modifying the server.py file to point to the desired RAG agent.
    3. Configuring Claude Desktop to point to the local MCP server.
  • This allows users to access the RAG agent's knowledge within the Claude Desktop environment.

13. Pricing and Takeaways

  • Individual components (parser, re-ranker, generate, LMUnit) are priced on a consumption basis (pay-per-token).
  • The RAG platform will also be consumption-based (number of documents ingested and queries).
  • Provisioned throughput is available for enterprise customers with specific latency or query volume requirements.
  • Key takeaways:
    • RAG can be treated as a managed service.
    • The platform offers a range of components and customization options.
    • Users can get started with the provided code and API keys.

14. Q&A Highlights

  • Who is doing the prompt engineering? Technical people are needed for advanced settings, while business users can experiment with simpler aspects like the system prompt.
  • Approach for different data types? Contextual AI has a customer machine learning engineering team to help with this.
  • Integrating with existing agents? Yes, the components can be integrated via JavaScript SDK.
  • How far will the $25 credit go? Try it out, and contact sales if needed.
  • Data sovereignty? VPC installation is supported, with Snowflake integration. AWS GovCloud is not yet supported.
  • Performance with millions of documents? The platform scales to handle large document volumes.
  • Repeatability of LMUnit scoring? The scoring is fairly deterministic with a random seed for repeatability.
  • Integration with Microsoft Copilot? Not directly integrated, but the API allows custom integrations.
  • Challenges in RAG? Document understanding, scalability, and structured data handling.
  • Public-facing web search? A managed solution for ingesting customer websites is on the roadmap.
  • Frequently updated content? Continuous ingestion pipelines are available.
  • Document-level permissions? An entitlements layer is being added to the platform.
  • Breakthroughs in RAG? Improvements in re-rankers and vision language models.
  • Replacing document parsers? The parser module aims to be a replaceable component.
  • Domain-specific language? Fine-tuning the grounded language model can help.
  • HIPAA compliance? Contextual AI is HIPAA certified.
  • Conflicting information in documents? The language model attempts to reason out the differences, but metadata can help prioritize the correct answer.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.