Key Concepts
RAG (Retrieval Augmented Generation), LLM (Large Language Model) evaluation, Ragas (framework for evaluating RAG systems), context precision, context recall, answer correctness, answer relevancy, faithfulness, vector embeddings, vector stores, Langchain, Ollama integration.
Ragas: Evaluating RAG Systems in Python
Introduction
The video introduces Ragas, a Python package designed to evaluate the quality of LLM responses in RAG systems. Ragas allows for quantitative assessment of LLM responses based on retrieved context, addressing the need for concrete metrics in evaluating dynamic and flexible language model outputs.
Setup and Dependencies
The video provides a step-by-step guide to setting up the environment for using Ragas:
- Install JupyterLab:
pip install jupyterlaborpip3 install jupyterlab - Run JupyterLab:
jupyter labin the desired directory. - Install Dependencies:
pip install faiss-cpuorpip install faiss-gpupip install openaipip install tiktokenpip install datasets(Hugging Face)pip install ragaspip install langchainpip install langchain-community
- A
requirements.txtfile will be provided in the GitHub repository for easy installation usingpip install -r requirements.txt.
Building a Basic RAG System (Optional)
The video demonstrates building a basic RAG system as a foundation for evaluation. This involves:
- Knowledge Base: Creating a set of documents (e.g., "Paris is the capital of France," "Mike loves the color pink").
- Embedding Model: Using
text-embedding-3-smallfrom OpenAI to embed text into a vector space. - Vector Store: Using FAISS (Facebook AI Similarity Search) to store and index the embeddings.
- Retrieval: Implementing a function to retrieve the most relevant context based on a user query using similarity search.
- Answer Generation: Using GPT-4 to generate an answer based on the retrieved context. The prompt instructs the model to "answer the user question only with facts found in the context."
Ragas Evaluation Metrics
Ragas focuses on evaluating the RAG system using several key metrics:
- Answer Correctness: How accurate is the answer provided by the LLM?
- Answer Relevancy: How relevant is the answer to the question asked?
- Faithfulness: How factually consistent is the response with the retrieved context? It measures how much the answer is based on the context. The formula is:
Faithfulness Score = (Number of claims in response supported by context) / (Total number of claims in response). - Context Precision: How relevant is the retrieved context to the question?
- Context Recall: How well does the retrieved context cover all the information needed to answer the question?
Evaluation Process with Ragas
The evaluation process involves:
- Crafting a Dataset: Creating a dataset of questions, ground truth answers (reference answers), retrieved context, and generated answers.
- Using the
evaluateFunction: Importing theevaluatefunction from Ragas and specifying the desired metrics. - Passing the Dataset: Providing the evaluation dataset to the
evaluatefunction. - Loading API Key: The code expects an
.envfile with the OpenAI API key defined asOPENAI_API_KEY=your_api_key.
Interpreting Evaluation Results
The video demonstrates how to interpret the evaluation results by manipulating the context and responses:
- High Scores: When the correct context is provided, the scores for correctness, relevancy, and faithfulness are high.
- Wrong Context: Providing irrelevant context leads to low correctness and relevancy scores, but potentially high faithfulness if the model avoids making unsupported claims.
- Correct Answer, Wrong Context: Hardcoding the correct answer with irrelevant context results in high correctness and relevancy, but low faithfulness, context precision, and context recall.
Ollama Integration
The video briefly touches on integrating Ragas with Ollama, a tool for running LLMs locally. This involves:
- Importing Necessary Modules: Importing modules from
langchain_community,langchain_llama, andragas. - Loading a Local LLM: Using
ChatOllamato load a local model (e.g., "Gwen 3 4 billion parameters"). - Wrapping the LLM: Using
LangchainLLMwrapper to make the LLM compatible with Ragas. - Wrapping the Embedding Model: Using
LangchainEmbeddingsWrapperto make the embedding model compatible with Ragas. - Running Evaluation: Running the evaluation process as before, but using the locally hosted LLM.
Conclusion
Ragas provides a valuable framework for evaluating RAG systems by quantifying the quality of LLM responses and the relevance of retrieved context. The video emphasizes the importance of understanding the individual metrics and how they relate to each other. While the video encountered some issues with Ollama integration due to resource constraints, it demonstrated the potential for using Ragas with locally hosted models. The presenter encourages viewers to provide feedback on the type of content they enjoy most, as the channel covers a range of topics including machine learning, Python tutorials, and backend development.
AI summaries can miss context or contain errors. Check important details against the original video.





