RAG — Replace Re-Ranking with Pruning For Fewer Hallucinations

Prompt EngineeringAbout 4 min readJul 8, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Retrieval Augmented Generation (RAG): A framework where a language model (LLM) retrieves relevant documents and uses them as context for generating responses.
  • Context Engineering: Optimizing the context provided to an LLM to improve the quality of its responses.
  • Hallucination: When an LLM generates incorrect or nonsensical information.
  • Re-ranking: Reordering retrieved documents based on their relevance to the user query.
  • Provenance: A context pruning method that removes irrelevant sentences from retrieved chunks while preserving local context.
  • Local GPT: A local GPT implementation by the video creator that incorporates the Prance model.

Context Engineering for RAG

  • Problem: Poor context provided to LLMs results in "garbage in, garbage out," leading to inaccurate or irrelevant responses.
  • Vanilla RAG System:
    • User query retrieves documents.
    • Top K documents are passed to the LLM.
    • Manual setting of Top K can introduce noise from irrelevant chunks.
  • Re-ranking: A common technique to reduce noise by evaluating the relevance of each chunk to the query.
  • Issue with Re-ranking: Even with re-ranking, individual chunks may contain irrelevant sentences, adding noise. Partial chunks are still problematic.

Practical Example: Deepseek Model Training Cost

  • Objective: To retrieve the total training cost of the Deepseek model from the Deepseek paper.
  • Initial Attempt: Using a hybrid search mechanism (dense vectors + full-text search) with top 20 chunks retrieved, re-ranking the top 10, and feeding them into the Qwen1.5-32B model. Chunk size of 512 tokens.
  • Result: The LLM incorrectly stated the training cost as 2.788 million GPU hours instead of $5.576 million, even though the correct number was present in the paper.
  • Analysis:
    • Large chunk size (512 tokens) introduced irrelevant information.
    • The relevant chunk was not among the top 6 displayed.
    • Chunks contained both relevant information (GPU hour count) and irrelevant information (table fragments, unrelated training mechanisms).

Provenance-Based Context Pruning

  • Solution: Implement a pruning phase to remove irrelevant sentences from the retrieved chunks, focusing on the most relevant parts of the text.
  • Result: After pruning, the LLM correctly identified the training cost as $5.576 million.
  • Benefits:
    • Reduces the number of tokens fed to the LLM (e.g., from 3000 to 500 tokens).
    • Improves accuracy by focusing on relevant information.

Provenance Explained

  • Paper Reference: "Provenance: Efficient and Robust Context Pruning for Retrieval Augmented Generation" (January 2024).
  • Core Idea: Identify and remove irrelevant information from retrieved chunks while maintaining the global context of the chunk.
  • Example: For the query "Can you eat pumpkins every day?", the sentence "Pumpkin is at its peak in the fall" is deemed irrelevant and removed.
  • Local Context Preservation: Provenance considers the context of surrounding sentences when determining relevance.
  • Example: In the sentence "It may also help lower blood pressure. However, if you eat too much, you may experience diarrhea from higher dose of fiber," the phrase "it" refers to pumpkin from the previous sentences. Provenance recognizes this connection and does not discard the sentence.
  • Sentence Scoring: Provenance assigns a relevance score to each sentence, enabling its use as a re-ranking model.

Performance and Implementation

  • Performance: Provenance achieves state-of-the-art performance on benchmark datasets, both with and without a re-ranker.
  • Hugging Face: The model is available on Hugging Face, along with a detailed blog post on training and inference.
  • Implementation Steps:
    1. Install NLTK.
    2. Provide your Hugging Face token.
    3. Use the model for pruning context.
  • Example: Using the Shephard Spy example, Provenance successfully extracts the sentence "On the bottom is ground meat" from a larger context, answering the question "What goes on the bottom of shepherd spy?".
  • Compression Rate: Provenance achieves about 80% compression in some cases.
  • Integration: Can be used as a secondary re-ranking step or to replace the re-ranker entirely.

Considerations

  • License: The current Provenance model has a non-commercial license.
  • Latency: Adding Provenance may increase latency by a few seconds depending on the text volume.

Synthesis/Conclusion

The Provenance context pruning technique offers a significant improvement to RAG systems by reducing hallucination and improving accuracy. By intelligently removing irrelevant information while preserving local context, Provenance ensures that LLMs receive cleaner and more focused input, leading to better results. The technique can be integrated into existing RAG pipelines either as a secondary re-ranking step or as a replacement for traditional re-rankers. While the current model has a non-commercial license, its open availability and strong performance suggest that it will become a crucial component of future RAG systems.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.