Key Concepts
- KAG (Cache Augmented Generation): A retrieval technique where documents are cached and augmented with user queries for LLM processing.
- RAG (Retrieval Augmented Generation): A system where data is ingested into a vector database, and relevant chunks are retrieved to augment LLM prompts.
- LLM (Large Language Model): AI models like OpenAI's GPT, Anthropic's Claude, and Google's Gemini.
- Vector Database: A database that stores data as vectors, enabling similarity searches. Examples include Quadrant, Pinecone, and Supabase.
- Embedding Model: A model that transforms text into numerical vector representations.
- Context Window: The amount of text an LLM can process in a single prompt.
- Prompt Caching: Storing prompts and their corresponding outputs to speed up future requests.
- Tokens: Units of text used by LLMs for processing.
- TTL (Time To Live): The duration for which cached data is stored.
KAG vs. RAG: A Detailed Comparison
1. Introduction to KAG and RAG
The video introduces KAG (Cache Augmented Generation) as a retrieval technique gaining traction due to the increasing context window lengths of LLMs. It contrasts KAG with the established RAG (Retrieval Augmented Generation) approach, demonstrating their implementation in n8n using OpenAI, Anthropic, and Google Gemini models.
2. KAG Demo with Google Gemini
- Process:
- A 179-page PDF document (Formula 1 technical regulations) is loaded from Google Drive.
- The text is extracted from the PDF.
- The file is sent to Google Gemini's cloud server to cache the contents.
- A unique cache ID is obtained, indicating the document is stored on Google Cloud servers.
- A query ("What are the rules around the gurnie?") is sent to Gemini along with the cache ID.
- Gemini retrieves the cached document and uses it to answer the query.
- Key Point: The entire document is available in the cache, allowing Gemini to provide a thorough answer.
- Technical Detail: The query is augmented with the cached document on the server side.
3. RAG Demo with OpenAI
- Process:
- The same 179-page PDF is loaded and its text extracted.
- The text is sent to a Quadrant vector store, where it's chunked and embedded.
- The query ("What are the rules around the gurnie?") is sent to the vector store.
- The vector store retrieves relevant chunks based on similarity.
- The retrieved chunks and the query are sent to OpenAI's GPT model.
- Key Point: The answer is based on a limited number of retrieved chunks, potentially missing context.
- Technical Detail: The document is split into 646 chunks, and the top K (initially 10, then increased to 30) most relevant chunks are retrieved.
4. Theory of RAG
- Ingestion Stage:
- Documents are split into chunks of specific sizes.
- Chunks are transformed into numerical vector representations using an embedding model.
- Vectors are stored in a vector database (e.g., Quadrant, Pinecone, Supabase).
- Maintenance:
- Requires a system to update or delete old data and vectors when documents change.
- Query and Retrieval Stage:
- The user's query is transformed into a vector using an embedding model.
- The vector database is queried to find the most similar vectors.
- The top K most similar chunks are retrieved and sent to the LLM along with the query.
- Flaws of RAG:
- Limited context due to chunking.
- Retrieval of irrelevant chunks.
- Mitigation Techniques:
- Contextual retrieval to improve embeddings.
- Reranking to reshuffle chunks.
5. KAG Architectures
- OpenAI and Anthropic (Prompt Caching):
- The document is sent with each query to the LLM.
- The LLM caches the prompt and document on the server side.
- Follow-up questions require sending the document again.
- Google Gemini (Cached Contents API):
- Documents are uploaded to Gemini's cloud server and assigned a unique cache ID.
- The cache ID is stored in a database.
- When a query is submitted, the cache ID is retrieved and sent to Gemini along with the query.
- Gemini loads the cached document and generates a response.
- Key Differences:
- OpenAI and Anthropic rely on prompt caching, while Gemini uses a dedicated caching API.
- Gemini offers more control over cache management and duration.
6. KAG Demo with OpenAI
- Process:
- The document is downloaded and its text extracted.
- The text and query are sent to OpenAI's GPT model.
- OpenAI caches the prompt and document.
- Key Point: OpenAI's prompt caching is easy to set up but offers limited control.
- Technical Detail: The system prompt must be structured with static content at the beginning and dynamic content (the query) at the end to ensure proper caching. Prompts are generally cached for 5-10 minutes.
7. KAG Demo with Anthropic
- Process:
- The document is downloaded and its text extracted.
- The text and query are sent to Anthropic's Claude model using the HTTP module.
- A
cache_controlparameter is included in the API call to enable caching.
- Key Point: Anthropic requires explicit cache control parameters.
- Technical Detail: The
cache_controlparameter is necessary for Anthropic to utilize its caching mechanism. The demo encountered rate limiting issues, highlighting a potential drawback of KAG.
8. KAG Demo with Google Gemini (Step-by-Step)
- Process:
- The workflow identifies new files in a folder and downloads them.
- The file type is determined (PDF or Google Doc).
- The text is extracted from the file.
- The token length is estimated (characters / 4).
- The file is encoded to base64.
- The encoded data, system instruction, and TTL (Time To Live) are sent to Gemini's cached contents endpoint.
- The endpoint returns a cache ID, which is stored in a database.
- When a query is submitted, the cache ID is retrieved from the database and sent to Gemini along with the query.
- Key Points:
- Gemini requires a separate workflow for uploading and caching documents.
- Cache IDs must be persisted in a database.
- A system is needed to manage cache expirations.
- Technical Details:
- Gemini's context caching requires documents to be at least 32,000 tokens.
- The TTL parameter specifies the duration for which the cache is stored.
- Gemini 2.0 will soon have cache enabled.
9. Comparison: RAG vs. KAG
| Feature | RAG "Accuracy vs. Relevance" | KAG produces higher accuracy because the entire thing is in context. RAG is limited by the problem of chunks being independent of each other. and I'll be comparing it against the Old Reliable rag to see what's a better fit for what situation so in your standard rag system there are two stages to it the first stage is to ingest or to import the data into the vector database and what I mean by that is you have a document or a series of documents let's say in a Google drive folder and then you run a job to split those documents into chunks and you can be quite specific around the size of those chunks and how you chunk a document then those independent chunks are sent into an embedding model and what this embedding model does is it transforms those chunks of text into numerical Vector representations of the information which are then stored and plotted within a vector database in this case I'm using quadrant but this could be pine cone or subab base so this is an important part of a rag system you need to get the data into the vector database and then you also need to maintain it within the database so if the data changes these Reg regulations for example if there's a new version of the regulations you'll need a system to delete out the old ones and and upsert or import the new ones and if data has been deleted you'll also need a way to actually delete out the vectors from the database so there can be quite a lot of Maintenance to this system design but once you have your vector database with all of your documents represented through these kind of dense Rich vectors The Next Step then is to query and retrieve data in this scenario the user submits a message or a query and two things happen that query is then sent through an embedding model similar to what happened here to generate these kind of dense numerical representations of that message those vectors those numbers are then sent into the vector database to query it because what we want to do is find the most similar vectors or numbers to these ones and the most similar ones are then returned in this top K that you see here these are different chunks or vectors essentially and they're sent into the llm at the same time the query itself is sent into the llm so in my example earlier of what are the rules of an F1 gurny that query is sent into the llm but then also all of the relevant chunks about F1 gurns are sent in as well so the llm has all of the information it needs to generate a response and rag is a great system design but it's not without its flaws so you saw earlier KAG produced a better output than rag for example because KAG had much more information in its context to actually generate a response there it was only the top 10 and then I changed it to the the top 30 results so it's only really seen a snippet of the document to actually form the basis of the answer another issue is that if you have a large array of documents sometimes the chunks that are retrieved are actually not relevant and it might be that they're maybe close matches to the specific query that were asked but it's in the wrong context so rag is great but it does have its issues and there are techniques like contextual retrieval where you can actually improve the embeddings that are created here or you can also use reranking to reshuffle these chunks to CH the best ones so there are various techniques to get around those issues and just to visualize that if I go to my quadrant collection this is the collection that I just created I imported 646 vectors for these F1 technical regulations and you can see all of the chunks that have been created and we've set the content length to a th000 characters plus an overlap between chunks so when the rag workflow produced its output it essentially just got 10 of these chunks that were the most relevant for the query to formulate the answer so let's compare this now to the KAG architecture and there are actually two versions that we can go through here in many ways KAG is essentially just prompt caching so if we look at the open Ai and the clawed version of this all that's happening is when a user submits a query the documents are also sent in with that query to the llm and the llm produces the output if there's a follow-up question the document is sent back in again so you're constantly sending in the documents in each request to open AI or into anthropic to produce the output the actual prompt caching happens on the llm side or on the server side with open Ai and entr Tropic it's able to realize that it's already received that before and it's been stored in the k cache within the Transformer and that's how the prompt caching Works within this system you need to continually send in the documents or it might be a chat history or a large system message or whatever it is whereas the Google Gemini version of KAG that I showed you earlier is very different so here you need to upload the data initially so you have a document or documents you run a job to import them into cach and what that does is it's going to return to you a cash ID that you need to save somewhere you need to save it to your own database and then when a user submits a query again like what is a gurnie on an F1 car you need to get that cach ID for the F1 regulations and then you need to send that query with the cash ID to Gemini and what Gemini will do then is it'll load up the cash with that ID and then it'll trigger the generation for the query so here all you're doing is you're sending the query and the cach ID you're not sending 170 odd pages of documentation each time so with this architecture there's definitely more setup involved but you have huge control then over how you manage this cash and the duration of the cach with Gemini is much longer than open Ai and anthropic it defaults to an hour I believe it can go to one or two days as well so let's jump back to this prompt caching version with open Ai and anthropic and let's test it out in our nadn workflows so here is our CAG with open AI workflow let me get my chat trigger I will drop it in here connect it up so now I've just sent it the the message it has downloaded the file extracted the text and then sent it with the query to open AI I'm hitting gbd4 mini here and on the right hand side you can see the total tokens used was around 5,000 this is a smaller version of the document now if I send another message then when this goes to open aai it's still going to send the full document but that was so much faster and if I click through here you can see on the right hand side cash tokens was 5,120 so even though we sent the full document you can actually see how many have been cached on the server side so we're not actually paying as much for those so this full breakdown of usage you don't actually get with the AI agent version so if I connect this up and then ask it the same question I still obviously get a response but if I click into the llm I'm not getting the number of cached tokens so I do actually think it is probably caching on the server side it's just not visible within nad's console but that's why I set this up just so that you can actually see it on screen so that is essentially the open aai version of KAG which is their prompt caching and it's very easy to set up in the sense that you don't actually really need to do anything the key though is you need to structure your system prompt in a certain way so if you look here you can't have any Dynamic content at the start of the prompt so here we have our standard system prompt the document is going to be static it's not going to change and then the dynamic element is the query at the bottom the reason for this is within open ai's own documentation they talk about the need to structure your prompts like this because it's based off a prefix so everything at the start of the prompt will be cached and then the minute something Dynamic is found that's where it'll be cut off so this whole cash lookup cash hit and cash Miss you'll need to kind of work through this if you really want to make sure that your prompts are being cached the other thing to say is prompts are generally cached for 5 to 10 minutes it can be up to an hour off peak but it's a lot lower than what you can actually get using that Google Gemini so as I said that's the easy version there's no real changes you need to make to your workflows if you're going the anthropic route you do actually need to make changes so let's bring over the chat interface to here let's load up the full regulations and let's trigger it so we're downloading the full RS now so we're calling claw we're not doing it using an n8
AI summaries can miss context or contain errors. Check important details against the original video.





