Deep Dive into Long Context
By Google for Developers
Share:
Key Concepts
- Tokens: Sub-word units used by LLMs, impacting how they process and understand text.
- Context Window: The amount of text (in tokens) an LLM can consider when generating a response.
- In-Weight Memory: Knowledge the LLM learned during pre-training.
- In-Context Memory: Knowledge explicitly provided to the LLM in the current prompt.
- RAG (Retrieval Augmented Generation): A technique to enhance LLMs by retrieving relevant information from a knowledge corpus and adding it to the context window.
- Context Caching: Reusing previously processed context to speed up and reduce the cost of subsequent queries.
- Needle in a Haystack: A benchmark task where the model must find a specific piece of information within a large context.
- Hard Distractors: Contextual elements that closely resemble the target information, making retrieval more challenging.
- Lost in the Middle Effect: A phenomenon where LLMs struggle to attend to information located in the middle of a long context.
Tokens and Their Significance
- A token is a unit of text slightly less than a word, including word parts and punctuation.
- While character-level generation has been explored, tokenization remains prevalent due to faster generation speeds.
- Tokenization can lead to "weirdness and complexity" as models perceive the world differently than humans.
- Example: Counting the number of "R"s in "strawberry" is difficult because the tokenizer might break the word into different parts.
- Whitespace is often prefixed to tokens, which can cause unusual effects during concatenation.
Context Windows: Types and Importance
- The context window comprises the tokens fed into the LLM, including the prompt, previous interactions, and uploaded files.
- LLMs have two knowledge sources: in-weight (pre-training) and in-context memory.
- In-context memory is easier to modify and update than in-weight memory.
- Context is crucial for:
- Updating obsolete facts.
- Providing private or personalized information.
- Supplying rare facts not well-represented in the pre-training data.
- Short context models require careful selection of information, while long context models allow for higher recall and coverage.
Retrieval Augmented Generation (RAG)
- RAG is an engineering technique that retrieves relevant information from a knowledge corpus before feeding it to the LLM.
- Process:
- Chunk the knowledge corpus into smaller textual units.
- Embed each chunk into a real-valued vector using an embedding model.
- Embed the query into a real-valued vector.
- Compare the query vector to the chunk vectors.
- Pack the closest chunks into the LLM's context.
- RAG is still relevant even with long context models because enterprise knowledge bases can contain billions of tokens.
- Long context and RAG can work synergistically, with long context increasing the recall of relevant information retrieved by RAG.
- The choice between RAG and long context depends on the application's latency requirements.
Scaling Long Context: Challenges and Future Directions
- The initial goal for 1.5 Pro was 1 million tokens, a 5x increase over the competition (around 200k tokens).
- A 2 million token model was released shortly after, representing a 10x increase.
- Inference tests were run at 10 million tokens with good quality, but the cost was prohibitive.
- "We could have shipped this model but it's pretty expensive to run this inference."
- Further scaling requires innovation, not just brute force.
- The cost of long context models is expected to decrease, leading to increased use of RAG with larger context windows.
Development and Evolution of Long Context Models
- The development of long context capabilities for 1.5 Pro was rapid and unexpected.
- The team worked "really hard" to ship the model quickly.
- Subsequent models (1.5 Flash, 2.0 Flash, 2.5 Pro) have seen improvements in quality at both 128k and 1 million token context sizes.
- 2.5 Pro demonstrates better benchmark results compared to strong baselines like GPT-4, Claude 3, and DeepSeek models.
Quality Evaluation and Challenges
- Evaluation is crucial for aligning and driving progress in LLM research.
- Single needle-in-a-haystack retrieval with easy distractors is a solved problem.
- Current challenges:
- Handling hard distractors (context that closely resembles the target information).
- Retrieving multiple needles.
- Realism in evaluations can compromise the ability to measure core long context capabilities.
- Tasks should integrate information over the whole context (synthesis) rather than just retrieval.
- Automatic evaluation metrics for synthesis tasks (e.g., summarization) are imperfect and can be noisy.
Long Context and Reasoning
- There's a deep connection between reasoning and long context.
- Improved next-token prediction with longer context can be interpreted as better reasoning.
- Long context is important for reasoning because it allows the model to make multiple logical jumps through the context.
- Feeding the output back into the input allows the model to overcome network depth limitations and perform harder tasks.
Long Output Generation
- There isn't a fundamental limitation to generating long outputs straight out of pre-training.
- Careful handling is required in post-training due to the end-of-sequence token.
- Short SFT data can lead the model to prematurely generate the end-of-sequence token.
- Reasoning is just one type of long output task; translation is another.
- Properly aligning the model is key to encouraging long output generation.
Developer Best Practices for Long Context
- Context Caching:
- Rely heavily on context caching to reduce costs and improve speed.
- Cache files uploaded by users (e.g., documents, videos, codebases).
- Place the question after the context to maximize caching benefits.
- Combination with RAG:
- Combine long context with RAG for knowledge bases containing billions of tokens.
- Use RAG even for shorter contexts when multiple needles need to be retrieved.
- Relevance:
- Avoid packing the context with irrelevant information.
- Prompting:
- Explicitly resolve contradictions between in-weight and in-context memory by using prompts like "based on the information above."
Fine-Tuning Considerations
- Fine-tuning on a knowledge corpus involves training the network to predict the next token.
- Limitations of fine-tuning:
- Requires hyperparameter tuning.
- Can lead to overfitting.
- May increase hallucinations.
- Advantages of fine-tuning:
- Cheaper and faster inference.
- Drawbacks of fine-tuning:
- Privacy implications (knowledge is cemented into the weights).
- Knowledge is not easy to update.
Future of Long Context (3-Year Outlook)
- Increased Quality: The quality of 1-2 million token context models will increase dramatically, maximizing retrieval tasks.
- Decreased Cost: The cost of long context will decrease, making 10 million token context windows a commodity.
- Coding Applications: 10 million token context windows will unlock incredible coding applications by allowing large projects to be included in the context.
- Superhuman Abilities: Long context will unlock superhuman abilities, such as processing information and connecting the dots more effectively than humans.
- "This thing is going to be incredible for coding applications."
Hardware and Inference Engineering
- Having powerful chips is not enough; talented inference engineers are crucial.
- The inference team's work was essential for delivering 1-2 million token context models.
Long Context and Agents
- Agents can be both consumers and suppliers of long context.
- Agents need long context to keep track of previous actions, observations, and the current state.
- Agents can automatically fetch and pack context using tool calls, reducing the tedium of manual context management.
Conclusion
Long context is a rapidly evolving area with the potential to unlock significant advancements in AI capabilities. While challenges remain in terms of cost, quality, and evaluation, the future looks promising, with expectations of increased quality, decreased cost, and the emergence of superhuman abilities, particularly in coding applications. The interplay between long context, RAG, reasoning, and agents will be crucial in shaping the future of AI systems.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development