Key Concepts
- Prompt Caching: A technique to store and reuse processed model inputs (KV vectors) to reduce latency and costs.
- Prefill vs. Decode: The two phases of LLM inference; prefill is compute-bound (processing input), while decode is memory-bandwidth bound (generating output).
- KV Cache: The storage of Key and Value vectors generated during the prefill phase.
- Multi-Head Latent Attention (MLA): An architectural innovation by DeepSeek that compresses the KV cache, allowing it to be stored on low-cost disks rather than expensive High Bandwidth Memory (HBM).
- Cache Invalidation: Actions that force a model to recompute the prefill phase, resulting in higher costs and latency.
1. The Economics of LLM Inference
While major AI labs (Google, OpenAI) are increasing prices, DeepSeek has implemented a 75% price cut on its V4 Pro model. This is not a subsidy but an architectural achievement.
- The Bottleneck: LLM requests consist of prefill (parallel processing of input) and decode (sequential token generation). Prefill is compute-bound, and decode is memory-bound.
- The DeepSeek Advantage: DeepSeek utilizes Multi-Head Latent Attention (MLA), which reduces the KV cache size by 93%. This allows the cache to be stored on distributed disk arrays rather than expensive HBM. Because the cache is small, streaming it from disk is faster than recomputing the prefill, enabling significantly lower pricing.
2. How Prompt Caching Works
- Mechanism: During prefill, the transformer generates Query, Key, and Value vectors. The Key and Value (KV) vectors describe how a token contributes to future attention. If a new request shares the same initial tokens as a previous one, the KV vectors are byte-identical and can be reused.
- Impact: Research (e.g., "Don't Break the Cache") indicates that prompt caching can save between 41% and 80% in costs across various agentic workflows.
3. Anthropic’s Framework for "Cloud Code"
To maximize cache hits, Anthropic organizes requests into a hierarchical structure:
- Layer 1 (Static): System prompts and tool definitions (globally cached).
- Layer 2 (Project):
cloud.mdand project context (cached per project). - Layer 3 (Session): Session-specific context (cached per session).
- Layer 4 (Conversation): The actual turn-by-turn dialogue (new on every request).
The Golden Rule: A change in a lower layer (e.g., conversation) preserves the cache of higher layers, but a change in a higher layer (e.g., system prompt) invalidates everything below it.
4. Five Common Cache-Busting Actions
To maintain cost efficiency, developers must avoid these triggers:
- Switching Models: Each model maintains its own cache; switching mid-session forces a full rebuild.
- Modifying Tools: Adding, removing, or updating tool parameters (or auto-reconnecting MCP servers) invalidates the cache.
- Dynamic System Prompts: Including timestamps or volatile metadata in the system prompt causes cache invalidation every time the data changes.
- Naive Compaction: Summarizing history via a separate API call with a different system prompt destroys the cache.
- Upgrading Software: New versions of agentic software often change system prompts or tool definitions, triggering a full cache rebuild.
5. Actionable Best Practices
- Use Messages, Not Prompt Changes: If you need to update the model on world changes (e.g., file edits or time), append a "system reminder" message rather than editing the system prompt.
- Cache-Friendly Design: Implement features like "Plan Mode" by toggling tools that are already defined, rather than swapping the entire toolset.
- Efficient Compaction: Perform compaction using the same system prompt and toolset as the parent conversation, appending the summary as a final user message.
- Strategic Commands: Use
/rewindto truncate conversations back to a previous state, allowing the next request to hit the existing cache entry.
Synthesis
The disparity in pricing between AI providers is largely driven by architectural choices regarding memory management. While DeepSeek has optimized the hardware/storage layer via MLA, developers must optimize the software layer through disciplined prompt engineering. By treating the system prompt as a static foundation and using message-based updates for dynamic information, builders can ensure high cache hit rates, significantly reducing both latency and operational costs.
AI summaries can miss context or contain errors. Check important details against the original video.





