Context Platform Engineering to Reduce Token Anxiety — Val Bercovici, WEKA

AI EngineerAbout 8 min readNov 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Context Platform Engineering: A discipline focused on optimizing the infrastructure and processes for managing and utilizing context within AI systems, particularly large language models (LLMs).
  • KV Cache (Key-Value Cache): A memory cache used by LLMs to store previously computed key-value pairs from attention mechanisms. This significantly speeds up token generation by avoiding redundant computations.
  • KV Cache Hit Rate: The percentage of times the LLM can retrieve a required key-value pair from the KV cache instead of recomputing it. A higher hit rate is crucial for production-grade AI agents.
  • Token Anxiety: The concern and limitations faced by developers due to token rate limits imposed by LLM providers, hindering productivity and efficient development.
  • Context Financial Engineering: The practice of strategically managing input and output tokens to balance costs, often involving "prompt arbitrage" based on token pricing and predicted cache usage.
  • Prompt Arbitrage: A form of context financial engineering where developers balance the cost of input and output tokens, considering new pricing models that focus on cache rights and reads.
  • Cadence Mismatch: The discrepancy between the slower human feedback loops for AI agents and the much faster iteration cycles of agent swarms and subtasks.
  • Token Storage Problem: The fundamental challenge of efficiently storing and retrieving tokens for LLMs, especially in dynamic and complex agentic workflows.
  • Service Level Agreements (SLAs): Contracts or commitments made by users when subscribing to token tiers or specifying token usage in agent instructions.
  • Service Level Objectives (SLOs): The actual performance targets delivered by the context platform to meet the user's SLAs.
  • Memory Tiering: The use of different types of memory (e.g., HBM, DRAM, NVMe-backed storage) with varying performance and capacity characteristics to store and access tokens.
  • Agent Swarms: A collection of AI agents working collaboratively on a task.
  • Agent Subtasks: Smaller, discrete tasks performed by individual agents within a larger agentic workflow.
  • Model Parallelism: Techniques for distributing LLM computations across multiple devices or models.
  • Disaggregated/Aggregated Pre-fill and Decode: Different strategies for handling the pre-fill (initial token generation) and decode (subsequent token generation) phases of LLM inference.
  • Load Generator: A tool used to simulate user traffic and test the performance of a system under various loads.
  • HBM (High Bandwidth Memory): A type of high-performance memory commonly used in GPUs.
  • DRAM (Dynamic Random-Access Memory): A common type of computer memory, generally slower and less dense than HBM.
  • NVMe (Non-Volatile Memory Express): A high-speed interface for solid-state drives (SSDs).

Open Sourcing the Context Platform Engineering Toolkit

Weta's Chief AI Officer, Valberkichi, and Kellen Fox, Head of Product Management, announced the open sourcing of their Context Platform Engineering Toolkit. This toolkit includes a sophisticated load generator developed by Kellen. Key features of the toolkit are:

  • Agent Swarm Configuration: Allows for the configuration of agent swarms and subtasks.
  • Specific SLOs: Enables setting and testing against specific Service Level Objectives.
  • Prompt Cycles: Supports deterministic and random prompt cycling.
  • Model Parallelism: Offers various model parallelism options.
  • Disaggregated/Aggregated Pre-fill and Decode: Provides options for optimizing these inference stages.
  • Memory Tiering: Includes important memory tiering options for efficient context management.

The toolkit is available on GitHub, and the team encourages users to download, experiment with, and provide feedback, as well as contribute to its development.

The Importance of KV Cache Hit Rate

A core insight driving the need for context platform engineering comes from the "Context Engineering" blog by Manis. They highlighted that KV cache hit rate is the single most important metric for production-grade AI agents.

Context platform engineering aims to dramatically simplify reaching maximum KV cache hit rates. This directly addresses the issue of token anxiety, where developers frequently encounter token rate limits. By engineering platforms that eliminate these limits, productivity can be significantly enhanced.

Context Financial Engineering vs. Context Platform Engineering

In the absence of robust context platform engineering, developers often resort to context financial engineering. This is fundamentally prompt arbitrage, where the cost of input and output tokens is balanced. This practice is complicated by new token pricing categories focusing on cache rights and reads, requiring developers to be "clairvoyant" in predicting future cache usage (e.g., for 5-minute or 1-hour time-to-live with Anthropic).

The speakers argue that applying context prompt engineering techniques is a superior approach to overcome token anxiety and prompt cash arbitrage compared to continuing with complex financial balancing.

The Cadence Mismatch and Token Storage Problem

A key challenge is the cadence mismatch between slow human feedback loops for agents and the high-iteration cadence of agent swarms and subtasks. These subtasks consume many tokens, often cachable, but the platform's ability to respond efficiently is limited.

Fundamentally, this is a token storage problem. The Service Level Agreements (SLAs) users sign for token tiers or commit to in agent instructions translate into Service Level Objectives (SLOs) delivered by the context platform. Kellen's research at Weta Labs revealed that when users subscribe to token tiers or pay for token rights, they are essentially purchasing KV cache slots in token storage.

Context platform engineering involves optimizing infrastructure and KV caching/memory tiers to meet these SLA requirements and deliver specific SLOs.

Visualizing Agentic Workflows and Cache Usage

Kellen Fox presented visualizations of typical agentic workflows:

  • Column Graph: Depicts new tokens (salmon), cachable tokens (gray), and output tokens (blue). Blue dots represent user responses, highlighting the slower human feedback loop.
  • Common Pattern: Agents consume context until a model or inference provider limit is reached, triggering a summarization phase. This summarization can lead to a loss of fidelity.
  • Prompt Breakdown: In agentic coding, user input is a small fraction of the total context. The majority consists of tool use and tool responses (e.g., bash commands, grep results).
  • Median Time Between Requests: For conversations, this can be 10-15 seconds, but the mean time can extend to minutes or hours due to human response times.
  • Multi-Agent Systems: Orchestrator agents spawn sub-agents for specific tasks. These sub-agents can be short-lived or persistent, impacting context endurance and overall context usage.
  • Cachable vs. Actual Cache Hits: While a significant portion of tokens might be theoretically cachable (gray in the visualization), the actual cache hit rate achieved in production is often lower.

The Impact of Low Cache Hit Rates

Low KV cache hit rates have significant consequences:

  • Increased Costs (API Users): Every "miss" (yellow in visualization) results in paying for input tokens again, potentially 10x more than if it were cached.
  • Rate Limit Issues (Subscription Users): Even with flat-rate subscriptions, low cache hit rates can lead to hitting rate limits faster due to increased token usage.

Temporal Analysis of Cache Performance

Kellen introduced a temporal view of cache performance:

  • Working Set vs. Time-to-Live (TTL): Graphs illustrate how the "working set" (tokens held in memory based on TTL) affects cache hit rates.
    • 1-minute TTL (Red): Shows thrashing and low hit rates when the time between requests exceeds 1 minute.
    • 5-minute TTL (Blue): Improves hit rates by riding out more cache hits.
    • 1-hour TTL (Green): Requires holding more tokens for longer but results in a better user experience and higher hit rates.
  • Refresh Rate: Visualized as the number of times a chunk of tokens is refreshed. A 1-minute TTL can lead to refreshing the same tokens 15-16 times, whereas longer TTLs approach a refresh rate of one.

Inference Provider Perspective and Token Storage

From an inference provider's standpoint, maintaining a high cache hit rate is critical to avoid "melting GPU clusters." They incentivize users to stay within optimal cache hit rate bands.

Token storage requires:

  • Sufficient Capacity: To hold an optimal amount of cache, reaching a point of diminishing returns.
  • Fast Ingestion: To store tokens rapidly without dropping them or blocking GPUs.
  • Rapid Fetching: To retrieve tokens quickly and avoid blocking GPUs, which are the primary resource.

Memory Tiering for Context Platforms

Different memory tiers are employed:

  • HBM: High-performance, ideal but not always feasible due to batching and cost.
  • DRAM: Common, but limited in size and tightly coupled with compute, making expansion difficult and potentially hurting performance (e.g., pooled DRAM).
  • Weta's Augmented Memory Grid: Leverages NVMe for denser storage (up to 1000x denser than DRAM) and optimized connectors between inference systems and their existing products. This provides capacity at DRAM speeds, allowing for sustained benefits.

Testing Methodology and Results

The open-sourced toolkit acts as an inference provider, aiming to keep load within SLOs (time to first token, output tokens per request). Testing can be done:

  • Deterministically: By sequentially processing prompts and observing performance drops when memory tiers overflow.
  • Realistically (Random Sampling): By increasing concurrent users over time and randomly sampling prompts, simulating a blend of HBM, DRAM, and other memory tiers.

Benchmark Results:

  • Comparison 1: HBM with Weta (Purple) vs. HBM + DRAM (Orange) vs. HBM + DRAM + Weta (Pinky)
    • Initially, all benefit from HBM.
    • As concurrent users increase, DRAM tiers overflow, causing significant drops in Orange and Pinky.
    • Weta's solution (Purple) also drops but less dramatically, maintaining higher concurrency and output tokens in steady state due to its capacity and DRAM speeds.
    • The Pinky color has the capacity but isn't fast enough to get data into the GPU for a difference.
  • Decode vs. Pre-fill Focus: Pre-fill focused roles show even better results for Weta, as GPUs are more efficient with large pre-fill batches.

Conclusion and Call to Action

The speakers reiterated their excitement about the open-sourced Context Platform Engineering Toolkit. They urged the audience to download, use, provide feedback, and contribute to the project. The ultimate goal is to collectively reduce token anxiety, minimize prompt cash arbitrage, and advance the field of context platform engineering. Links to relevant blogs and resources will be provided in the transcript section.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.