Andrej Karpathy's Math Proves Agent Skills Will Fail. Here's What to Build Instead.

By The AI Automators

Share:

Key Concepts

  • Agentic Workflows: Multi-step AI processes designed to execute complex tasks autonomously.
  • March of Nines: The concept that each incremental increase in reliability (e.g., from 99% to 99.9%) requires exponentially more engineering effort.
  • Agent Skills: Portable, self-contained units of domain knowledge and procedural logic (often markdown-based) that guide AI behavior.
  • Harness Engineering: The practice of building a software "scaffold" around an AI model to enforce deterministic behavior, validation, and reliability.
  • Sub-agents: Specialized, isolated AI instances triggered by an orchestrator to perform narrow tasks, allowing for context management and cost optimization.
  • Deterministic Rails: Constraints placed on AI systems to ensure they follow specific, validated paths rather than relying on probabilistic generation.

1. The Reliability Challenge in AI

The transition from simple AI tasks (writing blog posts) to complex business workflows (compliance audits, risk analysis) is hindered by the "march of nines." In a 10-step agentic workflow, a 90% success rate per step results in over six failures daily. To achieve enterprise-grade reliability, businesses must move beyond simple prompting and implement systems that guarantee performance.

2. Agent Skills vs. Harness Engineering

  • Agent Skills: These act as "instructions" or "prompts" baked into the workflow. While they improve performance, they are not inherently reliable because they rely on the model's ability to follow instructions without hallucinating or skipping steps.
  • Harness Engineering: This is the solution for high-stakes environments. By wrapping the AI in a "harness" (a software layer), developers can gate and validate outputs at every stage.
    • Example: Stripe’s "minions" system uses a harness to validate AI-generated code against a suite of 3 million tests before merging pull requests.

3. Case Study: Contract Review Harness

The video demonstrates a specialized harness for legal contract reviews, moving from a simple prompt-based approach to a structured, multi-phase system:

  • Phase-based Execution: The process is broken into eight deterministic phases (extraction, classification, user clarification, research, clause extraction, risk analysis, redline generation, and executive summary).
  • Virtual File System: Acts as a "scratchpad" where the agent saves progress. This provides resilience; if a process fails, it can be restarted from a specific phase using the saved state.
  • Sub-agent Orchestration: The main agent acts as a supervisor, delegating specific tasks (like clause-level risk analysis) to smaller, cheaper models (e.g., Gemini 2.5 Flash). This keeps the main context window clean and reduces costs.
  • Programmatic Output: Instead of asking the LLM to "write a document," the harness populates a pre-defined Word template, ensuring consistent formatting and structure every time.

4. 12 Pillars of Harness Design

To build a robust agentic harness, the following components are essential:

  1. Architecture: Choosing the right pattern (e.g., sequential, hierarchical, or DAG/graph-based).
  2. Planning: Implementing fixed plans (for deterministic workflows) or dynamic plans (where the AI adjusts steps based on the request).
  3. File System: Providing a workspace for the agent to read/write files, preventing context loss.
  4. Delegation: Using sub-agents for context isolation and parallel processing.
  5. Tool Calling & Guardrails: Restricting what tools an agent can access and requiring human-in-the-loop approval for sensitive actions.
  6. Memory: Utilizing short-term (markdown/scratchpad) and long-term (knowledge graphs) storage.
  7. State Management: Tracking the progress of a run (e.g., via a database table) to ensure the system knows exactly where it is in the workflow.
  8. Code Execution: Using secure, isolated sandboxes for the AI to run and test code.
  9. Context Management: Compacting and summarizing context to avoid "context rot" and token exhaustion.
  10. Prompt Engineering: Using specific, narrow prompts for sub-agents to maximize accuracy.
  11. Validation Loops: Implementing self-correction mechanisms where the agent tests its own output against requirements and iterates if it fails.
  12. Hybrid Approach: Combining the structure of a harness with the flexibility of agent skills for non-standard tasks.

5. Notable Quotes

  • "The march of nines: you can reach the first 90% of reliability with a strong build and a good demo. But each additional nine requires comparable engineering effort to achieve."
  • "The best approach is to create a specialized harness, where you can gate and validate the output of each stage to ensure it stays on track."

Synthesis

The future of AI in business lies in moving away from "chatting" with models and toward building deterministic agentic systems. By treating AI as a component within a software harness—rather than an autonomous black box—developers can achieve the reliability, observability, and scalability required for complex, multi-stage enterprise workflows. The key takeaway is that reliability is not a property of the model itself, but a result of the engineering framework built around it.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video