Hacking Subagents Into Codex CLI — Brian John, Betterup

AI EngineerAbout 6 min readNov 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Sub-agents: Independent instances of an AI agent that can be invoked by a main agent to perform specific tasks, offloading context and processing.
  • CodeX CLI: A command-line interface tool for interacting with AI models, particularly for code-related tasks.
  • Context Management: The process of efficiently handling and utilizing the limited token window of AI models. Sub-agents are crucial for this by processing information and returning only the essential results.
  • Wrapper Script: A script that orchestrates the execution of sub-agents, handling prompt construction, agent invocation, and result retrieval.
  • Permissions (CodeX Sandbox): Security mechanisms within CodeX that restrict agent access to file systems and credentials. Configuring these correctly is critical for sub-agent functionality.
  • Rollout Recorder: A logging feature in CodeX that can interfere with sub-agent execution by preventing file system access.
  • Agent's Rule of Two (Meta): A security framework for AI agents, considering input trustworthiness, access to sensitive systems/data, and the ability to change state or communicate externally.
  • agents.md: A configuration file in CodeX that defines available agents, their purpose, and how to invoke them.
  • Reasoning Effort: A parameter in agent configuration that dictates the computational resources and depth of processing for a given task (e.g., light, medium, high).
  • Asynchronous Execution: The ability of an agent to perform multiple tasks concurrently. CodeX, unlike Claude Code, does not support asynchronous execution, leading to slower performance.
  • Timeouts: A setting to prevent agent tasks from running indefinitely, especially for complex operations on large codebases.

Hacking Sub-Agents into CodeX CLI

This presentation details a method for integrating sub-agent functionality into the CodeX CLI, enabling more efficient context management and workflow flexibility. The speaker, Brian John, a Principal Fullstack Engineer at BetterUp, shares his experience in overcoming CodeX's sandbox limitations to achieve this.

Motivation for Sub-Agents

Brian John emphasizes the critical role of sub-agents in context management. By delegating tasks to sub-agents, the main agent can avoid consuming its limited token window with intermediate processing. The sub-agent performs its work, utilizes its own tokens, and returns only the final answer to the main agent, preserving the main agent's context for higher-level reasoning. This is particularly beneficial when working with large codebases.

Design and Implementation

The core design involves a parent CodeX session that executes a wrapper script. This script acts as an intermediary, determining which sub-agent to run, constructing the necessary prompt, and invoking the sub-agent via codeex exec.

  1. Child Agent Execution: A child CodeX process is launched to act as the sub-agent.
  2. Task Execution: The sub-agent responds to its prompt, performs its designated task.
  3. Result Output: The sub-agent writes its answer to a file.
  4. Result Retrieval: The wrapper script reads the output file and prints the result to standard output, making it accessible to the parent CodeX session.

Challenges and Solutions: CodeX Sandbox Permissions

The primary hurdle in implementing sub-agents within CodeX is its restrictive sandbox environment. Brian John highlights that achieving this with normal permissions, rather than using dangerously skip permissions, was a significant challenge.

  • Parent Permissions: Requires at least sandbox workspace to execute the codeex command.
  • Child Permissions:
    • sandbox workspace write: Essential for the sub-agent to write its output to a file.
    • Disabling Rollout Recorder: This logging mechanism, which restricts file system access to commands outside the workspace, must be disabled for the child process.
    • OpenAI Credentials: The sandbox prevents access to OpenAI credentials in the home directory.

Brian John spent considerable time identifying and configuring the minimum required permissions for both parent and child processes.

Security Considerations: Agent's Rule of Two

Brian John references Meta's "Agent's Rule of Two" paper, outlining three key security concerns for AI agents:

  1. Untrustworthy Input: Not a primary concern in this setup as the agents are controlled internally.
  2. Access to Sensitive Systems or Private Data: This is a relevant concern, especially when working with proprietary codebases.
  3. Ability to Change State or Communicate Externally: The sub-agents can change state (dependent on the system) and communicate externally to the OpenAI API.

While the risk is considered lower in this specific implementation due to controlled state changes and API communication, Brian John stresses that "lower risk does not mean no risk" and users must make their own risk assessments.

Configuring CodeX for Sub-Agents

To enable CodeX to recognize and utilize sub-agents, two key configuration steps are necessary:

  1. agents.md Configuration: This file informs CodeX how to invoke sub-agents. It specifies:

    • The command to run for a particular sub-agent.
    • When to trigger sub-agents (user request or deemed helpful).
    • A list of available sub-agents and their descriptions.
  2. Wrapper Script for Invocation: The agents.md configuration points to a wrapper script that handles the execution. This script is designed to:

    • Write the agent name and user query to separate files.
    • Execute a consistent command (codeex exec) to invoke the sub-agent. This consistency is crucial for CodeX's permissioning system, as it avoids repeated permission prompts for different command arguments.

Demo and Proof of Concept

Brian John presents a proof-of-concept repository with the following components:

  • Toy Agents: Simple agents defined with a name, reasoning_effort (light, medium, high), and a prompt. Examples include a word counter and a file writer.
  • Wrapper Script (agent_executor.py): A Python script (around 72 lines) that orchestrates the sub-agent execution. It takes inputs, calls an AgentExecutor class, and returns the output.
  • AgentExecutor Class: A small class that handles kicking off the child sub-agent with the correct permissions, reasoning effort, and disabling the rollout recorder.
  • CodeX Home Wrapper: A script that syncs CodeX home files into a subdirectory, allowing the agent to access them, and sets CodeX_HOME accordingly. It then launches CodeX, often in full auto mode (workspace write + approval on request).

Demo Walkthrough:

  1. Word Counter Agent:

    • The user queries CodeX to use the "work counter" sub-agent.
    • CodeX identifies the need to run agent_exec.
    • Agent name and query are written to files.
    • CodeX prompts for permission to run the command. The user grants permission and selects "don't ask again" to avoid repeated prompts.
    • The sub-agent executes serially (CodeX does not support asynchronous execution).
    • The result (word count) is printed to standard output and returned to CodeX.
  2. File Writer Agent:

    • The user queries CodeX to use the "file writer" sub-agent.
    • The process is similar to the word counter, but this time, no permission prompt appears due to the previous "don't ask again" selection.
    • A timeout of 600 seconds (10 minutes) is mentioned, as some agents can take significantly longer, especially on large codebases. Longer timeouts (e.g., 20 minutes) might be necessary.
    • The agent writes content to a file, which is then verified.

Performance Considerations

Brian John notes that CodeX's serial execution makes it slower than tools like Claude Code, which support asynchronous operations. However, he suggests this might be an intentional design choice, positioning CodeX as a more "hands-off, unattended" tool, while Claude Code is geared towards more iterative workflows. He finds this acceptable for his use cases.

Conclusion and Resources

The presentation concludes with Brian John sharing the URL to the open-source proof-of-concept repository and contact information for himself and BetterUp. He reiterates his interest in connecting with individuals passionate about working with LLMs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.