Build Hour: Agent RFT
By OpenAI
Key Concepts
- Agent RFT (Reinforcement Fine-Tuning): A method to improve agent performance by adjusting model weights based on a specified learning signal, teaching the model desired and undesired behaviors.
- Agents: Models capable of interacting with the outside world (via tools) to complete tasks autonomously.
- Tools: External functionalities (e.g., terminal, code interpreter, API calls) that agents can use to interact with their environment.
- Context Window: The memory or input space available to the model, which includes tool outputs.
- Prompt Engineering: Optimizing the input prompt to guide model behavior.
- Task Optimization: Simplifying tasks, adding guardrails, or adjusting tool availability to improve agent performance.
- Fine-Tuning: Training a model end-to-end on a specific task to achieve better performance.
- Reinforcement Learning (RL): A machine learning paradigm where an agent learns to make decisions by taking actions in an environment to maximize a cumulative reward.
- Reward Signal: A signal that indicates how well an agent is performing a task, used to guide the learning process.
- Latency: The time it takes for a model to produce an output, often influenced by the number of reasoning tokens and tool calls.
- Sample Efficiency: The ability of a model to learn effectively from a limited amount of training data.
- Distribution Shift: The difference between the data distribution during training and the data distribution during inference.
- Model Grader: A grader that uses a model (e.g., GPT-4) to evaluate the quality of an agent's output.
- Endpoint Grader: A grader that involves calling a custom endpoint to define a reward signal.
- String Grader: A brittle grader that relies on exact string matches to ground truth.
- F1 Score: A metric that balances precision and recall, often used to evaluate tasks where both false positives and false negatives are important.
- Compute Multiplier: A parameter that controls the amount of exploration an agent performs during training.
- Rollout: A single instance of an agent attempting to complete a task during training.
- Tool Calls per Rollout: The number of times an agent invokes tools within a single task completion attempt.
- Reasoning Tokens: Tokens used by the model to process information and make decisions.
- Domain Shift: The difference between the domain of the data the model was trained on and the domain of the task it's being applied to.
- Reward Hacking: A phenomenon where an agent exploits loopholes or unintended behaviors in the reward function to achieve high scores without genuinely improving performance.
Agent RFT: Enhancing Agent Performance Through Reinforcement Fine-Tuning
This build hour introduces Agent RFT (Reinforcement Fine-Tuning), a novel approach to significantly improve the performance of AI agents, particularly those that interact with external tools. Building upon previous discussions on agents and agent kits, Agent RFT focuses on fine-tuning models to better utilize tools and achieve superior task completion.
Introduction to Agent RFT
Agents are distinguished from regular models by their ability to interact with the outside world through tools to complete tasks autonomously. These tools can range from code interpreters and terminals for coding agents to internal software access for customer service agents. The outputs from these tool interactions are fed back into the agent's context window, allowing it to reason and decide on subsequent actions, potentially calling more tools.
While prompt engineering and task optimization (simplifying tasks, adding guardrails, adjusting tool sets) are effective initial steps for improving agent performance, fine-tuning offers a more profound level of optimization. Agent RFT specifically enables the fine-tuning of agents by allowing them to call tools during the training exploration process. This is achieved by adjusting the model's weights based on a custom learning signal (reward) that defines what constitutes good and less-than-good behavior.
Benefits of Agent RFT
Agent RFT offers several key advantages:
- Improved Reasoning and Tool Use: Enhances the agent's ability to reason over tool outputs and select the most effective tools to reach the optimal final answer.
- Sample Efficiency: Particularly valuable in domains with scarce training data, as it can learn effectively from fewer examples.
- Reduced Latency: By training the model to use fewer tool calls and reasoning tokens, Agent RFT can significantly decrease response times, leading to faster user experiences. This is achieved through an inherent penalty on token usage during training and the ability to set explicit tool call budgets.
- Domain Alignment: Helps bridge the distribution shift between the general domain of pre-trained models and the specific business context and tools of an application.
How Agent RFT Works: Technical Details
Agent RFT introduces two major updates to the existing RFT product:
- Tool Calling During Training: Models can now call user-provided tool endpoints during the training exploration phase.
- Custom Reward Signal Endpoint: Users can specify a custom reward signal via an endpoint, allowing for highly tailored training objectives.
During training, each agent rollout (a single attempt to complete a task) is assigned a unique identifier. This ID is attached to tool calls, enabling the user's system to track rollouts and manage state. When the final answer is generated, the user's grader is called, and all the context from the rollout, including tool calls, can be passed to it via the unique identifier for holistic grading. This process allows tool calls and grading to occur within the user's environment, mirroring production conditions and offering significant flexibility in shaping the model's policy.
Case Study: FinQA Financial QA with OpenAI
A compelling example of Agent RFT in action is the FinQA financial question-answering task. The original benchmark provided financial reports with questions. However, the task was modified to be significantly harder by providing only the question and requiring the agent to search through 2,800 financial reports using tools like semantic search, list (to view file system directories), and cat (to retrieve document content). A crucial constraint was to achieve the answer within 10 tool calls.
For grading, a model grader was used to avoid the brittleness of exact string matching and to allow for partial credit. This approach demonstrated how Agent RFT can train models to navigate complex file systems, identify relevant information, and perform numerical reasoning under strict constraints.
Demo and Training Process
The demo showcased the setup of a tool server using FastAPI, defining tools like semantic search (using embeddings and cosine similarity), list, and cat. The grader was implemented using GPT-4, providing flexibility in evaluating answers and awarding partial credit.
A baseline evaluation of GPT-5 on the FinQA task revealed significant variance in performance across samples, indicating room for improvement. The training process involved setting hyperparameters like epochs, batch size, and compute multiplier. The reward curve demonstrated rapid performance improvement within the first 10 steps, correlating with a decrease in tool calls and reasoning tokens. This suggests the model learned to use tools more efficiently.
Analysis of the training runs showed:
- Reduced Latency: A 5-second reduction in average response time (approximately 10%) and an 11 percentage point increase in average reward.
- Fewer Tool Calls: A reduction from an average of 6.9 tool calls per trace in the baseline to 4.2 in the fine-tuned model.
- Improved Efficiency: A significant drop in reasoning tokens from 2500 to 1500.
- Policy Shift: An analysis plot showed that a large fraction of data points achieved higher reward with fewer tool calls, indicating a more efficient and effective policy. The model also reduced repeated tool calls, demonstrating smarter tool usage.
Advice for Successful Agent RFT Implementation
Performance Optimization:
- Well-Specified and Constrained Task: Ensure a clear task with consensus on what constitutes a good answer, especially for subjective tasks.
- Non-Zero Baseline Performance: The model needs to have some capability to be right occasionally for RL to be effective.
- Improve Accuracy (IK): Analyze the variance in performance across multiple runs for each sample to identify areas for improvement.
- Quality Over Quantity: Focus on high-quality training data rather than sheer volume.
Infrastructure:
- Mirror Production Behavior: Host tools and graders in an environment that closely matches production to ensure seamless translation of improvements.
- Invest in Grader Design: Create a robust grader aligned with domain knowledge, difficult to game, and capable of providing nuanced feedback (e.g., partial credit).
- Limit Tool Call Output Length: Keep tool outputs concise to improve training speed and prevent model confusion.
Customer Spotlight: Cognition
Cognition, a company building Devon, an autonomous AI engineer, shared their experience using Agent RFT to optimize Devon's planning mode. Their goal was to reduce the time spent in planning to allow Devon to start making code edits faster.
They designed a task where Devon uses only read_file and shell tools to identify relevant files for editing, with the reward metric being F1 score to balance precision and recall. The results showed significant improvements:
- GPT-5 with 100 samples: Outperformed the base model.
- GPT-5 with 1000 samples: Achieved even greater gains.
- Reduced Back-and-Forths: The fine-tuned model reduced planning mode interactions from 8-10 to 4, effectively halving the time.
Cognition highlighted the importance of infrastructure, using isolated VMs for tool execution to prevent one rollout from affecting others, and the need for robust monitoring to handle infrastructure errors that could lead to zero rewards.
Other Success Stories
- Ambience (Healthcare): Improved an ICD-10 coding agent's F1 score from 0.52 to 0.57 and reduced latency by 18%, halving the number of responses exceeding their latency threshold.
- GenSpark (Slides Creation): Achieved an 88% improvement in bad cases for their slide creation agent by fine-tuning a reasoning model to harmonize output, focusing on both content and visual aesthetics.
- MacO (GPU Kernel Building): Used Agent RFT with as few as 100 PyTorch prompts to train GPT-5 to write performant GPU kernels for new hardware, achieving a 72% improvement in correctness and performance without needing code examples.
- Rogo (Financial Reasoning): Fine-tuned a model to summarize and present financial insights, using a custom LLM as an endpoint grader. This resulted in a 21% increase in core ML performance, lower hallucination rates, and fewer missing citations. Rogo also emphasized the importance of a watertight grader to prevent reward hacking.
When to Turn to Agent RFT
The recommended process for improving agent performance is:
- Build High-Quality Data: Ensure training and evaluation sets closely match production traffic.
- Establish Baseline Performance: Run baseline evaluations on models like GPT-5 to understand current performance.
- Optimize Without Fine-Tuning: Explore prompt improvements, infrastructure enhancements, and task harness adjustments.
- Apply Agent RFT: Once other optimizations are exhausted, use Agent RFT to fundamentally change model weights for end-to-end task and domain improvement.
Q&A Highlights
- Best Suited Tasks: Tasks with sufficient variance in the training data, allowing for exploration, and where the grading mechanism provides a nuanced signal (not purely binary) are ideal. Agent RFT is broadly applicable to any agent using out-of-distribution tools.
- Platform Updates: The Agent RFT platform now supports fine-tuning GPT-5 with tools and endpoint graders, offering significantly more flexibility than the initial May release. Observability and stability have also improved.
- Sample Efficiency: RL is sample efficient because the model generates its own training data through exploration. The strong prior of frontier models further enhances this efficiency.
- RL Training Objective: The core RL loss function remains consistent, but Agent RFT allows for more flexible reward functions (model graders, endpoint graders), enabling finer control over the model's policy.
- Alpha Endpoints: Access to alpha endpoints (like tool integrators) is typically managed through direct engagement with OpenAI account teams via an interest form.
- Continuous Learning: Currently, training is not continuous during inference. However, the model generates new trajectories during training, which are leveraged for objective computation.
The session concluded with resources for further engagement, including an interest form for Agent RFT and information on upcoming build hours.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Trump’s Tax Immunity Could Save Him More Than $600 Million
Forbes

Risk, Returns and everything in between | TEDxSVNIT 2026 | Ashu Bishnoi | TEDxSVNIT
TEDx Talks

Charter Communications CHTR Stock Explained!!!
Value Investing with Sven Carlin, Ph.D.

Retail, Cryptocurrency, Obituaries | Pointed News Quiz
Bloomberg Television

Micron expands US memory chip production amid AI demand surge
Fox Business

When Birds Speak The Nature Whisper | Dr. Bushra Nisar Khan | TEDxPunjab University
TEDx Talks