I Trained AI To "DOMINATE" Brackey's Game
By Nicholas Renotte
Key Concepts
- AI Agent (Baz): A character trained using reinforcement learning to navigate a game level.
- Godot Engine: The game development platform used to build the level.
- Stable Baselines 3 (SB3): A Python library for training reinforcement learning agents.
- GDO IRL: A library facilitating communication between Godot games and SB3.
- AI Controller Node: A Godot node used to integrate AI control into the game agent.
- Sensors: Mechanisms (e.g., ArrayCast, Position Sensor) allowing the AI agent to perceive its environment without visual input.
- Reward System: A crucial component of reinforcement learning that provides positive or negative feedback to guide the agent's behavior.
- Sync Node: A Godot node that opens a local server for communication with external training scripts.
- Parallel Environments: Running multiple instances of the game simultaneously to accelerate training.
n_stepsHyperparameter: In SB3, the number of steps collected from the environment before an algorithm update is performed.- Sparse Reward: A reward that is given infrequently, typically only upon completion of a major task or goal.
- Frame Stacking: A technique where multiple consecutive frames are provided to the agent as a single observation, allowing it to infer movement and velocity.
- Time Limit Wrapper: A mechanism that automatically resets the game environment if the agent gets stuck or exceeds a predefined time limit.
- Pre-trained Weights: Using a model that has already been trained on a similar task as a starting point for further training.
Project Overview and Initial Setup
The video details the process of training an AI agent named Baz to complete a specific level in a game built with the Godot Engine. The goal is for Baz to navigate obstacles and reach a flagpole, relying purely on sensory input rather than visual perception. The training is conducted using Stable Baselines 3 (SB3), a reinforcement learning library written in Python.
A key challenge was establishing communication between Godot and SB3. This was solved using GDO IRL, a library specifically designed for this purpose, which provides examples and solid documentation. The initial plan involved integrating this solution for Baz.
Initial Challenges and Goal Refinement
The creator admits to not fully reading the GDO IRL documentation due to a tendency to move too fast, leading to initial difficulties. This highlighted that significant changes were needed within the Godot game environment to support the reinforcement learning pipeline. Four key goals were identified:
- Hooking up Baz to the Agent: Integrating Godot's AI controller node into Baz's character.
- Implementing Sensors: Enabling Baz to perceive the environment without visual input. This involved adding an ArrayCast sensor to detect collisions with obstacles and a Position sensor to provide the coordinates of specified objects.
- Designing a Reward System: Creating a point-based system to incentivize Baz towards the flagpole.
- Adding a Sync Node: Integrating a Sync node to open a local server, allowing SB3 to communicate with the Godot game.
Reward System Details
The reward system was designed to guide Baz's behavior:
- Positive Rewards:
- A couple of points each time Baz achieves a new best distance towards the flagpole.
- 200 points for landing on a platform.
- 1 point for collecting a coin.
- 1,000 points for hitting the flagpole.
- Negative Rewards:
- -200 points for dying.
Training Script and Speed Optimization
A Python training script was developed to connect to the Godot game via the Sync node and allow the SB3 algorithm to train. Initial training speed was very slow (low FPS). To address this, three optimization strategies were implemented:
- In-game Speed Adjustment: Increasing the game's internal speed.
- Parallel Environments: Exporting the game and modifying the training script to run multiple game instances concurrently, significantly boosting frames per second (SPS).
n_stepsHyperparameter Tuning: Reducing then_stepshyperparameter (number of steps collected before an algorithm update) closer to the average game length. A largen_stepsvalue meant collecting too much information before updates, slowing down learning.
Debugging and Iterative Refinement
Despite optimizations, Baz struggled, particularly with a "moving platform of death." This led to a "second internal dilemma" where the creator, tired from "insane hours," resorted to using Claude (an LLM) with "crap prompts," which worsened the problems.
The strategy shifted to dialing down game difficulty: replacing a single long moving platform with two static ones and removing enemies. The hope was to transfer learning from easier levels to harder ones. Frame stacking was also suggested by Claude to help the agent learn movement, though the creator noted it had other issues. A risk identified was Baz memorizing the easier level, making the training effort wasted for harder levels.
The Reward System Flaw: A Major Takeaway
Baz showed progress on easier levels but developed an unexpected behavior: repeatedly jumping on platforms without progressing to the flagpole. This was not an anticipated "edge case." After trying various random hyperparameter options and even training for 10 million steps (roughly 5 hours with parallel environments), the issue persisted.
The breakthrough came when the creator realized a critical flaw in the reward system: Baz could gain 200 points every time he jumped on a platform, allowing him to accumulate millions of points without reaching the goal. The sparse reward of 1,000 points for the flagpole was insufficient to counteract this.
The solution was simple but crucial: adding a flag to ensure Baz could only receive the platform reward once per platform. This change immediately led to Baz learning to progress towards the flagpole. The creator emphasized this as one of the biggest takeaways: "Triple check the reward."
Conquering the "Platform of Death" and Time Limits
With the reward system fixed, Baz successfully learned to navigate static platforms. To tackle the "platform of death," pre-trained weights from the easier level were used as a starting point, and the game was re-exported with the challenging moving platform.
Another problem emerged: Baz would get stuck indefinitely, preventing further learning as no new frames were generated. This was identified as a time-related issue. The solution was to implement a time limit wrapper for each game, forcing a reset after 60 seconds if Baz got stuck.
After approximately 10 million steps of training with the time limit adjusted pipeline, Baz finally learned to consistently jump and hit the moving platform, successfully crossing the gap.
Conclusion and Main Takeaways
The project, spanning six weeks, successfully trained Baz to beat a challenging game level. The creator highlights several key learnings:
- Reinforcement learning is "part science and part art form."
- Perfect examples in papers and tutorials don't prepare you for real-world debugging.
- The importance of thoroughly checking the reward system in custom environments.
- The necessity of slowing down and taking a breath when encountering problems, rather than rushing or relying on LLMs with poor prompts.
- The value of time limits in preventing agents from getting stuck and facilitating continuous learning.
The journey demonstrated the iterative nature of reinforcement learning development, involving constant debugging, hyperparameter tuning, and creative problem-solving.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Every Kind of Volcano | SciShow Kids
SciShow Kids

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth