I found the new best coding agent.

By David Ondrej

Share:

Key Concepts

  • Droid: An AI coding agent developed by Factory AI, designed to be a more advanced and cost-effective alternative to tools like Code-Llama and Codex.
  • Factory AI: The company behind Droid.
  • Global Terminal Installation: Droid can be installed globally on a machine via a terminal command.
  • IDE Integration: Droid integrates with Integrated Development Environments (IDEs) like VS Code.
  • Token System: Droid utilizes a token-based system for model usage, with a generous free allocation.
  • Model Selection: Droid supports a wide range of AI models, including OpenAI (GPT-4, GPT-5), Anthropic (Sonnet 4.5, Haiku, Opus 4.1), and GLM 4.6.
  • GLM 4.6: An open-source model offered by Droid, noted for its performance and affordability.
  • Credit Consumption: Different models have varying credit costs, with Opus being highlighted as expensive.
  • Droid Modes (Shift+Tab):
    • Auto Off: All actions require explicit approval.
    • Spec: Planning mode.
    • Auto Low: Edits and read-only commands; aims for safety.
    • Auto Medium: Reversible commands; allows for slightly riskier but recoverable actions.
    • Auto High: "YOLO" mode; highest level of autonomy.
  • Ultra Deep Research Agent: A concept for an AI agent that performs extensive research by generating search queries, executing web searches, and synthesizing results.
  • Context Engineering/Management: Droid's ability to compress conversation history to save tokens.
  • Reasoning Effort: The ability to control the level of computational effort for model reasoning (e.g., for Sonnet 4.5).
  • Vectal: A platform offering access to various AI models and features like slash commands and task tagging.
  • Open Router: An API that aggregates access to multiple AI models, used by Droid.
  • Version Control (Git/GitHub): Emphasized as crucial for managing AI-generated code and preventing data loss.
  • AI Stack: The evolving landscape of AI tools and models, with rapid advancements and potential shifts in preferred solutions.
  • Multi-Agent Systems: The future trend of humans managing multiple specialized AI agents rather than relying on a single, monolithic AI.
  • Error Handling: Droid's capability to investigate and attempt to fix errors.
  • Context Cutoff Logic: The management of the amount of information a model can process at once.
  • System Prompt: The underlying instructions given to an AI model to guide its behavior and output.

Droid: A Powerful and Cost-Effective AI Coding Agent

This video introduces Droid, an AI coding agent developed by Factory AI, presented as a superior alternative to existing tools like Code-Llama and Codex, particularly in terms of user interface (UI) and cost-effectiveness.

Setup and Initial Impressions

The setup process for Droid is straightforward. Users can install it globally via a terminal command provided on the factory.ai website. Once installed, Droid integrates seamlessly with IDEs like VS Code. The UI is noted to be inspired by Code-Llama, which is considered a positive, suggesting a polished and smooth user experience. The presenter believes Droid might even surpass Code-Llama in polish.

Upon initial setup, users are prompted to create an account on the factory.ai website. A significant incentive is the 20 million tokens provided for free upon adding Droid to the cart, which is described as sufficient for building and deploying an entire SaaS application with tokens remaining.

Model Selection and Cost Efficiency

A key feature of Droid is its extensive model support. By typing /model, users can access a wide array of AI models, including:

  • OpenAI: GPT-4, GPT-5
  • Anthropic: Sonnet 4.5, Opus 4.1, Haiku
  • Droid Core: GLM 4.6 (an open-source model)

The presenter highlights that GLM 4.6 is the only open-source model available and performs exceptionally well, even outperforming Sonnet 4.5 on many benchmarks. Crucially, GLM 4.6 is also the most affordable option, indicated by the lowest multiplier for credit consumption. In contrast, Opus 4.1 is described as "retardedly expensive." The cost-effectiveness of Droid, combined with its model flexibility, is a major selling point.

Droid Modes and Control

Droid offers granular control over its operation through a Shift+Tab shortcut, which cycles through different modes:

  • Auto Off: The most restrictive mode, requiring explicit approval for all actions and preventing any terminal commands or risky changes.
  • Spec: A planning mode, allowing the AI to outline steps without execution.
  • Auto Low: Suitable for edits and read-only commands, aiming to avoid destructive actions like database deletion.
  • Auto Medium: Allows for reversible commands, enabling slightly riskier operations that can be undone.
  • Auto High: The most autonomous mode, referred to as "YOLO mode," where the AI has maximum freedom to act.

The presenter intends to use Auto Medium or Auto High for building a project from scratch to leverage Droid's full potential.

Building an "Ultra Deep Research" Agent

The presenter demonstrates Droid's capabilities by outlining the creation of an "ultra deep research" agent. The process involves:

  1. User Input: The user provides a topic.
  2. Reasoning Model Call: A model like Haiku 4.5 generates approximately 100 different search queries based on the topic.
  3. Asynchronous Web Search: Web search agents are dispatched concurrently to gather information.
  4. Result Aggregation: The results from the web searches are collected.
  5. Final Reasoning Model: Another reasoning model synthesizes the aggregated results into a high-signal report.

This process is likened to having a personal McKinsey analyst.

Droid's Advanced Research and Planning Capabilities

In a practical demonstration, the presenter provides a prompt for Droid. The AI's response is notable for its thoroughness. Unlike Code-Llama, which might immediately suggest code, Droid first performs three web search calls to verify the correct documentation for referenced technologies (Python, OpenRouter, Perplexity API). This pre-computation research phase is highlighted as a significant time-saver.

After its research, Droid suggests an architecture. The presenter then interacts with Droid to refine the plan, choosing to "keep iterating on spec." The ability to select different implementation versions (proceed with implementation, allow file edits read-only, medium, high) demonstrates Droid's customizability and the developers' deep understanding of the product ("dogfooding").

The presenter switches the model to Codex for implementation, noting that it's the "best model." Droid then compresses the conversation history to save tokens, showcasing its proper context engineering and management.

Handling Errors and Model Switching

An Error 400 status code occurs when attempting to simplify the plan with Codex. The presenter suspects it might be related to an AWS outage. To overcome this, the model is switched to Sonnet 4.5, and the "reasoning effort" is set to "high." This granular control over reasoning effort is a feature not found in Code-Llama, which offers more limited options like "think harder" or toggling thinking on/off. Droid's ability to manage these settings is presented as a significant advantage.

The presenter emphasizes the dynamic nature of the AI landscape, with new models like GPT-5, Opus 4.5, and Gemini 3 on the horizon. Droid's strength lies in its support for all these models, allowing users to adapt quickly.

Code Generation and Refinement

Droid generates an initial codebase, including files like config.py. The presenter identifies a need to update model references within the configuration (e.g., changing from 3.5 haiku to 4.5 haiku). This highlights the importance of verifying and correcting AI-generated configurations.

The presenter also stresses the critical importance of Git and GitHub fundamentals, especially when coding with AI. A dedicated 14-minute module within "the new society" is recommended for learning these essential skills, emphasizing that proper version control prevents catastrophic data loss.

The process of creating an .env file and a .gitignore file is demonstrated, underscoring best practices for managing API keys and avoiding accidental commits of sensitive information.

Advanced Use Cases and Debugging

The presenter then tasks Droid with investigating an error encountered during the application's execution. Simultaneously, another Droid instance is instructed to add more detailed progress print statements to the CLI output. This showcases the ability to run multiple Droid instances concurrently for different tasks.

The presenter observes Droid's architecture and its use of models available through OpenRouter, noting that Droid's value lies in its scaffolding, tools, and UI. The responsiveness of the Factory AI team (CTO following the presenter on Twitter) is seen as a positive indicator of the product's future.

While acknowledging Codex's current strength and reasoning capabilities, the presenter reiterates Droid's advantage in supporting a wide range of models, ensuring future compatibility.

Iterative Development and Research Enhancement

The "ultra deep research" agent is further refined. The presenter requests an improvement to the research process by adding an initial web search step to gather basic information before generating detailed search queries. This is done in Spec mode with a focus on minimal code changes.

The agent successfully generates research queries and performs web searches using the Perplexity API. However, issues arise with the summarization step, specifically regarding token limits and potential model confusion. The presenter notes that while Sonnet 4.5 has a 1 million token context window, the output was truncated, suggesting a potential configuration error or misunderstanding by the AI.

The presenter then uses Droid to investigate and fix the context cutoff logic, aiming to leverage the full 1 million token capacity. The ability to paste images directly into the Droid interface is highlighted as a superior feature compared to other agents.

Further iterations involve adjusting the summarization agent's prompt to make the report more concise and save it to an MD file with clear markdown formatting. The presenter also emphasizes the importance of committing changes frequently to Git, especially when working with AI, to mitigate risks of unintended modifications.

The agent is tasked with researching the AWS outage and providing protection strategies for startup founders. The generated report is detailed, covering the outage's impact, patterns, and recommended protection strategies like multi-AZ, multi-region, and multi-cloud approaches.

Conclusion and Future Outlook

Droid is praised for its powerful and versatile UI, surpassing both Code-Llama and Codex CLI in certain aspects, particularly its image handling capabilities. The presenter expresses regret for not trying Droid earlier.

The video concludes by encouraging viewers to subscribe for more in-depth content on AI coding tools and projects. The presenter differentiates their content by emphasizing their practical experience in building and scaling a real-world AI startup, contrasting it with channels that merely report news. The core takeaway is that Droid is a highly capable and user-friendly AI coding agent that supports a broad spectrum of models, making it a valuable tool for developers navigating the rapidly evolving AI landscape.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video