AI coding with Gemini CLI - Google's terminal agentic coding tool

Google for DevelopersAbout 5 min readSep 17, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Gemini CLI, Agent Loop, Context Awareness, Project-Oriented, Gemini md files, Tools (Google Search, Web Fetch, Save Memory), Sandboxing (Seatbelt, Docker, Podman), MCP (Microservice Communication Protocol) Tools, Extensions, Headless Mode, Telemetry (OpenTelemetry), Checkpointing, Chat Save/Restore, Context Compression, Node.js, React, Ink, System Prompts, Roadmap.

Gemini CLI Overview

The Gemini CLI is an interactive command-line tool designed to work with Gemini models, particularly Gemini 1.5 Flash Pro. Its key feature is an agent loop, enabling it to iterate on tasks using available tools until completion. It's context-aware, understanding the operating system (Mac, Windows, Linux), accessing the file system, and running binaries (ls, jq, FFmpeg). The CLI is project-oriented, intended for use within project directories like Git repositories, granting access to files within that directory.

Authentication and Setup

Upon initial launch, the CLI prompts for authentication. Users can log in with Google for 1,000 free API calls daily using 1.5 Flash and Pro. Alternatively, a Gemini API key can be used for unrestricted 1.5 Pro access, or a Vertex AI account for centralized models and billing.

Gemini md Files

Gemini md files are pre-canned instructions for projects, promoting consistent engineering practices. They can be placed in the top-level directory or subdirectories (frontend, backend) for specific instructions. A global directory in the home directory can store instructions applicable across all projects. These files contain information on building, testing (test frameworks, mocking, React testing), Git usage, and engineering style guides.

Example: A Gemini md file might specify the use of npm run preflight for building and running the project, details about test frameworks, and guidelines for class construction and code style.

Built-in Tools

The CLI includes tools for file editing, searching, reading, and making changes. Notable tools include:

  • Google Search: Uses the Gemini API's Google Search feature to find relevant information online.
  • Web Fetch: Grabs data from the internet using the Google Search Index or, as a fallback, tools like curl or wget.
  • Save Memory: Appends information to Gemini md files for future use, allowing the model to remember specific tasks or style preferences.

Example: Using "Save Memory" to store the preference for conventional commit syntax.

Editing and Interaction

The CLI allows users to ask questions about the codebase and make changes. It generates diffs for review, offering options to accept ("yes"), always allow edits, modify in an external editor (VS Code, Vim), or reject ("no") to provide feedback. Built-in commands (starting with /) are available, with help accessible via /help. Shell mode (accessed with !) executes commands in the local shell, including the output in the context. Files can be referenced using @ syntax, including their content in the prompt. The CLI supports multimodal inputs (PDFs, images, videos, audio).

Example: Asking "How is the tty window title set?" to understand the code and then requesting "Can you update this title to include a Unicode star?"

Sandboxing

Sandboxing provides security by limiting the CLI's access to files and tools. The default operation includes soft safeguards, but a sandbox (using gemini sandbox) offers stronger protection. On macOS, it uses the Seatbelt system, while Docker and Podman are supported on other systems. Custom Docker images can be configured. Settings can be specified via command-line flags or the gemini settings.json file.

Example: Running gemini sandbox to enable the macOS Seatbelt sandbox.

MCP (Microservice Communication Protocol) Tools

MCP tools extend the CLI's capabilities. Examples include:

  • Context7: Retrieves API documentation on demand.
  • Google Slides: Generates slide presentations using the Google Slides API.
  • Custom Extension: Connects Gemini API docs directly via an MCP server.

MCP tools are configured in gemini settings.json, allowing the use of binaries like npx, uvx, and pipx to download and run remote packages.

Example: Configuring Context7 with npx and arguments in gemini settings.json.

Extensions

Extensions bundle configuration, including metadata, SMTP server definitions, and context files. They can include custom / commands and are shareable.

Example: Creating an extension for a personal Git repository with specific API documentation instructions.

Headless Mode

Headless mode allows running prompts directly from the command line (e.g., summarize my last week's Git logs and write an email to email.msg). The yolo mode automatically approves tool permissions. This is useful for Cron jobs.

Example: Using headless mode to generate a weekly email summary of Git logs.

Telemetry

The CLI supports writing telemetry data to an OpenTelemetry (OTel) endpoint. This includes prompts, model details, status, latency, and token usage.

Example: Configuring the CLI to write telemetry data to a local JSON file.

Checkpointing

Checkpointing creates a shadow Git repository for backups. Before each tool call (e.g., write file), a checkpoint is created. The restore command allows reverting to a previous state.

Example: Using checkpointing to revert a change made by the model that replaced the readme with jokes.

Chat Save/Restore

The chat save and chat resume commands allow saving and restoring chat sessions, enabling users to jump back to specific points in a conversation.

Example: Saving a chat session with ideas for refactoring documentation and then resuming it later.

Context Compression

The CLI has a context compression feature that reduces the token count when the context window limit is reached. It can be triggered automatically or manually using the compress command.

Example: Using compress to reduce the token count in a long chat session.

Architecture

The CLI is built with Node.js and React, using the Ink library to render web components as a terminal interface. It has a core package (agent loop, command logic) and a CLI package (interface).

System Prompts

The CLI uses system prompts to guide the agent's behavior. The main system prompt defines the agent's role, mandates (e.g., not making assumptions), and workflow. The first turn context prompt provides the model with information about the user's system (OS, directory structure).

Example: The main system prompt includes instructions on using Three.js for 3D games and Kotlin Multiplatform or Flutter for mobile apps.

Open Source and Roadmap

The Gemini CLI is open source on GitHub. The roadmap is public, with high-level themes like improving agent performance and automation.

Conclusion

The Gemini CLI is a powerful and versatile tool for interacting with Gemini models. Its agent loop, context awareness, and project-oriented design enable complex tasks. Features like Gemini md files, built-in tools, sandboxing, MCP tools, extensions, headless mode, telemetry, checkpointing, and chat save/restore enhance its functionality and security. The open-source nature and public roadmap foster community involvement and continuous improvement.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.