Google just destroyed all open-source models (Gemma 4)
By David Ondrej
Key Concepts
- Gemma 4: Google’s latest open-source, multimodal AI model series.
- Local AI: Running models directly on personal hardware (laptop/phone) rather than cloud servers.
- Dense vs. Mixture of Experts (MoE): Architectural differences in parameter activation; MoE is sparse and more efficient.
- Parameters: The internal variables a model learns; higher counts generally correlate with higher intelligence.
- Inference: The process of running a trained AI model to generate predictions or text.
- VRAM/RAM: Critical hardware resources for local AI; VRAM is dedicated to the GPU, while Apple Silicon uses unified memory.
- Ollama: A tool for running and managing local LLMs via the command line.
- MCP (Model Context Protocol): A standard for connecting AI agents to external data sources like databases.
1. Overview of Gemma 4
Gemma 4 is Google’s new open-source model series designed to be highly efficient. Despite its relatively small size (26B–31B parameters), it competes with massive models like Claude 3.5 Sonnet (1.1 trillion parameters). It is multimodal, capable of processing audio, images, and video.
- Model Variants:
- E2B & E4B: "Effective" parameter models designed for mobile devices.
- 26B (MoE): A sparse architecture where only specific "experts" are activated per query, making it faster and more efficient.
- 31B (Dense): A dense architecture where all parameters are active, offering more predictable behavior but requiring more computational power.
2. Technical Advantages of Local AI
- Privacy: Data remains on the user's device, bypassing cloud-based data harvesting.
- Cost-Efficiency: Eliminates recurring subscription fees (e.g., ChatGPT Plus).
- Accessibility: No rate limits or internet dependency; models function offline.
3. Implementation and Setup
Laptop Setup (Ollama)
- Installation: Download Ollama from
ollama.comusing the provided one-liner terminal command. - Execution: Use
ollama run gemma4:31bin the terminal to download and initiate the model. - Interface: Users can interact via the terminal or the Ollama desktop app for a GUI-based chat experience.
Mobile Setup (Google AI Edge Gallery)
- App: Download the "Google AI Edge Gallery" from the App Store or Google Play.
- Model Selection: Choose the E2B or E4B models within the app.
- Usage: The app initializes the model locally, allowing for offline AI interaction on smartphones (e.g., iPhone 16 Pro Max).
4. Integrating AI Agents with Databases
The video highlights Supabase as a solution for AI agents needing real-time data access.
- Problem: Connecting agents to databases usually requires complex API wiring and "glue code."
- Solution: Supabase provides an MCP (Model Context Protocol) out of the box, allowing AI tools (like Cursor or Hermes) to query PostgreSQL databases directly via a single connection string.
- Benefit: Embeddings, authentication, and file storage are unified in one instance, simplifying vector similarity searches.
5. Real-World Applications & Performance
- Coding & UI Generation: Gemma 4 demonstrates high proficiency in replicating web designs and generating functional code components from reference images.
- Hermes Agent Integration: By setting a custom endpoint (
localhost:11434/v1), users can power the Hermes agent with Gemma 4, creating a fully local, autonomous coding assistant. - Performance Note: While local models are highly capable, they may be slower than cloud-hosted supercomputers when handling complex, multi-step agentic tasks.
6. Notable Quotes
- "The big AI companies want you to ignore local models because if you can run AI on your laptop, their business model falls apart."
- "We're entering an era where capable AI models are no longer locked behind the paywall, but instead they can run on the devices you already have."
Synthesis
Gemma 4 represents a paradigm shift in the AI industry by bringing trillion-parameter-level intelligence to consumer hardware. By leveraging local execution tools like Ollama and database integration platforms like Supabase, developers and power users can build private, cost-effective, and offline-capable AI agents. The transition from cloud-dependent AI to local, device-native AI is not only possible but increasingly competitive with the industry's largest proprietary models.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Stanford CS153 Frontier Systems | Building the Frontier Ecosystem
Stanford Online

'Things are going to be okay, in Canada and the U.S.': Thorne
BNN Bloomberg

I'M OUT: The $11 Trillion AI Bubble is Breaking!
Steven Van Metre

South Korea bets big on AI with nearly a trillion dollars of investment • FRANCE 24 English
FRANCE 24 English

The Bubble is Bursting... (Emergency Update)
Bravos Research

The AI Bubble Just Ended - Without Popping
Heresy Financial

AI Market Volatility, Europe Heat Wave, Venezuela Quakes Damage | Bloomberg This Weekend: June 27
Bloomberg Television