Google Gemma 3 Beats DeepSeek V3: 100% FREE and Local!

Mervin PraisonAbout 4 min readMar 25, 2025Watch original
THE SUMMARYAI-generated

Jemma 3: A Capable Open-Source Model

Key Concepts:

  • Jemma 3: A family of open-weight language models by Google.
  • Parameter Size: The number of parameters in a model, influencing its complexity and performance.
  • Ollama: A tool for running language models locally.
  • Chainlit: A Python library for creating chatbot user interfaces.
  • Prais Agents: A framework for building AI agents.
  • AI Agents: Autonomous entities that can perform tasks using language models and tools.
  • Function Calling: The ability of a language model to call external functions or APIs.
  • Context Window: The amount of text a language model can consider when generating output.
  • Quantization: A technique for reducing the size of a model by using lower-precision numbers.
  • LM Arena: A platform for evaluating language models through human preference testing.
  • RAG (Retrieval-Augmented Generation): An AI framework that combines a pre-trained language model with an information retrieval system to generate more accurate and contextually relevant responses.

Overview of Jemma 3

Jemma 3 is presented as a highly capable open-weight language model that can be run on a single GPU or TPU. It is available in four different versions: 27 billion, 12 billion, 4 billion, and 1 billion parameter models. It supports multimodal (image, text, short videos) and multilingual (140 languages out-of-box, 35 languages fully supported) capabilities, a 128,000 token context window, and function calling.

Performance and Comparison

  • Jemma 3 (27B) outperforms DeepSeek V3 (7B) in preliminary human preference evaluations on LM Arena's leaderboard.
  • In the Chatbot Arena Elo SC ranking, Jemma 3 (27B) is positioned in the top 10, surpassing models like Qwen 2.5 Max and slightly trailing 01.
  • The model's performance is notable considering its relatively small size compared to other models with similar capabilities.
  • Quantization allows for high performance even with compressed models.

Local Execution with Ollama

Step-by-step process:

  1. Installation:

    • pip install ollama
    • pip install chainlit
    • Download Ollama from ollama.com.
  2. Model Download:

    • ollama pull jemma3 (downloads the 4 billion parameter model).
  3. Basic Chatbot Implementation (app.py):

    import ollama
    
    response = ollama.chat(model='jemma3', messages=[
        {
            'role': 'user',
            'content': 'Give me a meal plan for today.',
        },
    ])
    print(response['message']['content'])
    
    • This code snippet demonstrates a basic interaction with the Jemma 3 model, requesting a meal plan and printing the response.
  4. Running the script:

    • python app.py
  5. Chatbot User Interface with Chainlit (ui.py):

    import chainlit as cl
    import ollama
    
    @cl.on_message
    async def main(message: str):
        response = ollama.chat(model='jemma3', messages=[
            {
                'role': 'user',
                'content': message,
            },
        ])
        await cl.Message(content=response['message']['content']).send()
    
    @cl.on_chat_start
    async def start():
        await cl.Message(content="Hello! I'm Jemma 3. How can I help you today?").send()
    
    • This code creates a simple chatbot interface using Chainlit, allowing users to interact with the Jemma 3 model through a chat window.
  6. Running the Chainlit app:

    • chainlit run ui.py

AI Agent Creation with Prais Agents

Step-by-step process:

  1. Installation:

    • pip install prais-agents[llm] (includes Ollama support).
  2. Basic Agent Implementation (app.py):

    from prais_agents.agents import Agent
    
    agent = Agent(
        instruction="You are a helpful assistant.",
        llm="ollama/jemma3"
    )
    
    agent.start("Why is the sky blue?")
    
    • This code creates a basic AI agent using the Prais Agents framework, instructing it to be a helpful assistant and using the Jemma 3 model.
  3. Running the script:

    • python app.py
  4. Adding Internet Search Tool:

    from prais_agents.agents import Agent
    from prais_tools.internet_search import InternetSearch
    
    agent1 = Agent(
        instruction="You are a helpful assistant that writes LinkedIn posts.",
        llm="ollama/jemma3",
        tools=[InternetSearch()]
    )
    
    agent2 = Agent(
        instruction="You are a helpful assistant that writes Tweets based on LinkedIn posts.",
        llm="ollama/jemma3"
    )
    
    agent1.start("Write about Donald Trump 2025 election.")
    linkedin_post = agent1.memory.get_last_user_message()
    agent2.start(f"Write a Tweet based on this LinkedIn post: {linkedin_post}")
    
    • This code demonstrates how to integrate an internet search tool into an AI agent, allowing it to access and utilize information from the web.
    • Two agents are created: one to write a LinkedIn post and another to write a tweet based on the LinkedIn post.

Key Arguments and Perspectives

  • The presenter emphasizes the accessibility and cost-effectiveness of Jemma 3, highlighting its ability to run locally on a computer without requiring an internet connection.
  • The model's open-weight nature is presented as a significant advantage, allowing for greater transparency and customization.
  • The presenter showcases the potential of Jemma 3 for creating AI agents that can automate various tasks, such as content creation and information retrieval.

Conclusion

Jemma 3 is presented as a powerful and accessible open-weight language model that offers a compelling alternative to larger, more resource-intensive models. Its ability to run locally, combined with its multimodal and multilingual capabilities, makes it a versatile tool for a wide range of applications, including chatbot development and AI agent creation. The presenter encourages viewers to explore the model and share their experiences in the comments. The video also recommends watching another video on creating 100% local RAG AI agents based on custom data.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.