Grok 4 Just Beat Every AI Model!

Mervin PraisonAbout 4 min readJul 10, 2025Watch original
THE SUMMARYAI-generated

Grok 4: Overview, API Usage, and Agent Creation

Key Concepts:

  • Grok 4: XAI's latest large language model (LLM) with improved intelligence, coding capabilities, and context window.
  • Context Window: The amount of text or data the model can consider at once (256,000 tokens for Grok 4).
  • Function Calling: The ability of the model to use external tools or APIs.
  • Structured Outputs: The ability of the model to generate data in a specific format (e.g., JSON).
  • Reasoning: The model's ability to think through problems and provide logical solutions.
  • XAI SDK: Software Development Kit provided by XAI for interacting with their models.
  • Praise AI Agents: A library for creating AI agents that can use LLMs.
  • Agentic Behavior: The ability of the model to act autonomously and use tools to achieve a goal.
  • Benchmarks: Standardized tests used to evaluate the performance of LLMs (e.g., GPQA, VendingBench).
  • Token: A unit of text used for pricing and model input/output.

Grok 4 Overview

  • Capabilities: Grok 4 is presented as a highly intelligent model, surpassing even Claude 3 Pro in some areas. It excels in coding tasks and demonstrates strong reasoning abilities.
  • Performance:
    • Top performer in GPQA benchmark AIME25 LCB HMMT and USA.
    • Number one for coding.
    • Number two for output speed.
    • Tops the math index.
    • Grok 4 heavy is the top performing humanity lost exam.
    • Grok 4 is the top beating Claude Opus 4 in vending bench.
  • Variations: Three versions exist: Grok 4 (standard), Grok 4 Heavy (top-performing), and Grok 4 without tools.
  • Pricing: $3 for input and $15 for output per million tokens. Cached input is $0.75 per million tokens.
  • Access: Available through a paid subscription on grok.com ($300/year or $30/month for Grok 4, $300/month for Grok 4 Heavy).

Using the Grok 4 API with XAI SDK

  • Installation:

    1. Open terminal.
    2. pip install xai-sdk
    3. export XAI_API_KEY=<your_api_key> (API key obtained from XAI website).
  • Code Example (app.py):

    from xai import client
    from xai.client import Chat
    
    client = Client(api_key="YOUR_API_KEY")
    
    chat = client.chat.create(
        model="grok",
        messages=[
            {"role": "system", "content": "You are a PhD level mathematician."},
            {"role": "user", "content": "Solve this equation: 2x + 3 = 7"}
        ]
    )
    
    print(chat.response)
    
  • Explanation: The code initializes the XAI client, creates a chat session with the "grok" model, provides a system instruction (role-playing as a mathematician), asks a question, and prints the response.

  • Running the Code: python app.py in the terminal.

Creating AI Agents with Praise AI Agents

  • Installation: pip install praise-ai-agents llm

  • Code Example:

    from praise_ai_agents import Agent
    from llm import Instruction
    
    agent = Agent(
        instruction=Instruction("You are a helpful assistant."),
        llm="grok"
    )
    
    response = agent.ask("What is the capital of France?")
    print(response)
    
  • Explanation: This code imports the necessary libraries, creates an agent with a specific instruction and LLM ("grok"), and then asks a question.

  • Enhancements:

    • Multiple agents can be added for collaboration.
    • MCP tools can be integrated with a single line of code to enhance agent capabilities.
  • OpenAI SDK Integration: Possible by providing the Grok API base URL and API key.

Grok 4 Testing and Examples

  • Misguided Attention Test: Grok 4 correctly identified the ethical dilemma in a modified trolley problem.
  • Python Expert Level Challenge (Vel Q): The model provided a response, but the testing system indicated a failure.
  • Simplify Josephus Task: The model successfully solved the task.
  • Far Sequence Task: The model generated code, but it initially produced a syntax error due to the testing system's Python version. Grok 4 was able to identify the Python version issue and correct the code.
  • Safety Test (Breaking into a Car): The model refused to provide instructions for illegal activities and suggested legitimate alternatives (e.g., calling a locksmith).

Observations and Insights

  • Agentic Behavior: The Grok interface exhibits agentic behavior, utilizing tools and internet search to answer questions.
  • Python Version Awareness: Grok 4 can identify and adapt to different Python environments.
  • Safety Measures: The model incorporates safety measures to prevent misuse.
  • Limitations: The testing system used in the video may have limitations (e.g., outdated Python version) that can affect the results.

Conclusion

Grok 4 is a powerful LLM with strong capabilities in coding, reasoning, and general knowledge. The XAI SDK and Praise AI Agents library provide convenient ways to access and utilize Grok 4 in various applications. While some limitations were observed during testing, the model demonstrated impressive performance and adaptability. The presenter encourages viewers to try Grok 4 and share their experiences.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.