Vercel AI SDK Masterclass: From Fundamentals to Deep Research

AI EngineerAbout 7 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI SDK Primitives: generateText, generateObject, streamText, streamObject
  • Unified Interface: Ability to switch between language models with minimal code changes.
  • Tools/Function Calling: Extending language models to interact with the outside world.
  • Max Steps: Allowing the model to autonomously choose the number of steps in a process.
  • Structured Outputs: Generating data in a predefined format using generateObject and Zod schemas.
  • Deep Research Clone: Building an application that performs in-depth research and generates a report.
  • Recursion: Using a function to call itself to achieve a desired outcome.

Generate Text Function

  • The generateText function is a core primitive for calling large language models and generating text.
  • It accepts a prompt or an array of messages as input. Messages have a role (e.g., "user") and content.
  • The AI SDK provides a unified interface, allowing easy switching between language models by changing the model parameter.
  • Example:
    import { generateText } from '@vercel/ai';
    
    async function main() {
      const text = await generateText({
        model: 'openai/gpt-4o-mini',
        prompt: 'Hello world',
      });
      console.log(text);
    }
    main();
    
  • Models can be selected from different providers (e.g., OpenAI, Perplexity, Google) by importing the respective provider and specifying the model name.
  • Some models, like Perplexity, provide sources for their responses, accessible via the sources property.
  • Example of switching to Perplexity:
    import { generateText } from '@vercel/ai';
    import { perplexity } from '@vercel/ai/providers/perplexity';
    
    async function main() {
      const text = await generateText({
        model: perplexity({ model: 'sonar-pro' }),
        prompt: 'When was the AI Engineer Summit in 2025?',
      });
      console.log(text);
      console.log(text.sources);
    }
    main();
    

Tools/Function Calling

  • Tools enable language models to interact with the outside world and perform actions.
  • The model receives a list of available tools with their names, descriptions, and required data.
  • If the model decides to use a tool, it generates a "tool call" containing the tool's name and arguments.
  • The developer is responsible for parsing the tool call, executing the tool, and handling the results.
  • The AI SDK simplifies tool usage with the tool utility function, providing type safety between parameters and arguments.
  • Example: Adding two numbers using a tool:
    import { generateText, tool } from '@vercel/ai';
    
    async function main() {
      const result = await generateText({
        model: 'openai/gpt-4o-mini',
        prompt: "What's 10 + 5?",
        tools: {
          addNumbers: tool({
            description: 'Adds two numbers together',
            parameters: {
              num1: { type: 'number' },
              num2: { type: 'number' },
            },
            execute: async ({ num1, num2 }) => {
              return num1 + num2;
            },
          }),
        },
        maxSteps: 3,
      });
      console.log(result.steps);
      console.log(result.text);
    }
    main();
    
  • maxSteps allows the model to iteratively use tools and incorporate the results into its response.
  • The AI SDK automatically invokes the execute function of a tool and returns the result in a toolResults array.
  • The maxSteps property enables multi-step agents, where the model can autonomously choose the next step in the process.
  • Example: Getting weather in two cities and adding the temperatures:
    • The model infers latitude and longitude from the city name.
    • Parallel tool calls are used to get weather data for both cities.
    • The addNumbers tool is then used to sum the temperatures.

Structured Outputs

  • Structured outputs allow generating data in a predefined format.
  • Two methods: generateText with the experimental output option and the generateObject function.
  • Zod is a TypeScript validation library used to define schemas for structured outputs.
  • Example: Using generateText with experimental output:
    import { generateText, output } from '@vercel/ai';
    import { z } from 'zod';
    
    const schema = z.object({
      sum: z.number(),
    });
    
    async function main() {
      const result = await generateText({
        model: 'openai/gpt-4o-mini',
        prompt: "What's 10 + 5?",
        tools: {
          addNumbers: tool({
            description: 'Adds two numbers together',
            parameters: {
              num1: { type: 'number' },
              num2: { type: 'number' },
            },
            execute: async ({ num1, num2 }) => {
              return num1 + num2;
            },
          }),
        },
        maxSteps: 3,
        output: {
          object: schema,
        },
      });
      console.log(result.experimental_output);
    }
    main();
    
  • The generateObject function is dedicated to structured outputs and is considered a "workhorse" function.
  • Example: Generating definitions for AI agents:
    import { generateObject } from '@vercel/ai';
    import { z } from 'zod';
    
    const schema = z.object({
      definitions: z.array(z.string().describe("I want the language model to use as much jargon as possible. It should be completely incoherent.")),
    });
    
    async function main() {
      const obj = await generateObject({
        model: 'openai/gpt-4o-mini',
        prompt: 'Please come up with 10 definitions for AI agents',
        schema: schema,
      });
      console.log(obj);
    }
    main();
    
  • The describe function in Zod allows providing detailed instructions for specific values in the schema.

Deep Research Clone

  • The goal is to build a simplified version of a deep research tool in Node.js.
  • The workflow involves:
    1. Taking an input query.
    2. Generating subqueries.
    3. Searching the web for relevant results for each subquery.
    4. Analyzing the results for learnings and follow-up questions.
    5. Recursively repeating the process with follow-up questions.
  • Depth: The number of levels to go down in the research process.
  • Breadth: The number of subqueries to generate at each level.
  • The first step is to build a function to generate search queries:
    import { generateObject } from '@vercel/ai';
    import { z } from 'zod';
    
    const mainModel = 'openai/gpt-4o-mini';
    
    async function generateSearchQueries(query: string, n = 3) {
      const schema = z.object({
        queries: z.array(z.string()).min(1).max(5),
      });
    
      const obj = await generateObject({
        model: mainModel,
        prompt: `Generate ${n} search queries for the following query: ${query}`,
        schema: schema,
      });
      return obj.queries;
    }
    
  • The next step is to search the web for relevant results using a service like Exa.
  • A function is created to search the web with Exa, filtering the results to include only relevant information.
  • The exa-js package is used to interact with the Exa API.
  • The liveCrawl option ensures that the results are up-to-date.
  • The most complex part is analyzing the search results for learnings and follow-up questions.
  • This is achieved using generateText with two tools:
    • searchWeb: Searches the web for information about a given query.
    • evaluate: Evaluates the relevance of the search results.
  • The evaluate tool uses generateObject with enum mode to determine if the search result is relevant or irrelevant.
  • If the evaluation is irrelevant, the model is instructed to search again with a more specific query.
  • A search and process function is created to search the web and determine if the results are relevant.
  • A generate learnings function is created to analyze the search results for learnings and follow-up questions.
  • The generate learnings function uses generateObject to generate a learning and follow-up questions from the search result.
  • Recursion is introduced to the process to allow for deeper research.
  • A deep research function is created to handle the recursion.
  • A global state variable is created to store the accumulated research.
  • The deep research function is updated to update the store as it iterates through the levels of recursion.
  • The deep research function is called recursively, decrementing the depth and breadth.
  • A check is added to the evaluate tool to ensure that the same source is not used twice.
  • A generate report function is created to synthesize all of the accumulated research and put it into a report.
  • The generate report function uses generateText to generate the report.
  • A system prompt is added to the generate report function to provide guidance on what the report should look like.
  • The final report is written to a markdown file.

Conclusion

The session provides a comprehensive overview of building agents with the AI SDK, covering fundamental concepts like generating text, using tools, and generating structured outputs. It demonstrates how to combine these primitives to build a more complex application, a deep research clone, showcasing the power and flexibility of the AI SDK for creating intelligent and autonomous systems. The use of Zod for schema validation and the emphasis on type safety throughout the examples highlight the importance of robust development practices when working with language models.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.