AI SDK Version 5 on TanStack Start!

Jack HerringtonAbout 4 min readAug 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Verscell AI SDK version 5: Upgraded AI development kit with polished APIs and new features.
  • Tanstack Start: A full-stack framework for building web applications.
  • Speech Transcription: Converting spoken audio into written text using AI.
  • Text-to-Speech (TTS): Converting written text into spoken audio using AI.
  • Tools: Functions or APIs that the AI can use to interact with external data or services.
  • UI Messages vs. Model Messages: Distinction between the structure of messages displayed in the UI and the format required by the AI model.
  • Stop When: New functionality in V5 that allows for more flexible control over when the AI stops processing tool invocations.
  • React Media Recorder: A React library for handling microphone input and audio recording.
  • OpenAI Whisper: An OpenAI model for speech transcription.
  • OpenAI TTS1: An OpenAI model for text-to-speech conversion.

Creating a Tanstack Start App with AI Chat Example

  • The video demonstrates how to create a new Tanstack Start application using the create-start-app command.
  • The --example tanstack-chat flag is used to set up a pre-built AI chat example.
  • The example features a guitar recommendation system, showcasing the use of AI tools to access and interact with a database.
  • The application requires setting up environment variables for API keys, specifically for Anthropic and OpenAI.

Backend Implementation (API Endpoint)

  • The backend API endpoint (/api/demo/chat) handles incoming chat requests.
  • It uses the streamText function from the AI SDK to connect to the Anthropic model.
  • Key Change (V5): The code converts UI messages to model messages using a cleaner, more distinct approach.
  • Key Change (V5): The maxSteps parameter is replaced with the stopWhen functionality, allowing for more flexible control over tool invocations.
  • The toUIMessageStreamResponse function converts model messages back to UI messages for display.
  • The API returns a raw streamed response, which integrates seamlessly with Tanstack Start.

Tools Implementation

  • The example uses two tools: getGuitars and recommendGuitar.
  • Key Change (V5): The parameters field in tool definitions is renamed to inputSchema. An outputSchema is also available.
  • The recommendGuitar tool is a UI tool that instructs the AI to display a specific guitar card.
  • The tool's execute function marshals the input ID into the output.

Frontend Implementation (UI)

  • The useChat hook is used to manage the chat interface.
  • Key Change (V5): The useChat hook no longer tracks the input field; this must be managed separately using useState.
  • Key Change (V5): The useChat hook now uses a transport mechanism, allowing for different methods of communicating with the AI (e.g., fetch, server functions, local OpenAI calls). The default is defaultChatTransport which uses the API route.
  • The messages component formats the chat messages for display.
  • Key Change (V5): Messages are now modeled as parts, including text and tool calls.
  • The code iterates through the message parts and renders them accordingly, using Markdown for text and displaying guitar recommendations based on tool call outputs.
  • Key Change (V5): The sendMessage function is now separate from the form post, providing more control over message sending.

Adding Speech Transcription

  • An API route (/api/transcribe) is created to handle audio transcription.
  • The route uses the OpenAI Whisper model via the AI SDK's transcribe function (experimental).
  • The route receives audio data as form data, converts it to a U8 int array, and passes it to the transcribe function.
  • A useTranscribe hook is created to wrap the react-media-recorder hook and handle audio recording and transcription.
  • The hook provides startRecording and stopRecording functions, and a transcribeAudio function that sends the audio data to the API.
  • A TranscribeButton component is created to provide a UI for starting and stopping the recording.
  • The transcribed text is appended to the existing input text in the chat interface.

Adding Text-to-Speech

  • An API route (/api/tts) is created to handle text-to-speech conversion.
  • The route uses the OpenAI TTS1 model via the AI SDK's generateSpeech function (experimental).
  • The route receives text as input and returns an MP3 audio file as a U8 int array.
  • A useTextToSpeech hook is created to handle the text-to-speech conversion.
  • The hook provides a convertTextToSpeech function that sends the text to the API and plays the resulting audio.
  • A TTSButton component is created to provide a UI for converting text to speech.
  • The TTSButton is added to the message formatting logic, allowing users to hear the AI's responses.

Conclusion

The video demonstrates the significant upgrades in Verscell's AI SDK version 5, particularly in handling UI/model message distinctions, providing flexible tool invocation control with stopWhen, and offering new transport mechanisms in the useChat hook. The addition of speech transcription and text-to-speech functionalities, leveraging OpenAI's Whisper and TTS1 models, showcases the SDK's expanded multimedia capabilities. The integration with Tanstack Start allows for a seamless development experience, making it easier to build sophisticated AI-powered applications. The key takeaway is that V5 offers a more refined and powerful set of tools for building AI-driven chat applications with enhanced control and flexibility.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.