Key Concepts
- Verscell AI SDK version 5: Upgraded AI development kit with polished APIs and new features.
- Tanstack Start: A full-stack framework for building web applications.
- Speech Transcription: Converting spoken audio into written text using AI.
- Text-to-Speech (TTS): Converting written text into spoken audio using AI.
- Tools: Functions or APIs that the AI can use to interact with external data or services.
- UI Messages vs. Model Messages: Distinction between the structure of messages displayed in the UI and the format required by the AI model.
- Stop When: New functionality in V5 that allows for more flexible control over when the AI stops processing tool invocations.
- React Media Recorder: A React library for handling microphone input and audio recording.
- OpenAI Whisper: An OpenAI model for speech transcription.
- OpenAI TTS1: An OpenAI model for text-to-speech conversion.
Creating a Tanstack Start App with AI Chat Example
- The video demonstrates how to create a new Tanstack Start application using the
create-start-appcommand. - The
--example tanstack-chatflag is used to set up a pre-built AI chat example. - The example features a guitar recommendation system, showcasing the use of AI tools to access and interact with a database.
- The application requires setting up environment variables for API keys, specifically for Anthropic and OpenAI.
Backend Implementation (API Endpoint)
- The backend API endpoint (
/api/demo/chat) handles incoming chat requests. - It uses the
streamTextfunction from the AI SDK to connect to the Anthropic model. - Key Change (V5): The code converts UI messages to model messages using a cleaner, more distinct approach.
- Key Change (V5): The
maxStepsparameter is replaced with thestopWhenfunctionality, allowing for more flexible control over tool invocations. - The
toUIMessageStreamResponsefunction converts model messages back to UI messages for display. - The API returns a raw streamed response, which integrates seamlessly with Tanstack Start.
Tools Implementation
- The example uses two tools:
getGuitarsandrecommendGuitar. - Key Change (V5): The
parametersfield in tool definitions is renamed toinputSchema. AnoutputSchemais also available. - The
recommendGuitartool is a UI tool that instructs the AI to display a specific guitar card. - The tool's
executefunction marshals the input ID into the output.
Frontend Implementation (UI)
- The
useChathook is used to manage the chat interface. - Key Change (V5): The
useChathook no longer tracks the input field; this must be managed separately usinguseState. - Key Change (V5): The
useChathook now uses atransportmechanism, allowing for different methods of communicating with the AI (e.g., fetch, server functions, local OpenAI calls). The default isdefaultChatTransportwhich uses the API route. - The
messagescomponent formats the chat messages for display. - Key Change (V5): Messages are now modeled as parts, including text and tool calls.
- The code iterates through the message parts and renders them accordingly, using Markdown for text and displaying guitar recommendations based on tool call outputs.
- Key Change (V5): The
sendMessagefunction is now separate from the form post, providing more control over message sending.
Adding Speech Transcription
- An API route (
/api/transcribe) is created to handle audio transcription. - The route uses the OpenAI Whisper model via the AI SDK's
transcribefunction (experimental). - The route receives audio data as form data, converts it to a U8 int array, and passes it to the
transcribefunction. - A
useTranscribehook is created to wrap thereact-media-recorderhook and handle audio recording and transcription. - The hook provides
startRecordingandstopRecordingfunctions, and atranscribeAudiofunction that sends the audio data to the API. - A
TranscribeButtoncomponent is created to provide a UI for starting and stopping the recording. - The transcribed text is appended to the existing input text in the chat interface.
Adding Text-to-Speech
- An API route (
/api/tts) is created to handle text-to-speech conversion. - The route uses the OpenAI TTS1 model via the AI SDK's
generateSpeechfunction (experimental). - The route receives text as input and returns an MP3 audio file as a U8 int array.
- A
useTextToSpeechhook is created to handle the text-to-speech conversion. - The hook provides a
convertTextToSpeechfunction that sends the text to the API and plays the resulting audio. - A
TTSButtoncomponent is created to provide a UI for converting text to speech. - The
TTSButtonis added to the message formatting logic, allowing users to hear the AI's responses.
Conclusion
The video demonstrates the significant upgrades in Verscell's AI SDK version 5, particularly in handling UI/model message distinctions, providing flexible tool invocation control with stopWhen, and offering new transport mechanisms in the useChat hook. The addition of speech transcription and text-to-speech functionalities, leveraging OpenAI's Whisper and TTS1 models, showcases the SDK's expanded multimedia capabilities. The integration with Tanstack Start allows for a seamless development experience, making it easier to build sophisticated AI-powered applications. The key takeaway is that V5 offers a more refined and powerful set of tools for building AI-driven chat applications with enhanced control and flexibility.
AI summaries can miss context or contain errors. Check important details against the original video.





