Add Gemini Live agents to your video conferencing with Fishjam
By Google for Developers
Key Concepts
- Fishjam: A low-latency video conferencing and live-streaming API.
- Fishjam Agents: A feature allowing backend services to join a video conference as a standard participant.
- Gemini 3.1 Flash Live: A multimodal AI model capable of real-time audio and video processing.
- Multimodal Live API: Google’s interface for streaming text, audio, and video to/from AI models via WebSockets.
- tRPC: A framework for end-to-end typesafe APIs, used here for backend-frontend communication.
- Codec Settings: Specific configurations required to ensure audio/video compatibility between the AI model and the conferencing server.
1. Architecture and Integration Overview
The application functions as a real-time speech-to-speech bridge. The architecture involves:
- Frontend: A React application using Fishjam SDK hooks to manage room state, device toggling, and peer tracking.
- Backend: A Fastify server with tRPC endpoints that manages room creation, peer token generation, and the Gemini session lifecycle.
- Communication: The system uses WebSockets to stream audio and video data between the Fishjam conference room and the Gemini Live API.
2. Step-by-Step Implementation Process
- Initialization: Install
Fishjam Server SDKandGoogle GenAI SDK. - Room Management: Create a room via the Fishjam API and map the room name to a unique ID in the backend cache.
- Agent Creation:
- Initialize the agent in the room with specific codec settings to match Gemini’s input requirements.
- Create tracks for the agent to send/receive audio.
- Data Streaming:
- Audio: Listen to
trackDataevents from Fishjam, convert audio chunks to Base64, and send them to Gemini. Conversely, convert Gemini’s Base64 audio responses into JavaScript buffers to play back in the room. - Video: Capture frames from the user's video track at 1 FPS and send them to Gemini for visual analysis.
- Audio: Listen to
- Cleanup: Implement a function to remove listeners, delete tracks, and disconnect the session to prevent resource leaks.
3. Key Technical Features
- Interruption Handling: The system monitors Gemini’s output; if the model signals an interruption, the backend flushes the current audio buffer in the Fishjam room to stop playback immediately.
- Selective Subscriptions: While the demo uses a global audio listener, the speaker notes that
Fishjam Selective Subscriptions APIcan be used for more granular control over which participants the AI hears. - System Prompting: The agent’s behavior is controlled via system instructions. The presenter demonstrated "guardrails" by restricting the agent to only discuss bananas, effectively ignoring or redirecting queries about other topics (e.g., oranges or calculus).
- Tool Use: The integration enables Google Search and custom function calling, such as the
disconnectfunction, which allows the AI to terminate its own session.
4. Notable Quotes
- "We are building a real-time speech-to-speech agent using Fishjam and Google's Multimodal Live API. To achieve our goal, what we need to do is basically take the audio from Fishjam and send it to Gemini, wait for Gemini's response, and put the audio back in the Fishjam conference room." — Adrian Czerwiec
- "The codec settings for its output need to match the exact settings that Gemini expects on its input." — Adrian Czerwiec (emphasizing the importance of technical alignment).
5. Real-World Application
The presenter demonstrated the agent's capabilities through:
- Contextual Awareness: The agent accurately described the presenter's physical environment (purple lighting, plants, "Software Mansion" sign).
- Dynamic Information Retrieval: Successfully performed a live Google search for weather in Kraków.
- Multilingual Support: Demonstrated real-time translation and conversation in Polish and German.
- Constraint Enforcement: Showcased how system prompts can be used to create specialized AI personas (e.g., the "banana-only" agent).
Synthesis/Conclusion
The integration of Fishjam and Gemini 3.1 Flash Live provides a robust framework for developers to inject AI participants into video conferences. By leveraging WebSockets for low-latency communication and utilizing Fishjam’s agent feature, developers can create interactive, multimodal AI agents that see, hear, and speak in real-time. The process is highly accessible via the Fishjam sandbox, requiring only basic API keys and a structured backend to handle the audio/video stream synchronization.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

99% Follow Goals, Only 1% Do this
Him-eesh Madaan

Why Does This Guy Appear In Kids Videos?
sphynx

NVIDIA Monopoly is DEAD | OPEN-SOURCE Chips Are HERE!
Hefty LLM

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

¿Trabajas en Oficina? EL ERROR que comete el 99% con Julieta Manzano | Martha Debayle
Martha Debayle

How East India Company Captured India | Nitish Rajput | Hindi
Nitish Rajput @

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED