“Forget AI Agents, VOICE AGENTS are the future” – Jannis Moore

David OndrejAbout 12 min readMay 5, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Voice AI Agents: AI systems primarily interacting through voice, seen as the next evolution of AI agents.
  • AI Employees: Autonomous AI agents performing specific job roles within a business.
  • Emotional Intelligence (in Voice AI): The ability of voice AI to understand and respond to human emotion conveyed through voice, moving beyond simple text transcription.
  • Orchestration Layer: The technical system managing the flow of information in voice AI (Speech-to-Text -> LLM -> Text-to-Speech).
  • Multimodal Models: AI models capable of processing and generating information across multiple modalities (text, voice, image, video).
  • Agent-to-Agent (A2A) Communication / Multi-Agent Systems (MAS): Frameworks allowing different AI agents to communicate and collaborate.
  • State Management: Enabling AI agents to maintain context and memory across interactions or sessions.
  • Process Mining: Technology (like Celonis) used to analyze workflows and identify patterns, potentially applicable to optimizing AI agent tasks.
  • VIP Coding (Visually Integrated Programming): AI-assisted coding or low-code/no-code development, enabling faster prototyping and building.
  • Technical Fundamentals: Basic understanding of concepts like APIs, Requests, Webhooks, and JSON for building AI solutions.
  • Niche Selection: The strategy of focusing on a specific industry or problem to build a targeted AI solution.
  • Open Source AI: AI models and tools with publicly available code or weights (e.g., Sesame TTS).
  • AGI (Artificial General Intelligence): Hypothetical future AI with human-like cognitive abilities across diverse tasks.
  • Societal Impacts: Effects of AI on jobs, privacy, security, bias, government control, and decentralization.

The Future is Voice AI

Janice, a German entrepreneur with experience building thousands of AI agents, argues that voice AI agents represent the next significant evolution in artificial intelligence. She states, "They're the most powerful representation of ourselves," predicting they will become indistinguishable from humans within the year. The core advantage of voice is its speed and naturalness for communication compared to text (estimated 200-300 words per minute speaking vs. 40-50 typing). This allows users to interact with technology more efficiently, potentially commanding an "AI employee army" to perform tasks automatically.

AI Employees and Practical Applications

Janice's company builds voice AI for businesses and has deployed its first "digital employee." This AI agent automates the process of creating and distributing community updates and social media posts.

  • Process: Janice sends a Telegram message with the topic (e.g., "OpenAI released XYZ"), resource URL, and instructions.
  • AI Actions: The agent performs ideation, research (using the URL), content creation tailored to different platforms, posting/publishing, and then asks for confirmation, allowing for corrections.
  • Key Principle: Janice emphasizes starting AI deployment with highly repetitive tasks due to the current variability and need for control in AI systems. The goal is to outsource tasks that don't require constant human oversight.

The Technology Behind Voice AI

The effectiveness of voice AI depends on both the underlying Large Language Model (LLM) and the voice synthesis (Text-to-Speech, TTS).

  • LLMs: While important, the specific LLM choice (e.g., OpenAI's GPT-4.0 or 4.1) depends on the task complexity, especially for "advanced reasoning" involving multiple "tool calls" (features/actions the AI can perform). GPT-4.1 is noted as particularly good for instruction following.
  • Voice Quality & Emotion: Sounding natural and conveying emotion is crucial. Early TTS like ElevenLabs was good but lacked "emotional intelligence."
  • Orchestration Layer: The current process involves Speech-to-Text (STT) -> LLM -> TTS. Janice explains this "loses all of the emotional knowledge" during the text conversion stages.
  • The Breakthrough: New methods are emerging to bypass text conversion, potentially "tokenizing the emotion" directly from speech, sending these emotional tokens through the system, and decoding them back into emotionally nuanced speech. Janice predicts, "that's going to be the future of 2025 specifically emotional intelligence which becomes just such a massively big market." This unlocks use cases like AI therapists or agents responding differently based on the user's tone (angry vs. chill).

Overcoming Voice AI Limitations

A potential bottleneck for voice interaction is its unsuitability for public environments. Solutions discussed include:

  • Whispering: Using models like OpenAI's Whisper to interact quietly.
  • Delegation: Giving a brief voice command, which then triggers AI agents to communicate and coordinate amongst themselves to complete the task.

The Future Vision: AI Agent Armies and Interoperability

The discussion envisions a future where individuals have numerous AI agents working for them, delegated by a master agent.

  • Concept: An "AI employee army" managed via voice commands during simple activities like a walk.
  • Interoperability: Technologies like Multi-Agent Collaboration Platform (MCP) and Agent-to-Agent (A2A) communication are key but currently lack robust standards.
  • Challenge: Giving a single agent access to too many tools (e.g., thousands of Zapier integrations) can lead to confusion, hallucinations, and suboptimal performance.
  • Needed Advancements:
    • State Management: Agents need sessions to store temporary information and pass context effectively between tasks, improving stability and enabling more general AI capabilities beyond predefined repetitive tasks. OpenAI's Assistant API does this partially, but cross-agent state management is needed.
    • Process Mining for Agents: Systems (analogous to Celonis for human workflows) are needed to monitor agent performance, identify successful pathways, and allow for feedback (reinforcement learning) to optimize multi-agent systems dynamically. This mirrors how synapses strengthen in the brain.

Web Interaction and Tool Integration

Instead of websites needing specific AI versions, the focus will be on tools and services (CRMs, accounting software) having standardized APIs or servers (like MCP or A2A) that agents can interact with. Agents can already scrape information from standard websites effectively.

Common Business Use Cases for Voice AI

Janice identifies three consistently effective use cases across industries:

  1. Lead Reactivation: Re-engaging old or cold leads. (Case Study: A client campaign generated $500,000 from 20,000-30,000 automated calls).
  2. Customer Support: Handling inquiries and providing assistance.
  3. Lead Qualification: Screening potential leads to determine fit. The market is shifting from exploring the technology to building niche Software-as-a-Service (SaaS) solutions targeting these specific, high-value problems.

How to Start Building and Selling Voice AI

Janice outlines a methodology for aspiring AI entrepreneurs:

  1. Identify Pain Points: Leverage existing industry knowledge or research (using AI/Google) to find significant problems businesses face, particularly those involving phone communication.
  2. Market the Problem: Focus outreach on the specific pain point and the quantifiable value of solving it, not just the "fancy AI tech."
  3. Outreach: Use cold emailing, direct outreach, and leverage personal networks. Specificity is key; vague promises of "automation" are ineffective.
  4. Build & Gather Data: Create the initial solution and meticulously collect data on its performance. Janice quotes: "In God we trust, for the rest we have data."
  5. Quantify Value: Calculate the tangible benefits (e.g., hours saved, cost reduction, revenue generated) compared to the status quo (human labor).
  6. Package the Offer: Price the solution based on the demonstrated value, often as a percentage of the savings or earnings generated for the client.
  • Example (Dental Office): Frame the AI not as replacing the beloved front desk staff, but as augmenting them – removing repetitive and negative calls, allowing them to focus on higher-value tasks, thus increasing job satisfaction. This empathetic approach is more effective for SMBs than threatening replacement.

Advancements in LLMs: GPT-4.1

Both speakers note GPT-4.1 (specifically the API version, likely referring to gpt-4o or similar recent models) as a significant step forward due to its superior instruction-following capabilities and large context window (1 million tokens mentioned). It's very literal, requiring precise prompting. Prompting strategies often need adjustment with each new major model release.

Tools and Platforms for Building Voice Agents

The choice of platform depends on technical expertise:

  • Non-Technical: Voiceflow (mentioned as "Synflow" in transcript) is accessible but offers less granular control, potentially limiting success rates for complex agents.
  • Slightly Technical (Recommended): Vapi, Retell AI provide more control and higher potential success rates. They don't require coding but necessitate understanding web concepts.

Demystifying Technical Concepts for AI Builders

Many beginners fear technical aspects. Janice argues:

  • JSON vs. Coding: JSON is structured data, a way to format information for machines, not complex programming logic. It's relatively easy to learn.
  • AI Assistance: Tools like ChatGPT (especially GPT-4o with its advanced reasoning and tool use) and AI-integrated IDEs like Cursor significantly lower the barrier to entry for technical tasks, even writing necessary code snippets. Platforms like E2B allow AI to execute generated code.
  • Value of Basics: Understanding fundamental concepts is crucial even if not coding deeply.

Essential Technical Concepts

Key terms to understand for building web-connected AI agents:

  • Request: The fundamental action of asking for data from a server (like opening a website).
  • API (Application Programming Interface): A structured way for software systems to communicate, often involving requests that return data in formats like JSON.
  • Webhook: An automated message sent from one system to another when a specific event occurs (event-driven API).
  • JSON (JavaScript Object Notation): A standard text-based format for representing structured data.
  • Voice AI Specific: Orchestration Layer, STT (Speech-to-Text), TTS (Text-to-Speech).

The Evolution of Voice AI Architecture

The current STT -> LLM -> TTS "orchestration layer" is evolving to eliminate the text conversion steps, which lose emotional information. The future involves direct speech-to-token (including emotional data) and token-to-speech processing, likely using multimodal models, preserving the nuances of human communication.

Should You Learn to Code in the Age of AI?

  • Janice's Advice: Depends on motivation. If you genuinely enjoy development, yes – there's value in understanding systems deeply, fixing legacy code, or developing new protocols. If purely for money, there are likely faster paths (like building AI solutions using existing tools).

The Role of VIP Coding / AI-Assisted Development

Visual Integrated Programming (VIP Coding) or AI-assisted development is valuable, despite criticism.

  • Benefits: Enables rapid prototyping, building functional MVPs (David mentions his own VIP-coded startup reaching 5 figures/month), and crucially, learning.
  • Key to Success: Actively engage with the process, try to understand the underlying concepts (e.g., code structure, modularity), rather than blindly accepting AI suggestions. This improves outcomes and builds valuable knowledge.
  • Importance for Founders: Even non-technical founders benefit from a high-level understanding to manage developers effectively and make informed decisions.

Multimodal Models and Their Importance

Models processing multiple data types (text, voice, image, video) like Google's Gemini or potentially Llama 3 are becoming increasingly important. They simplify interactions but raise concerns about potential monopolies and ecosystem lock-in by major players like Google integrating AI deeply into their existing product suites (Workspace, etc.).

Privacy, Security, and the Risks of Centralized AI

Significant concerns were raised about data privacy and control:

  • Data Collection: Tech companies increasingly act like governments, demanding user data and ID.
  • Government Access: Centralized AI agents, knowing users intimately, could become targets for government surveillance, potentially revealing personal weaknesses and information. AI companies may be legally compelled to cooperate.
  • Security Risks: Installing unknown software (like MCP/A2A servers) without understanding is dangerous. Basic cybersecurity (password managers, skepticism) is crucial. Voice cloning scams using AI are becoming sophisticated.

Open Source AI: Potential and Challenges

Open source offers transparency but faces hurdles:

  • Incomplete Openness: Many models are "open weight" but lack open training data, hindering bias analysis and replication. This may be due to protecting proprietary data, fear of misuse, or legal concerns over scraped data.
  • Sesame TTS: An exciting open-source emotive TTS model, demonstrating the potential of new techniques. However, only a smaller version was released, and the open nature means safety features (like watermarking) can be easily removed, facilitating misuse (e.g., undetectable voice scams).

AI Bias, Misinformation, and Societal Impact

AI models inherit and can amplify biases present in their training data.

  • Risks: Subtle manipulation of opinions (political bias), perpetuation of misinformation (e.g., flawed health narratives on cholesterol, meat, influenced by historical lobbying).
  • Need for Critical Thinking: Users must be aware of potential biases and not blindly trust AI outputs, especially on critical topics.

AI Making People Dumber/Smarter?

AI could widen the gap:

  • Smarter: Engaged users leverage AI as a tool to learn faster, become more productive, and gain "superpowers."
  • Dumber/Dependent: Passive users might outsource critical thinking and basic decision-making to AI, becoming overly reliant and susceptible to manipulation or bias. The danger lies in losing the ability to think independently.

AI Agents in 12 Months: Predictions

Janice expects:

  • More capable "AI employees" handling complex, repetitive tasks within businesses.
  • Emergence of "agentic workforces" composed of specialized agents, possibly from different vendors.
  • Leaner human teams focused on management, oversight, and strategic decision-making.
  • Potential development of "manager agents" that monitor and optimize the performance of other agents.

The Future of Work: Lean Teams and Solo Entrepreneurs

The speakers agree with Naval Ravikant's idea of a future with many more small (1-3 person) or solo companies. AI can automate the "boring," repetitive tasks common in many profitable businesses, allowing individuals or small teams to operate complex ventures. The human role shifts towards creativity, strategy, decision-making, and judgment. AI handling bureaucracy (contracts, taxes) is a highly anticipated development.

Is AGI Near?

Janice believes true AGI is not close (estimating 5-10 years). Key missing pieces include robust mechanisms for:

  • Refining AI actions based on feedback (understanding "good" vs. "bad" outcomes).
  • Hierarchical control systems where agents manage and learn from sub-agents.
  • The ability for AI systems to autonomously deploy and train new agents based on learned objectives.

System Vulnerability and Personal Preparedness

Concerns about societal reliance on fragile systems (power grid, internet) were discussed. Personal preparedness (solar power, food, water, security) and building resilient, location-independent businesses are seen as prudent strategies.

Thoughts on Europe's Trajectory

Both speakers expressed concern about the direction of Western Europe, citing high taxes, excessive bureaucracy, counterproductive policies (e.g., energy, immigration), and a decline in safety and traditional values. Eastern Europe was viewed more favorably for its culture and lifestyle, despite potential political issues. They argued for policies that incentivize wealth creation and work, rather than dependency. The power of changing one's environment (moving countries) was highlighted as a key strategy for entrepreneurs seeking freedom and better opportunities (e.g., moving to Dubai).

Niche Selection Nuance

While focusing on a niche is helpful for clarity and purpose, Janice advises beginners not to be overly rigid. Gathering data and experience across initial projects is paramount. Once a specific problem/solution proves successful and data supports it, then doubling down on that niche becomes highly effective for crafting strong offers and scaling. Consistency and action are more important than finding the "perfect" niche initially.

Conclusion/Synthesis

The conversation paints a picture of rapid advancement in AI, particularly with voice agents poised to become integral tools due to their natural interaction style and increasing emotional intelligence. While the technology (LLMs like GPT-4.1, new voice architectures, multimodal capabilities) is evolving quickly, significant challenges remain in areas like agent interoperability, state management, security, privacy, and mitigating bias. Entrepreneurs are encouraged to focus on solving specific, high-value business problems (like lead reactivation or qualification) using AI, leveraging data to refine their offerings, and understanding basic technical concepts even when using AI-assisted tools. The rise of AI is expected to shift the nature of work towards smaller, more agile teams and individuals focused on high-level strategy and decision-making, while automating repetitive tasks. However, this transition brings societal risks related to job displacement, potential misuse of powerful AI tools (like voice cloning), centralized control, and the amplification of misinformation, necessitating critical thinking, robust security practices, and a push for transparency. AGI is considered further out (5-10 years), requiring breakthroughs in agent learning and self-improvement.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.