Investors Called This Entrepreneur's Idea 'Ridiculous.' Now They've Raised $100 Million

By Forbes

Share:

Key Concepts

  • Video as a Data Format: Emphasizing video's ubiquity (90% of global data) but often overlooked analytical potential.
  • Temporal Understanding: The ability to comprehend events and context within video over time, crucial for human-like understanding.
  • Multimodal AI: AI systems that process and integrate information from multiple modalities, such as visual, audio (including non-speech sounds), and language, simultaneously.
  • Foundation Models: Large-scale AI models trained on vast datasets that can be adapted for a wide range of downstream tasks.
  • Semantic Search: Searching for content based on meaning and context rather than just keywords.
  • Vector Embeddings: Numerical representations of data (like video segments) that capture their semantic meaning, allowing for efficient similarity search and retrieval.
  • Agentic Workflow: An automated process where an AI agent performs a series of tasks, often in response to a prompt, to achieve a specific goal.
  • Prompt-Based Highlight Creation: Using natural language prompts to instruct AI to generate specific video highlights.
  • Go-to-Market (GTM): The strategy a company uses to bring a product or service to market.
  • Large Language Models (LLMs): AI models trained on massive text datasets, primarily focused on language understanding and generation.
  • Computer Vision: A field of AI that enables computers to "see" and interpret visual information from images and videos.
  • Speech Recognition: The process of converting spoken language into text.
  • Horizontal Platform: A technology platform designed to be applicable across many different industries or use cases, rather than being specialized for one.
  • API (Application Programming Interface): A set of rules and protocols that allows different software applications to communicate with each other.
  • Proof of Concept (POC): A small-scale project to demonstrate the feasibility of a concept or idea.
  • Flywheel (Startup Context): A business model where different parts of the business reinforce each other, creating a self-sustaining growth loop.
  • Hyperscalers: Large cloud providers (e.g., Microsoft, AWS) with immense computing resources.

The Overlooked Power of Video Data and Its Challenges

So Young, Co-founder of 12 Labs, highlighted that video, despite constituting 90% of all data in the world and being integral to daily communication, storytelling, and work, is often a "forgotten data format" in terms of analysis and utilization. Video is crucial across diverse industries including media, sports, entertainment, advertising, the creator economy, law enforcement (evidence management, investigations), healthcare, automotive, and enterprise knowledge.

The core problem lies in traditional video understanding technologies. These typically rely on:

  1. Frame Understanding: Analyzing individual frames (e.g., 24-30 frames per second) to identify objects, which misses temporal context.
  2. Speech Recognition: Extracting transcripts to analyze dialogue, which ignores visual and non-linguistic audio cues.

Neither approach is sufficient because human comprehension of the world is inherently temporal (understanding changes and movement over time) and multimodal (integrating visual, auditory – including non-language sounds like clapping – and linguistic information). Large Language Models (LLMs), even today, cannot fully grasp video because language is a distinct data format; video is redundant, multimodal, and requires temporal understanding.

Current enterprise practices for managing video assets are largely manual, involving watching footage, taking notes, and tagging (often file-based, lacking in-scene or temporal context), with tags stored in spreadsheets. This makes searching for specific moments incredibly inefficient, akin to having "no control F for video." For example, a sports organization trying to find a touchdown with crowd applause and a specific logo in the background cannot easily describe and pinpoint that exact segment using traditional methods.


12 Labs' Multimodal Video Understanding Solution

12 Labs addresses this challenge by building multimodal video understanding foundation models. These models enable:

  • Semantic Search and Retrieval: Users can search across vast amounts of video data using natural language queries, finding exact scenes and moments.
  • Communication and Conversation with Video: The AI creates a "memory" of the footage, allowing users to ask questions and reason with the video content.

12 Labs' Foundation Models: Moringo and Pegasus

12 Labs employs two primary models:

  1. Moringo: This model creates the initial "memory" of video by seeing, hearing, and understanding its context. This memory is stored as vector embeddings, which capture the holistic context of the video and facilitate efficient search and retrieval of specific moments.
  2. Pegasus: This model builds upon Moringo's memory, enabling reasoning capabilities. It allows users to ask complex questions about the video, such as "describe what happens in this clip," "what's funny about this video," or "create a highlight from it," and receive intelligent, context-aware responses.

These two models work in tandem to provide a comprehensive solution for various video-centric use cases. 12 Labs provides this technology via APIs, allowing customers and partners to integrate and build on their platform.


Real-World Applications and Case Studies

The technology has significant applications, particularly in media, entertainment, and advertising:

Content Management Lifecycle

  • Content Creation: Transforming raw footage into a narrative becomes a "search problem." Creatives can quickly find relevant moments across vast archives to iterate on stories, a challenge faced by individual creators and large production houses alike.
  • Distribution: Enhances personalized content recommendations by understanding actual user behavior and preferences, leading to more relevant content discovery.
  • Content Repackaging & Monetization: Enables repurposing historical content into more digestible, personalized, and scalable formats for viewers.

Case Study: Maple Leaf Sports Entertainment (MLSE)

Maple Leaf Sports Entertainment (MLSE), one of Canada's largest sports entities (owning the Toronto Raptors and Maple Leafs), faced the challenge of efficiently accessing their extensive historical archives and game footage to create personalized content for a rapidly growing fan base.

MLSE implemented 12 Labs' video understanding technology to build an agentic workflow for prompt-based highlight creation. This allowed creative teams to type natural language prompts (e.g., "Vince Carter's 30th anniversary highlight") and instantly visualize potential highlights in real-time during meetings.

Impact:

  • Efficiency: The time required to create sports highlights was drastically reduced from 16 hours or days to just 9 minutes.
  • Scalability: Enabled MLSE to "do more as an organization" and scale their storytelling capabilities.
  • Personalization: Facilitated the creation of content that resonates with specific audiences, acting as a "content machine" for engagement.

This was highlighted as "one of the first use cases of true agent workflows built on top of video understanding," demonstrating how organizations can scale and innovate their content strategies. 12 Labs' platform is horizontal, meaning its AI can comprehend any type of footage, from animation and sports to dash cam videos.


The Founding Journey of 12 Labs

12 Labs was founded by five co-founders, who met as researchers and scientists building AI for language and image data. They identified video as an unsolved problem due to the lack of suitable technology. The company began building its technology in 2020 and officially incorporated in 2021.

Overcoming Investor Skepticism

In the early days, the concept of building "general purpose foundation models" that were multimodal (understanding sound, visual, and video) and offered via APIs was met with skepticism from investors, especially when LLMs were still nascent (pre-GPT3). The common startup advice was to focus on a narrow problem and verticalize, but 12 Labs pursued a horizontal, infrastructural AI approach.

Their breakthrough came when their CTO, Aiden (a Forbes 30 Under 30 recipient), used $50,000 in AWS compute credits (a startup grant) to train a "baby version" of their Moringo model. He then entered the ICCV Value Challenge, a prestigious AI video competition hosted by Microsoft, and won first place, outperforming Microsoft's state-of-the-art model. This victory attracted their first investor, Index Ventures, who had a strong thesis on multimodal technology, and later Dr. Fei-Fei Li, who became an investor and advisor. 12 Labs has since raised over $100 million.

Landing the First Customer

Given video's ubiquity, finding the first customer was challenging because "anyone could be your customer." 12 Labs proactively sought out potential clients. Their first customer, a manager for one of the largest sports leagues, was identified at a conference where he was coincidentally presenting on the very challenges of video tagging and archive management that 12 Labs aimed to solve.

The 12 Labs team prepared a customized demo using the organization's own YouTube videos. They approached the manager directly after his panel, conducting a demo on an iPad. This led to a $5,000 Proof of Concept (POC), which, while small, allowed 12 Labs to learn extensively from the customer as a "design partner," shaping their product and technology. This early partnership was crucial for validating their approach and exploring broader market applicability.


Advice for AI Startups in a Big Player Landscape

So Young offered key advice for small AI startups competing with "hyperscalers" like OpenAI and Anthropic:

  • Understand Your Problem Space Deeply: Startups must have a profound understanding of the specific problem they are solving.
  • Minimize Communication Cycles: Leverage the agility of a small team to rapidly integrate market feedback into product development, creating a tight "flywheel" between market, product, data, and research.
  • Focus and Prioritization: Large companies often have many products and internal silos, making it difficult to match a startup's speed, focus, and prioritization.
  • Dedicated Capital: Startups invest all their capital into solving one specific problem, a level of singular dedication that larger organizations cannot replicate across their diverse portfolios.
  • Vocalize and Specialize: Clearly articulate the problem being solved and relentlessly pursue that specific goal. This focus becomes the startup's unique strength.

11 Labs vs. 12 Labs

Addressing a common question, So Young clarified that 12 Labs was founded first, approximately a year before 11 Labs. They maintain a good relationship with the 11 Labs team and even co-hosted a "23 Labs multimodal AI hackathon" (combining their numbers) to celebrate multimodal AI. Another hackathon is planned in New York focusing on advertising technology.


Conclusion

12 Labs is at the forefront of solving the critical challenge of video understanding, transforming video from a "forgotten data format" into an actionable source of insight. By developing multimodal foundation models like Moringo and Pegasus, they enable semantic search, reasoning, and automated content creation at scale. Their success with clients like Maple Leaf Sports Entertainment demonstrates significant efficiency gains and new opportunities for personalization and engagement. The company's journey highlights the importance of deep problem understanding, agile execution, and unwavering conviction for startups aiming to innovate in the AI landscape, even against larger players.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video