Building Alice’s Brain: an AI Sales Rep that Learns Like a Human - Sherwood & Satwik, 11x

AI EngineerAbout 6 min readJul 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • SDR (Sales Development Representative): Entry-level sales role focused on sourcing leads, engaging them, and booking meetings.
  • Seller: The company using 11X's Alice to sell their products or services.
  • Lead: The potential customer being contacted by Alice.
  • Knowledge Base: A centralized repository of seller information that Alice uses to personalize emails.
  • Parsing: Converting non-text resources (e.g., PDFs, videos) into text for LLMs.
  • Chunking: Splitting large text documents into smaller, semantically meaningful units for embedding.
  • Vector Database: A database optimized for similarity search using vector embeddings.
  • RAG (Retrieval Augmented Generation): A framework for enhancing LLM outputs by retrieving information from external sources.
  • Deep Research Rag: An advanced RAG technique using agents to plan and execute multiple retrieval steps.
  • Hallucinations: Instances where an AI generates incorrect or fabricated information.

1. Introduction: Building Alice's Brain

  • 11X builds digital workers for go-to-market organizations, including Alice (AI SDR) and Julian (voice agent).
  • The talk focuses on Alice's brain, which is the knowledge base that powers her email generation.
  • Alice sends about 50,000 emails daily, compared to 20-50 for a human SDR, and runs campaigns for about 300 organizations.
  • The goal is to enable Alice to proactively pull context about the seller, rather than relying on manual input.

2. The Role of an SDR

  • SDRs source leads, engage them across channels, and book meetings.
  • Key metrics: positive replies and meetings booked.
  • A significant part of the job involves writing emails.
  • Example: An email written by Alice is shown.

3. Defining Seller and Lead

  • Seller: The company selling through Alice (11X's customer).
  • Lead: The person being sold to.
  • The seller provides context (products, case studies) to Alice, who then personalizes emails for leads.

4. Requirements for Alice's Success

  • Know the Seller: Products, services, case studies, pain points, value propositions, Ideal Customer Profile (ICP).
  • Know the Lead: Role, responsibilities, concerns, solutions tried, pain points, company.
  • The talk focuses on how Alice knows the seller.

5. The Old Approach: Manual Library

  • Sellers manually input information about their business into a "library."
  • Users entered details about products, services, pain points, and solutions.
  • These descriptions were used as context for Alice's emails.
  • Problems:
    • Tedious and cumbersome user experience.
    • Onboarding friction due to the need to fill out the library.
    • Suboptimal emails due to either too few or too many offers in the context.

6. The New Approach: Knowledge Base Overview

  • Alice proactively pulls context about the seller.
  • The knowledge base is a centralized repository for seller information.
  • Users upload source materials (marketing materials, case studies, sales calls, press releases).
  • Resources are categorized into documents, images, websites, and media (audio/video).

7. Knowledge Base Architecture

  • User uploads a resource to the client.
  • The resource is saved to an S3 bucket and sent to the backend.
  • The backend creates resources in the database and triggers jobs based on resource type and vendor.
  • Vendors asynchronously parse the resource and send a webhook with the parsed artifact.
  • The parsed artifact is stored in the database and upserted to Pinecone (vector database) for embedding.
  • UI is updated, and the agent can query Pinecone for the stored information.

8. Pipeline Step 1: Parsing

  • Parsing converts non-text resources into text (specifically, markdown).
  • Markdown contains structural information and formatting that is semantically meaningful.
  • The process makes non-text information legible to large language models (LLMs).
  • Five resource types (documents, images, websites, audio, video) are parsed into markdown.
  • 11X chose to work with vendors for parsing instead of building it from scratch due to complexity and time constraints.

9. Parsing Vendor Selection

  • Requirements:
    • Support for necessary resource types.
    • Markdown output.
    • Webhook support.
  • Initially, accuracy and comprehensiveness were not prioritized, assuming vendors were within a reasonable range.
  • Cost optimization was planned for after production.

10. Parsing Vendor Choices

  • Documents and Images: Llama Parse (Llama Index product)
    • Supported the most file types.
    • Excellent support from Jerry and the Llama Index team.
    • Example: Converting a 11X sales deck PDF into markdown.
  • Websites: Firecrawl
    • Familiarity from previous projects.
    • Tavi's crawl endpoint was still in development.
    • Example: Converting a homepage into markdown.
  • Audio and Video: Cloudglue
    • Supported both audio and video.
    • Extracted information from the video itself, not just transcribing audio.
    • Example: Converting YouTube videos and MP4 files into markdown.

11. Pipeline Step 2: Chunking

  • Splitting long markdown documents into semantic entities for embedding in the vector database.
  • Preserving the structure of the markdown (e.g., titles vs. paragraphs).
  • Chunking strategies:
    • Splitting on tokens.
    • Splitting on sentences.
    • Splitting on markdown headers.
    • Using LLMs to split the document.
  • Considerations:
    • What logical units to preserve.
    • What to extract during retrieval.
    • Whether to use different strategies for different resource types.
    • Expected query and retrieval strategies.
  • 11X used a combination of splitting on markdown headers, sentences, and tokens in a waterfall approach.

12. Pipeline Step 3: Storage

  • Storage technologies for RAG:
    • Graph databases.
    • Document databases.
    • Relational databases.
    • Key-value stores.
    • Object storage (e.g., S3).
    • Vector databases (used by 11X for similarity search).
  • 11X chose Pinecone (vector database) because:
    • Well-known solution.
    • Cloud-hosted.
    • Easy to get started with great guides and SDKs.
    • Bundled embedding models.
    • Excellent customer support.

13. Pipeline Step 4: Retrieval

  • Evolution of RAG techniques:
    • Traditional RAG (pulling information and enriching the system prompt).
    • Agentic RAG (using tools for information retrieval within an agentic flow).
    • Deep Research RAG (using deep research agents to plan and execute retrieval steps).
  • 11X built a deep research agent using Leta (cloud agent provider).
  • The agent receives lead information, creates a plan with retrieval steps, summarizes results, and generates a Q&A answer.

14. Pipeline Step 5: Visualization

  • Reassuring customers that Alice understands their business.
  • Interactive 3D visualization of the knowledge base in the product.
  • Vectors from Pinecone are projected down to three dimensions and rendered as nodes.
  • Users can click on nodes to view the associated chunk.
  • The UI shows resources and allows users to interrogate Alice about the knowledge base using a Leta-powered agent.
  • The knowledge base content shows up as a Q&A during campaign creation, with retrieved chunks displayed.

15. Conclusion and Lessons Learned

  • The knowledge base was a revolutionary project that improved the user experience.
  • Lessons learned:
    • RAG is complex.
    • Get to production before benchmarking and improving.
    • Lean on vendors for expertise and support.
  • Future plans:
    • Track and address hallucinations.
    • Evaluate parsing vendors on accuracy and completeness.
    • Experiment with hybrid RAG (graph database + vector database).
    • Reduce costs across the pipeline.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.