THE SUMMARYAI-generated
Key Concepts
- SDR (Sales Development Representative): Entry-level sales role focused on sourcing leads, engaging them, and booking meetings.
- Seller: The company using 11X's Alice to sell their products or services.
- Lead: The potential customer being contacted by Alice.
- Knowledge Base: A centralized repository of seller information that Alice uses to personalize emails.
- Parsing: Converting non-text resources (e.g., PDFs, videos) into text for LLMs.
- Chunking: Splitting large text documents into smaller, semantically meaningful units for embedding.
- Vector Database: A database optimized for similarity search using vector embeddings.
- RAG (Retrieval Augmented Generation): A framework for enhancing LLM outputs by retrieving information from external sources.
- Deep Research Rag: An advanced RAG technique using agents to plan and execute multiple retrieval steps.
- Hallucinations: Instances where an AI generates incorrect or fabricated information.
1. Introduction: Building Alice's Brain
- 11X builds digital workers for go-to-market organizations, including Alice (AI SDR) and Julian (voice agent).
- The talk focuses on Alice's brain, which is the knowledge base that powers her email generation.
- Alice sends about 50,000 emails daily, compared to 20-50 for a human SDR, and runs campaigns for about 300 organizations.
- The goal is to enable Alice to proactively pull context about the seller, rather than relying on manual input.
2. The Role of an SDR
- SDRs source leads, engage them across channels, and book meetings.
- Key metrics: positive replies and meetings booked.
- A significant part of the job involves writing emails.
- Example: An email written by Alice is shown.
3. Defining Seller and Lead
- Seller: The company selling through Alice (11X's customer).
- Lead: The person being sold to.
- The seller provides context (products, case studies) to Alice, who then personalizes emails for leads.
4. Requirements for Alice's Success
- Know the Seller: Products, services, case studies, pain points, value propositions, Ideal Customer Profile (ICP).
- Know the Lead: Role, responsibilities, concerns, solutions tried, pain points, company.
- The talk focuses on how Alice knows the seller.
5. The Old Approach: Manual Library
- Sellers manually input information about their business into a "library."
- Users entered details about products, services, pain points, and solutions.
- These descriptions were used as context for Alice's emails.
- Problems:
- Tedious and cumbersome user experience.
- Onboarding friction due to the need to fill out the library.
- Suboptimal emails due to either too few or too many offers in the context.
6. The New Approach: Knowledge Base Overview
- Alice proactively pulls context about the seller.
- The knowledge base is a centralized repository for seller information.
- Users upload source materials (marketing materials, case studies, sales calls, press releases).
- Resources are categorized into documents, images, websites, and media (audio/video).
7. Knowledge Base Architecture
- User uploads a resource to the client.
- The resource is saved to an S3 bucket and sent to the backend.
- The backend creates resources in the database and triggers jobs based on resource type and vendor.
- Vendors asynchronously parse the resource and send a webhook with the parsed artifact.
- The parsed artifact is stored in the database and upserted to Pinecone (vector database) for embedding.
- UI is updated, and the agent can query Pinecone for the stored information.
8. Pipeline Step 1: Parsing
- Parsing converts non-text resources into text (specifically, markdown).
- Markdown contains structural information and formatting that is semantically meaningful.
- The process makes non-text information legible to large language models (LLMs).
- Five resource types (documents, images, websites, audio, video) are parsed into markdown.
- 11X chose to work with vendors for parsing instead of building it from scratch due to complexity and time constraints.
9. Parsing Vendor Selection
- Requirements:
- Support for necessary resource types.
- Markdown output.
- Webhook support.
- Initially, accuracy and comprehensiveness were not prioritized, assuming vendors were within a reasonable range.
- Cost optimization was planned for after production.
10. Parsing Vendor Choices
- Documents and Images: Llama Parse (Llama Index product)
- Supported the most file types.
- Excellent support from Jerry and the Llama Index team.
- Example: Converting a 11X sales deck PDF into markdown.
- Websites: Firecrawl
- Familiarity from previous projects.
- Tavi's crawl endpoint was still in development.
- Example: Converting a homepage into markdown.
- Audio and Video: Cloudglue
- Supported both audio and video.
- Extracted information from the video itself, not just transcribing audio.
- Example: Converting YouTube videos and MP4 files into markdown.
11. Pipeline Step 2: Chunking
- Splitting long markdown documents into semantic entities for embedding in the vector database.
- Preserving the structure of the markdown (e.g., titles vs. paragraphs).
- Chunking strategies:
- Splitting on tokens.
- Splitting on sentences.
- Splitting on markdown headers.
- Using LLMs to split the document.
- Considerations:
- What logical units to preserve.
- What to extract during retrieval.
- Whether to use different strategies for different resource types.
- Expected query and retrieval strategies.
- 11X used a combination of splitting on markdown headers, sentences, and tokens in a waterfall approach.
12. Pipeline Step 3: Storage
- Storage technologies for RAG:
- Graph databases.
- Document databases.
- Relational databases.
- Key-value stores.
- Object storage (e.g., S3).
- Vector databases (used by 11X for similarity search).
- 11X chose Pinecone (vector database) because:
- Well-known solution.
- Cloud-hosted.
- Easy to get started with great guides and SDKs.
- Bundled embedding models.
- Excellent customer support.
13. Pipeline Step 4: Retrieval
- Evolution of RAG techniques:
- Traditional RAG (pulling information and enriching the system prompt).
- Agentic RAG (using tools for information retrieval within an agentic flow).
- Deep Research RAG (using deep research agents to plan and execute retrieval steps).
- 11X built a deep research agent using Leta (cloud agent provider).
- The agent receives lead information, creates a plan with retrieval steps, summarizes results, and generates a Q&A answer.
14. Pipeline Step 5: Visualization
- Reassuring customers that Alice understands their business.
- Interactive 3D visualization of the knowledge base in the product.
- Vectors from Pinecone are projected down to three dimensions and rendered as nodes.
- Users can click on nodes to view the associated chunk.
- The UI shows resources and allows users to interrogate Alice about the knowledge base using a Leta-powered agent.
- The knowledge base content shows up as a Q&A during campaign creation, with retrieved chunks displayed.
15. Conclusion and Lessons Learned
- The knowledge base was a revolutionary project that improved the user experience.
- Lessons learned:
- RAG is complex.
- Get to production before benchmarking and improving.
- Lean on vendors for expertise and support.
- Future plans:
- Track and address hallucinations.
- Evaluate parsing vendors on accuracy and completeness.
- Experiment with hybrid RAG (graph database + vector database).
- Reduce costs across the pipeline.
AI summaries can miss context or contain errors. Check important details against the original video.





