Key Concepts
Generative AI, Multimodal Search, Semantic Understanding, Vector Embeddings, Prompt Engineering, Evals, Context Management, Model Economics, AWS Bedrock, AWS SageMaker, Neuron SDK, Tranium, Inferentia, Intelligent Document Processing, Generative UI.
Kalin's Approach to Generative AI: Building Cool Stuff
Kalin is a company that builds custom applications and solutions for clients ranging from Fortune 500 companies to startups. They focus on building cool things for their customers, leveraging technologies like generative AI. They emphasize that generative AI is not a magic bullet and that a practical, hands-on approach is crucial.
Examples of Kalin's Work:
- Brainbox AI: An agent for decarbonizing the built environment by managing HVAC systems in buildings. This project was recognized as one of the Times 100 best inventions.
- Simmons: Water management and conservation solutions using AI.
- Nature Footage: Multimodal search and semantic understanding of stock footage videos.
Multimodal Search and Semantic Understanding of Videos
Kalin is working with Nature Footage to index and make their stock footage searchable. This involves:
- Leveraging Nova Pro models to generate understandings, timestamps, and features of the videos.
- Storing all data in Elasticsearch.
- Building a pooling embedding by taking frame samples and pooling the embeddings of those frames.
- Using Titan v2 multimodal embeddings for text-based image search.
Real-Time Sports Footage Processing Architecture
Kalin processes sports footage in real-time and in batch archival. The architecture involves:
- Data Splitting: Separating audio and video.
- Audio Transcription: Generating transcriptions from the audio. A simple hack for finding highlights is to use
ffmpegto get an amplitude spectrograph of the audio and look for audience cheering. - Embedding Generation: Creating embeddings from both the text and the video.
- Behavior Identification: Identifying specific behaviors with a certain vector and confidence.
- Database Storage: Storing the identified behaviors in a database (Postgres PGVector, OpenSearch).
- Push Notifications: Sending push notifications to end-users via Amazon End User Messaging (SNS) when specific events (e.g., a three-pointer) are detected.
Key Insight: Annotating static camera angles with simple visual cues (e.g., a blue line on the three-pointer line) and asking targeted questions significantly improves video understanding model performance. SAM 2 (from Meta) can be used for automated annotation.
Randall Hunt's Background and Kalin's Services
Randall Hunt, the speaker, shares his background:
- Started with hacking and video games.
- Worked at NASA on physics projects.
- Joined Tenen (later MongoDB) before its IPO.
- Led the CI/CD team at SpaceX.
- Spent a long time at AWS, building various technologies.
Kalin offers services ranging from building chatbots and co-pilots to AI agents. They address customer needs in three main areas:
- Self-Service Productivity Tools: Building custom applications on top of existing tools, often for institutions with specific requirements.
- Automating Business Functions: Using intelligent document processing (IDP) with generative AI and custom classifiers to improve efficiency in processes like receipt and bill of lading processing.
- Monetization: Adding new AI-powered features to existing SaaS platforms.
Building a Moat in AI: Inputs, Outputs, and Evals
The speaker emphasizes the importance of defining the inputs and outputs of an AI system as the most fundamental aspect. He references a slide (possibly from Jason Louu or DSPY) that strategically identifies the specifications for building a moat in your business.
Key Steps:
- Define Inputs and Outputs: Clearly specify what data goes into the system and what results are expected.
- Evals Layer: Implement a robust evaluation process (Evals) to prove the system's reliability and performance.
- System Architecture: Design the overall architecture of the system.
- LLMs and Tools: Select the appropriate LLMs and tools, recognizing that these are incidental and subject to change.
The speaker channels Steve Ballmer's famous "developers" chant, replacing it with "evals, evals, evals, evals" to highlight the critical role of evaluation.
AWS Infrastructure for AI
Kalin builds AI solutions on AWS, utilizing services like:
- Bedrock: A service for accessing various pre-trained models.
- SageMaker: A comprehensive machine learning platform (comes at a compute premium).
- EKS/EC2: Alternative compute options for running models.
- Tranium and Inferentia: Custom silicon chips offering a 60% price-performance improvement over Nvidia GPUs. These require the Neuron SDK.
- Vector Stores: Postgres (preferred), OpenSearch, and MemoryDB (Redis). Redis offers extremely fast vector search but is expensive due to its RAM requirements.
Important Note: Amazon has reduced the prices of P4 and P5 instances by up to 40%, making GPUs more accessible.
Lessons Learned from Building AI Systems
The speaker shares key lessons learned from Kalin's experience:
- Eval and Embeddings are Not Enough: Understanding access patterns and user behavior is crucial.
- Embeddings Alone Don't Make a Great Query System: Faceted search and filters require additional tools like OpenSearch and Postgres.
- Speed Matters: Slow inference negatively impacts UX. Caching and UI enhancements can mitigate this.
- Know Your End Customer: Tailor solutions to the specific needs and constraints of the users.
- Prompt Engineering is Key: As models improve, prompt engineering becomes increasingly effective.
- Know Your Economics: Ensure that inference costs are sustainable.
Evals and Metrics
The speaker emphasizes the importance of creating effective evals. He suggests starting with a "vibe check" and then iteratively refining the data and tests. Metrics don't always need to be complex scores; a simple boolean (true/false) indicating success or failure can be sufficient.
UX and Context Management
UX is critical for the success of AI applications. Generative UI, as demonstrated with Cloud Zero, allows for dynamic and personalized rendering of information. Context management is also essential for differentiating applications. Injecting user-specific information (e.g., browsing history, cookies) into the model can lead to more strategic inferences.
Real-World Examples and Customer Insights
- Cloud Zero: Kalin built a chatbot for interacting with AWS infrastructure and is now using generative UI to render charts dynamically.
- Hospital System: Initially built a voice bot for nurses, but switched to a chat interface based on user feedback.
- Customer in Remote Areas: Instead of sending large PDF files, Kalin sends screenshots of relevant pages to accommodate low connectivity.
Prompt Engineering and Economics
The speaker briefly touches on prompt engineering best practices and emphasizes the importance of optimizing output tokens and costs. He also highlights the benefits of prompt caching, tool usage, and batch processing (e.g., batch on Bedrock offers a 50% discount). Optimizing context by minimizing irrelevant information is crucial for improving model performance and reducing costs.
Conclusion
Kalin focuses on building practical AI solutions for a variety of clients. They emphasize the importance of understanding the problem, defining clear inputs and outputs, implementing robust evaluation processes, and optimizing for both performance and cost. They leverage AWS services and custom silicon to build scalable and efficient AI applications. The key takeaways are the importance of a practical approach, the power of prompt engineering, and the need to prioritize user experience and economic viability.
AI summaries can miss context or contain errors. Check important details against the original video.