Exclusive: Inside the Best AI Model for Coding and Writing | Scott White (Anthropic)
By Peter Yang
Key Concepts
- Hybrid Reasoning Model: A model capable of both instant responses and step-by-step thinking for complex tasks.
- Extended Thinking: A mode where the model dedicates more computational resources and time to solve a problem.
- Dogfooding: Using one's own products or services to test and improve them.
- Evals: Evaluations used to measure the performance and quality of model outputs, especially in AI product development.
- Model CICD: Continuous Integration and Continuous Delivery for AI models, ensuring consistent performance and quality.
- Talent Stack Collapse: The blurring of traditional role boundaries, where individuals can perform tasks across multiple disciplines.
- Constitutional AI: Training AI systems to adhere to a set of principles or "constitution" to ensure ethical and safe behavior.
- Model Features: Features closely integrated with the model to drive specific outcomes.
Claude 3.7 Sonnet: A Hybrid Reasoning Model
Anthropic's Claude 3.7 Sonnet is introduced as a "hybrid reasoning model," designed to mimic human thinking by offering both quick, intuitive responses and slower, more deliberate analysis for complex problems. This is achieved through toggles in the user interface and API, allowing users to choose between instant responses and "Extended Thinking."
- Thinking Fast and Slow: The model is designed with two modes of operation, similar to the human brain's "thinking fast and slow" concept.
- API Control: API customers can adjust the "budget versus thinking frontier," controlling how long and hard the model thinks.
- Practical Use Cases: Claude 3.7 is optimized for real-world tasks, particularly in the workplace, rather than competition-style math or computer science problems.
- Software Engineering Focus: The model excels at production software engineering and coding, benefiting both Anthropic's internal development and its customers.
- Visualization Capabilities: Surprisingly strong at creating visualizations, aiding in product development and communication.
Training the Model for Real-World Use Cases
The process of training Claude to excel in practical use cases is described as similar to traditional product development.
- Strong Vision: Having a clear vision of the problems the model should solve and the target customers.
- Dogfooding: Internally using early versions of the model to identify and address shortcomings.
- Customer Feedback Loops: Maintaining close contact with API customers and internal users to gather feedback and refine the model.
- Prioritization: Focusing on specific reinforcement learning environments and data relevant to key use cases.
- Product Development Analogy: The overall process is likened to product development, with the model being a new type of product component.
- Andy Rackliffe Quote: "The only thing that matters is if the dogs are eating the dog food," emphasizing the importance of internal adoption.
Practical Applications and Use Cases
Scott discusses how he uses Claude in his daily work, highlighting the differences between the standard model and the extended thinking mode.
- Standard Claude: Used for document creation, summarization, PRDs, and evaluations.
- Extended Thinking: Employed for complex tasks involving multiple documents and connecting disparate information.
- Claude Code: A research preview product for coding, where extended thinking is particularly valuable.
- Bug Fix Example: A real-world example of quickly fixing a UI contrast issue, demonstrating the model's impact on internal workflows.
- Visualization Use: Creating flowcharts, user journey charts, and even full React components for product specifications.
The Role of Personality in AI Models
The discussion shifts to the importance of personality in AI models, particularly Claude's "niceness" and collaborative nature.
- Commoditization Concerns: Addressing concerns about AI models becoming commodities.
- Multi-Model Approach: Businesses are increasingly adopting a multi-model approach, using different models for different tasks.
- Model Personality: Claude's personality is seen as a key differentiator, driven by specific character traits like curiosity, truth-seeking, and open-mindedness.
- Amanda Asal: The "Claude character leader" who guides the model's personality development.
- Customer Relationships: A pleasant personality fosters stronger customer relationships, influencing adoption decisions for companies like Intercom, DoorDash, and Lyft.
Collaboration with Research Teams
Scott explains how the product team collaborates with research teams at Anthropic.
- Four-Legged Stool: Traditional product development (product, design, engineering) is now a four-legged stool with research as the fourth leg.
- Shared Vision: A consistent vision for the product and company is crucial for aligning research and product efforts.
- Early Risk Mitigation: Building early prototypes and maintaining a portfolio of research bets.
- Feedback Loops: Embedding researchers in product teams to pipeline learnings and address gaps.
- Model Features: Building features that are closely integrated with the model, requiring close collaboration with research.
- Single-Threaded Ownership: Researchers act as product experts, translating product needs into reinforcement learning environments.
Building a New Product with Claude: The "Styles" Feature
Scott walks through the process of building a new product, using the "Styles" feature as an example.
- Product Spec: Creating a product specification based on customer feedback.
- Evals Development: Building a set of evaluations (evals) to measure the product's performance against specific use cases.
- Use Case Examples: Tailoring technical explanations, output format templates, and learning goals.
- Eval Framework: Translating use cases into an eval framework that Claude can understand and grade.
- Continuous Integration: Using evals as part of a continuous integration process to ensure features are not broken by new model updates.
- User Preferences Example: The "Styles" product includes "user preferences" (custom instructions) to guide Claude's output.
- Regression Testing: Establishing evals as part of a regression set to maintain quality over time.
Evals: Practical Examples
Scott provides practical examples of how to create evals for the "Styles" product.
- Explanatory Style Eval: Creating prompts and outputs that demonstrate the desired explanatory style.
- Ground Truth vs. Synthetic Evals: Using ground truth to verify accuracy or using Claude to evaluate whether the output fulfills the intended purpose.
- Concise Style Eval: Asking Claude to grade whether the output is concise while maintaining clarity.
- AI Evaluating Itself: Using AI to evaluate its own performance, with human oversight.
Mike Krieger's Influence
Mike Krieger's influence on Anthropic's product development is discussed.
- Fewer Things Better: Krieger has pushed the team to focus on fewer, more impactful initiatives.
- Pragmatism: Emphasizing the practicality of the model and product.
- Indispensable Product: Making the product indispensable for internal use, aligning with customer needs.
- Bottoms-Up Energy: Fostering a culture of bottoms-up ideation and iteration.
Advice for Aspiring AI Product Managers
Scott offers advice for individuals interested in joining the product team at Anthropic.
- Mission-Driven: Prioritizing the company's mission and commitment to AI safety.
- Relentlessness: Demonstrating energy, urgency, and a relentless drive.
- One-Team Mentality: Embracing collaboration and interconnectedness across teams.
- Customer Centricity: Focusing on customer needs and working closely with sales teams (for Enterprise roles).
- Data Familiarity: Having a strong understanding of data and iteration cycles (for growth roles).
- Eval Upskilling: Developing skills in creating and using evals for model feature development.
Claude's Future: From Assistant to Collaborator
Scott previews Anthropic's vision for Claude's future, as outlined in a chart showing the evolution from "Claude Assist 2024" to "Claude Collaborates 2025" and "Claude Pioneers 2027."
- Capable Assistant: Claude currently feels like a capable assistant that requires specific guidance.
- Capable Collaborator: The goal is to transform Claude into a capable collaborator that can take on meaningful tasks and delegate work.
- Key Dimensions: Improving Claude's knowledge, communication abilities, and ability to take action.
- Meaningful Strides: Making significant progress in these dimensions to move beyond being just an assistant.
- Solving Biggest Problems: Helping users solve their most pressing problems and meaningfully taking things off their plate.
Conclusion
The interview provides a detailed look into Anthropic's approach to building AI products, emphasizing the importance of a hybrid reasoning model, dogfooding, customer feedback, collaboration with research, and a focus on practical use cases. Scott's insights offer valuable guidance for product managers looking to navigate the evolving landscape of AI product development. The key takeaway is the shift from AI as a simple assistant to a capable collaborator, capable of understanding and addressing complex problems with minimal guidance.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development