AI in Healthcare Series: Accelerating the AI Revolution in Medicine, with Peter Lee, Microsoft

Stanford OnlineAbout 6 min readJul 30, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Large Language Models (LLMs): AI models like Grok, GPT-4, Gemini, and others, used for various tasks including reasoning, problem-solving, and collaboration.
  • Pre-training vs. Post-training: Pre-training refers to the initial training of a model on a massive dataset, while post-training involves fine-tuning or adapting the model for specific tasks or applications.
  • Inference Time Compute: The computational resources used when the model is actively generating outputs or making predictions.
  • Benchmarks: Standardized tests used to evaluate the performance of AI models on specific tasks.
  • Evaluation Frameworks: Methodologies and metrics used to assess the effectiveness and usefulness of AI models in real-world scenarios.
  • AI Agents: Autonomous entities that can interact with their environment and perform tasks, often in collaboration with humans or other agents.
  • Sequential Diagnosis: An approach to medical diagnosis where an AI model starts with minimal information and iteratively asks questions, orders tests, and makes referrals to arrive at a diagnosis.
  • Ambient Listening AI: AI systems that passively listen to conversations and extract relevant information, such as clinical notes.
  • Quality Measures: Metrics used to assess the quality of healthcare services and outcomes, often tied to revenue for healthcare organizations.
  • Fire Mandates: Fast Healthcare Interoperability Resources, data standard mandates.

Grok Launch and Model Performance

  • The podcast discusses the recent launch of Grok, a new LLM, and its performance benchmarks.
  • Grok's benchmarks show it outperforming models like GPT-3.5 Pro and Gemini 2.5 Pro on multimodal tasks and specific exams (ARC, AGI, etc.).
  • The CEO of Grok has stepped down due to differences on alignment and training.
  • The discussion acknowledges the trend of new models with more compute and better performance continually emerging.

Post-Training and Inference Time Compute

  • The focus in AI research is shifting towards post-training and inference time compute due to the increasing difficulty of achieving breakthroughs in pre-training.
  • Pre-training is becoming a high-stakes endeavor, similar to committing to a new silicon processor architecture.
  • A quote from OpenAI suggests that a few seconds of "thinking" time for a model can be equivalent to scaling up its size significantly.
  • The quality of the pre-trained base model is crucial for achieving good results in reasoning paradigms.

Evaluation and Benchmarking Challenges

  • There's a concern that the AI community is "chasing benchmarks," leading to models optimized for specific tests rather than real-world usefulness.
  • The "Goodhart's Law" is mentioned: "When a measure becomes a target, it ceases to be a good measure."
  • Microsoft Research is experimenting with an evaluation approach called Adele, which uses ideas from psychometrics to evaluate AI models.
  • The ultimate goal is to create AI collaborators that can work effectively with humans, requiring evaluation methods that assess qualities like listening skills, asking the right questions, and knowing when to seek help.

AI Agents and Healthcare Applications

  • The discussion highlights the potential of AI agents to assist with complex tasks in healthcare, such as participating in tumor board meetings.
  • The Healthcare Agent Orchestrator developed by Matt Lungren is mentioned as an example of an agent that can facilitate the use of other AI models and assist in meetings.
  • Current limitations of AI agents include the inability to proactively participate in conversations.
  • The importance of building AI systems that can complete long tasks without getting off track is emphasized.

Narrow Models vs. Super Agents

  • The conversation explores the debate between using one super-smart model to do everything versus using multiple specialized agents.
  • The current narrative seems to favor multiple agents that are more specialized and can work together to complete tasks.
  • The MAI (Medical AI) team at Microsoft is working on sequential diagnosis, where an AI model starts with minimal information and iteratively asks questions, orders tests, and makes referrals to arrive at a diagnosis.
  • The evaluation framework for sequential diagnosis includes penalties for the costs of tests and referrals, encouraging the model to be economically reasonable.
  • The sequential diagnosis model has shown promising results, outperforming human doctors in some scenarios, but further research is needed to address limitations.

Ambient Listening AI and Revenue Cycle Management

  • The discussion touches on the potential of ambient listening AI to improve revenue cycle management in healthcare.
  • The example of a dermatologist upcoding a clinical note to justify a treatment is used to illustrate the potential for both optimistic and pessimistic outcomes with automated coding.
  • The possibility of using AI to audit and monitor coding practices at a system level is explored.

Quality Measures and the Burden on Providers

  • Connecting ambient listening AI to quality measures is seen as a way to improve revenues for healthcare organizations and justify the cost of these tools.
  • The podcast emphasizes the importance of reducing the burden on healthcare providers, who are often overloaded with cognitive, effort, and cost burdens due to technology and policy mandates.
  • The ambient documentation space is "on fire" because it saves clinicians time, which is a major benefit.

Chatty EHR and Patient Engagement

  • The Chatty EHR project at Stanford is mentioned as an example of connecting AI directly to electronic health records to improve patient engagement.
  • Patients are increasingly using AI models on their own and expect their doctors to be literate on the topic.
  • The potential for patients to have a normal conversation with their chart through a patient portal is discussed, making complex medical information more accessible.

The AI Revolution in Medicine Podcast Series

  • Peter Lee and his colleagues have created a podcast series to discuss the real-world impact of generative AI in medicine.
  • The series features interviews with experts from various backgrounds, including clinical, business, and technology.
  • The podcast explores open questions and emerging themes in the field, such as the potential for AI to change the structure of medical specialties.

The Future of Medical Specialties

  • The discussion explores the potential for AI to either increase or decrease the number of medical specialties.
  • Historically, technology has increased the number of specialties, but AI could potentially enable general practitioners to handle a wider range of medical knowledge.
  • The possibility of AI models understanding agriculture in different regions and triangulating across those practices to become a superhuman aronomist is mentioned as an analogy.
  • The podcast participants express uncertainty about which way the field will go, but emphasize the importance of thinking about these issues now.

Shifting Value in Healthcare

  • The discussion touches on the potential for AI to shift value in healthcare from diagnosis to prevention and intervention.
  • A chart is presented showing the shifting of value from diagnosis to intervention, but the podcast participants suggest that value may shift even further to the right, towards changes in coding practices and finding high-value patients and procedures.
  • The importance of understanding the bigger picture of healthcare, including prevention and intervention, is emphasized.

Conclusion

The podcast explores the rapid advancements in AI and their potential impact on healthcare. Key takeaways include the shift towards post-training and inference time compute, the challenges of evaluating AI models, the potential of AI agents to assist with complex tasks, the debate between narrow models and super agents, the importance of reducing the burden on healthcare providers, and the potential for AI to change the structure of medical specialties. The participants emphasize the need for ongoing research, collaboration, and thoughtful consideration of the ethical and economic implications of AI in healthcare.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.