AI In Healthcare Series: Leveraging GPT-5, Cosmos, and Predictive Models for Better Outcomes

Stanford OnlineAbout 6 min readSep 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • GPT-5: The latest iteration of OpenAI's large language model.
  • USMLE: United States Medical Licensing Examination, a benchmark for AI in healthcare.
  • Thinking/Pro Models: Higher quality, reasoning-focused versions of GPT models with increased latency.
  • Model Selection: Choosing the appropriate AI model for a specific task.
  • Specialized Models: AI models trained for specific domains or tasks (e.g., image models, healthcare-specific models).
  • Benchmarks: Standardized tests used to evaluate the performance of AI models.
  • Non-deterministic: LLMs do not always produce the same output for the same input.
  • Reinforcement Learning: A type of machine learning where an agent learns to make decisions by receiving rewards or penalties.
  • User Experience (UX): The overall experience of a person using a product such as a website or computer application, especially in terms of how easy or pleasing it is to use.
  • Cosmos: A collaborative effort across health systems using Epic software to build a large, de-identified dataset for AI research.
  • De-skilling: The loss of skills due to reliance on AI or automation.
  • Human-AI Interaction: The study and design of interactions between humans and AI systems.
  • MyChart: Epic's patient-facing application for accessing medical records, scheduling appointments, etc.
  • Information Asymmetry: Unequal access to information between physicians and patients.
  • Visit Agenda: A pre-visit plan created collaboratively by patients and providers to focus the appointment.
  • Real-World Evidence: Data collected from real-world healthcare settings used to inform clinical decisions.

GPT-5 and the State of AI Models

  • Initial Impressions: Seth Hayne uses his son's "Magic deck" test as a personal benchmark, finding that GPT-5 still hasn't mastered it. He primarily uses the "thinking" and "pro" models for higher quality output, accepting increased latency.
  • Rate of Change: Justin Norden notes that the improvement from GPT-4 to GPT-5 doesn't feel like a "Death Star giant leap forward." He highlights the emergence of specialized models like Gemini's Nano Banana.
  • Model Choice: Matt Lungren preferred having a choice of models, including the intermediate "4.5" model, which he found better for writing. He believes casual users may be more impressed with GPT-5's "intelligent responses."
  • Benchmark Limitations: Lungren feels benchmarks are becoming less useful for evaluating real-world applications. Hayne points out that current benchmarks are based on the education system and don't reflect the later stages of a clinician's career.
  • Multiple Choice Issues: Norden references a study from Nigam Shah's lab showing that even minor changes to multiple-choice questions can significantly impact LLM performance. He also notes the non-deterministic nature of LLMs.

Bridging the Gap: From Models to Real-World Healthcare Applications

  • Software Development vs. Research: Hayne emphasizes that academic research often doesn't replicate the checks and balances used in building real-world software. He calls for more dialogue between industry and academia.
  • Asking Meaningful Questions: Hayne stresses the importance of understanding the problems physicians face through "immersion" (spending time with them). He notes that questions can be diagnostic or administrative.
  • Building the Right Pipeline: Hayne discusses the need for robust back-end structures for monitoring and improving AI-generated outputs (summaries, notes, SQL queries).
  • New UX Patterns: Hayne highlights the need to explore new UX patterns for situations where longer model thinking time improves outcomes.
  • Minimizing LLM Noise: Norden argues that reliable, consistent software should minimize LLM usage to reduce noise. He emphasizes the importance of context, data, audit controls, monitoring, governance, and evaluation.
  • Model vs. Application: Hayne clarifies that the model is not the same as the end-user application. He emphasizes the need for guardrails and checks to ensure informed decisions.
  • Multi-Agent/Tool Story: Lungren envisions a future where AI agents can fetch the right tools and information, saving users time and mental effort.
  • Right Tool for the Right Job: Hayne argues that AI should not always be used, even if it's a generalizable framework.

Cosmos and Healthcare-Specific Models

  • Reevaluating the Task: Lungren discusses the need to reevaluate tasks and address gaps in natural language processing for healthcare, considering the language of health, longitudinal records, and prediction.
  • Epic's History: Hayne explains that the name "Epic" comes from classic Greek poems, and the company has always focused on the patient's story.
  • Chronological Patient Stories: Epic's approach involves using medical events, interventions, observations, and time intervals in chronological order to build patient stories.
  • Transformer Architecture: The same transformer architecture used in natural language processing is applied to predict future events in a patient's health journey.
  • Data Set Size: The largest model is trained on approximately 8 billion encounters, with potential for growth using the Cosmos community dataset (2-3 times larger).
  • Prediction Horizon: The model can predict both near-term and long-term events (3-10 years), which can help manage chronic diseases and address capacity problems.
  • UX Exploration: The team is exploring how to visualize simulated future trajectories from a UX perspective.
  • Collaboration: The research is a collaboration with Yale and Microsoft.

De-skilling and the Future of Human-AI Interaction

  • De-skilling Concerns: Norden raises concerns about the de-skilling of users due to reliance on AI, referencing a paper showing that physicians perform worse on colonoscopies after using AI for polyp detection and then having it removed.
  • Acceptable De-skilling: Lungren questions what areas in healthcare are acceptable for de-skilling, where technology can be reliably used.
  • Human-AI Combo: Lungren emphasizes the need for a human-AI interaction model where the combination is better than either alone.
  • Quality of Care in 20 Years: Hayne believes that the quality of care will be unequivocally better in 20 years due to real-world evidence, personalized medicine, and new tools.
  • Interim State: Hayne highlights the need to navigate the interim state, learning how to best use AI tools. He uses the example of "vibe coding" tools that allow developers to create functioning prototypes more quickly.
  • Time Scale of Change: Norden questions the time scale of these changes, considering the increasing use of consumer tools by patients.
  • Integrated Software: Hayne discusses Epic's work on integrating MyChart (patient-facing) and physician-facing experiences to answer consumer questions grounded in medical records and facilitate handoffs to care teams.
  • Last Hand on the Doorknob Moment: Lungren describes the common scenario where patients ask their most important question as the appointment ends. He sees the potential for AI to address this by allowing patients to tee up questions in advance.
  • Information Asymmetry: Lungren highlights the information asymmetry between physicians and patients and the need to level the playing field.
  • Visit Agenda: Epic is building towards a real-time experience driven by the conversation in the exam room, using a pre-visit agenda to guide the visit and bring in real-world evidence.

Epic UGM and Societal Challenges

  • Societal Questions: Hayne emphasizes the importance of addressing societal challenges facing health systems, such as physician shortages and an aging demographic.
  • Scaling Meaningfully: He stresses the need to scale solutions that help with access challenges and answer billing questions.
  • Epic UGM: Lungren describes the Epic UGM conference as a major event in healthcare, bringing together a community to discuss the future of the field.

Conclusion

The discussion highlights the current state of AI models, particularly GPT-5, and the challenges of translating their capabilities into real-world healthcare applications. Key themes include the need for careful model selection, robust software development practices, and a focus on user experience. The conversation also explores the potential of healthcare-specific models like Cosmos and the importance of addressing societal challenges facing health systems. The future vision includes a more integrated and collaborative patient-physician experience, driven by AI and grounded in real-world evidence. The panelists acknowledge the potential for de-skilling but emphasize the importance of human-AI interaction models that enhance overall performance and improve the quality of care.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.