Generative AI for Healthcare (Part 1): Demystifying Large Language Models

By Unknown Author

Share:

Generative AI for Healthcare: Part 1 - What is an LLM Really?

Key Concepts:

  • Generative AI
  • Large Language Models (LLMs)
  • Epochs of AI in Healthcare (Rules-based, Machine Learning, Generative AI)
  • Tokenization
  • Static Embeddings
  • Vectors and Vector Space
  • Cosine Similarity
  • Context-Aware Embeddings
  • Transformer Architecture
  • Self-Attention
  • Temperature
  • Pre-training
  • Supervised Fine-tuning
  • Reinforcement Learning with Human Feedback (RLHF)
  • Reward Model
  • LLM as a Judge
  • Test-Time Scaling
  • Reasoning Models
  • Parameters
  • Compute (Petaflop per Second Days)
  • Training Data
  • Loss
  • Back Propagation
  • Gradient Descent

1. Introduction and Motivation

  • Generative AI is rapidly reshaping healthcare, but healthcare professionals lack accessible educational resources to understand and effectively use these technologies.
  • The video aims to provide a clear, comprehensive resource for healthcare professionals to understand generative AI fundamentals and tips.
  • The speakers, Dong and Shivam, are clinical informaticists and physicians at Stanford Medicine focused on deploying generative AI in clinical settings.
  • They are independent contractors for Greenlight, contributing to OpenAI safety initiatives, and Shivam was previously a consultant for Glass Health.
  • The content offers a high-level overview of complex technical concepts, streamlined for clarity and based on expert consensus.
  • The video primarily focuses on OpenAI's models like ChatGPT, but the core concepts apply to all LLMs.

2. The Challenges of Prompting

  • Prompting AI models effectively is challenging due to:
    • Difficult AI Literature: Original AI papers like "Attention is All You Need" are technically complex and hard for clinicians to understand.
    • Lack of Healthcare-Specific Resources: Existing prompting resources are either too basic or too technical, lacking practical detail for healthcare applications.
    • Rapid Pace of AI Progress: The exponential growth in AI research (e.g., PubMed articles on AI/ML increased from 272 in 2014 to over 20,000 in 2024) makes it difficult to stay updated.

3. Key Takeaways

  • The video aims to:
    • Contextualize generative AI within the broader framework of AI and healthcare.
    • Provide an intuitive understanding of how LLMs work.
    • Explain the pre-training and post-training phases that shape model behavior.

4. A Brief History: The Three Epochs of AI in Healthcare

  • Based on a framework from Michael Howell and Karen DeSalvo, the history of healthcare AI is divided into three epochs:
    • Epoch 1 (1970s): Symbolic AI and Probabilistic Models (Rules-Based AI):
      • Examples: Clippy, TurboTax, most video game AI, clinical decision support tools (e.g., drug interaction alerts), risk calculators (e.g., MDCalc), automated billing algorithms.
      • Characteristics: Logic-based, non-adaptive, doesn't learn from new data, hardcoded by experts for specific tasks.
    • Epoch 2 (2010s): Deep Learning (Traditional Machine Learning):
      • Examples: Facial recognition, targeted advertising, computer vision (e.g., self-driving cars), automated EKG/STEMI detection, AI deterioration models (e.g., sepsis risk scores), radiology abnormality detection.
      • Characteristics: Trained on massive labeled datasets, learns from data for specific tasks, often low in interpretability (black box models).
    • Epoch 3 (2017-Present): Large Language Models and Generative AI:
      • Examples: ChatGPT, AI summarization tools (e.g., Amazon customer review summaries), customer service chatbots (e.g., Bank of America's Erica), image generation (e.g., DALL-E, Midjourney, Sora), clinical knowledge retrieval (e.g., OpenEvidence, ClinicalKey AI), chart summarization, automated note drafting, ambient dictation.
      • Characteristics: General purpose, generative capabilities, pre-trained on enormous datasets, multimodal, often opaque in interpretability.

5. How LLMs Work: Anatomy and Physiology

  • Input Prompt and Tokenization:
    • The input prompt is broken down into smaller pieces called tokens (roughly equivalent to single words).
    • Example: "Please help me draft a short but professional letter to a patient explaining why they don't need an MRI for their seasonal allergies" is tokenized into 25 tokens.
  • Static Embeddings:
    • Each token is associated with a static embedding, a vector that captures the meaning of the word.
    • A vector is a list of numbers; an embedding is a high-dimensional vector representing a word's meaning.
    • Each number in the embedding represents the strength of an association between the token and a particular concept, sentiment, or idea (ranging from -1 to 1).
    • Example: In a simplified 3-dimensional vector space (legs, tail, speaks), the angular distance (cosine similarity) between "cat" and "dog" is smaller than between "cat" and "human," reflecting their semantic similarity.
    • Word2Vec example: Subtracting the vector for "man" from "woman" yields a vector representing "gender." Applying this vector to "king" results in a vector close to "queen."
    • Modern LLMs use thousands of dimensions (e.g., GPT-3 used around 12,000 dimensions per token embedding).
    • Static embeddings are learned during the pre-training phase.
  • Transformer Architecture and Self-Attention:
    • The transformer architecture, particularly the self-attention mechanism, is crucial for modern LLMs.
    • Self-attention allows the model to weigh the importance of other tokens in the sentence, dynamically adjusting the representation of each word based on context.
    • This transforms the static embedding into a context-aware embedding.
  • Generating Output:
    • The model takes all the context-aware embeddings and predicts the next most likely token in the sequence.
    • The "temperature" setting controls the randomness and creativity of the output.
      • Higher temperature leads to more creative but potentially unexpected outputs.
      • Lower temperature leads to more predictable outputs.
    • The generated token is appended to the sequence, and the process repeats until the output is complete.
    • The output isn't just the answer; it becomes part of the reasoning path.
    • Prompting the model to write out its reasoning step-by-step (chain of thought) can improve the quality of answers.
  • Visualizing LLM Navigation in Vector Space:
    • The LLM navigates through a vector space, with the goal of reaching the "solution space" where a good answer resides.
    • Each token generated is an opportunity to guide the output towards the ideal solution space.

6. Training LLMs: From GPT-1 to Reasoning Models

  • Attention is All You Need (2017): Introduced the transformer architecture and self-attention.
  • GPT-1 (2018):
    • 117 million parameters.
    • Trained on the "Books Corpus" (4.6 GB of text).
    • Compute: 1 petaflop per second day.
    • Example output for "Which antibiotics are first line for treating inpatient community acquired pneumonia assuming no risk factors?": "I don't know the woman said but I'll ask good now what else do you know about this plague."
  • GPT-2 (2019):
    • 1.5 billion parameters.
    • Trained on 40 GB of text scraped from web pages linked to in Reddit posts with >3 upvotes.
    • Compute: Estimated 600 petaflop per second days.
    • Example output for the same prompt: Repetitive questions, not answering the question directly.
  • Scaling Laws for Neural Language Models (2020):
    • Outlined three key levers in model training: compute, dataset size, and number of parameters.
    • Optimal performance requires scaling all three together.
  • GPT-3 (2020):
    • 175 billion parameters.
    • Trained on WebText2 (400 billion tokens, 570 GB).
    • Compute: Estimated 3600 petaflop per second days.
    • Example output for the same prompt (using an equivalent open-source model): "I am a nurse i am starting to develop a cough as far as I know it's an easy to treat cough i am not sure."
  • GPT-3.5 (ChatGPT) (2022):
    • Similar parameters, training data, and compute to GPT-3.
    • Breakthrough was in post-training techniques:
      • Supervised Fine-tuning: Training on human-curated input-output pairs to learn appropriate combinations. Instruction fine-tuning trains the model to follow instructions.
      • Reinforcement Learning with Human Feedback (RLHF): Generating multiple outputs and having humans rank them by overall quality. This helps improve factual accuracy and align behavior with human values.
      • Human reviewers are used to rank outputs, but this is being scaled using reward models (machine learning models that predict human preferences) and LLMs as judges (using LLMs with guiding principles to rank outputs).
    • Example output for the same prompt: "First line antibiotics for treating inpatient community acquired pneumonia in patients without risk factors include a combo of betalactone such as septrixone or septoxim plus a macroly like a zithro."
  • The Shift to Test-Time Scaling (2024-Present):
    • Jensen Wong (Nvidia CEO) introduced the idea of three distinct scaling laws: pre-training, post-training, and test-time scaling.
    • Ilia Sutskever (OpenAI co-founder) stated that pre-training as we know it will end due to running out of high-quality internet data.
    • Test-time scaling involves giving the model more compute at inference time to reason through a problem.
    • OpenAI's GPT-0 series (reasoning models) showed significant performance gains on challenging benchmarks like ARC AGI by increasing test-time compute.
    • Sam Altman (OpenAI CEO) signaled a shift towards scaling and optimizing test-time reasoning.

7. What is an LLM Really?

  • The vast majority of humanity's collective data output (estimated at 100 trillion GB) is distilled into a much smaller form: the parameters of a model (weights, biases, and embeddings).
  • For GPT-4, this compressed representation takes up around 3,500 GB (a compression factor of 30 billion).
  • The training process for GPT-4 cost OpenAI an estimated $100 million and 7,200 megawatt hours of electricity.
  • This model contains a distillation of a part of the human experience, capturing our collective understanding and interpretation of reality, including knowledge, reasoning, and thinking.
  • This model representation of our reality is in a portable format.

8. Conclusion

  • LLMs are powerful tools that can help us better understand the world and serve as building blocks for new technologies.
  • The next video will focus on evidence-based methods for effective prompting.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video