Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP - Dan Mason

AI EngineerAbout 9 min readJun 23, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agent Workflows: Using LLMs to control application flow.
  • Langraph: A framework for building and visualizing agent workflows.
  • MCP (Model Control Plane): A system for managing and controlling LLMs.
  • Blueprints: Structured documents defining treatment protocols and approved language.
  • Anchors: Key events or milestones in a treatment timeline.
  • State Management: Maintaining and updating patient-specific data throughout the interaction.
  • LLM as a Judge: Using an LLM to evaluate the performance and complexity of another LLM's decisions.
  • Confidence Scoring: Assessing the LLM's certainty and the complexity of a situation to determine if human review is needed.
  • Evals: Evaluation processes to measure the performance and accuracy of the system.

1. Introduction and Stride's AI Focus

  • The speaker introduces themselves and Stride, a custom software consultancy specializing in AI.
  • Stride uses AI for code generation (unit test creation, maintenance), modernization of legacy codebases (e.g., early 2000s .NET), and agent workflows.
  • The presentation focuses on agent workflows, specifically a healthcare use case.
  • The speaker emphasizes that this is "the way we did it," and welcomes feedback and alternative approaches.

2. Case Study: Aila Science and Early Pregnancy Loss Treatment

  • Client: Aila Science, a women's health institution focused on early pregnancy loss (miscarriage) treatment.
  • Problem: Patients receive medication at the hospital/clinic and self-administer it at home. This is a traumatic time, making it difficult to track medication schedules and answer questions.
  • Existing System: A text message-based system with humans manually pushing buttons to send pre-approved messages.
  • Limitations of Existing System:
    • Difficult to scale the human team to serve more patients.
    • Inflexible for supporting new treatments.
  • Solution: Rebuild the system with an LLM at the core to make it more flexible and capable.
  • Key Features of the New System:
    • Flexible blueprint and knowledge base for defining new treatments.
    • Medically approved language.
    • Self-evaluation function to catch complicated situations and surface them for human review.
  • Disclaimer: The presentation shows the client's actual code, with some redactions.
  • Custom Software: Stride built custom software due to specific client requirements, but used off-the-shelf tools (Langraph, Langchain, Langsmith) where possible.
  • Hosting Constraints: The system intersects with patient data, requiring adherence to HIPAA and other privacy requirements.
  • Hybrid System: Humans are very much in the loop, preserving human judgment.
  • Capacity Increase: Estimated 10x increase in the number of people that can be served.
  • New Treatments: New treatments and workflows can be supported without writing more code.

3. System Architecture and Components

  • Virtual Operations Associate: Assesses the state of a conversation with a patient and determines the best response (text message, questions, actions).
  • Evaluator Agent: An LLM as a judge that evaluates the virtual operations associate's decisions, assessing correctness and complexity.
  • Tools: A mix of MCP (Model Control Plane) and state management tools.
    • MCP: Looks at local files or goes across the wire to a larger software system.
    • State Management: Maintains the state of the conversation in real-time.
  • Stack:
    • LLM Piece: Python and a Langraph container.
    • Text Message Gateway, Database, Dashboard: Node, React, MongoDB, Twilio, hosted in AWS.
  • Eval System:
    • Custom harness that pulls data out of Langsmith, processes it, and runs it through PromptFu.
    • PromptFu's LLM rubric is used to evaluate the system.
  • Team: Two software engineers, one designer, and the speaker.
    • The software engineers maintain the core system (gateway, dashboard, text message stuff, database).
    • The speaker maintains and built everything on the Langraph side, with AI assistance.
  • Code Generation: The speaker used Klein to write the Python code and mostly hand-coded the prompts.

4. Detailed Walkthrough of the System in Action

  • Dashboard: The operations associates (humans) use a dashboard to monitor conversations and approve/provide feedback on the LLM's decisions.
  • "Needs Attention" Feature: The system flags conversations that require human review.
  • Rationale: The LLM provides a rationale for its decisions, explaining why it said a given thing at a given time.
  • Time Zone Detection: The system automatically detects the patient's time zone.
  • Anchor Points: The system sets anchor points to track key events in the treatment timeline.
  • Flexibility: The system is flexible and can handle patients who skip ahead or provide information out of order.
  • State Management: The system maintains the state of the conversation, including messages, anchors, and treatment phase.
  • Langsmith Integration: The entire conversation is logged in Langsmith, allowing for detailed debugging and analysis.
  • Confidence Scoring: The system assigns a confidence score to each decision, based on factors such as correctness, complexity, and potential risks.
  • Human Review: If the confidence score is below a certain threshold, the conversation is flagged for human review.
  • Feedback Mechanism: Humans can provide feedback on the LLM's decisions, which is used to improve the system.
  • Blueprint Structure: The blueprints are structured documents that define the treatment protocol and approved language.
  • Triage/Knowledge Base: If the blueprint doesn't address a patient's question, the system consults a triage/knowledge base.
  • Model Selection: Claude was chosen for its steerability, transparency, and flexible hosting.
  • Prompt Injection: The system is designed to prevent prompt injection attacks by obscuring patient data and limiting the LLM's access to external resources.
  • Data Sensitivity: The system is designed to handle sensitive patient data in a secure and compliant manner.
  • Model Learning: The system does not use patient data to train the LLM, but the team does use human feedback to improve the prompts and guidelines.
  • Caching: The system uses caching to reduce the cost and latency of LLM calls.
  • Tool Calling: The LLM uses tool calling to interact with external systems and perform specific tasks.
  • Retries: The system has a retry mechanism to handle errors and malformed tool calls.
  • Evals: The system uses a custom eval harness to measure the performance and accuracy of the system.

5. Key Arguments and Perspectives

  • Transparency and Control: The speaker emphasizes the importance of transparency and control in agent workflows.
  • Human-in-the-Loop: The speaker argues that humans should be in the loop to provide feedback and handle complex situations.
  • Flexibility: The speaker argues that the system should be flexible enough to handle new treatments and workflows without requiring code changes.
  • Scalability: The speaker argues that the system should be scalable to handle a large number of patients.
  • Cost-Effectiveness: The speaker argues that the system should be cost-effective, taking into account the cost of LLM calls and human review.
  • Model Selection: The speaker argues that the choice of LLM should be based on factors such as steerability, transparency, and flexible hosting.
  • Fine-Tuning: The speaker argues that fine-tuning is not always necessary, as LLMs are constantly improving.
  • Prompt Engineering: The speaker emphasizes the importance of prompt engineering in building effective agent workflows.
  • Testing and Evaluation: The speaker emphasizes the importance of testing and evaluation in ensuring the quality and accuracy of the system.

6. Notable Quotes

  • "That's dumb. You should do that better." - The speaker encourages feedback and alternative approaches.
  • "Hey this is how this thing works You can see it goes from here to here There's loops here Like this is where we're doing our um you know our our evaluation of the process and here's where humans come in." - Describing the ease of explaining Langraph to clients.
  • "I'm sorry I can't answer that question call 911 go to your doctor whatever it is" - Example of how the system handles situations outside its scope.

7. Technical Terms and Concepts

  • LLM (Large Language Model): A type of AI model that can generate human-like text.
  • Agent: An LLM that controls the flow of an application.
  • Langraph: A framework for building and visualizing agent workflows.
  • MCP (Model Control Plane): A system for managing and controlling LLMs.
  • Blueprints: Structured documents defining treatment protocols and approved language.
  • Anchors: Key events or milestones in a treatment timeline.
  • State Management: Maintaining and updating patient-specific data throughout the interaction.
  • LLM as a Judge: Using an LLM to evaluate the performance and complexity of another LLM's decisions.
  • Confidence Scoring: Assessing the LLM's certainty and the complexity of a situation to determine if human review is needed.
  • Evals: Evaluation processes to measure the performance and accuracy of the system.
  • Prompt Engineering: The process of designing and optimizing prompts for LLMs.
  • Tool Calling: The ability of an LLM to interact with external systems and perform specific tasks.
  • Prompt Injection: A type of attack where an attacker tries to manipulate an LLM by injecting malicious prompts.
  • HIPAA (Health Insurance Portability and Accountability Act): A US law that protects the privacy of patient data.
  • VPC (Virtual Private Cloud): A private network within a public cloud.
  • Terraform: An infrastructure-as-code tool.
  • GitHub Actions: A continuous integration and continuous delivery (CI/CD) platform.
  • TDD (Test-Driven Development): A software development process where tests are written before the code.

8. Logical Connections

  • The introduction sets the stage by explaining Stride's expertise in AI and the focus on agent workflows.
  • The case study provides a real-world example of how agent workflows can be used to solve a specific problem in healthcare.
  • The system architecture section explains the different components of the system and how they work together.
  • The detailed walkthrough shows the system in action, demonstrating its flexibility and capabilities.
  • The key arguments and perspectives section summarizes the speaker's main points and provides supporting evidence.
  • The technical terms and concepts section defines the key terms used in the presentation.

9. Data, Research Findings, and Statistics

  • Estimated 10x increase in the number of people that can be served with the new system.
  • The most complicated conversation seen was something like 150 texts.
  • Average cost to generate a single message is somewhere in the 15 to 20 cent range.

10. Synthesis/Conclusion

The presentation provides a detailed overview of how Stride built an agent workflow for Aila Science to improve the treatment of early pregnancy loss. The system uses an LLM to automate text message-based interactions with patients, providing personalized support and guidance. The system is designed to be flexible, scalable, and cost-effective, while also ensuring patient safety and privacy. The speaker emphasizes the importance of transparency, control, and human-in-the-loop in building effective agent workflows. The presentation also highlights the challenges and trade-offs involved in building such a system, such as model selection, prompt engineering, and testing and evaluation.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.