How Intuit uses LLMs to explain taxes to millions of taxpayers - Jaspreet Singh, Intuit

AI EngineerAbout 5 min readJul 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Generative AI in tax preparation
  • Large Language Models (LLMs) for tax assistance
  • Regulatory compliance and data security in AI
  • Prompt engineering and model fine-tuning
  • Evaluation frameworks for LLM performance
  • Retrieval-Augmented Generation (RAG) and Graph RAG
  • LLM as a Judge
  • Latency and scalability challenges with LLMs
  • Human-in-the-loop for AI development

Introduction

Jaspit, a senior staff engineer at Intuit, discusses how Intuit uses LLMs to enhance TurboTax and help users understand their taxes better. TurboTax successfully processed 44 million tax returns for tax year 23. The goal is to provide users with high confidence in their tax filings and ensure they receive the best possible deductions.

Intuit Geni and GenOS

Intuit Geni experiences are built on GenOS, a proprietary generative OS platform. GenOS addresses the limitations of out-of-the-box tooling, particularly in regulatory compliance, safety, and security, which are critical in the tax domain. GenOS is designed for large-scale deployment within Intuit. It includes GenUI (UI side), an orchestrator for managing different LLM solutions, and Intuit Assist, the overall experience powered by GenOS.

LLM Implementation in TurboTax

Scalable Solutions

Intuit aims to build scalable solutions to support millions of customers.

Prompt-Based Solutions

The initial approach involved prompt-based solutions to understand users' tax situations. For example, explaining the constituents of a tax refund (deductions, credits, etc.).

Model Selection

Claude is the primary production model, backed by a multi-million dollar contract. OpenAI models are used for question answering.

Query Types

  • Static Queries: Prepared statements for known information, such as a tax refund summary.
  • Dynamic Queries: Answering user questions about their tax situation (e.g., "Can I deduct my dog?").

Model Iteration

The models are constantly changing. The models are being iterated on with newer versions.

RAG and Graph RAG

RAG and Graph RAG solutions are used to leverage Intuit's proprietary tax information and tax engines, improving the accuracy of answers.

Fine-Tuning

A pilot project involved fine-tuning Claude for static queries. While the quality was good, it was deemed too specialized.

Evaluation

Evaluation is a critical component, ensuring quality throughout the development lifecycle and in production.

Key Pillars of LLM Implementation

  • Scalability: Building solutions that can handle millions of users.
  • Tax Accuracy: Ensuring the accuracy of tax-related information.
  • Relevancy: Providing relevant answers to user questions.
  • Coherence: Ensuring the clarity and understandability of explanations.
  • Human Domain Experts: Tax analysts provide expertise, decode IRS changes, and assist with evaluations.
  • Phased Evaluation System: Manual evaluations are conducted initially, followed by automated evaluations.
  • Prompt Engineering: Tax analysts serve as prompt engineers, allowing data scientists and ML engineers to focus on quality and metrics.
  • LLM as a Judge: Using LLMs to evaluate the performance of other LLMs.

Fine-Tuning Details

  • Fine-tuning was performed on Claude 3 Haiku, powered by AWS Bedrock.
  • The goal was to improve response quality, reduce the need for instructions, and decrease latency.
  • Only consented user data was used, adhering to regulations (7 to 616 regulations).

Evaluation Metrics and Tools

  • Accuracy: Ensuring tax accuracy.
  • Relevancy: Ensuring the information is relevant to the user's query.
  • Coherence: Ensuring the response is clear and understandable.
  • Manual Evaluations: Conducted by tax experts.
  • Automated Evaluations: Using LLM as a Judge and in-house tooling for automated prompt changing.
  • AWS Ground Truth: Used for creating golden datasets for evaluation.
  • Broad Monitoring: Monitoring LLM performance with real users in real time.

Model Changes

The move from Anthropic Claude Instant to Claude Haiku for tax year 24 required significant effort and was validated through clear evaluation processes.

Major Learnings

  • Contract Costs: LLM contracts are expensive, and long-term contracts offer better pricing but create vendor lock-in.
  • Vendor Lock-in: Prompts are a form of lock-in, and upgrading models even within the same vendor can be challenging.
  • Latency: LLM latency (3-10 seconds) is significantly higher than backend services, especially with complex tax situations. Latency increases during peak filing times (e.g., tax day).
  • Product Design: Product design must account for latency, including fallback mechanisms and user experience considerations.
  • Evaluations: Evaluations are crucial for launching and maintaining LLM-powered features. Clear guidelines and golden datasets are essential.

Q&A Highlights

  • Evaluation Types: Manual evaluations with tax experts for initial development, auto-evaluations for minor prompt tweaks, and re-evaluation for major changes (e.g., new tax year).
  • LLM Interactions: Question answering covers both product-related questions (how to use TurboTax) and tax-situation questions (deduction eligibility).
  • Number Verification: TurboTax uses a proprietary tax knowledge engine for calculations. LLMs do not perform calculations. ML models and safety guardrails prevent hallucinated numbers in explanations.
  • RAG vs. Graph RAG: Graph RAG provides better response quality and personalized answers compared to regular RAG.
  • Answer Generation: Answers for static questions are generated using the tax knowledge engine and prompts crafted by tax experts. Evaluations ensure accuracy.

Conclusion

Intuit leverages LLMs to enhance TurboTax, focusing on accuracy, scalability, and regulatory compliance. The implementation involves a combination of prompt engineering, model fine-tuning, and rigorous evaluation processes, with human experts playing a crucial role. Key challenges include managing contract costs, addressing latency issues, and ensuring data security. The insights shared provide a comprehensive overview of how Intuit is using generative AI to improve the tax preparation experience for millions of users.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.