Evaluating Domain Specific LLMs for Real World Finance — Waseem Alshikh, Writer

AI EngineerAbout 5 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Domain-Specific Models vs. General Models: The core question is whether to focus on building specialized models for specific domains (e.g., finance, medicine) or rely on general-purpose models.
  • Accuracy vs. Grounding: The distinction between a model's ability to provide a correct answer (accuracy) and its ability to ground that answer in the provided context (grounding).
  • Hallucination: When a model generates information that is not supported by the provided context or real-world knowledge.
  • Robustness: The ability of a model to perform well even when presented with noisy or incomplete input data.
  • FAIL Benchmark: A custom evaluation framework designed to assess model performance in real-world scenarios with query and context failures.
  • RAG (Retrieval-Augmented Generation): A framework that combines information retrieval with text generation to improve the accuracy and relevance of generated text.

Evaluation of Domain-Specific vs. General Models

Background and Motivation

  • In early 2024, the company observed that general-purpose language models (LLMs) were achieving high accuracy (80-90%) on general benchmarks.
  • This raised the question of whether it was still necessary to invest in building domain-specific models, or if fine-tuning general models would suffice.
  • To answer this question, the company created a benchmark called FAIL to evaluate model performance in real-world scenarios.

The FAIL Benchmark

  • Goal: To evaluate how well models perform when faced with common issues like misspelled queries, incomplete information, and irrelevant context.
  • Categories of Failures:
    • Query Failure: Issues related to the input query itself.
      • Misspelling Queries: Queries with spelling errors.
      • Incomplete Queries: Queries missing key information.
      • Out of Domain Queries: Queries that are outside the scope of the model's expertise.
    • Context Failure: Issues related to the context provided to the model.
      • Messing Context: Asking questions about context that doesn't exist in the provided information.
      • OCR Error: Errors introduced during the optical character recognition (OCR) process when converting physical documents to text.
      • Irrelevant Context: Providing the model with a completely unrelated document.
  • Data Diversity: The benchmark includes a diverse set of financial service-specific data.
  • Open Source: The white paper, data, evaluation set, and leaderboard are all available on GitHub and Hugging Face.
  • Evaluation Metrics:
    • Correct Answer: Whether the model provides the correct answer.
    • Context Grounding: Whether the model's answer is grounded in the provided context.

Model Selection and Evaluation Process

  • A group of models was selected for evaluation, including both chat models and "thinking" models.
  • The models were evaluated using the FAIL benchmark, and the results were analyzed based on the two key metrics: correct answer and context grounding.

Results and Analysis

Key Findings

  • Thinking Models and Refusal to Answer: Thinking models generally don't refuse to answer questions, even when given wrong context or data.
  • Hallucination: When given wrong context, these models often fail to follow the context and provide answers that are not grounded in the provided information, leading to higher hallucination rates.
  • Grounding Issues: Models struggle with grounding, especially in tasks like text generation and question answering.
  • Smaller Models vs. Larger Models: Smaller models sometimes perform better than larger, more complex models in terms of grounding.
  • Chain of Thought: The data suggests that "thinking" models may not be truly thinking in domain-specific tasks, as their hallucination rates are high.
  • Gap Between Robustness and Grounding: There is a significant gap between a model's robustness (ability to handle noisy input) and its ability to ground its answers in the correct context.
  • Best Model Performance: Even the best-performing models still have a significant error rate (around 20%) in terms of providing completely wrong answers.

Specific Examples

  • The presenter highlights the performance of models like O1, O3, and B-fan, noting that they perform well when faced with misspelled queries, incomplete information, or out-of-domain queries.
  • However, when it comes to grounding, these models perform significantly worse, with grounding accuracy dropping by 50-60% in some cases.

Data and Statistics

  • The presenter mentions that the white paper includes details on the amount of data and tokens used in the benchmark.
  • The data shows that even with the best models, around 20% of requests result in completely wrong answers.

Conclusion and Implications

Need for Full-Stack Solutions

  • The presenter concludes that, based on the current state of technology and models, a full-stack approach is necessary for reliable utilization of LLMs.
  • This includes RAG systems, grounding techniques, and guardrails to ensure that models provide accurate and contextually relevant answers.

Continued Need for Domain-Specific Models

  • The presenter answers the initial question by stating that there is still a need to build and continue developing domain-specific models.
  • While general models are improving in accuracy, their ability to ground their answers in the correct context is still significantly behind.
  • "Even accuracy is keep growing but the grounding the context for following all the context correctly it's still way way way behind from everything we see today in the market"

Synthesis

The video presents a compelling argument for the continued importance of domain-specific models, even as general-purpose LLMs improve. The FAIL benchmark highlights the critical distinction between accuracy and grounding, demonstrating that models can provide correct answers without necessarily understanding or utilizing the provided context. This can lead to high hallucination rates and unreliable results, particularly in sensitive domains like finance. The presenter advocates for a full-stack approach that combines RAG, grounding techniques, and guardrails to mitigate these issues and ensure the responsible use of LLMs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.