ChatGPT vs Claude vs Gemini: The BEST AI for Research (Hallucination Test Results)

By Andy Stapleton

Share:

Key Concepts

  • First-Order Hallucinations: AI models providing references that do not exist.
  • Second-Order Hallucinations: AI models citing a claim with a reference, but the claim is not actually supported by the content of that reference, or the reference is not cited for the proper reasons.
  • Plausibility Machines: AI models that generate outputs that appear convincing and real, even if they are inaccurate.
  • Web Search/Deep Research Features: AI model functionalities that allow them to access and process information from the internet or specialized databases, impacting their accuracy in providing references.
  • Specialized Research Tools: AI-powered platforms designed specifically for academic and research purposes, offering more reliable citation and information retrieval.

Testing Methodology and Criteria

The study aimed to evaluate the truthfulness of popular large language models (LLMs) for academic research by testing their ability to provide accurate references and ensure the cited claims are supported by the referenced content.

Success and Failure Criteria:

  1. Accurate References (First-Order Hallucinations): Does the AI provide a reference that actually exists?
  2. Supported Claims (Second-Order Hallucinations): Does the cited reference actually support the claim being made, and is it cited for the proper reasons?

The testing involved using specific prompts designed to stress-test the LLMs, asking them to provide actual answers, real quotations, and format them in an APA bibliography style.

Model Testing and Results

The following LLMs were tested across different platforms (ChatGPT, Claude, Gemini):

  • ChatGPT: Tested various models including GPT-5 Thinking, GPT-5 Auto, and GPT-5 Agent.
  • Claude: Tested models like Sonnet 4 and Opus 4.1.
  • Gemini: Tested models including Flash 2.5 Pro and Flash 2.5.

Important Note: The study found that paying for premium versions of these models did not necessarily guarantee improved accuracy in providing references.

First-Order Hallucinations: Do References Actually Exist?

  • Overall Averages:

    • ChatGPT: Over 60% correct responses.
    • Claude: Approximately 56% correct responses.
    • Gemini: Only 20% correct responses.
  • Detailed Model Performance (First-Order Hallucinations):

    • ChatGPT:
      • GPT-5 Thinking (with web search enabled): By far the best performer.
      • GPT-5 Auto + Deep Research: Provided real references when deep research or web search was enabled.
    • Claude:
      • Sonnet 4 + Research: Achieved a 100% success rate in providing real references.
      • Opus 4.1: Performed poorly, with none of the provided references actually existing.
    • Gemini:
      • Flash 2.5 Pro (paid version) + Deep Research: None of the references existed.
      • Flash 2.5: None of the references existed.
      • Other Gemini models: Only 40% of references existed.

Conclusion on First-Order Hallucinations: ChatGPT, particularly GPT-5 with web search or deep research enabled, was the most reliable for generating existing references. Gemini performed significantly worse.

Second-Order Hallucinations: Does the Citation Support the Claim?

  • Overall Averages (across all models):

    • ChatGPT: Just under 50% of citations contained the information they were cited for.
    • Claude: Just over 40% of citations contained the information they were cited for.
    • Gemini: 0% success rate. Gemini was unable to provide any references where the content supported the claim being cited.
  • Detailed Model Performance (Second-Order Hallucinations):

    • ChatGPT:
      • GPT-5 Thinking (with deep research or web search): Performed best.
      • GPT-5 Auto + Deep Research: Performed well.
      • GPT-5 Agent: 0% success rate.
    • Claude:
      • Had approximately 40-50% success rate in citing references where the claim was supported.
    • Gemini:
      • Had a 0% success rate.

Conclusion on Second-Order Hallucinations: Gemini failed entirely in this regard. ChatGPT performed better than Claude, but still struggled with over half of its citations not supporting the claims.

Overall Ranking and Recommendations

Based on both first and second-order hallucination tests, the models are ranked from least to most successful:

  1. Most Successful: ChatGPT-5 Thinking + Web Search. This combination offered the best chance of obtaining real references that contained appropriate information for their citations.
  2. Highly Recommended: ChatGPT-5 Auto + Deep Research or Web Search.
  3. Mixed Batch: Various Claude models fell into this category.
  4. Worst Performers: Models that achieved 0% for both real references and supported claims. This included some Gemini models and the ChatGPT-5 Agent.

Key Take-Home Messages:

  • Payment Does Not Guarantee Accuracy: Paying for premium AI models does not automatically make them more accurate, as demonstrated by Gemini's poor performance in this test.
  • Gemini Struggled: Gemini performed poorly across the board in this specific academic research context.
  • Common Failure Mode: A frequent issue was LLMs citing information from the introduction of a paper, essentially referencing what another source cited within that paper, rather than primary sources.
  • Plausibility Machines: LLMs are highly effective "plausibility machines," generating outputs that seem convincing and real. This necessitates manual verification of every claim and reference.
  • Safe Workflow: The only reliable method is to generate content with an LLM and then manually trace every claim to its specific PDF and page number for verification.

Alternative Tools for Academic Research

The presenter strongly advises against using general LLMs for gathering references due to their unreliability. Instead, they recommend specialized AI tools designed for academia:

  1. Elicit: This tool uses real papers and verifies them in the background, ensuring that provided references exist and contain the cited information.
  2. Scispace: Described as a "powerhouse," Scispace allows users to search for papers, create literature reviews, and is built upon real-world references.
  3. Consensus: Recommended for quickly finding "yes or no" answers from specific research fields.

Final Recommendation: LLMs are excellent for language manipulation but not for reliable information gathering. For academic research, utilize specialized tools like Elicit, Scispace, and Consensus, and always verify information manually. The channel will continue to stress-test these tools for academic research.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video