ChatGPT’s Hallucination Problem For Research TESTED!
By Andy Stapleton
Key Concepts
- First-Order Hallucinations: Refers to whether a cited academic reference (paper, study) actually exists.
- Second-Order Hallucinations: Refers to whether the specific claim or content for which a paper is cited is actually present within that paper. This tests the accuracy of the citation's justification.
- Claim Citation Match: The degree to which the content cited by the LLM matches the actual content of the referenced paper.
- Deep Research: A specific feature or mode within ChatGPT that involves a more thorough search for academic papers and information, leading to better accuracy.
- Web Search: A general web search feature, found to be less effective than Deep Research for academic referencing.
- ChatGPT5 Auto: A model configuration that automatically selects the best underlying model for a given task.
- ChatGPT5 Instant: A faster, less thorough model configuration.
- ChatGPT5 Thinking: A model configuration that involves more processing time, potentially leading to better results.
- ChatGPT5 Agent: An advanced model configuration designed to perform complex tasks by acting as an "agent" in the digital world.
Introduction: The Challenge of AI-Generated References
Researchers frequently question the reliability of references provided by large language models (LLMs) like ChatGPT, specifically whether they are real and accurately reflect the cited content. This video presents a stress test of the latest ChatGPT models to identify which ones provide trustworthy academic references and which are prone to "hallucinations." The study focuses on two types of hallucinations: first-order (does the reference exist?) and second-order (does the reference support the claim?).
Methodology: Stress Testing ChatGPT Models
The research team designed an experiment to evaluate the accuracy of references generated by various ChatGPT models.
Types of Hallucinations Tested:
- First-Order Hallucinations: Assesses if the provided reference (e.g., a research paper) genuinely exists. Previous studies on this channel have touched upon this.
- Second-Order Hallucinations: This is a more critical and less explored area. It evaluates whether the specific claim or reason for which a paper is cited is actually contained within that paper. An LLM might cite a real paper but attribute content to it that isn't there, or misrepresent its findings. This is crucial because a real reference with incorrect content is still misleading.
Evaluation Matrix:
- Accurate References: A "correct response" means the reference exists.
- Claim Citation Match: The response is "supported by the content of the paper," meaning the claim made by the LLM is genuinely found in the cited paper.
Models and Features Tested:
The following ChatGPT5 configurations were tested:
- Auto: Automatically selects the best model.
- Instant: A quicker response model.
- Thinking: A model that involves more processing.
- Agent: An advanced, autonomous agent model.
These models were tested with additional "sprinkles" (features):
- Deep Research: A specialized search mode for academic databases.
- Web Search: A general internet search.
Example Prompt Structure:
The research team used a specific prompt to elicit detailed responses and references: "My research team wanted to know three things: [specific research question, e.g., 'understand the behavior between X and Y']. Provide me with three things:
- A summary answer.
- The exact quotation where you found that information.
- An APA bibliography of the studies you provided."
This prompt was designed to thoroughly test the LLMs' ability to summarize, extract direct quotes, and provide accurate, appropriately cited references.
Results: First-Order Hallucinations (Reference Existence)
This section addresses whether the papers cited by the LLMs actually exist. Five randomly selected papers from each model's output were checked.
- ChatGPT5 Auto + Deep Research: Achieved 100% correct responses (5 out of 5 papers existed). This configuration consistently performed well.
- ChatGPT5 Instant + Web Search: Only two out of five papers were correct. This configuration was largely unreliable.
- ChatGPT5 Thinking + Web Search: Performed moderately.
- ChatGPT5 Thinking + Deep Research: Showed improvement over Thinking + Web Search, indicating the benefit of "Deep Research."
- ChatGPT5 Agent: Performed poorly, providing only one real reference out of five.
Key Finding: The "Deep Research" feature significantly improves the likelihood of generating real references. ChatGPT5 Auto, when allowed to select the best model, combined with Deep Research, was the most reliable for finding existing papers.
Results: Second-Order Hallucinations (Claim-Citation Match)
This section investigates the more critical aspect: whether the content cited by the LLM is actually present in the referenced paper.
- ChatGPT5 Auto + Deep Research: Performed the best among all tested configurations, demonstrating the highest accuracy in matching claims to paper content. However, it was "still not perfect," indicating some inaccuracies remained.
- ChatGPT5 Instant (with or without Deep Research): Consistently performed poorly in this regard.
- ChatGPT5 Thinking (with or without Deep Research): Performed better than Instant but not as well as Auto + Deep Research.
- ChatGPT5 Agent: Produced "not a correct result at all." Despite some of its references existing (as per first-order testing), the claims for which they were cited were entirely fabricated or misrepresented within the papers. This was a significant surprise, given the "Agent" model's advanced reputation.
Key Finding: "Deep Research" consistently enhances the accuracy of claim-citation matches. ChatGPT5 Auto with Deep Research is the most reliable, though still imperfect. The "Agent" model is particularly unreliable for ensuring content accuracy within cited papers.
Key Takeaways and Recommendations
The study provides clear, actionable insights for researchers using ChatGPT for academic work:
- Always Turn On Deep Research: This feature is crucial for both ensuring references exist and that the cited content is accurate. It consistently leads to better responses across models.
- Use ChatGPT5 Auto: This configuration, especially when paired with Deep Research, delivered the most reliable results for both first and second-order hallucinations. It appears to intelligently select the best underlying model for academic tasks.
- Avoid ChatGPT5 Agent for Referencing: Despite its advanced capabilities, the Agent model proved to be highly unreliable for generating accurate academic references and claim-citation matches. It frequently fabricated the reasons for citations, even if the papers themselves existed.
- Consider ChatGPT5 Thinking + Deep Research: If Auto is not available or preferred, Thinking with Deep Research is a viable, though slightly less optimal, alternative.
- Stay Away from ChatGPT5 Instant: This model consistently performed poorly in all aspects of reference accuracy.
Visual Summary (Scatter Graph Analogy): The ideal scenario is to have "real references" and a "correct claim citation match." The results indicate that configurations like "Auto + Deep Research" and "Thinking + Deep Research" fall into this desired quadrant, while "Agent" is far from it.
Conclusion
While large language models offer significant potential for research, their current ability to provide reliable academic references varies drastically. The study unequivocally demonstrates that ChatGPT5 Auto combined with the "Deep Research" feature is the most effective configuration for generating trustworthy references, minimizing both first and second-order hallucinations. Researchers should exercise caution, always verify AI-generated references, and specifically avoid general-use agents like ChatGPT5 Agent for critical academic tasks. The presenter also hints at future tests of other specific AI tools for academia, suggesting that specialized tools might offer even greater accuracy.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth

Nvidia Crushes Earnings again — What Jensen Huang sees next for AI
CGTN America

Samsung union suspends strike after reaching tentative pay deal • FRANCE 24 English
FRANCE 24 English

OH SH*T! The Banks are Dumping AI Loans!
Steven Van Metre