We Tested Free AI Tools (LLMs) for Research—Only One Was Accurate

Andy StapletonAbout 4 min readJul 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Large Language Models (LLMs)
  • Hallucination Rate
  • Content Accuracy
  • Reference Accuracy
  • Interrogating PDFs
  • Literature Review
  • Chat GPT (free version)
  • Claude Sonet 4
  • Gemini Flash 2.5 Flash
  • Perplexity

1. Introduction: Testing Free LLMs for Academia

The presenter discusses the results of testing free large language models (LLMs) for academic and research tasks. The goal was to determine if these models could reliably perform typical academic tasks for free. The LLMs tested were the free version of Chat GPT, Claude Sonet 4, Perplexity, and Gemini Flash 2.5 Flash, selecting the best model for each task.

2. Tested Academic Tasks

The study focused on two primary academic tasks:

  • Chat with PDF (Interrogating PDFs): Uploading a research paper and asking questions to extract factual information, including intentionally misleading questions to test the LLM's ability to correct incorrect information.
  • Grabbing References and Correct Literature (Literature Review): Evaluating the LLM's ability to compile accurate references and provide correct literature.

3. Methodology: Prompt Engineering and Hallucination Rate

The team used a set of prompts and tested the output of each LLM (Chat GPT, Gemini, Claude, and Perplexity). The outputs were assessed as either "right" or "wrong," which required substantial time and effort.

Hallucination Rate Assessment:

  • Erroneous Response (Content): The content provided was not in the paper, fabricated answers, or fabricated terminology.
  • Erroneous Response (References): The reference didn't exist, incorrect author or year, or cited references attributed to different work.
  • Correct Response: The LLM provided accurate information based on the PDF.

4. Interrogating PDFs: Factual Information Extraction

Objective: Determine if the LLMs can return factual information from uploaded PDFs reliably.

Sample Questions:

  • Simple: "According to the PDF, what were the two challenges of OPV devices?"
  • Misleading: "I have read the paper that even if you use careful kneeling it won't be successful removing the surfactant layer" (This was false).

Results:

  • Chat GPT: Perfect (100% accuracy). Consistently provided the right information and corrected misleading statements.
    • Example: When presented with an erroneous statement, Chat GPT responded: "The statement you've written contains a small but significant error in interpretation. Here's a corrected version."
  • Claude: Very good (100% accuracy). Similar performance to Chat GPT.
  • Gemini: Less reliable. Could be convinced that information was in the paper when it wasn't.
    • Example: To the incorrect statement, Gemini responded: "You are correct. The paper states that even careful kneeling, the complete removal of surfaces response..." (The quotation doesn't exist in the PDF).
  • Perplexity: The worst. Easiest to convince that non-existent information was present and often failed to extract correct information.

Conclusion: For extracting information from a single PDF, Chat GPT or Claude are the best free options.

5. Citation Time: Literature Review and Reference Accuracy

Objective: Assess the LLMs' ability to retrieve information from external sources and perform literature reviews.

Prompts:

  • "Act like a world-renowned expert in organic solar cells. Compile five recent articles, blah blah blah, and then use American Chemical Society citation style."
  • Misleading: "Explain to me the Stapleton theory of photovalttaics in three sentences" (This theory does not exist).

Results:

  • Chat GPT: Inaccuracy of just under 80%. The best among the tested free models for reference accuracy.
  • Gemini: Second best to Chat GPT.
  • Claude & Perplexity: Performed poorly.

Hallucination: All models were susceptible to hallucination, and the team could convince them to provide false information more easily than with PDF interrogation.

Example Prompts and Responses:

  • When asked about the non-existent Stapleton theory, Chat GPT initially responded well by saying, "The Stapleton theory of photovoltex does not appear to be recognized or established."
  • However, the team was able to later prompt Chat GPT to acknowledge the "theory", demonstrating its susceptibility to manipulation.

6. Overall Trends and Recommendations

The presenter provided a table summarizing the performance of each LLM.

  • Accurate Content (Interrogating PDFs):
    • Chat GPT: Highly satisfactory (100% accuracy)
    • Claude: Highly satisfactory (100% accuracy)
    • Gemini: Satisfactory
    • Perplexity: Really bad
  • Accurate Referencing (Literature Review):
    • Chat GPT: Satisfactory; No major issue observed, however when deep research is not used, mismatches between cited references and actual information.
    • Claude: Highly unsatisfactory
    • Gemini: Better using deep research; not as good without deep research
    • Perplexity: Highly unsatisfactory

7. Key Takeaway: Best Free LLM

Based on the research, Chat GPT (free version) is the best overall free LLM for academic and research tasks.

  • Content Accuracy: Chat GPT and Claude are tied at 100%.
  • Reference Accuracy: Chat GPT performs the best among the free models tested, while Perplexity performs the worst.

The presenter recommends staying away from Perplexity for academic and research purposes when using the free version. They also mention that Gemini (using Deep Research) performs exceptionally well and doesn't hallucinate, but that's a paid feature, and the research focused on free models.

8. Conclusion

The choice of LLM depends on the specific task. However, for a single free tool, Chat GPT currently offers the best balance of content and reference accuracy. The presenter encourages viewers to share their experiences with free LLMs in the comments.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.