Which AI Wins for Marketers? We Tested Them All

By Authority Hacker Podcast

Share:

Key Concepts

  • LLMs (Large Language Models): AI models used for various tasks like copywriting, data analysis, and customer support.
  • Thinking Models vs. Non-Thinking Models: Thinking models (e.g., DeepSeek, Gemini 2.0 Flash Thinking) engage in internal monologues to break down problems, leading to potentially better but more expensive and slower results. Non-thinking models are faster and cheaper but may lack depth.
  • Context Size: The amount of information an LLM can process in a single prompt, measured in tokens (approximately 0.75 words per token).
  • API (Application Programming Interface): A way to access and use LLMs programmatically, often with usage-based pricing.
  • Token: A unit of measurement for the input and output of LLMs, roughly equivalent to 0.75 words.
  • Hallucination: When an LLM generates incorrect or fabricated information.
  • Model Agnostic: The practice of using different AI models for different tasks, rather than relying on a single model for everything.

Test Setup and Methodology

  • Typing Mind: A platform used to test various LLMs by plugging in API keys from different providers (OpenAI, Anthropic, Google, DeepSeek, etc.).
  • Real-World Marketing Tasks: The models were tested on tasks such as copywriting, content brainstorming, data transformation, summarization, light coding, campaign planning, data analysis, and customer support chatbot development.
  • Challenging Prompts: The prompts were designed to be complex and specific to differentiate the models' performance.
  • Cost Analysis: The cost of using each model was tracked to evaluate the quality-to-cost ratio.

Copywriting Test

  • Task: Write a sales email with a cynical, tech-savvy persona for an AI marketing course.
  • Prompt: A detailed prompt outlining the persona, product, problem, solution, benefits, USPs, call to action, and tone.
  • Key Findings:
    • Claude 3.7 Sonnet (with reasoning) performed the best, producing engaging and persuasive copy. Cost: $0.043.
    • GPT 4.5 was good but not as effective as Claude 3.7 Sonnet, especially considering its high cost. Cost: $0.13.
    • DeepSeek offered good value for the price, but the copy was slightly cliché. Cost: $0.0037.
    • GPT-4o mini was generic and lacked personality.
    • Anthropic models have been significantly better at copywriting and marketing-based tasks.
  • Conclusion: Switching to Claude can significantly improve the quality of marketing copywriting.

Content Brainstorming Test

  • Task: Generate 10 YouTube video ideas for a channel focusing on AI and automation for entrepreneurs and marketers.
  • Prompt: A prompt explaining the channel's focus, target audience, and desired output format.
  • Key Findings:
    • Gemini models (Google's models) performed poorly, generating generic and unappealing titles.
    • DeepSeek was surprisingly good, suggesting real tools and concepts. Cost: $0.002.
    • GPT-4o mini was boring and corporate.
    • Grok 3 generated some good ideas but lacked specific details.
    • Claude 3.7 Sonnet picked up on "make money" video trends and generated decent ideas.
    • GPT 4.5 was expensive but provided more thought-out ideas with specific tool recommendations. Cost: $0.32.
    • GPT-4 was significantly worse than other models.
  • Conclusion: DeepSeek, Grok 3, and Claude 3.7 Sonnet understood YouTube trends better than others. GPT 4.5 was valuable for its specific tool recommendations.

Data Transformation Test

  • Task: Prepare candidate profiles for a recruiter by combining data from a JSON file with a job specification.
  • Prompt: A prompt instructing the model to highlight information that directly matches job requirements and maintain consistent formatting.
  • Key Findings:
    • DeepSeek created a visually appealing profile but was inconsistent in formatting.
    • Gemini 2.0 Flash Thinking was verbose but provided decent information.
    • GPT-4o mini was too short and lacked analysis.
    • Gemini 2.0 Flash offered the best value for money due to its low cost and decent performance. Cost: $0.0025.
  • Conclusion: Gemini 2.0 Flash is a cost-effective option for automation tasks.

Summarization Test

  • Task: Read a book ("The Business of Being a Housewife") and create a step-by-step SOP for making a family budget.
  • Prompt: A prompt instructing the model to identify the process in the book and prepare a concise and clear SOP.
  • Key Findings:
    • DeepSeek got distracted and focused too much on food budgeting.
    • Gemini 2.0 Flash Thinking was too verbose.
    • GPT-4o mini hallucinated quotes.
    • Gemini 2.0 Flash was decent and included templates.
    • GPT-4o mini made up quotes and was not as good as Gemini 2.0 Flash.
  • Conclusion: Gemini 2.0 Flash is a clear winner for small models, outperforming GPT-4o mini.

Light Coding Test

  • Task: Write an app script to add a menu item in Google Sheets that allows users to set an API key for Gemini 2.0 Flash and call the API.
  • Prompt: A detailed prompt providing information about the API and desired functionality.
  • Key Findings:
    • DeepSeek failed completely.
    • Gemini 2.0 Flash Thinking and GPT-4o mini worked after receiving error messages.
    • Grok 3 failed.
    • Claude 3.7 Sonnet got it right on the first prompt.
  • Conclusion: Claude 3.7 Sonnet is the best model for light coding tasks.

Campaign Planning Test

  • Task: Develop a high-level marketing plan for the launch of an AI marketing course.
  • Prompt: A large prompt providing context about the company, target audience, and past marketing efforts.
  • Key Findings:
    • GPT-4o mini and Claude wanted to transition the traditional SEO audience to AI.
    • Other models suggested separate campaigns for SEO and AI audiences.
    • Thinking models were better than creative models for complex queries.
    • GPT 4.5 was nothing special. Cost: $0.78.
  • Conclusion: Thinking models are better for marketing planning. Claude was the preferred model due to its balance of creativity and marketing knowledge.

Data Analysis Test

  • Task: Analyze Facebook Ads data from a past launch and provide insights for improvement.
  • Prompt: A prompt instructing the model to identify areas for improvement based on the provided data.
  • Key Findings:
    • Gemini wrote a lot of stuff that was hard to use.
    • GPT-4o mini and Claude were the best.
    • GPT-4o mini provided the most actionable insights.
    • Claude told you what to do, but the formatting was less clear.
    • Grok was okay and provided tables.
  • Conclusion: GPT-4o mini is the best for data analysis due to its actionable insights.

Customer Support Chatbot Test

  • Task: Build a customer support chatbot that can answer FAQs, empathize with users, and escalate to human support when necessary.
  • Prompt: A large prompt providing FAQ links, login instructions, and payment plan details.
  • Key Findings:
    • All models (Gemini 2.0 Flash, GPT-4o mini, Claude) performed similarly, providing brief answers and directing users to links.
    • The setup may have been flawed, as the models interpreted the task as replacing a search engine for help documents.
  • Conclusion: The performance difference between models was minimal, suggesting that using a smaller model like Gemini 2.0 Flash or GPT-4o mini is more cost-effective.

Overall Conclusions

  • Model Agnosticism: It is crucial to be model agnostic and use different LLMs for different tasks to maximize results and cost-effectiveness.
  • Claude's Strengths: Claude excels at copywriting and marketing-related tasks.
  • Gemini 2.0 Flash's Value: Gemini 2.0 Flash offers incredible value for basic automation tasks due to its low cost and decent performance.
  • DeepSeek's Versatility: DeepSeek is a versatile model with a good balance of capabilities, but its API can be unreliable.
  • GPT-4o mini's Analytical Prowess: GPT-4o mini is effective for data analysis tasks.
  • The Importance of Prompt Engineering: Prompt engineering is crucial to get the best results from LLMs.
  • Cost Considerations: The cost of using different models can vary significantly, especially for large-scale automation tasks.
  • Google's Competitive Edge: Google's control over both hardware and software allows them to offer competitive pricing for their Gemini models.
  • The Future of AI: The AI landscape is constantly evolving, and it is essential to stay updated on new models and their capabilities.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video