Which AI Wins for Marketers? We Tested Them All
By Authority Hacker Podcast
Share:
Key Concepts
- LLMs (Large Language Models): AI models used for various tasks like copywriting, data analysis, and customer support.
- Thinking Models vs. Non-Thinking Models: Thinking models (e.g., DeepSeek, Gemini 2.0 Flash Thinking) engage in internal monologues to break down problems, leading to potentially better but more expensive and slower results. Non-thinking models are faster and cheaper but may lack depth.
- Context Size: The amount of information an LLM can process in a single prompt, measured in tokens (approximately 0.75 words per token).
- API (Application Programming Interface): A way to access and use LLMs programmatically, often with usage-based pricing.
- Token: A unit of measurement for the input and output of LLMs, roughly equivalent to 0.75 words.
- Hallucination: When an LLM generates incorrect or fabricated information.
- Model Agnostic: The practice of using different AI models for different tasks, rather than relying on a single model for everything.
Test Setup and Methodology
- Typing Mind: A platform used to test various LLMs by plugging in API keys from different providers (OpenAI, Anthropic, Google, DeepSeek, etc.).
- Real-World Marketing Tasks: The models were tested on tasks such as copywriting, content brainstorming, data transformation, summarization, light coding, campaign planning, data analysis, and customer support chatbot development.
- Challenging Prompts: The prompts were designed to be complex and specific to differentiate the models' performance.
- Cost Analysis: The cost of using each model was tracked to evaluate the quality-to-cost ratio.
Copywriting Test
- Task: Write a sales email with a cynical, tech-savvy persona for an AI marketing course.
- Prompt: A detailed prompt outlining the persona, product, problem, solution, benefits, USPs, call to action, and tone.
- Key Findings:
- Claude 3.7 Sonnet (with reasoning) performed the best, producing engaging and persuasive copy. Cost: $0.043.
- GPT 4.5 was good but not as effective as Claude 3.7 Sonnet, especially considering its high cost. Cost: $0.13.
- DeepSeek offered good value for the price, but the copy was slightly cliché. Cost: $0.0037.
- GPT-4o mini was generic and lacked personality.
- Anthropic models have been significantly better at copywriting and marketing-based tasks.
- Conclusion: Switching to Claude can significantly improve the quality of marketing copywriting.
Content Brainstorming Test
- Task: Generate 10 YouTube video ideas for a channel focusing on AI and automation for entrepreneurs and marketers.
- Prompt: A prompt explaining the channel's focus, target audience, and desired output format.
- Key Findings:
- Gemini models (Google's models) performed poorly, generating generic and unappealing titles.
- DeepSeek was surprisingly good, suggesting real tools and concepts. Cost: $0.002.
- GPT-4o mini was boring and corporate.
- Grok 3 generated some good ideas but lacked specific details.
- Claude 3.7 Sonnet picked up on "make money" video trends and generated decent ideas.
- GPT 4.5 was expensive but provided more thought-out ideas with specific tool recommendations. Cost: $0.32.
- GPT-4 was significantly worse than other models.
- Conclusion: DeepSeek, Grok 3, and Claude 3.7 Sonnet understood YouTube trends better than others. GPT 4.5 was valuable for its specific tool recommendations.
Data Transformation Test
- Task: Prepare candidate profiles for a recruiter by combining data from a JSON file with a job specification.
- Prompt: A prompt instructing the model to highlight information that directly matches job requirements and maintain consistent formatting.
- Key Findings:
- DeepSeek created a visually appealing profile but was inconsistent in formatting.
- Gemini 2.0 Flash Thinking was verbose but provided decent information.
- GPT-4o mini was too short and lacked analysis.
- Gemini 2.0 Flash offered the best value for money due to its low cost and decent performance. Cost: $0.0025.
- Conclusion: Gemini 2.0 Flash is a cost-effective option for automation tasks.
Summarization Test
- Task: Read a book ("The Business of Being a Housewife") and create a step-by-step SOP for making a family budget.
- Prompt: A prompt instructing the model to identify the process in the book and prepare a concise and clear SOP.
- Key Findings:
- DeepSeek got distracted and focused too much on food budgeting.
- Gemini 2.0 Flash Thinking was too verbose.
- GPT-4o mini hallucinated quotes.
- Gemini 2.0 Flash was decent and included templates.
- GPT-4o mini made up quotes and was not as good as Gemini 2.0 Flash.
- Conclusion: Gemini 2.0 Flash is a clear winner for small models, outperforming GPT-4o mini.
Light Coding Test
- Task: Write an app script to add a menu item in Google Sheets that allows users to set an API key for Gemini 2.0 Flash and call the API.
- Prompt: A detailed prompt providing information about the API and desired functionality.
- Key Findings:
- DeepSeek failed completely.
- Gemini 2.0 Flash Thinking and GPT-4o mini worked after receiving error messages.
- Grok 3 failed.
- Claude 3.7 Sonnet got it right on the first prompt.
- Conclusion: Claude 3.7 Sonnet is the best model for light coding tasks.
Campaign Planning Test
- Task: Develop a high-level marketing plan for the launch of an AI marketing course.
- Prompt: A large prompt providing context about the company, target audience, and past marketing efforts.
- Key Findings:
- GPT-4o mini and Claude wanted to transition the traditional SEO audience to AI.
- Other models suggested separate campaigns for SEO and AI audiences.
- Thinking models were better than creative models for complex queries.
- GPT 4.5 was nothing special. Cost: $0.78.
- Conclusion: Thinking models are better for marketing planning. Claude was the preferred model due to its balance of creativity and marketing knowledge.
Data Analysis Test
- Task: Analyze Facebook Ads data from a past launch and provide insights for improvement.
- Prompt: A prompt instructing the model to identify areas for improvement based on the provided data.
- Key Findings:
- Gemini wrote a lot of stuff that was hard to use.
- GPT-4o mini and Claude were the best.
- GPT-4o mini provided the most actionable insights.
- Claude told you what to do, but the formatting was less clear.
- Grok was okay and provided tables.
- Conclusion: GPT-4o mini is the best for data analysis due to its actionable insights.
Customer Support Chatbot Test
- Task: Build a customer support chatbot that can answer FAQs, empathize with users, and escalate to human support when necessary.
- Prompt: A large prompt providing FAQ links, login instructions, and payment plan details.
- Key Findings:
- All models (Gemini 2.0 Flash, GPT-4o mini, Claude) performed similarly, providing brief answers and directing users to links.
- The setup may have been flawed, as the models interpreted the task as replacing a search engine for help documents.
- Conclusion: The performance difference between models was minimal, suggesting that using a smaller model like Gemini 2.0 Flash or GPT-4o mini is more cost-effective.
Overall Conclusions
- Model Agnosticism: It is crucial to be model agnostic and use different LLMs for different tasks to maximize results and cost-effectiveness.
- Claude's Strengths: Claude excels at copywriting and marketing-related tasks.
- Gemini 2.0 Flash's Value: Gemini 2.0 Flash offers incredible value for basic automation tasks due to its low cost and decent performance.
- DeepSeek's Versatility: DeepSeek is a versatile model with a good balance of capabilities, but its API can be unreliable.
- GPT-4o mini's Analytical Prowess: GPT-4o mini is effective for data analysis tasks.
- The Importance of Prompt Engineering: Prompt engineering is crucial to get the best results from LLMs.
- Cost Considerations: The cost of using different models can vary significantly, especially for large-scale automation tasks.
- Google's Competitive Edge: Google's control over both hardware and software allows them to offer competitive pricing for their Gemini models.
- The Future of AI: The AI landscape is constantly evolving, and it is essential to stay updated on new models and their capabilities.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development