The 2025 AI Engineering Report — Barr Yaron, Amplify

AI EngineerAbout 4 min readAug 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • AI Engineering: The application of engineering principles to the development, deployment, and maintenance of AI systems.
  • LLMs (Large Language Models): AI models trained on vast amounts of text data, capable of generating human-like text, translating languages, and answering questions.
  • RAG (Retrieval-Augmented Generation): A technique for enhancing LLMs by retrieving relevant information from external knowledge sources and incorporating it into the generated output.
  • Fine-tuning: The process of further training a pre-trained model on a specific dataset to improve its performance on a particular task.
  • LoRA/Q-LoRA (Low-Rank Adaptation): Parameter-efficient fine-tuning methods that reduce the number of trainable parameters.
  • DPO (Direct Preference Optimization): A reinforcement learning technique for fine-tuning LLMs based on human preferences.
  • Supervised Fine-tuning: Fine-tuning a model using labeled data.
  • AI Agents: Systems where an LLM controls the core decision-making or workflow.
  • Observability: The ability to monitor and understand the internal state of a system based on its outputs.
  • Vector Database: A database optimized for storing and querying vector embeddings, which represent data points in a high-dimensional space.

Survey Overview and Demographics

  • The "2025 State of AI Engineering Survey" was conducted to understand the current AI engineering landscape.
  • 500 respondents participated, primarily identifying as engineers (software or AI).
  • Many seasoned software engineers are relatively new to AI, with nearly half having three years or less of AI experience, and 10% starting in the past year.
  • The term "AI Engineering" has seen a surge in interest since late 2022, coinciding with the launch of ChatGPT.

LLM Usage and Applications

  • Over half of the respondents use LLMs for both internal and external applications.
  • OpenAI models dominate external use cases, holding three of the top five spots and half of the top 10.
  • Top use cases include code generation/intelligence and writing/content generation.
  • 94% of LLM users employ them for at least two use cases, and 82% for at least three, indicating diverse applications.

Model Customization and Fine-tuning

  • RAG is the most popular customization method (70% of respondents).
  • Fine-tuning is more prevalent than expected.
  • 40% of fine-tuners use LoRA or Q-LoRA, indicating a preference for parameter-efficient methods.
  • Other fine-tuning techniques include DPO, reinforcement fine-tuning, and supervised fine-tuning.

Model and Prompt Management

  • Over 50% of respondents update their models at least monthly, with 17% updating weekly.
  • Prompt updates are even more frequent, with 70% updating monthly and 10% updating daily.
  • 31% of respondents lack any prompt management system.

Multimodal Adoption

  • Image, video, and audio usage lag behind text usage, indicating a "multimodal production gap."
  • Audio has the highest intent to adopt among those not currently using it (37%).

AI Agents

  • AI agents are defined as systems where an LLM controls the core decision-making or workflow.
  • 80% of respondents report LLMs working well, but only 20% say the same about agents.
  • Most respondents plan to use agents eventually.
  • The majority of agents in production have write access, often with a human in the loop.

Monitoring and Observability

  • 60% of respondents use standard observability methods to monitor their AI systems.
  • Over 50% rely on offline evaluation.
  • Human review remains the most popular method for evaluating model accuracy and quality.
  • Most respondents rely on internal metrics for monitoring model usage.

Vector Databases

  • 65% of respondents use a dedicated vector database.
  • 35% self-host their vector database, while 30% use a third-party provider.

Other Findings

  • Most respondents believe AI agents should disclose that they are AI.
  • Respondents are willing to pay more for faster inference time, but not by a wide margin.
  • Most believe transformer-based models will still be dominant in 2030.
  • The majority think open-source and closed-source models will converge.
  • The mean guess for the percentage of US Gen Z population that will have AI girlfriends/boyfriends is 26%.
  • Evaluation is the most painful aspect of AI engineering today.

Popular Resources

  • The survey identified the top 10 podcasts and newsletters that AI engineers actively learn from.
  • Swix is listed as both a popular newsletter and podcast for Latent Space.

Conclusion

The survey provides a snapshot of the rapidly evolving AI engineering landscape. Key takeaways include the widespread adoption of LLMs, the importance of RAG and fine-tuning, the challenges of prompt management, the emerging interest in AI agents, and the critical need for robust monitoring and evaluation practices. The results highlight the dynamism of the field and the ongoing efforts to translate cutting-edge research into practical applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.