AI Engineer Paris 2025 (Day 2)

AI EngineerAbout 12 min readSep 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Engineering: Building AI applications, including content engineering, VLM, VIP, coding, MCP, and agents.
  • Agentic AI Stack: A more sophisticated AI stack including front-end frameworks, diverse APIs, agents, and databases with AI capabilities.
  • AI Agents: Software systems using AI to pursue goals, complete tasks, reason, plan, and learn autonomously.
  • Secure Sandbox Environments: Isolated environments for executing untrusted code generated by AI agents.
  • COAB: A serverless platform providing infrastructure for agents and inference, offering diverse hardware and secure sandbox environments.
  • Scale to Zero: A technique to reduce costs by scaling down idle workloads to zero, mitigating cold starts with memory snapshots.
  • Flux Model Family: A state-of-the-art image model family from Black Forest Labs, combining text-to-image generation and image editing.
  • Latent Flow Matching: An algorithm used in Flux models, involving latent generative modeling and flow matching to generate images.
  • Latent Generative Modeling: A technique to find a representation of an image that only contains the details that matter to human perception.
  • Flow Matching: An algorithm to find a vector field that maps from a simple distribution to a complex distribution of images.
  • Adversarial Diffusion Distillation: An algorithm to reduce the number of numerical integration steps in image generation, making it faster.
  • MCP (Meta-Control Protocol): A protocol for tool calling and agent interaction, enabling dynamic tool discovery and secure communication.
  • Computer User Agents: Agents that control a computer through the graphical user interface, mimicking human actions.
  • Policy Model: A visual language model (VLM) used by computer user agents to output actions based on tasks and memory.
  • Localizer Model: A model used by computer user agents to output the coordinates of elements for actions like clicking.
  • Specialization Frontier: The point where specialized models outperform generalist models in specific tasks.
  • System Prompt Learning: A paradigm for LLM learning where the system prompt is updated based on feedback from the environment.
  • Prompt Optimization: Techniques to improve the performance of prompts, including traditional RL, metaprompting, and prompt learning.
  • Attention Mechanism: A fundamental component of transformer models that allows the model to focus on relevant parts of the input.
  • Attention D: A technology from ZML that runs the attention mechanism on the CPU instead of the GPU, enabling larger models and unlimited context.
  • Kernel Bypass: A technique to directly access the network card, reducing latency for network communication.
  • Notebook Llama: An open-source alternative to Notebook LM, built with Llama Index and Llama Cloud, providing tools for document processing and analysis.
  • Llama Pars: A Llama Cloud product for parsing complex documents with tables and images.
  • Llama Extract: A Llama Cloud product for extracting structured information from documents based on predefined schemas.
  • Workflows: A Llama Index abstraction for defining the logical flow of an application, allowing for controlled agent behavior.
  • Gemini: A natively multimodal model from Google DeepMind, capable of understanding and outputting multiple modalities.
  • VO3: A video generation model from Google DeepMind, capable of creating photorealistic videos with advanced features.
  • Genie 3: An open worlds model from Google DeepMind, allowing users to explore and interact with generated environments.
  • Gemini Nano: A small model from Google DeepMind, designed to run on mobile devices and browsers.
  • Gemma: An open model family from Google DeepMind, offering smaller and more efficient models.
  • Full Duplex Conversation: A conversational AI system where the model always speaks and always listens, allowing for natural interruptions and overlaps.
  • Audio Language Model: A model that processes and generates audio, using techniques like Mimi codec to compress audio into tokens.
  • Multistream Modeling: An architecture for modeling audio tasks, using multiple streams in parallel to represent different audio sources.

Main Topics and Key Points

Day 2 Opening Remarks (Ralph Jabri & Yan Leger)

  • Ralph Jabri, the MC, highlights the success of day one and introduces the lineup for day two, covering agents, MCP, open models, and generative media.
  • Yan Leger, CEO of COB, discusses the conference's growth to five tracks and over 35 speakers.
  • All sessions are recorded and available on YouTube.
  • COB is a serverless platform simplifying application deployments.

The State of AI Engineering: Managing State in AI Engineering (Emil Eifrem, Neo4j)

  • Context Engineering: The process of providing the best inputs to LLMs to get the best outputs.
  • Three Main Sources of State:
    • Rag Corpus
    • Agentic Memory
    • Application
  • Evolution of Application Architecture:
    • Pre-2022: Simple CRUD apps with UI, backend, database, and object storage.
    • 2023: Chatbots on top of an orchestration layer, using vector databases for unstructured data.
    • 2025: Real applications with embedded AI features, agents, and a complex data layer.
  • The Data Layer Mess: Vector databases adding structured data, relational databases adding unstructured data, leading to a fragmented landscape.
  • Four Properties of a Kick-Ass Data Layer for AI Applications:
    1. Handle Structured, Unstructured, and Semistructured Data: Store, retrieve, index, and handle transactional scope across all three types.
    2. Extract Entities from Unstructured Data: Consistently and reliably extract entities using named entity recognition and entity resolution.
    3. Link Entities Across Persistent Agentic Memory and Application Data: Link entities between agentic memory, Rag corpus, and application state for performance, reduced complexity, and reliability.
    4. Disambiguate Between First-Party and Derived Data: Differentiate between data collected directly from users and data derived from the Rag corpus for proper handling.
  • Demo: Shows how Neo4j can extract entities from a Wikipedia page and store both raw data and extracted entities in a graph.
  • Agentic Memory: Intrinsically graph-oriented, with examples like Zap, Mezzero, Cognney, and MCP's initial implementation.
  • Build vs. Buy: Enterprises are increasingly building AI solutions rather than buying them due to the promise of higher quality and the desire to rationalize software ecosystems.
  • Innovation in AI: Shifting from Silicon Valley to Europe, with Paris and Stockholm gaining traction.

Agents are the New Microservices (Tushar Jain, Docker)

  • Agents are the new microservices, requiring standardized packaging, trusted catalogs, and easy sharing.
  • C Agent: An open-source agent builder that packages agents as OC artifacts and shares them via an OCR registry or hub.
  • MCP Catalog: A catalog of trusted and verified MCP servers, similar to Docker Official Images.
  • Docker Desktop Tooling: Easy discovery and configuration of MCP servers, with security and containerization.
  • Demo: Shows how to easily connect MCP servers to clients in Docker Desktop, automating workflows and providing security.

Lessons Learned from Running an MCP Server at GitHub Scale (Martin Woodward, GitHub)

  • MCP is more than just tool calling; it includes resources, prompts, sampling, root, and elicitation.
  • Tools are not the answer; too many tools can confuse the LLM.
  • Dynamic tool discovery is essential to reduce the number of tools available to the LLM.
  • Installing MCPs is a pain, and local MCP installations are rarely upgraded.
  • Remote MCP servers are easier to upgrade and scale, but require good authentication.
  • Password-based authentication is bad; OAUTH support is key.
  • MCPs are pointless without discoverability; an open-source MCP registry is needed.
  • GitHub has created an open-source MCP registry for discoverability.

AI Continuously Redefining Cloud Infrastructure (Yan Leger, COB)

  • AI engineering has evolved from LLM-backed GPUs to content engineering, VLM, VIP, coding, MCP, and agents.
  • The agentic AI stack is more sophisticated, requiring a mix of GPU, CPU, and accelerators.
  • Agentic workloads require secure sandbox environments, performance, efficiency, and deployment speed.
  • COAB provides a global serverless platform for agents and inference, with diverse hardware and secure sandboxes.
  • Demo: Shows how to execute sandbox code using the COAB MCP server.
  • Agentic workloads need to create thousands of secure sandboxes daily with subsecond starts.
  • COAB leverages virtualization with cloud hypervisor to isolate containers.
  • Networking is a major bottleneck, mitigated by preemptively starting machines and caching images.
  • Scale to zero and autoscaling are used to increase efficiency, with memory snapshotting to reduce cold start times.
  • Predicts that Nvidia's GPU monopoly will fall, similar to what happened with Intel in the CPU market.

How to Unify Text to Image Generation and Image Editing (Andreas Blattmann, Black Forest Labs)

  • Image editing is as important as image generation, allowing for iteration and control.
  • Flux Context combines text-to-image generation and editing in one model, enabling character consistency, style reference, and local/global editing.
  • Image generation uses a prompt to describe a scene, while image editing uses an instruction text prompt to describe how to change the initial image.
  • Latent Flow Matching:
    • Latent Generative Modeling: Finds a representation of an image that only contains the details that matter to human perception.
    • Flow Matching: Finds a vector field that maps from a simple distribution to a complex distribution of images.
  • Flux models use a general transformer model conditioned on a text prompt and a context image.
  • Adversarial Diffusion Distillation: An algorithm to reduce the number of numerical integration steps, making image generation faster.
  • Latent Adversarial Diffusion Distillation: Applies adversarial diffusion distillation to the latent space, further reducing compute efforts.
  • Demo: Shows how to edit an image of a football club logo, changing its style and context in real time.

Building Computer User Agents (Lie Jiang, H)

  • Computer user agents control a computer through the graphical user interface, mimicking human actions.
  • APIs and MCPs won't work for the long tail of tasks due to inertia, business models, and the intelligence built into UIs.
  • Computer user agents use a policy model (VLM) to output actions and a localizer model to output coordinates.
  • Specialized models are more efficient than generalist models, as demonstrated by AlphaZero's performance in chess compared to GPT-5.
  • H's OLO models are state-of-the-art for UI localization.
  • Training OLO involves supervised fine-tuning on UI localization examples and reinforcement learning to optimize for task execution success.
  • Open weights models build trust, start conversations, verify performance, and drive innovation.
  • Next steps include extending to desktop and mobile, annotating use cases, and optimizing end-to-end task execution.
  • H is building a surfer showcase with a portal for launching agents and doing inference.

Vibe Coding with Data (Andreas Kollegger, Neo4j)

  • Multi-agent systems are used to break down complex problems, allowing for specialized models.
  • A multi-agent system can help build a knowledge graph by interacting with the user, finding data, modeling the data, and building the graph.
  • Demo: Shows how a multi-agent system can create a bill of materials graph from data files, enabling root cause analysis.

The State of Open LLMs in 2025 (VB, Hugging Face)

  • Open LLMs are competitive with proprietary models in terms of performance, with several open models in the top 10.
  • Open LLMs are easy to use, with options for serverless APIs, managed deployments, and self-deployment.
  • Serverless APIs offer a familiar SDK for interacting with open LLMs.
  • Managed deployments provide a simple way to deploy open LLMs with various GPU options.
  • Self-deployment offers maximum control over the model and data.
  • Having multiple deployment options increases optionality and reduces reliance on a single provider.
  • Trends in the Open LLM Landscape:
    • Reasoning Revolution: Open models now offer chain-of-thought reasoning, with the ability to distill reasoning to smaller models.
    • Large Context: Open LLMs now support 128K, 256K, and even 1 million context windows.
    • Lower Cost: The cost of using open LLMs has decreased due to software, hardware, and architecture optimizations.
  • Open LLMs have become easier to use, with standards for chat templates and native support for quantization.
  • Proprietary models still win in general reasoning, end-to-end multimodal, and safety/jailbreak scaffolding.
  • A playbook for trying open models is to pick a simple project, swap the proprietary model with an open model, evaluate, tune the prompt, and iterate.
  • Future trends include smaller and domain-specific models, effort-based reasoning, better quantization schemes, and sparse models.

Prompt Learning: Evolving System Instructions in Real World Environments (Aparna Dinakaran, Arise AI)

  • System prompts are key to building effective agents and should be consistently updated.
  • System prompt learning uses English feedback to improve the prompt, inspired by Karpathy's tweet.
  • The process involves taking data, the original prompt, and passing it into a meta prompt to generate a new prompt.
  • Case Study: Klein Coding Agent:
    • Initial benchmark showed 31% accuracy on Sweetbench light.
    • Error analysis revealed common failure patterns.
    • Prompt learning with 5-10 loops improved accuracy by 15%.
    • Significant improvements were seen in harder problem categories like salient translation and error detection.
  • Collecting errors and running evals is important to understand what to fix.
  • Online evals are more important than offline evals because they use real production data.
  • Prompt learning is also effective in other domains like structured JSON webpage generation and support query classification.
  • Future trend: Automated updates to system prompts based on data and evals.

Attention D: Pushing the Limits of What's Possible Beyond GPUs (Steve Morren, ZML)

  • The attention mechanism in transformer models has quadratic complexity, limiting context window size.
  • The softmax function in attention creates sparse outputs, suggesting that it can be modeled as a graph problem.
  • Modeling attention as a graph problem allows it to run in log(n) complexity, paving the way for unlimited context.
  • GPUs are bad at branching, which is required for graph-based attention.
  • CPUs are good at branching, but the question is whether they can do it fast enough.
  • ZML's Attention D runs the attention mechanism on the CPU instead of the GPU, freeing up GPU memory and enabling larger models.
  • Kernel bypass with DPDK is used to reduce latency for network communication.
  • Demo: Shows a 32B model running on a 32GB GPU with Attention D, which was previously impossible.

Scaling Realtime Voice AI (Neil Zeghidour, Qoutai)

  • Qoutai is a nonprofit AI research lab focused on open research and open science, particularly in multimodal LLMs.
  • The goal is to scale real-time voice AI for interactive and high-volume applications like gaming, robotics, and personalized media.
  • Key aspects of voice AI include synthesis, transcription, translation, transformation, and conversational experience.
  • Quality factors include fidelity, voice design, emotions, flow, and latency.
  • Scalability requires either large-scale cloud generation or small-scale on-device generation.
  • Moshi: The first full-duplex conversational AI, using a multistream architecture to allow the model to always speak and always listen.
  • Mimi Codec: Compresses audio into tokens, allowing audio to be processed by LLMs with similar efficiency to text.
  • Multistream modeling uses two streams in parallel to model different audio sources, enabling tasks like translation and transcription.
  • Qoutai's models are highly accurate and fast, with high throughput for real-time applications.
  • A demo shows how to create a new conversational experience by uploading a voice sample and writing a personality for the LLM.
  • Qoutai is looking for collaborations to incorporate its models into products, particularly for large-scale generation use cases.

Synthesis/Conclusion of the Main Takeaways

The AI Engineer Paris conference highlighted the rapid advancements and evolving landscape of AI engineering. Key takeaways include the increasing sophistication of AI stacks, the importance of managing state in AI applications, the shift towards building custom AI solutions, and the growing role of open-source models. The conference also showcased innovative techniques for improving model performance, such as prompt learning and specialized architectures, and emphasized the need for scalable and efficient infrastructure to support the next generation of AI applications. The presentations underscored the importance of community collaboration and the potential for AI to transform various industries, from image editing and robotics to voice AI and personalized media.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.