AIE CODE 2025: AI Leadership ft Anthropic, OpenAI, McKinsey, Bloomberg, Google Deepmind, and Tenex

By AI Engineer

Share:

Here's a comprehensive summary of the provided YouTube video transcript:

Key Concepts

  • AI Engineering: The core theme of the conference, focusing on the development and application of AI in software engineering.
  • Coding Agents: AI systems designed to assist or automate coding tasks, including code generation, debugging, refactoring, and more.
  • Vibe Coding: A term used to describe the new paradigm of interacting with AI for coding, emphasizing an iterative, conversational approach rather than traditional manual coding.
  • Harnesses: The interface layer that allows AI models to interact with users, code, and tools, acting as the "special sauce" for AI products.
  • Context Management: Crucial for AI agents to effectively process and utilize information, including managing context windows, memory, and tool interactions.
  • Proactive Agents: AI agents that can anticipate user needs and take action without explicit instruction, aiming to reduce developer cognitive load.
  • AI Native Workflows & Roles: Shifting organizational structures and processes to fully integrate AI into the software development lifecycle, moving beyond point solutions.
  • Developer Experience (DevX): The overall experience of software developers, with a focus on how AI tools can improve productivity, reduce friction, and enhance job satisfaction.
  • ROI Measurement: The challenge of quantifying the return on investment for AI in software engineering, moving beyond simple usage metrics to focus on engineering and business outcomes.
  • Prompt Engineering: The art and science of crafting effective prompts to elicit desired behaviors from AI models.
  • Agent Orchestration: The process of coordinating multiple AI agents to work together to achieve complex tasks.
  • Model Behavior: The deliberate design and shaping of AI model outputs and interactions to align with product goals and user expectations.
  • AI Security: Addressing vulnerabilities like prompt injections and ensuring the safe and trustworthy use of AI in software development.
  • Compounding Engineering: A new paradigm where each AI-assisted feature makes subsequent features easier to build, creating a virtuous cycle of improvement.
  • Quality Gateways: Implementing automated checks and processes to ensure the quality and reliability of AI-generated code.

Summary of Presentations

Alex Lieberman (Host & Co-founder of Morning Brew, Managing Partner of 10X)

Alex Lieberman, co-founder of Morning Brew and managing partner of 10X, opened the AI Engineer Code Summit 2025. He highlighted his transition from a newsletter business to the AI frontier, co-founding 10X to help companies with AI transformation. He emphasized the conference's goal to provide a look back at the industry's progress and a tactical view of its future, featuring speakers from labs, startups, academia, and Fortune 500 companies. He also acknowledged the sponsors, including Google DeepMind (presenting sponsor), Anthropic (platinum sponsor), and others.

Caitlyn Les (Head of Engineering, Clawed Developer Platform at Anthropic)

Caitlyn Les from Anthropic discussed how Anthropic is evolving its platform to help developers build powerful agentic systems using Claude. She outlined three key areas:

  1. Harnessing Claude's Capabilities: Exposing customizable features like "thinking time" (allowing Claude to reason longer for complex tasks, with a token budget) and reliable tool use (both built-in and custom tools with defined schemas). Claude Code leverages these for debugging and quick answers.
  2. Managing Claude's Context Window: Addressing the complexity of context management for coding agents. This includes:
    • MCP (Model Context Protocol): A standardized way for agents to interact with external systems (e.g., GitHub, Sentry).
    • Memory Tool: Keeping context outside the window and pulling it in when needed (e.g., storing codebase patterns or Git workflow preferences).
    • Context Editing: Clearing irrelevant information from the window (e.g., old tool results). Combining memory and context editing showed a 39% performance bump in internal evals. Anthropic is also expanding context windows (up to 1 million tokens) and teaching Claude to better understand its context.
  3. Giving Claude a Computer: Enabling Claude to autonomously write and run code. This involves providing infrastructure for secure, sandboxed environments, container orchestration, and session persistence. Key primitives include the Code Execution Tool (allowing Claude to write and run code securely on Anthropic's servers) and Agent Skills (folders of scripts and resources Claude can access). For Claude Code, a web design skill example was given. Anthropic is hiring for roles in developer products and DevRel.

Mikuel Katasta (President and Head of AI at Replit)

Mikuel Katasta from Replit discussed building coding agents for non-technical users, focusing on autonomy as a Northstar. He presented a third dimension to the agent landscape: autonomy for non-technical users, contrasting supervised autonomy (like Tesla FSD, requiring user oversight) with the "Whimo" experience Replit aims for, where users are in the back seat and don't need technical expertise. He argued that autonomy shouldn't be conflated with long runtimes or seen as a vanity metric, but rather as maximizing "reducible runtime" where the agent can accomplish tasks without user technical decisions. The three pillars of autonomy are:

  1. Frontier Model Capabilities: Leveraging the baseline intelligence of LLMs.
  2. Verification: Ensuring local correctness at every step to prevent compounding errors. Replit found over 30% of individual features were broken on first generation by AI, leading to a focus on autonomous testing. They use browser-based testing, simulating UI interactions via DOM abstractions and programmatic interactions with the application (database, logs, APIs, UI clicks). They write Playwright code directly, which is LLM-amenable, expressive, and reusable, making it an order of magnitude cheaper and faster than computer vision approaches.
  3. Context Management: While long context windows are helpful, Replit emphasizes managing context effectively through offloading to the codebase (documentation, plan descriptions), file systems, and sub-agent orchestration. Sub-agent orchestration significantly improves memory compression and separation of concerns.

Lisa Orur (Engineering Leader at Zapier)

Lisa Orur from Zapier shared Zapier's AI agent journey, focusing on empowering their support team to ship code and address "app erosion" (API changes causing reliability issues).

  • Experiment 1: Empowering Support to Fix Bugs: Started two years ago, focusing on four target apps, with engineering review for merge requests.
  • Experiment 2: AI for App Erosion (Project Scout):
    • Discovery: Identified pain points like context gathering (API docs, internet searches) as significant time sinks.
    • API Development: Built APIs for diagnosis (LLM-based context gathering) and unit test case finding (search-based).
    • Embedding Tools: Realized tools need to be embedded in engineers' workflows. MCP (Model Context Protocol) helped embed tools within IDEs like Cursor.
    • Scout Agent: Developed an agent to orchestrate tools, starting with the support team for small, emergent bugs. The flow involves issue categorization, fixability assessment, merge request generation, and then review/adjustment by support.
    • Implementation: Kicked off by Zaps, embedding into support workflows, using a GitLab CI/CD pipeline (plan, execute, validate phases) with Scout APIs and Cursor SDK.
    • Impact: Scout generates 40% of support team's app fixes, doubling some support team members' velocity (from 1-2 tickets/week to 3-4). It also reduces friction in the triage flow and allows engineering to focus on more complex tasks. Support team members involved are transitioning to engineering roles.

Steve Yaggi (Engineering Leader at Sourcegraph) & Jean Kim (Author & Researcher at IT Revolution)

Steve Yaggi and Jean Kim discussed the future of tools, predicting the "death of the IDE" by 2026. They argued that current tools like Cloud Code, while popular, are too complex and have cognitive overhead. They likened current AI coding tools to drills or saws, contrasting them with future CNC machines for precision. They believe all code will be written by "giant grinding machines" overseen by engineers who don't look at code directly.

  • Challenges with Current Tools: Cognitive overhead, AI "lying, cheating, and stealing" (generating incorrect code), and developer resistance, particularly from senior engineers.
  • The "Diver Metaphor": Criticized the "bigger tank" approach (larger context windows) for AI agents, advocating instead for task decomposition and specialized agents (product manager diver, coding diver, review diver, etc.) rather than a single "bigger diver."
  • "Vibe Coding": Defined as anything where code is not typed by hand, emphasizing the iterative conversation with AI. They noted that AI will likely reshape technology organizations and the economy significantly.
  • Productivity Gains: Cited examples like Booking.com seeing double-digit productivity increases and Travelopia replacing a legacy application in six weeks with a small team. Dr. Top Pal at Fidelity built an application in five days using vibe coding.
  • Trust in AI: Research shows that the longer people use AI tools, the more they trust them, suggesting practice and skill development are key.
  • Hot Take: Steve declared that engineers still using IDEs after January 1st are "bad engineers."
  • Backlash: Acknowledged a real backlash against AI adoption, with many engineers refusing to use it.

Bill Chen & Brian Fioa (Applied AI Team at OpenAI)

Bill Chen and Brian Fioa from OpenAI discussed building coding agents, focusing on CodeX. They highlighted the rapid evolution of the field and the challenge of constantly adapting agents to new models.

  • Anatomy of a Coding Agent: Composed of three parts: User Interface (CLI, IDE, cloud agent), Model (GPT-5.1, CodeX Max), and Harness (prompts, tools, agent loop).
  • The Harness: The interface layer to the model, managing multi-turn interactions, tool calls, and user intent interpretation. Building a good harness is challenging due to tool compatibility (AV), prompt tuning, latency, context window management, and API changes.
  • Model Behavior: A combination of "intelligence plus habit." OpenAI trains models with habits like planning, gathering context, and testing. Over-prompting can hinder performance.
  • CodeX: An agent for everywhere you code (VS Code plugin, CLI, cloud, ChatGPT). It turns specs into runnable code, navigates repos, runs commands, and reviews PRs. The CodeX harness manages complex tasks like parallel tool calls, security, sandboxing, prompt forwarding, permissions, compaction, and MCP support.
  • Patterns for Building Agents:
    • Harness as the New Abstraction Layer: Reduces reliance on model upgrades, allowing focus on product differentiation.
    • CodeX as an SDK: Enables programmatic use via TypeScript or Python, integration into CI/CD pipelines (GitHub actions), and use within other agents (e.g., agents that make other tools).
    • Examples: Zed wrapping CodeX for an IDE interface, GitHub integrating directly with CodeX via an SDK.
  • Future of CodeX: Models will improve, handle longer horizon tasks unsupervised, and raise the trust ceiling. The SDK will evolve to support these capabilities, enabling agents to learn and adapt.

Martin Harrison & Natasha Mania (McKinsey & Company)

Martin Harrison and Natasha Mania from McKinsey's Software X practice discussed the people and operating model aspects of AI in software development. They noted a disconnect between AI's potential and the 5-10% productivity gains many enterprises are seeing, attributing this to bottlenecks in collaboration, manual code reviews, and amplified tech debt.

  • Bottlenecks: Uneven impact of AI across tasks and individuals, making work allocation difficult. Manual code reviews are increasing despite automation elsewhere. AI can amplify tech debt and complexity.
  • AI-Native Workflows & Roles: Top performers are seven times more likely to have AI-native workflows (scaling AI across the SDLC) and six times more likely to have AI-native roles (smaller pods, consolidated roles like "product builders" with full-stack fluency).
  • Client Study: A study with a bank showed positive results from team-level interventions: assigning stories using agents, co-creating prototypes with agents, reorganizing squads by workflow, and using agents for cross-repository impact analysis. This led to a 51% increase in code merges and improved efficiency.
  • Talent Model Shifts: Moving from "two-pizza teams" to smaller pods (3-5 individuals) with consolidated roles. Product Managers are creating code prototypes directly.
  • Scaling Change Management: Crucial for large organizations, involving getting many small things right (communication, incentives, upskilling). A reset with hands-on upskilling and measurement systems was needed for a tech company that saw usage drop off.
  • Measurement System: Emphasized the need for a holistic system capturing inputs (tools, upskilling), outputs (adoption, velocity, capacity), and outcomes (developer NPS, code quality, MTTR, time to revenue, cost reduction).
  • Key Takeaways: Start now, embrace human change, find the right model, set bold ambitions, and focus on the journey.

Jaor Dennis Blanch (Stanford Researcher)

Jaor Dennis Blanch, a Stanford researcher, presented findings on the impact of AI on software engineering productivity, using a machine learning model to replicate human expert evaluations of code commits.

  • AI Productivity Gains: A study of 46 teams using AI vs. 46 similar teams not using AI showed a median net productivity gain of ~10% as of July, with a widening gap between top and bottom performers.
  • Factors Driving Gains:
    • AI Usage (Tokens): Correlation is loose; quality matters more than quantity. A "death valley" effect was observed around 10 million tokens/month, where usage seemed to decrease productivity.
    • Environment Cleanliness Index: A composite score (tests, types, documentation, modularity, code quality) showed a decent correlation (R-squared of 0.40) with AI productivity gains. Investing in codebase hygiene is crucial.
  • AI Engineering Practices Benchmark: A tool to scan codebases for AI fingerprints, categorizing usage from Level 0 (no AI) to Level 4 (agentic orchestration).
  • Measuring AI ROI: Proposed a framework focusing on engineering outcomes (primary metric: engineering output via ML model) and guardrail metrics (rework, quality, tech debt, people/DevOps). Usage metrics (telemetry, experience sampling, surveys) are also important.
  • Case Study: A company saw a 14% increase in PRs after AI adoption, but code quality decreased by 9% and rework increased by 2.5x, suggesting a potential negative ROI if not managed properly. The key takeaway is that measuring impact beyond simple PR counts is essential.

Edidomar Freiedman (CEO of Kodto)

Edidomar Freiedman, CEO of Kodto, discussed the state of AI code quality, highlighting the hype versus reality. He noted that while AI adoption is high (82% daily/weekly use), developers have serious quality concerns due to a lack of frameworks for measuring and managing quality.

  • AI Adoption & Concerns: High adoption rates (82% weekly/monthly use), but 67% of developers have quality concerns. AI influences up to 80% of code for some developers.
  • Quality Dimensions: Problems exist across the SDLC (planning, development, review, testing, deployment) and at code level (security, inefficiency) and process level (learning, verification, guardrails).
  • Impact: 42% of developers spend more time fixing AI-generated issues, leading to 35% project delays. Some reports indicate 3x more security incidents with AI-generated code.
  • Solutions:
    • Testing: Heavily using AI for testing doubles trust in AI-generated code.
    • Code Review: AI code review tools can block PRs based on quality gates (e.g., test coverage), improving quality (double quality gain) and productivity (47% improvement). Kodto scans 1 million PRs/month, finding 17% with high severity issues.
    • Context: Developers distrust AI due to poor context (80% of the time). Better context (code, versioning, PR history, logs, standards) is crucial for AI quality. Kodto's context engine is a key technology.
  • Recommendations: Invest in automated quality gateways, parallel agents for background quality workflows, intelligent code review/testing, and living documentation. The future involves agents orchestrating quality and verification workflows dynamically. Quality is a competitive edge.

Olive Song (Senior Researcher at Minimax)

Olive Song from Minimax introduced Minimax M2, an open-weight, 10 billion parameter model designed for coding and agentic tasks.

  • Minimax M2: Highly cost-efficient, top-ranked in intelligence and agent benchmarks among open-source models. High community adoption (most downloads, top 3 token usage on OpenRouter).
  • Key Characteristics & Training:
    • Code Experience: Supported by scaled environments and "expert developers" as reward models, leading to full-stack, multilingual capabilities.
    • Long Horizon Tasks: Achieved through "interleaved thinking" (thinking interleaved with tool calling) and reinforcement learning, allowing adaptation to noisy environments and automation of complex workflows.
    • Robust Generalization: Supported by data pipeline perturbations, enabling generalization to various agent scaffolds and operational spaces.
    • Multi-Agent Scalability: Enabled by its small size and cost-effectiveness, supporting parallel agent tasks.
  • Future: Development of M2.1, M3, with focus on better coding, memory, context management, proactive AI, vertical experts, and multimodal integration.

Kath Corvec (Director of Product at Google Labs)

Kath Corvec from Google Labs discussed Jules, a proactive asynchronous autonomous coding agent from Google Labs, aiming to reduce developer cognitive load.

  • The Problem: Developers are serial processors, carrying mental load when managing asynchronous agents. Proactivity is needed to reduce context switching costs (up to 40% of productive time).
  • Jules's Vision: Agents should behave like collaborators, not command-line utilities, understanding context, anticipating needs, and intervening at the right moment.
  • Proactive System Ingredients: Observation (continuous understanding of code, patterns, workflow), Personalization (learning user habits, preferences, ignored code), Timeliness (intervening at the right moment), and Seamless Workflow Integration (working within existing tools like terminals, repos, IDEs).
  • Jules's Proactivity Levels:
    • Level 1 (Attentive Sous Chef): Detects and fixes issues like missing tests, unused dependencies, unsafe patterns automatically.
    • Level 2 (Kitchen Manager): Becomes contextually aware of the entire project, learning user workflows and frameworks.
    • Level 3 (Convergence): Understands context and consequence, integrating with other agents (Stitch for design, Insights for data) to propose improvements across boundaries.
  • Jules Features: Memory (editable agent memories), Critic Agent (adversarial code review), Verification (Playwright scripts, screenshots), To-Do Bot (proactively working on future tasks), Best Practices integration, Environment Setup, and Just-in-Time Context.
  • Demo: Showcased Jules indexing a codebase, identifying to-dos and best practices, and allowing manual or automated initiation of tasks.
  • Analogy: Compares proactive AI to Google Nest learning user habits or the human body anticipating falls.
  • Future Vision: Moving from code awareness to system awareness, with agents collaborating across the full project lifecycle. The challenge is to invent the future of software development, questioning old ways of building software.

ASAF Board (Engineering Leader at Northwestern Mutual)

ASAF Board from Northwestern Mutual shared their journey in building GenBI (Generative AI + Business Intelligence) and the challenges of balancing innovation with risk aversion in a large enterprise.

  • GenBI: An agent that helps answer business questions with data, aiming for data democratization.
  • Northwestern Mutual Context: A 160-year-old, risk-averse financial services company with significant data, resources, and talent, but also a strong emphasis on stability ("don't f up").
  • Challenges: No prior precedent for GenBI, preference for using messy real data, building trust (with users and leadership), and securing budget for an unproven concept.
  • Approach:
    • Real Data: Used messy, real data to understand production complexities and involve end-users early.
    • Building Trust: Employed a "crawl, walk, run" approach, starting with BI experts, then business managers, before considering executives. Focused on delivering verified information from existing reports initially, rather than generating new SQL.
    • Incremental Process: Broke down the project into 6-week sprints with tangible deliverables, allowing for continuous learning, risk control, and leadership buy-in. Phases included natural language to SQL, metadata understanding, semantic search, data pivoting, and eventually full NBI agent capabilities.
  • Productization: Each step (rag agent, metadata agent, SQL agent) could be packaged as a product, demonstrating incremental business value (e.g., automating 80% of BI team's report-finding tasks, proving metadata value through AB tests).
  • Future: Evaluating tools like DataBricks Genie, enriching metadata, and exploring usage-based pricing models for SAS in the GenAI era.

Lei Jang (Head of Technology Infrastructure Engineering at Bloomberg)

Lei Jang from Bloomberg discussed their experience deploying AI within Bloomberg's engineering organization, focusing on evolving the codebase and handling incidents.

  • Bloomberg Context: 9,000+ engineers, large codebase (billions of lines, largest JavaScript codebase), significant AI research and product teams, private network, and contributions to open source (Envoy, AI Gateways).
  • AI for Coding: Started ~2 years ago, initially seeing benefits in greenfield projects (proof-of-concept, test rollouts) but limitations in broader applications.
  • Key Initiatives:
    • Uplift Agents: Broadly scanning the codebase to identify and apply patches, facing challenges with deterministic verification, increased PR open rates, and longer merge times.
    • Incident Response Agents: Developing agents to analyze codebase, telemetry, feature flags, and call traces unbiasedly to assist in troubleshooting.
  • Platform Principles & Pay Path: Emphasized providing a "golden path" with enablement teams, making the "right thing easy" and the "wrong thing hard." This includes a gateway for model experimentation, a tool discovery hub (MCP directory), a platform-as-a-service for tool creation/deployment, and easy proof-of-concept capabilities.
  • Adoption & Change Management: Incorporated AI coding into onboarding training to drive adoption and create change agents. Leveraged existing "champ" and "guild" programs for AI communities to foster shared learning and inner-sourcing. Addressing leadership adoption gaps through dedicated workshops.
  • Cost Function Shift: AI changes the cost function of software engineering, making some work cheaper and enabling engineers and leaders to focus on higher-value activities and strategic questions about quality and purpose.

Samir Motti (Head of AI Engineering at The Browser Company)

Samir Motti from The Browser Company discussed reimagining the browser around AI-native experiences with their new browser, DIA.

  • Browser Company Journey: Started with Arc (incremental improvement), then evolved to DIA (AI-native browser) based on the thesis that AI will transform internet usage and the browser itself.
  • Lessons Learned:
    • Optimizing Tools & Process for Iteration: Moved from rudimentary prompt editors to integrating all tools (prompts, context, models, parameters) into the product itself, enabling 10x speed improvements and broader team access for ideation and refinement. This includes tools for prototyping, evals, data collection, and "hill climbing" (optimizing AI products).
    • Jepa Mechanism: A sample-efficient way to improve LLM systems without RL or fine-tuning, using prompt seeding, task execution, scoring, PA selection, and LLM reflection for prompt mutation.
    • Model Behavior as Craft & Discipline: Treating model behavior (style, tone, response shaping) as a product design element. This involves behavior design, data collection, and model steering (prompting, model selection, context management). The process is iterative, with feedback loops from internal and external users.
    • AI Security (Prompt Injections): Crucial for browsers due to their access to private data, exposure to untrusted content, and external communication capabilities. Strategies include tagging untrusted context, separating data and instructions, and building security into the product experience (e.g., confirmation steps for autofill, event scheduling, email writing).
  • Future Vision: Embracing technology shifts with conviction, recognizing that AI is not just a product evolution but a company one, impacting hiring, training, communication, and collaboration.

Max Canet Alexander (Executive Distinguished Engineer at Capital One)

Max Canet Alexander from Capital One discussed essential investments for AI-assisted organizations, emphasizing that "what's good for humans is good for AI."

  • No Regrets Investments:
    • Standardize Development Environments: Use industry-standard tools (package managers, linters) that align with AI training sets. Avoid obscure languages or custom tooling that fights the training data.
    • CLI/API Access: Agents need CLIs or APIs for actions. Prefer text-based interactions over orchestrating browsers unless absolutely necessary.
    • Validation: Invest in objective, deterministic validation (tests, linters) to provide clear error messages for agents. Address legacy codebases that lack testability.
    • Structure: Agents work better on well-structured codebases. Improve codebase structure and testability.
    • Documentation: Write down external context, intentions, and the "why" behind code, as agents cannot access tribal knowledge or implicit information.
    • Code Review Velocity: Make individual responses faster, improve code review quality through apprenticeship, distribute review load, and set SLOs.
    • Code Review Quality: Maintain a high bar for code quality, rejecting suboptimal code to prevent a vicious cycle of decreasing productivity.
  • Addressing Fear: Emphasize psychological safety, transparency about AI's role (augmentation, not replacement), clear intents, and proactive communication.
  • Metrics: Focus on core developer experience metrics (speed and quality) like PR throughput, change failure rate, code quality, change confidence, and maintainability, rather than just AI utilization. Utilize telemetry, experience sampling, and effective surveys.
  • Employee Success: Provide education and time to learn AI skills. Tie AI adoption to employee success and competitive edge.
  • Bottlenecks: Identify and fix actual SDLC bottlenecks (context switching, meetings, legacy code modernization) rather than just focusing on AI-assisted code generation. Morgan Stanley's Dev Gen AI for mainframe modernization and Zapier's hiring strategy based on AI-driven engineer effectiveness were cited.

Dan Shipper (Founder of Every)

Dan Shipper, founder of Every, shared dispatches from the future on building an AI-native company, emphasizing the 10x difference in organizations with 100% AI adoption.

  • Every's Model: A small team (15 people) running four software products with high MRR growth, low funding, and 99% of code written by AI agents. Each app is built by a single developer.
  • Compounding Engineering: A paradigm where each feature makes the next easier to build, achieved through a four-step loop: Plan, Delegate, Assess, Codify.
    • Plan: Detailed planning for agents.
    • Delegate: Telling the agent to perform the task.
    • Assess: Evaluating agent output using tests, agent feedback, code review, etc.
    • Codify: Turning learnings into explicit prompts for reuse across the organization.
  • Second-Order Effects:
    • Tacet Code Sharing: Easier to learn from and reuse code/processes from other developers via agents.
    • New Hire Productivity: Agents set up environments and provide guidance, making new hires productive on day one.
    • Expert Freelancer Collaboration: Lower startup cost for external experts to contribute specific functionalities.
    • Cross-Product Commits: Developers easily contribute fixes or improvements to other products within the company.
    • Stack Flexibility: AI facilitates translation between different languages and frameworks, reducing the need for strict standardization.
    • Manager Commits: AI enables managers with fractured attention to contribute code by investigating bugs and submitting PRs.
  • Key Takeaway: 100% AI adoption creates a fundamentally different operational model, enabling single developers to build complex products and fostering cross-organizational collaboration through compounding engineering.

N.L.W. (Host of AI Daily Brief, CEO of Super Intelligent)

N.L.W. discussed the status of enterprise AI adoption and ROI findings from a study of ~2500 use cases.

  • Enterprise Adoption: Growing, with significant inflection in coding/software engineering. Agents are seeing meaningful uptake, with 42% of large enterprises having production agents (up from 11% in Q1).
  • ROI Challenges: Traditional impact metrics struggle with AI. Expectations for ROI realization are high and accelerating.
  • Study Findings (ROI Survey):
    • Overall ROI: 44.3% see modest ROI, 37.6% see high ROI. Only 5% see negative ROI. Expectations are very optimistic.
    • Impact Categories: Time savings (35%), increased output, quality improvement, and new capabilities are dominant.
    • Time Savings: Clusters around 5-10 hours/week, representing significant win-backs of work weeks per year.
    • Organizational Size: Differences observed, with mid-sized organizations (200-1000 people) focusing more on increasing output.
    • Role: Seuites/leaders are less focused on time savings and more on new capabilities and transformational impact.
    • Risk Reduction: Lowest primary benefit category (3.4%), but highest likelihood of transformational impact (25%).
    • Coding/Software Use Cases: Higher ROI and lower negative ROI than average.
    • Systematic Adoption: Organizations submitting more use cases and thinking cross-organizationally see better ROI.
  • Future Focus: Moving from generic impact conversations to systematic experiments, with a focus on automation and agentic use cases showing significantly higher ROI.

Arman Hezarki (Co-founder & Managing Partner at 10X)

Arman Hezarki from 10X presented on compensating engineers like salespeople to scale output, not overhead.

  • The Problem: Traditional hourly or salary+bonus models may not incentivize engineers to leverage AI effectively. Equity-based models have risks, especially for non-Google-like companies.
  • 10X Model: Engineers are compensated based on story points completed, directly linking pay to output. This is counterbalanced by strategists compensated on Net Revenue (customer happiness) and rigorous internal QA/client approval processes.
  • Case Studies:
    • Billboard Company: Built an AI moderation model with 96% accuracy in two weeks, paid per story point.
    • Retailer Devices: Developed five parallel AI models (heat mapping, Q detection, theft detection) for low-power devices, again paid per story point.
  • Risks & Mitigation:
    • Inflated Story Points: Mitigated by strategists scoping work and rigorous review processes.
    • Quality Drop: Mitigated by hiring the right people and robust QA.
    • Sharp Elbows: Mitigated by hiring the right people and ensuring strategists are involved in client satisfaction.
  • Core Principle: "AI makes people look like crazy mirrors." It amplifies existing attributes. Hiring the right people is paramount, making everything else easier.

Justin Rio (Deputy CTO at DX)

Justin Rio from DX discussed effective leadership in AI-enhanced organizations, focusing on reducing fear, improving metrics, and enabling employee success.

  • Current Impact Volatility: GenAI impact varies widely, with studies showing both productivity increases and decreases. A "induced flow" can make engineers feel productive even when data shows otherwise.
  • Key Factors for Positive Impact: Clear AI policies, time for learning and experimentation (not just providing materials), and effective measurement systems. Top companies see positive impacts on KPIs like change confidence, code quality, and change failure rate.
  • Reducing Fear: Emphasize psychological safety, transparency about AI's role (augmentation, not replacement), clear intents, and proactive communication.
  • Metrics: Focus on core developer experience metrics (speed and quality) like PR throughput, change failure rate, code quality, change confidence, and maintainability. Utilize telemetry, experience sampling, and effective surveys. DX's DXAI Measurement Framework normalizes metrics into utilization, impact, and cost.
  • Enabling Employee Success: Provide education, time to learn, and tie AI adoption to employee success. Distribute guides, establish feedback loops for system prompts, and manage model temperature (creativity vs. determinism).
  • Unblocking Usage: Leverage self-hosted/private models, partner with compliance early, and think creatively around barriers.
  • Integrating Across SDLC: Identify and fix actual SDLC bottlenecks (context switching, meetings, legacy code modernization) rather than just focusing on AI-assisted code generation.

Mel Lutzky (Co-founder & CEO of Graphite)

Mel Lutzky, co-founder and CEO of Graphite, discussed AI-powered code review and the broader development process.

  • Graphite's Mission: Applying AI to the entire development process, making code review as quick as possible.
  • The Challenge: Writing code is only the first step; testing, review, merging, and deployment often take as long or longer.
  • Graphite's Solution: An AI agent integrated into pull requests, aiming to provide a "2025" code review experience, moving beyond 2015 practices.
  • Sponsorship: Graphite sponsored the afterparty and tomorrow's event at Public Records, aiming to showcase New York experiences for attendees.

Alex Lieberman (Host & Co-founder of Morning Brew, Managing Partner of 10X)

Alex Lieberman provided closing housekeeping and thanked attendees, speakers, and the production team. He encouraged networking during the break and afterparty, suggesting "hottest takes" on AI or debating "is a hot dog a sandwich?" as conversation starters. He announced that his co-founder, Arman, would be speaking later about paying engineers like salespeople. He also reminded attendees about tomorrow's engineering track sessions (MC'd by Jed from Google) and the leadership brunch for those with leadership passes only. He thanked Graphite for sponsoring the afterparty.

N.L.W. (Host of AI Daily Brief, CEO of Super Intelligent)

N.L.W. discussed the status of enterprise AI adoption and ROI findings from a study of ~2500 use cases.

  • Enterprise Adoption: Growing, with significant inflection in coding/software engineering. Agents are seeing meaningful uptake, with 42% of large enterprises having production agents (up from 11% in Q1).
  • ROI Challenges: Traditional impact metrics struggle with AI. Expectations for ROI realization are high and accelerating.
  • Study Findings (ROI Survey):
    • Overall ROI: 44.3% see modest ROI, 37.6% see high ROI. Only 5% see negative ROI. Expectations are very optimistic.
    • Impact Categories: Time savings (35%), increased output, quality improvement, and new capabilities are dominant.
    • Time Savings: Clusters around 5-10 hours/week, representing significant win-backs of work weeks per year.
    • Organizational Size: Differences observed, with mid-sized organizations (200-1000 people) focusing more on increasing output.
    • Role: Seuites/leaders are less focused on time savings and more on new capabilities and transformational impact.
    • Risk Reduction: Lowest primary benefit category (3.4%), but highest likelihood of transformational impact (25%).
    • Coding/Software Use Cases: Higher ROI and lower negative ROI than average.
    • Systematic Adoption: Organizations submitting more use cases and thinking cross-organizationally see better ROI.
  • Future Focus: Moving from generic impact conversations to systematic experiments, with a focus on automation and agentic use cases showing significantly higher ROI.

Arman Hezarki (Co-founder & Managing Partner at 10X)

Arman Hezarki from 10X presented on compensating engineers like salespeople to scale output, not overhead.

  • The Problem: Traditional hourly or salary+bonus models may not incentivize engineers to leverage AI effectively. Equity-based models have risks, especially for non-Google-like companies.
  • 10X Model: Engineers are compensated based on story points completed, directly linking pay to output. This is counterbalanced by strategists compensated on Net Revenue (customer happiness) and rigorous internal QA/client approval processes.
  • Case Studies:
    • Billboard Company: Built an AI moderation model with 96% accuracy in two weeks, paid per story point.
    • Retailer Devices: Developed five parallel AI models (heat mapping, Q detection, theft detection) for low-power devices, again paid per story point.
  • Risks & Mitigation:
    • Inflated Story Points: Mitigated by strategists scoping work and rigorous review processes.
    • Quality Drop: Mitigated by hiring the right people and robust QA.
    • Sharp Elbows: Mitigated by hiring the right people and ensuring strategists are involved in client satisfaction.
  • Core Principle: "AI makes people look like crazy mirrors." It amplifies existing attributes. Hiring the right people is paramount, making everything else easier.

Justin Rio (Deputy CTO at DX)

Justin Rio from DX discussed effective leadership in AI-enhanced organizations, focusing on reducing fear, improving metrics, and enabling employee success.

  • Current Impact Volatility: GenAI impact varies widely, with studies showing both productivity increases and decreases. A "induced flow" can make engineers feel productive even when data shows otherwise.
  • Key Factors for Positive Impact: Clear AI policies, time for learning and experimentation (not just providing materials), and effective measurement systems. Top companies see positive impacts on KPIs like change confidence, code quality, and change failure rate.
  • Reducing Fear: Emphasize psychological safety, transparency about AI's role (augmentation, not replacement), clear intents, and proactive communication.
  • Metrics: Focus on core developer experience metrics (speed and quality) like PR throughput, change failure rate, code quality, change confidence, and maintainability. Utilize telemetry, experience sampling, and effective surveys. DX's DXAI Measurement Framework normalizes metrics into utilization, impact, and cost.
  • Enabling Employee Success: Provide education, time to learn, and tie AI adoption to employee success. Distribute guides, establish feedback loops for system prompts, and manage model temperature (creativity vs. determinism).
  • Unblocking Usage: Leverage self-hosted/private models, partner with compliance early, and think creatively around barriers.
  • Integrating Across SDLC: Identify and fix actual SDLC bottlenecks (context switching, meetings, legacy code modernization) rather than just focusing on AI-assisted code generation.

Dan Shipper (Founder of Every)

Dan Shipper, founder of Every, shared dispatches from the future on building an AI-native company, emphasizing the 10x difference in organizations with 100% AI adoption.

  • Every's Model: A small team (15 people) running four software products with high MRR growth, low funding, and 99% of code written by AI agents. Each app is built by a single developer.
  • Compounding Engineering: A paradigm where each feature makes the next easier to build, achieved through a four-step loop: Plan, Delegate, Assess, Codify.
    • Plan: Detailed planning for agents.
    • Delegate: Telling the agent to perform the task.
    • Assess: Evaluating agent output using tests, agent feedback, code review, etc.
    • Codify: Turning learnings into explicit prompts for reuse across the organization.
  • Second-Order Effects:
    • Tacet Code Sharing: Easier to learn from and reuse code/processes from other developers via agents.
    • New Hire Productivity: Agents set up environments and provide guidance, making new hires productive on day one.
    • Expert Freelancer Collaboration: Lower startup cost for external experts to contribute specific functionalities.
    • Cross-Product Commits: Developers easily contribute fixes or improvements to other products within the company.
    • Stack Flexibility: AI facilitates translation between different languages and frameworks, reducing the need for strict standardization.
    • Manager Commits: AI enables managers with fractured attention to contribute code by investigating bugs and submitting PRs.
  • Key Takeaway: 100% AI adoption creates a fundamentally different operational model, enabling single developers to build complex products and fostering cross-organizational collaboration through compounding engineering.

Conclusion/Synthesis

The AI Engineer Code Summit 2025 showcased a rapidly evolving landscape where AI is not just a tool but a fundamental shift in how software is conceived, built, and maintained. Key themes revolved around the practical application of AI agents, the critical importance of context management and robust harnesses, and the need for organizations to adapt their workflows, roles, and compensation models to leverage AI effectively. Presentations highlighted the potential for significant productivity gains, but also underscored the challenges related to quality, security, measurement, and the human element of adoption. The consensus was that a proactive, iterative, and quality-focused approach, coupled with a willingness to experiment and adapt, is essential for navigating this transformation and unlocking the true potential of AI in engineering. The future of software development is increasingly agentic, collaborative, and driven by intelligent systems that augment human capabilities.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video