Coding, Agents, and the SDLC

By South Park Commons

Share:

Key Concepts

  • Minus One to Zero: A period of exploration for technologists, particularly in AI.
  • Agentic Development Environments/Frameworks: Next-generation IDEs that leverage AI agents.
  • Cloud Code: AI models (like Claude) used for code generation and development.
  • Harness: A combination of model, tools, memory, and other methods that teaches an LLM about its environment and goals, making AI products effective.
  • Context Engineering: The discipline of structuring inputs (prompts, tools, memory) for LLMs to achieve desired outcomes.
  • Organizational Memory: Pre-existing knowledge and processes within a company.
  • Agent Memory: The internal memory and learning capabilities of an AI agent.
  • Synchronicity vs. Asynchronicity: The debate on whether AI-assisted coding workflows should be real-time and interactive or operate in the background.
  • AI Quality Engineer: A new role focused on debugging and observing AI agent outputs, often non-technical subject matter experts.
  • Token-to-Value Ratio: The economic justification for the cost of training and running large language models based on the value they create.

The Evolving Software Development Life Cycle with AI Agents

This discussion, originating from South Park Commons, explores the transformative impact of AI agents and code generation on the software development life cycle. It features insights from founders building in this space: Charlie (Conductor), Ben Hilac (Raindrop), Oliver (Mesa), and Harrison (LangChain).

The Shift Beyond the IDE and New Development Paradigms

Charlie introduces Conductor, a Mac app designed to run cloud code in parallel, aiming to build "whatever comes after the IDE." He notes that coding is moving "a level higher than the IDE," driven by the power of cloud code models like Claude, which enabled the vision for a "next generation of IDEs."

Ben Hilac from Raindrop focuses on "Sentry for AI products," monitoring agent behavior and performance in real-world scenarios. He highlights the distinct nature of building coding agents compared to traditional software.

Oliver from Mesa addresses collaborative bottlenecks in large companies, such as pull requests and code reviews. He believes solving these issues can dramatically improve development speed, potentially leading to a "living codebase" where actions occur autonomously, even without human intervention.

Harrison from LangChain aims to simplify building intelligent agents, believing LLMs will make applications more "agentic" and "smart." LangChain focuses on developer tools to create reliable agents.

The Divergence of AI-Written vs. Human-Written Code

A significant "unspoken tailwind" identified by Oliver is the profound difference between AI-generated code and human-written code.

  • AI code failure modes can be catastrophic, with entire files being "completely wrong" or unnecessary.
  • Human code failure modes are generally smaller.
  • This leads to an attribution problem in tools like GitHub, where it's unclear if code was written by AI or a human, complicating code reviews.
  • Examples: AI might remove intentional code (requiring do not change this line comments) or critical human-written comments, while AI-generated comments are often "garbage."
  • Proposed Solution: Similar to auto-generated proto-buffs, there's a need for clear labeling like "autogenerated, do not modify."
  • Granularity of Control: Teams are becoming more restrictive with AI in sensitive parts of the codebase (e.g., core infrastructure, migration files) compared to new front-end files. There's a desire for CI-level locking of files or directories.

The Product-Infrastructure Blur and Model Volatility

Harrison discusses Open Suite, an autonomous coding agent built on LangGraph, primarily to test LangGraph's infrastructure. He views infra and product as separate, with core infra not AI-written. A key challenge is making cloud code proficient at writing and updating LangGraph code.

Oliver shares a friend's "existential fear" for dev tool builders: the unpredictable nature of model providers. An example is Anthropic's Claude 3.7 upgrade breaking a system built on 3.5 because the new model was "RL'd to use these very specific tools," necessitating a complete internal product update. This highlights the difficulty of adding novel capabilities not "baked in" by model providers.

Ben emphasizes the rapid deprecation of models ("six months later it's gone") and the lack of clear upgrade paths, likening it to a "different brain." He finds this absurd from a traditional software engineering perspective (e.g., imagining Stripe doing this). However, he argues that the first-mover advantage (e.g., Cursor) necessitates "surfing the wave" and quick adaptation despite these risks.

The Agent as Product and the Human-AI Review Loop

Oliver questions if agents will remain "the product" as foundation models converge capabilities. He notes that he could build a code review agent in a weekend that outperforms market leaders by delegating most work to a foundation model.

Charlie believes there will be a "different level" of product, with UI/UX around agents becoming crucial (as Conductor aims to do). He foresees a future where managing multiple agents requires a distinct UX.

Ben describes an evolving review cycle:

  1. Humans write, humans review.
  2. Humans + AI write, humans + AI review.
  3. AI writes, AI reviews (e.g., Reptile).
  4. Current/Future: AI writes, humans review. AI is good at generating code but not at judging it. Conductor's future vision is to become a "reviewing work app" or "superhuman inbox," where humans provide final judgment on AI-generated work, with agents operating "under the hood."

Harrison introduces the open-source "agent inbox" concept, designed to manage numerous agents, especially those triggered by background processes. This system, resembling an email inbox or customer support board, allows humans to intervene when agents are "stuck" or have questions, mirroring human-to-human management.

The "Harness" and the Model-Harness Split

Ben's notable quote defines the "harness": "Models have actually become really malleable. So LLM harnesses that teach a model about its environment and its goals have outsized impact. It's this harness, a combination of model, tools, memory, and other methods that makes AI products actually good."

Charlie envisions a future where humans remain in the loop, reviewing at increasingly higher levels. He recounts an experiment where he proxied GPT-5 into cloud code; despite the hype, GPT-5 performed "atrociously" within cloud code's harness. He asked Baris, the cloud code creator, about the model-harness split, who estimated it as 70% model, 30% harness. This implies constant work to adapt the model to the harness, with much work potentially "wasted" with new model releases.

Oliver observes that models are "RL'd for specific tool calls" with precise names. Using different names or tools with similar functions leads to "terrible" performance, indicating that models are highly tuned to their specific environments.

Harrison adds that models are not truly portable or interchangeable; prompts and context engineering methods often don't transfer between models.

Context Engineering as a Discipline and the Human-AI Interface

Harrison argues that the best context engineers deeply understand the problem and "put themselves in the agent's position," as models are not "mind-readers." This understanding informs prompts, tools, and workflow design.

Charlie expresses concern that treating AIs as humans might be a "local minima" because AIs have different strengths and weaknesses. He questions if a human-centric approach to teaching AIs is optimal.

Ben points to DSPy as evidence against a purely human-centric approach, where the system generates optimal, often "gibberish" prompts that outperform human-crafted ones. He also notes that while there was a trend towards simple, plain language prompts, the importance of harnesses suggests that deep context engineering remains vital. He counters Paul Graham's AGI definition (telling it what you want in a sentence) by stating that even telling a person what you want is hard, and highly instructible models like GPT-5 can take typos verbatim, lacking human-like judgment.

Harrison highlights a trend where models/harnesses are doing more of their own context engineering, using file systems to manage context rather than relying solely on prompt stuffing, which increases flexibility.

Organizational Memory vs. Agent Memory

The discussion delves into the convergence or divergence of organizational memory and agent memory.

  • Ben states that much of Raindrop's harness work is a substitute for long-term memory, fearing that future model providers offering online learning or long-term memory could render their code irrelevant.
  • Ben also emphasizes that memory requires good judgment (when to save, when to revoke), a capability models currently lack. He observes that models often misinterpret preferences or save temporary preferences as permanent.
  • Harrison views memory as "reflection on top of" raw data, with current efforts focused on crafting prompts to guide what to remember.
  • Behavioral changes desired by users (e.g., "be more strict") are difficult to capture in simple memory or system prompts.
  • Human-in-the-loop workflows are crucial for empowering memory through feedback.
  • Types of Memory: Learning through interactions (human feedback) and integrating existing organizational context (via RAG or distillation).
  • Raindrop's approach to organizational memory involves defining "good" or "bad" semantic signals, refining them through user feedback, and asking "why" changes were made to extrapolate rules. They found LLMs are good at formulating questions but bad at deciding when to ask them (Deep Research example).
  • Charlie describes using cloud.md as their primary, albeit "lame," method for model memory, noting that Claude often needs explicit instructions (e.g., against useRef or fallbacks). He prefers to be the "arbiter" of cloud.md rather than letting Claude update it autonomously.
  • Oliver's approach stores rules/memory as an array in a database, allowing users to define when these rules are pulled into context (e.g., specific files, PR authors). Rules can be injected directly into the system prompt or exposed as tools for the LM to decide when to use them, offering more granular control.

Synchronicity vs. Asynchronicity in Coding Tools

The panel discusses the dichotomy between synchronous (real-time, interactive) and asynchronous (background) AI-assisted coding.

  • Charlie notes a shift in the "meta" from fully asynchronous (e.g., Devin) back to more "coding in the loop" (synchronous). He likens it to coding with humans (intern vs. pair programming), suggesting no single answer. Conductor is moving towards more asynchronous work, with a "founder mode" overview/inbox, while retaining synchronous chat UI options.
  • Ben references Sam Altman's Starcraft analogy (CEO + direct reports + "tons of agents" as doers). He finds himself reverting to autocomplete for hard work, using chat for questions/review, and letting AI write front-end code. For low-stakes, non-critical tasks (e.g., adding a button), he "fires it off to Devin" and checks PRs later.
  • Oliver is mostly back to writing backend code himself, using AI for front-end. He sees AI not as a "junior engineer" but as a tool for asking questions to deeply understand a system before implementing it.
  • Need for Higher-Level Tools: Oliver highlights the need for agents to operate on "larger building blocks" (e.g., "move function" instead of line-by-line rewriting) rather than just micro-level code generation. This is an "underrated tailwind."

New Roles and Screening for AI-Native Talent

The discussion concludes with how programming jobs are changing and how to screen for talent in this new AI-native world.

  • Harrison foresees a new role: the "AI quality engineer," often non-technical subject matter experts, who observe and debug agent outputs (similar to code review for specific domains).
  • Dario (Anthropic) believes coding is the biggest use case for LLM APIs due to model capabilities and the developer-centric user base in Silicon Valley, though a "long tail" of other use cases (finance, medicine, legal) exists.
  • Ben suggests that coding's token-to-value ratio makes it economically viable for high token costs, especially in lower-risk domains compared to healthcare or legal.
  • Harrison adds that agents are best suited for producing "first drafts" after long periods of work (e.g., coding, deep research).
  • Hiring for AI-Native Founders:
    • Charlie: Looks for people "anxious that you are not burning tokens 24/7," excited about "alien intelligence." Excitement is paramount; other skills can be taught.
    • Harrison: For infrastructure roles, seeks thoughtful individuals with scaling experience. Resumes and personal projects are less indicative; GitHub stars (indicating actual usage) are a high signal. He prioritizes quick phone calls, asking "What thing have you made that you're the most proud of?"
    • Oliver: Emphasizes bringing people on-site quickly to assess culture and excitement. He believes CS fundamentals are becoming more important, not less, as engineers need to understand systems at a high level and explain their code deeply. He warns against the trap of older developers relying on outdated knowledge of model limitations, stressing the need to constantly retest model capabilities.
    • Ben: Prioritizes "understanding" in candidates, as AI makes it easy to produce work without deep comprehension. GitHub stars are a good indicator of understanding a problem well enough to build and explain a solution.

Conclusion

The conversation underscores that AI is fundamentally reshaping software development, moving beyond traditional IDEs to agentic environments. While AI offers immense productivity gains, it introduces new challenges related to code attribution, model volatility, the need for sophisticated "harnesses," and the critical role of human judgment in reviewing and guiding AI agents. The future of software engineering will likely involve a blend of synchronous and asynchronous AI interactions, with new roles emerging for "AI quality engineers" and a renewed emphasis on fundamental understanding and adaptability for human developers. The ability to effectively manage and integrate AI's capabilities while mitigating its unique failure modes will be key to unlocking its full potential.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video