Stanford CS547 HCI Seminar | Autumn 2025 | HCI Systems in the Age of AI Code Generation
By Unknown Author
Key Concepts
- HCI Systems Research: Human-Computer Interaction research focusing on building and evaluating interactive systems.
- AI Code Generation: Using artificial intelligence, particularly Large Language Models (LLMs), to automatically produce code.
- Augmented Development Environments: Rethinking tools and workflows to enhance programmer capabilities using AI.
- Alterable Applications: Interactive applications that can dynamically change their functionality at runtime through embedded code generation.
- Program Synthesis: A traditional approach to code generation involving searching a space of valid programs, often domain-specific.
- Token Prediction: The mechanism by which LLMs generate text, including code, by predicting the next most probable token.
- Problem-Space Exploration: The initial phase of design focused on understanding and reframing the problem, rather than immediately seeking solutions.
- Design Space Exploration: The process of generating and evaluating multiple potential solutions to a problem.
- Trigger Action Programming (TAP): End-user automation environments (e.g., If This, Then That) linking simple triggers to actions.
- Procedural Content Generation (PCG): Algorithmic generation of game content (textures, models, worlds).
- Procedural Behavior Generation: Algorithmic generation of interactive behaviors within a game or application at runtime.
- Disposable Code: Code written for epistemic purposes, to explore or figure something out, not intended for long-term, critical use.
- Durable Code: Robust, well-engineered code intended for long-term use in critical systems.
- Document-Grounded Alignment: Using a stable, shared document (e.g., requirements doc) to maintain context and agreement in a design process.
- RAG (Retrieval-Augmented Generation): A technique where LLMs retrieve information from an external knowledge base to inform their generation.
Introduction: HCI Systems Research in the Age of AI Code Generation
The presentation focuses on how to conduct HCI systems research in the era of AI code generation, drawing on two years of the speaker's group investigating systems that integrate code generation. The core question is: "How will we interact with AI code generation?"
The generative AI wave significantly impacted the HCI research community around UIST 2023. Since then, there has been a surge in papers presenting applications made with pre-trained models. Meta-analyses, such as a 2024 paper by Park, Forlizzi, and Zimmerman, characterize model capabilities and application areas. A key finding is that programming support is a major application, with systems transforming text to programs, programs to text, answering code questions, and ranking code.
Industry has rapidly adopted this, with tools like Copilot, Claude Code, CodeWhisperer, Ghostwriter, Kite, Tabnine, Cursor, and Windsurf. Vendors claim impressive productivity gains (e.g., 10x, 1,000 stories, 280,000 hours saved). However, the HCI and software engineering communities offer a more nuanced view, suggesting current code-generation tools are best suited for self-sufficient, expert software developers, drawing a parallel to the limited success of MOOCs for non-expert learners.
Initial LLM UIs were "thin wrappers" (e.g., ChatGPT generating HTML). More commonly now, LLMs are integrated into development environments, accessing codebases and directly modifying editor states. An alternative vision, proposed at UIST 2023, is to use code generation as a building block inside interactive applications, where the user's goal is not to write code but to accomplish another task. An example is ReactGenie (Jackie Yang's work from Stanford), where a user speaks a desired action, and ReactGenie synthesizes and executes UI automation code in a domain-specific language (DSL) under the hood. This raises questions about correctness when no programmer reviews the generated code.
These two directions—augmented development environments and alterable applications—form the main themes of the talk. The speaker notes that these ideas are not entirely new, echoing previous waves of code generation in HCI that leveraged program synthesis. While program synthesis involved searching a constrained space of valid programs (requiring narrow domains and small programs), LLMs use token prediction. Intellectual parallels exist, such as empowering programmers (e.g., Armando Solar-Lezama's Sketch from 2013 for writing logic with "holes") and transducers under the hood (e.g., AutomataTutor for generating DSL programs to provide hints in theory of computation classes).
Theme 1: Augmented Development Environments
This theme explores how AI code generation can enhance the capabilities of programmers, moving beyond simple code snippets.
Project 1: PAIL – Beyond Code Generation (CHI 2025, led by JD)
PAIL targets experienced programmers, aiming to assist them beyond just generating function bodies. The dominant vision of AI creation—expressing a desire and getting a solution—is challenged. PAIL questions whether users know what to ask for and if there are better ways to achieve goals. Current code generation assumes a specific problem definition and provides one solution, often neglecting problem framing and alternative solutions.
PAIL shifts the perspective to programming as a design activity. Designers focus on task framing and explore solution spaces, understanding trade-offs (e.g., British Design Council's Double Diamond model). PAIL aims to backtrack from a specific problem to problem-space exploration, helping users investigate different problem statements and alternative implementations, and then converge on decisions. This focuses on program design rather than just code production.
Example: A parent programming for their child faces higher-level design questions (e.g., "How do I maintain attention?", "Is this pedagogically relevant?") that PAIL aims to support.
PAIL Tool Functionality:
- Users start with an application request.
- Instead of immediate code generation, PAIL encourages abstracting up to establish goals and requirements.
- It generates a shared document that captures design goals and requirements. This document is interactive, allowing users to agree on requirements, providing a stable representation of design decisions.
- The system proactively generates alternatives and rationales for design choices (e.g., target age demographic for an interactive application).
- Under the hood, LLMs simulate execution and hypothetically assess the implications of design choices, providing "information scent" about potential user outcomes.
- Users can iteratively go back and forth between design decisions and code generation/execution.
Goals of PAIL:
- Support the program design process.
- Broaden design space exploration.
- Overcome communication limits by using a stable, document-grounded representation.
- Help users consider high-level design decisions.
- Track grounding explicitly to mitigate mismatched metaphors and expectations.
Evaluation (11 participants):
- Successes: Accelerated design and iteration cycles, broadened design space exploration, generated alternatives were useful, document-grounded alignment was effective for persistence of requirements.
- Challenges: Information quantity quickly became overwhelming, leading to diffused attention. Users struggled to maintain awareness of multiple artifacts and regain mental models after significant changes. The speaker notes that we've crossed a "Rubicon" where we can generate things faster than we can think about them, creating an opposite problem to traditional design where implementation was too costly early on.
Takeaways: PAIL compresses the design process by supporting iteration across multiple abstraction levels and broader exploration. The key challenge is managing and directing information flow and user retention.
Project 2: Generative Trigger Action Programming with Ply (UIST, led by Tim and Hila)
This project shifts the target audience from experienced programmers to end-user programming. It asks: "What would working with code generation look like if you never showed the code?"
Context: Trigger Action Programming (TAP), exemplified by "If This, Then That," allows users to link simple triggers to actions (e.g., "Turn the lights on when I press the button"). However, complex conditions (e.g., "double-press," "fade slowly to green when triple-press") require developers to rewrite components. To make TAP more powerful, intermediate code is needed to consume simple events, transform them, and interact with components. LLMs are ideal for generating this logic.
Challenge: How can end users have confidence in generated code if they never see it?
Ply's Approach:
- Decomposition into "Layers": Users are encouraged to create higher-order sensors and actuators by defining an input, a one-phrase natural language behavior specification, and an observable output.
- Automated Visualization: Ply uses code generation to automatically create appropriate visualizations of what each "box of behavior" does.
- Automated Control Interfaces: Ply generates control interfaces that allow users to parameterize the behavior. The names of these parameters serve as a signal for users to assess if the code is doing the right thing.
Demo Example:
- Creating a counter using left and right buttons.
- Visualizing the sensor's current state.
- Changing generated parameters (e.g., counter step size to three).
- Using chat to modify behavior (e.g., "reset to zero when the middle button is pressed").
- Creating a sensor that cycles through languages (English, Spanish, Chinese) using top and bottom buttons, with Ply generating the configuration interface for complex parameters.
Core Ideas in Ply:
- Layers: Enable users to create higher-order sensors and actuators (e.g., pointing a camera at a rice cooker to interpret "cooking," "warm," or "off" states).
- UI for Events: Create user interfaces to show events dispatched by sensors.
- Customization: Let users customize sensor and actuator behavior.
Discussion on "End-User Programmer": The term refers to the order of magnitude more people who engage in computational thinking for work or hobbies but are not trained software engineers. A key challenge, acknowledged by the speaker, is that program decomposition (breaking down problems) is a very hard task, and Ply currently provides the substrate for expressing such skills rather than proactively helping users develop them.
Summary of Theme 1: Code generation can support higher-level tasks in software design and development, enabling program design, understanding trade-offs, and providing control over generated code through interfaces and visualizations.
Theme 2: Alterable Applications
This theme explores wiring code generation directly into applications, allowing them to change functionality at runtime.
Project 3: GROMIT – Procedural Behavior Generation (UIST 2024, led by Nicholas)
GROMIT investigates extending procedural content generation (PCG) in games (e.g., Minecraft's generated worlds, textures, models) to procedural behavior generation at runtime. The concept is an "infinite brewing stand" where new, developer-unforeseen behaviors can be created on the fly.
GROMIT's Approach: An LLM-powered runtime behavior generation system for Unity.
- A user in the game triggers a generation action (e.g., using one object on another where no interaction is defined).
- GROMIT takes the scene hierarchy and associated metadata as input.
- It generates a script that performs the desired action.
- The script is compiled and executed in the game at runtime. The name "GROMIT" references the Wallace and Gromit scene where Gromit lays tracks as the train moves, symbolizing on-the-fly creation.
Demo Example (Escape Room):
- A player is stuck in an escape room, needing a key.
- Some interactions (e.g., key opening door) are pre-defined.
- The player tries to use a torch on a bookshelf (an undefined interaction).
- GROMIT is invoked, creating a prompt based on the interaction ("bookshelf is an old wooden bookshelf filled with books," "torch is a flaming torch").
- The LLM generates the behavior: "the torch should burn down the bookshelf" (implemented by deleting the bookshelf).
- This new behavior is cached, allowing the player to burn down other bookshelves, revealing a hidden key and enabling escape.
Evaluation (13 game developers interviewed):
- Technical Feasibility: GROMIT works frequently without crashing.
- Impact on Game Development:
- Concerns: Developers expressed significant concern about "overreach," questioning their identity and vision if game experiences are dynamically generated. They worried about defining enough constraints and guardrails to maintain a quality experience.
- Excitement: Developers were enthusiastic about using GROMIT during development (e.g., in alpha testing). It could be a tool to collect data on player actions, cueing developers on where to allocate resources. This allows for an "open alpha sandbox" where player-generated behaviors (even hallucinations or crashes) are acceptable, as the goal is to inform the creation of a final, polished experience.
Disposable vs. Durable Code: This concept, shared by Armando Fox (UC Berkeley), suggests that AI code generation might be excellent for producing "epistemic" or disposable code—code written to figure something out. It's less suited for durable code—the robust, critical systems that run for a long time—unless paired with strong engineering practices and expert review. GROMIT's promise lies in generating disposable code to prototype and test novel behaviors, which developers can then make durable.
Takeaways: Runtime behavior generation is possible with current LLMs. While unchecked generation can lead to undesirable outcomes, its greatest promise lies in prototyping and testing novel behaviors that are ultimately refined and made durable by developers, often with expert supervision.
Project 4: CARE – Context-Aware Rehabilitation Exercises (Briefly Discussed)
CARE, a collaborative project with Stanford neurology, aims to use code generation to automatically create tailored prescriptions for physical rehabilitation, specifically for stroke patients.
Problem: Physical therapists currently provide written worksheets or use pre-authored software with limited customization (e.g., dropdown boxes for repetitions). This lacks personalization for individual patient needs.
CARE's Approach:
- A physical therapist writes down flexible instructions based on a patient's movement limitations.
- Code generation is used to automatically generate tailored exercise software for that specific patient.
- The system generates Python code, but conceptually, it's generating a program in a domain-specific language (DSL) focused on rehabilitation.
Safety and Constraints: Unlike games where crashes are less critical, rehabilitation involves patient well-being. To mitigate risks, CARE heavily constrains the code generation. An API was designed in collaboration with physical therapists, defining specific "building blocks" for movements. The LLM is instructed to only generate combinations of these pre-approved blocks, making the process much safer than generating arbitrary code.
Reflections on Systems HCI in the Age of Generative AI
The speaker raises a meta-question: "Does the Systems HCI playbook still work in the age of generative AI?"
The Traditional Systems HCI Playbook:
- Motivate a problem, make the audience care.
- Show that naive solutions are insufficient.
- Introduce a novel technical insight (often from a PhD student).
- Build a system around that insight.
- Evaluate the system with users to validate performance on "people metrics."
- Alternatively: Understand unmet user needs, build a novel system, evaluate with users.
Argument for Continued Relevance:
- Systems HCI's focus on the intersection of technical contributions and people is more needed than ever.
- It prevents disciplines (HCI, ML) from retreating into isolated corners (HCI black-boxing technology, ML focusing only on models/data without user context).
- The "systems thinking lens" is crucial for understanding larger interactions that cross boundaries of concern.
Argument for Non-Applicability/Challenges:
- Technical Insight: The traditional model of a single PhD student developing a novel technical insight is challenged. The resources (e.g., thousands of GPUs) required for large pre-trained models are not commonly available in research departments.
- Unique Leverage: The same LLM modules are often available to anyone, anywhere, at zero cost. This questions what unique leverage researchers have to ask and answer unique questions.
- "Rearranging Deck Chairs": Merely being "downstream consumers of pre-baked models feels a bit like rearranging the deck chairs on a boat that someone else built instead of building your own boat."
- Counter-argument: HCI research has always relied on layers built by others (hardware, OS, I/O devices). However, small teams historically could change these layers (e.g., custom hardware at UIST, alternative interactions at CHI).
- Current State: There's a "flood of LLM applications," many of which may not be durable.
The speaker concludes by raising these questions for graduate students, emphasizing that while there are no easy answers, these are critical considerations for future research projects.
Synthesis and Conclusion
The presentation thoroughly explores the evolving landscape of HCI systems research in the age of AI code generation, highlighting two primary avenues: augmenting developer environments and creating alterable applications. Projects like PAIL demonstrate how AI can elevate programming to a design activity, fostering problem-space and design-space exploration through stable, document-grounded representations, though managing information overload remains a challenge. Ply showcases how code generation can empower end-users in trigger-action programming by generating invisible logic, visualizations, and control interfaces, emphasizing the need for robust decomposition strategies. GROMIT illustrates the potential of runtime procedural behavior generation in games, particularly for prototyping "disposable code" during development, while CARE exemplifies the critical need for constrained, domain-specific code generation in high-stakes applications like rehabilitation.
Ultimately, the speaker prompts a critical reflection on the traditional Systems HCI playbook. While the interdisciplinary focus of Systems HCI is more vital than ever for bridging technical and human concerns, the nature of technical contributions and unique research leverage is shifting with the widespread availability of powerful, pre-trained AI models. The distinction between "disposable" and "durable" code offers a valuable framework for understanding where AI code generation currently provides the most impactful and safe contributions, often in scenarios with expert supervision or for exploratory purposes. The future of HCI systems research will likely involve navigating these complexities, focusing on how to effectively integrate, control, and evaluate AI-generated artifacts within human-centered workflows and applications.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

How the hometown humiliation of Putin marks a turning point for Ukraine | DW News
DW News

Shocking video shows moment paramedics are hit by Israel in 'double-tap' strike
Sky News

Every Kind of Volcano | SciShow Kids
SciShow Kids

Putin Xi, To Catch a Castro, Red Carpet Rebellion • FRANCE 24 English
FRANCE 24 English

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Pokemon goes prehistoric at Chicago's Field Museum
Reuters

Trump's supporters furious over Trump smartphone scam.
ABC News In-depth