Useful General Intelligence — Danielle Perszyk, Amazon AGI

AI EngineerAbout 4 min readAug 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts:

  • Controlled Hallucination: The brain's process of making predictions, taking in sensory information, and reconciling errors, essential for perception and understanding.
  • Technosocial Co-evolution: The reciprocal relationship between technology and human cognition, where technology shapes our brains and behavior.
  • Representational Alignment: The ability to reproduce the contents of our minds to better cooperate, a key aspect of human intelligence.
  • Cognitive Technologies: Tools that stabilize our thinking, reorganize our brains, and control our hallucinations, such as language, writing, and computers.
  • Computer Use Agents: AI agents that can interact with computer UIs like humans, seeing pixels and performing actions.
  • Nova Act: A research preview of an agent that combines a specialized version of Amazon Nova with an SDK to allow developers to build and deploy agents.
  • Models of Mind: The ability to infer the existence of other minds, a crucial component of general intelligence.

1. The Reliability of Our Minds and the Role of Hallucinations

  • The speaker, Danielle, a cognitive scientist at Amazon AGI SF lab, emphasizes that our brains are "prediction machines" that operate through "controlled hallucination."
  • Our brains don't have direct access to reality; they make predictions, take in sensory information, and reconcile errors.
  • Understanding language involves activating all meanings of a word, including the concept of hallucination.
  • Hallucinations are necessary for AI to go beyond the data and be flexible, like human intelligence. The key is to control them.

2. The Vision for Agents: Augmenting Human Intelligence

  • The talk challenges the standard vision of AGI, which focuses on making AI smarter and giving it more agency.
  • The speaker advocates for building AI that makes humans smarter and gives them more agency, complementing rather than replacing human intelligence.
  • Drawing a parallel to Douglas Engelbart's vision, the goal is to augment human intelligence through technosocial co-evolution.
  • Automation can lead to augmentation by freeing up attention, but it can also reduce agency if not carefully controlled.

3. Nova Act: A Research Preview of a Computer Use Agent

  • Nova Act is presented as a step towards building AI that unhobbles humans by meeting models and builders where they are.
  • It aims to make it frictionless for users to get started with agents.
  • Nova Act uses the browser as a tool, as most websites lack APIs.
  • It combines a specialized version of Amazon's foundation model Nova, trained for high reliability on UI tasks, with an SDK.
  • Developers can use "act calls" to translate natural language into actions on the screen.
  • A demo is shown where Nova Act is used to find an apartment, extract data, and calculate biking distance to a Caltrain station using Python integrations.
  • The underlying model is continuously improved and shipped every few weeks.

4. The Challenges of Computer Use and the Importance of Alignment

  • Even basic computer interactions, like interpreting icons, are deceptively challenging for AI.
  • Agents need to explore and learn through reinforcement learning (RL) to discover new ways of using computers.
  • It's critical that agents' perception of the digital world is aligned with our own.
  • Current agents are often LLM wrappers that lack an environment to ground their interactions.
  • Computer use agents, like Nova Act, can see pixels and interact with UIs, providing a form of embodiment.
  • The approach focuses on making the smallest units of interaction reliable and giving granular control.

5. The Evolution of Intelligence: From Social Cognition to Language

  • The talk highlights the co-evolution of humans and technology, tracing it back to the development of social cognition and language.
  • Representational alignment, the ability to reproduce the contents of our minds, was a key evolutionary adaptation.
  • Language co-evolved with our models of minds, integrating communication and representation.
  • Models of mind became the original placeholder concept, enabling generalization.
  • Language triggered a series of flywheels, leading to the development of cognitive technologies.

6. The Future of Agents: Models of Mind and a Common Language

  • To make agents truly reliable, they will eventually need models of our minds.
  • This requires a common language for humans and computers, including a model of our shared environment and intuitive interfaces.
  • Human-agent interaction data is needed to advance the models.
  • Useful products will motivate people to use the agents, leading to collective intelligence.
  • Nova Act is presented as the primitives for a cognitive technology that aligns agents and humans' representations.

7. Conclusion

  • The talk concludes by emphasizing the importance of building agents that augment human intelligence and give us more agency.
  • This requires focusing on representational alignment, developing models of mind, and creating a common language for humans and computers.
  • By building useful products and collecting human-agent interaction data, we can collectively build useful general intelligence.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.