New Nanobananas, OpenAI’s n8n, ChatGPT apps, Sora 2 PRO, Google’s free app builder - AI NEWS
By AI Search
Key Concepts
- Apps in ChatGPT: Integration of third-party applications directly within the ChatGPT interface.
- Apps SDK: A developer tool for building custom applications that can be embedded and called within ChatGPT.
- Agent Kit: A suite of tools from OpenAI for creating AI agent-centric workflows and integrating them into products.
- Agent Builder: A drag-and-drop interface within Agent Kit for constructing AI agent workflows.
- Chatkit: A tool within Agent Kit for deploying and displaying AI agents in front-end applications.
- GPT-5 Pro: OpenAI's advanced reasoning model, now with API access, noted for high performance in challenging benchmarks.
- Sora 2 API: Access to OpenAI's revolutionary video generation model, capable of creating scenes with dialogue and sound effects.
- Luminina Demu: A new state-of-the-art image generator and editor offering text-to-image and text-based image manipulation, including ControlNet-like features.
- Apriel 1.5 15B Thinker: A small (15 billion parameters) yet highly performant open-source language model, excelling in reasoning tasks and fitting on consumer GPUs.
- Paper 2 Video: An AI tool that converts scientific papers into narrated presentation videos, complete with slides, voice cloning, and avatar narration.
- Mimix: An AI capable of mixing and matching different fictional characters and styles within the same video.
- ChronoEdit: Nvidia's AI image editor that applies edits by reasoning through how changes would occur over time, ensuring physical consistency.
- Codemender: Google DeepMind's AI agent designed to autonomously detect and fix software vulnerabilities, including zero-day exploits.
- Humanoid Robots: Advancements in physical AI, including waterproof/dustproof models (DR2), home-environment robots (Figure 03), and robots with expressive gaits (Orca).
- Ling 1 Trillion: Alibaba's Ant Group's large language model (1 trillion total parameters, 15 billion active) demonstrating state-of-the-art performance in reasoning and coding.
- Google Opal: A free, drag-and-drop workflow builder for creating AI-powered applications, offering automated workflow generation from prompts.
- Grock Imagine v0.9: A free video generator with native audio, offering fast video creation from images or text prompts.
- Tiny Recursion Model (TRM): Samsung's ultra-small (7 million parameters) model that achieves state-of-the-art reasoning performance on ARC AGI benchmarks by using a recursive, iterative problem-solving approach.
- ARC AGI: A benchmark testing an AI model's ability to learn and understand new patterns from unseen data.
OpenAI DevDay Updates
OpenAI's recent DevDay introduced several significant updates:
- Apps in ChatGPT: Users can now directly call and interact with third-party applications like Canva, Spotify, Zillow, and Expedia within the ChatGPT interface. For example, ChatGPT can create a custom Spotify playlist or find homes for sale via Zillow. A notable feature allows uploading a document and using Canva to automatically generate a presentation deck. This functionality is available for free users outside the EU, with EU access expected soon.
- Apps SDK: A developer tool built on MCP (Multi-Context Protocol) enables developers to create custom apps with unique interfaces and logic that AI models can interact with. This SDK allows embedding and calling these apps within ChatGPT, potentially reaching over 800 million users, presenting a "golden opportunity" for early adopters.
- Agent Kit: This set of tools facilitates the creation of workflows with AI agents and their integration into existing products.
- Agent Builder: A drag-and-drop interface for constructing agent-centric workflows, similar to Zapier or N8N. The speaker clarifies that Agent Builder is focused on the OpenAI ecosystem, calling only OpenAI agents, unlike N8N which is open-source, more flexible, and integrates with numerous third-party services and different AI models.
- Chatkit: Used to embed the built agents into web or app frontends after their logic has been defined in Agent Builder.
- GPT-5 Pro API Access: OpenAI has released API access to GPT-5 Pro, their most advanced reasoning model. Benchmarks show it saturates competitive math tasks (Python), performs well in Frontier Math (previously under 30%), and has improved scientific knowledge (GPQA Diamond). However, its price is significantly higher at $120 per output token, compared to $10 for GPT-5, for what are sometimes marginal improvements.
- Sora 2 API Access: The revolutionary video generator, Sora 2, known for its world understanding, ability to generate scenes with scripted dialogue and sound effects from simple prompts, is now accessible via API. Sora 2 Pro offers 1080p resolution and up to 15 seconds of video, though all Sora 2 Pro outputs currently include a watermark.
- GPT Realtime Mini: A new, lower-latency, and cheaper version of GPT Realtime Voice, claiming comparable quality to its predecessor.
- GPT Image 1 Mini: A smaller, more cost-effective, and faster version of the GPT image model.
New State-of-the-Art Image Editor: Luminina Demu
Luminina Demu is a new image generator and editor capable of both text-to-image generation and text-based image editing, akin to Nano Banana.
- Generation Capabilities: It can perfectly follow complex prompts for generating images, including detailed scenes with multiple elements, text, specific lighting conditions (e.g., slatted sunlight), and anime styles.
- Editing Capabilities:
- Text-based edits: Examples include adding a bike to a scene, transforming a photo into a book illustration, or changing a plain wall to a brick wall.
- Style Transfer: Applies the style from a reference photo to an input image.
- ControlNet-like features: Can generate images from input edge maps, depth maps, or pose skeletons.
- Inpainting/Outpainting: Fills in blanked-out sections of an image and expands image edges to create wider panoramic views.
- Performance: Benchmarks indicate Luminina Demu achieves the highest average scores compared to other image generators and editors.
- Availability: The models are released on Hugging Face, with a GitHub repository providing instructions for local installation and execution. The total model size is around 17 GB, potentially usable with 12 GB VRAM with CPU offloading.
Tiny but Super Performant Model: Apriel 1.5 15B Thinker
Samsung has released Apriel 1.5 15B Thinker, a remarkably small yet powerful model.
- Size and Performance: With only 15 billion parameters, it is significantly smaller (1/10th or less) than leading models like DeepSeek Car 1, Kimik 2, or Claude 4.5 Sonnet. Despite its size, it achieves an intelligence score of 52 on the artificial analysis leaderboard, matching DeepSeek Car 1 and outperforming Kimik 2 and Claude 4.5 Sonnet in challenging reasoning tasks. Its small size allows it to run on most consumer-grade GPUs.
- Efficiency: It occupies the "upper left quadrant" on a performance vs. model size chart, indicating high performance with high efficiency.
- Training Methodology:
- Mid-training/Continual Pre-training: Trained on a vast, diverse dataset including math, coding, science problems, and logical puzzles to build strong reasoning skills.
- Supervised Fine-tuning: Further refined using over 2 million high-quality text samples covering math, science, coding, and conversational use cases.
- Key Distinction: Notably, it achieves state-of-the-art performance without using reinforcement learning, a common technique for other thinking models like DeepSeek.
- Availability: It is open-source, with instructions for local download and execution available on Hugging Face.
Paper to Video AI: Paper 2 Video
This AI tool automates the creation of narrated presentation videos from scientific articles.
- Functionality: It takes a scientific paper, a user's face photo, and a few seconds of their voice as input. It then generates a full presentation video, including slides, a transcript, subtitles, and an avatar of the user narrating the content in a corner.
- Examples: Demos include a presentation on "Paper 2 Video" itself, Jeffrey Hinton explaining the "Forward Forward Algorithm," and Jensen Huang explaining a "Sauna" paper.
- Step-by-step Process:
- Slide Builder: Processes the paper to create presentation slides.
- Subtitle Builder: Generates the transcript and subtitles.
- Audio Generator: Clones the person's voice and makes the avatar speak the transcript.
- Cursor Builder: Synchronizes subtitles with the narration.
- Full Output: Compiles all components into the final video.
- Availability: All components are released, with a GitHub repository providing installation and running instructions. A significant minimum requirement is 48 GB of VRAM.
Video Character Mixing AI: Mimix
Mimix is an AI that allows mixing and matching different fictional characters or styles within the same video, similar to Sora 2.
- Capabilities: It can combine characters like Mr. Bean into a Tom and Jerry episode or have Mr. Bean interact with a realistic Ice Bear. It can even mix 2D and realistic characters in the same video.
- Performance: It reportedly outperforms other video models in preserving the appearance of characters.
- Availability: A technical paper and GitHub repository are available, but the code has not yet been released.
Nvidia's ChronoEdit: Temporal Reasoning for Image Editing
Nvidia's ChronoEdit is an AI that edits images using text prompts, similar to Nano Banana, but with a unique temporal reasoning approach.
- Functionality: It can perform various edits, such as changing a character's view, transforming a sketch into an anime battle scene, removing glasses, or turning a photo into a professional portrait while maintaining face consistency. It also exhibits ControlNet-like abilities, such as extracting edge maps.
- Temporal Reasoning: The core innovation is its ability to apply edits by reasoning how they would play out over time. For example, if asked to move a camera forward, it simulates the scene's evolution to generate the final output. This process ensures physical consistency in the edited images.
- Applications: This capability makes ChronoEdit potentially useful for generating synthetic data to train humanoid robots (e.g., showing a robot picking up objects) or autonomous driving systems (e.g., moving cars, making turns).
- Process:
- Reference Image Input: Takes the original image.
- Temporal Reasoning Stage: Imagines a short video of how the desired edit would happen.
- Editing Frame Generation Stage: Uses tokens from the previous step to refine the edited image, ensuring realism and adherence to physical laws.
- Availability: Nvidia plans to release the code for ChronoEdit.
Google DeepMind's Codemender
Codemender is an AI agent developed by Google DeepMind to autonomously detect and fix vulnerable code.
- Functionality: It identifies software vulnerabilities, including new zero-day exploits, more efficiently than manual methods.
- Impact: In the past six months, Codemender has upstreamed 72 security fixes to open-source projects, some as large as 4.5 million lines of code.
- Process:
- Code Analysis: Spots vulnerabilities.
- Reasoning and Patching: Writes the fix or patch.
- Validation: Runs the fix through a validator to ensure it works and doesn't break other features.
- Self-Correction: If validation fails, it self-corrects and retries.
- Human Review: Validated patches are compiled for human review, then submitted to the code repository.
- Concerns: The speaker notes the potential for misuse, where a similar agent could be used by hacking teams to exploit vulnerabilities rather than fix them.
- Availability: Codemender has been announced but is not yet released for public use.
Humanoid Robot News
Several new humanoid robot demos highlight diverse advancements:
- Deep Robotics DR2: This new robot is certified waterproof and dustproof (IP67), enabling reliable operation in outdoor and industrial conditions, including rain, humidity, and dust. It can operate continuously in temperatures from -20°C to 55°C. It integrates a multi-sensor system (lidar, depth cameras, wide-angle cameras) for object detection and path planning, can climb stairs/slopes, and carry loads up to 20 kg.
- Figure 03: Designed for safe operation in home environments, featuring soft materials and multi-density foam protection. Demos show it performing household chores like serving water, cleaning, washing dishes (though primarily rinsing), loading/folding clothes, and even operating a washing machine. It also shows potential for receptionist or delivery roles. The speaker notes these are carefully planned and staged demos with hard cuts, not continuous autonomous operation.
- Cyan Robotics Orca: This robot focuses on "expressive gaits" or walking positions, demonstrating happy, tired, angry, nervous, and "monkey king" walks, with corresponding facial expressions. The speaker views this as primarily a marketing stunt.
Alibaba's Ling 1 Trillion
Alibaba's Ant Group, through its Inclusion AI team, has released Ling 1 Trillion, a significant new model.
- Size: It boasts 1 trillion total parameters with 15 billion active parameters.
- Performance: Ling 1 Trillion outperforms leading open-source models like DeepSeek 3.1 Terminus, as well as Kimik 2, GPT-5, and Gemini 2.5 Pro (low-inc version) in math and reasoning benchmarks, particularly excelling in ARC AGI (which tests learning new patterns). It also shows superior coding performance across various benchmarks.
- Efficiency: The model demonstrates high performance relative to the number of tokens used, placing it in the "upper left quadrant" for efficiency.
- Availability: The model is released on Hugging Face, including the full model for download and usage instructions.
Google's Opal Workflow Builder
Google offers Opal, a free workflow builder for AI-powered applications, similar to OpenAI's Agent Kit but focused on apps rather than agents.
- Functionality: It features an intuitive drag-and-drop interface with nodes to define app steps and information flow. It can automatically generate complex workflows from a simple text prompt.
- Availability: Access has been expanded to 15 additional countries, and it is currently free to use.
- Demo Example (Blog Post Generator):
- User prompts for a blog post generator (research, photos, full blog).
- Opal automatically creates a workflow: User input -> Conduct deep research -> Generate blog post text -> Generate image prompt (highly detailed) -> Gemini 2.5 Flash or Imagen 4 for image generation -> Compile formatted blog post.
- Users can customize or add nodes, choosing from Google models (LLMs, image generators, video, music, text-to-speech).
- The output can be generated as a web page, saved to Google Docs, a presentation, or Google Sheets.
- Other Workflow Ideas: Personalized research reports, YouTube video analysis to create quizzes.
Grock Imagine v0.9: Free Video Generator with Audio
Grock has released Imagine version 0.9, a video generator with native audio, offering a free alternative to Sora 2 and V3.
- Functionality: It generates videos from text prompts or uploaded images, with audio natively built in.
- Quality: While not matching the quality of Sora or V3, it is completely free to use.
- Process: Users type a prompt to create a photo, select an image, then click "make video." Generation is relatively fast, taking about 1 minute.
- Examples: Generating a video of "two kung fu masters fighting on a city rooftop" or animating an uploaded image of "Will Smith eating spaghetti" with custom dialogue.
- Availability: Free to use via grock.com/immagine or the Grock app.
Samsung's Tiny Recursion Model (TRM)
Samsung has introduced the Tiny Recursion Model (TRM), a groundbreaking model for its efficiency and performance.
- Size and Performance: With only 7 million parameters (not billions), TRM scores higher on the challenging ARC AGI 1 and ARC AGI 2 benchmarks than models tens of thousands of times larger, such as DeepSeek R1, 03 Mini, and Gemini 2.5 Pro. It also successfully solves mazes and Sudoku puzzles where larger state-of-the-art models fail.
- ARC AGI Significance: ARC AGI tests an AI's ability to learn and understand new patterns from data it has never seen during training, a historically difficult task for AI models.
- Methodology (Recursive Process): Instead of solving a problem in one go, TRM solves it iteratively. It starts with an initial guess, then repeatedly refines its answer using a tiny neural network, learning from each previous attempt. This recursive process allows the model to "think deeply" and reason effectively without requiring a massive neural network with billions or trillions of parameters.
- Availability: A technical paper detailing the model is available.
Synthesis and Conclusion
This week in AI demonstrates an unprecedented pace of innovation across multiple domains. Key themes include:
- Agentic AI and Workflow Automation: OpenAI's Agent Kit and Google's Opal highlight a strong move towards AI agents and intuitive workflow builders, empowering both developers and end-users to create sophisticated AI-powered applications and automate complex tasks.
- Efficiency and Accessibility: Models like Samsung's Apriel 1.5 15B Thinker and the Tiny Recursion Model (TRM) showcase a significant breakthrough in achieving state-of-the-art performance with dramatically smaller parameter counts, making advanced AI more accessible on consumer-grade hardware.
- Multimodal Generative AI: Advancements in image and video generation continue with Luminina Demu offering advanced editing, Mimix enabling character mixing, and Grock Imagine providing a free video generator with audio. OpenAI's Sora 2 API access further solidifies the move towards integrating high-quality video generation into applications.
- Specialized AI for Practical Applications: Tools like Paper 2 Video automate academic presentations, ChronoEdit offers physically consistent image editing for synthetic data generation, and Codemender provides autonomous vulnerability patching, demonstrating AI's growing utility in specific, high-value tasks.
- Humanoid Robotics Progress: While still in early stages, new humanoid robot demos from Deep Robotics, Figure, and Cyan Robotics hint at future applications in industrial, home, and even expressive roles, pushing the boundaries of physical AI.
Overall, the landscape of AI is rapidly expanding, with a clear focus on making powerful AI more efficient, accessible, and integrated into daily workflows and specialized applications, while also exploring new frontiers in multimodal generation and physical embodiment.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Why Does This Guy Appear In Kids Videos?
sphynx

TIC en las Organizaciones - Electiva Complementaria II Unisimon
Julieth Güell S

How to Tame Your Advice Monster | Michael Bungay Stanier | TED
TED

Margaret Heffernan: Why it's time to forget the pecking order at work
TED

The importance of psychological safety: Amy Edmondson
The King's Fund

What Is Psychological Safety?
Harvard Business Review

13-Conflict Management: Listening in Conflict
Deliberate Development