NVIDIA’s New AI: Impossible Video Game Animations!

Two Minute PapersAbout 3 min readMay 12, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • GENMO: NVIDIA's AI model for "everything to motion," encompassing video to motion, text to motion, and audio to motion.
  • Motion Transfer: Transferring movements from a recorded video to a virtual character.
  • Text Prompting: Using text to guide and modify the AI's motion generation.
  • Audio Input: Incorporating music to influence the generated motion.
  • Keyframes: Specific poses at defined points in time used as constraints for motion generation.
  • Breakpoints: Transition points where different input types (video, text, keyframes) are seamlessly integrated.
  • Diffusion Backbone: The underlying architecture of the AI model, requiring denoising steps.
  • SLAM (Simultaneous Localization and Mapping): An off-the-shelf method used to extract camera position and direction from videos.
  • Pose Estimation: Marking up the video with a doll in 2D.

1. Introduction to GENMO

  • NVIDIA has released a new AI work called GENMO, which is described as "everything to motion."
  • GENMO goes beyond text-to-motion, incorporating video and audio inputs.

2. Capabilities of GENMO

  • Video to Motion: GENMO can learn movements from a recorded video and transfer them to a virtual character.
    • Example: The AI can analyze a video of someone climbing stairs and replicate the motion on a 3D character.
  • Text to Motion: Users can add text prompts to guide the AI's motion.
    • Example: Adding a text prompt to make the character perform a lunge.
  • Audio to Motion: GENMO can incorporate music as an input to influence the generated motion.
  • Combining Inputs: GENMO can seamlessly weave together different input types (video, text, keyframes) at breakpoints.
    • Example: Transitioning from an initial video to lunges based on text prompts, while maintaining the style of the initial motion.

3. Examples and Demonstrations

  • Invisible Stairs: The AI is shown climbing invisible stairs, demonstrating its ability to understand and react to the environment.
  • Keyframe Integration: The AI is tasked with hitting specific poses (keyframes) at defined points in time.
  • Real Dancing: GENMO is tested with real dance movements, including cha-cha-cha, and performs impressively.
  • Monkey Mimicry: The AI is shown mimicking human actions, such as typing on a keyboard, with realistic movements.

4. Technical Details and Limitations

  • GENMO relies on an off-the-shelf SLAM method to extract camera position and direction from videos.
  • The AI model has a heavy diffusion backbone and requires 5 denoising steps.
  • Limitations:
    • Only handles full-body motion.
    • No facial gestures or hand articulation.
    • Not an end-to-end solution, as it relies on SLAM.

5. Editing and Timing Adjustments

  • Users can edit the timings of the generated motion.
  • The AI re-does the animation from scratch to ensure seamless transitions after edits.

6. Significance and Potential Applications

  • GENMO is described as an "absolutely fantastic AI contribution to computer games and virtual worlds."
  • It has the potential to revolutionize character animation and virtual world creation.

7. Conclusion

  • GENMO represents a significant advancement in AI-driven motion generation.
  • While it has limitations, its capabilities are impressive and hold great promise for future applications.
  • The presenter expresses hope that NVIDIA will release the source code for this work.
  • The presenter thanks the viewers for watching Two Minute Papers.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.