New AI video editor, Bytedance's VEO, new top 3D generator, new open-source AI beats DeepSeek

AI SearchAbout 6 min readJun 22, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

Text-aware image restoration (TAIR), Hunyan 3D 2.1, Part Tracker (Nvidia), Lauraedit, Miniax M1, Hya O2, Immerse Gen, MidJourney V1, Interact Human (ByteDance), Polaris, Align Your Flow (Nvidia), Diffusion Transformer, Mixture of Experts (MoE), Reinforcement Learning, Distillation.

Text-Aware Image Restoration (TAIR)

  • Main Topic: AI for fixing blurry or damaged images, especially those with text.
  • Key Points:
    • TAIR (Text-aware image restoration) sharpens and makes text more readable in blurry images.
    • It is faster and more effective than existing alternatives like Divir.
    • The AI can accurately reconstruct text even when barely discernible in the original image.
  • Examples: Restoring images of Louis Vuitton logos, airplane lettering, and movie posters.
  • Technical Details:
    • Uses a combination of a diffusion transformer and a text recognition component.
    • The text recognition component detects and recognizes text, aiding in image restoration.
  • Availability: Data sets and pipeline are available on Hugging Face and Google Drive, but model weights are not yet released.

Hunyan 3D 2.1

  • Main Topic: Open-source 3D model generator.
  • Key Points:
    • Generates 3D models from 2D images, including shape and texture (albedo, metallic properties, roughness).
    • Completely free and open-source.
    • Outperforms other leading 3D model generators like Trippo and Trellis.
  • Examples: Generating 3D models of a wardrobe, a cat, a golden truck with jade, and a character.
  • Hugging Face Space: A free online space is available for testing.
  • Local Installation: Models can be downloaded from Hugging Face for local use with a provided Gradio graphical interface.
  • Benchmark Scores: Hunyan 3D 2.1 beats other leading 3D model generators in quality assessments.

Part Tracker (Nvidia)

  • Main Topic: AI tool for creating segmented 3D models from single images.
  • Key Points:
    • Reconstructs 3D objects and segments them into meaningful parts for editing and animation.
    • Works with various images and art styles.
    • Allows for editing or animating individual parts of the 3D model.
  • Examples: Segmenting a toast, spaceship, mini golf course, car, treasure chest, and head with hair.
  • Comparison: Superior to Hollow PartrA in detecting and segmenting meaningful parts.
  • Hugging Face Space: A free online space is available for testing.
  • Local Installation: Models and code are available on GitHub for local use.

Lauraedit

  • Main Topic: AI tool for editing videos by modifying the first frame.
  • Key Points:
    • Applies changes made to the first frame of a video to the rest of the video.
    • Preserves details of unchanged elements in the video.
    • Can use additional frames to guide video generation further.
  • Examples: Swapping a watch, replacing headphones with a notebook, and changing flower colors in a time-lapse.
  • Comparison: Outperforms other video editors like Vase (Alibaba) and Cling in preserving pose and motion.
  • Technical Details: Uses mask-aware LoRA fine-tuning to identify and edit specific parts of the video.
  • Availability: GitHub repo with instructions for local installation and a Gradio graphical interface. ComfyUI integration is in development.

Miniax M1 and Hya O2

  • Main Topic: Miniax releases a new video model (Hya O2) and a large language model (M1).
  • Hya O2:
    • Insanely good video model with excellent prompt adherence, physics understanding, and overall coherence.
  • Miniax M1:
    • A large language model (LLM) comparable to Gemini and GPT.
    • Completely free and open-source (Apache 2 license).
    • 456 billion parameter hybrid mixture of experts model.
    • Supports a context length of 1 million tokens (approximately 700,000 words).
    • Performance is on par with the best closed-source models.
    • API usage is significantly cheaper than other models like Gemini 2.5 Pro and DeepSeek R1.
  • Examples: Diagnosing a complex medical problem with web search and reasoning.
  • Availability: Models are available on Hugging Face for local installation.

Immerse Gen

  • Main Topic: AI for creating detailed and realistic 3D virtual reality environments from text prompts.
  • Key Points:
    • Generates high-quality, detailed VR environments.
    • Supports various scenes and artistic styles.
    • Creates coherent scenes with minimal warping or errors.
  • Examples: Generating islands, lakes, futuristic cities, deserts, anime rooms, winter landscapes, and sunset beaches.
  • Technical Details:
    • Uses a base terrain, terrain-conditioned texturing, lightweight 3D assets, RGBA texture synthesis, and dynamic elements.
  • Availability: Code is coming soon, but no GitHub repo has been released yet.

MidJourney V1

  • Main Topic: MidJourney's first AI video model.
  • Key Points:
    • Image-to-video generation only (no text-to-video).
    • 5-second videos at 480p resolution.
    • Unique aesthetic vibe.
    • Requires a paid subscription (at least $8/month).
  • Limitations: Low resolution, no sound generation, and lower quality than other state-of-the-art video generators for high-action scenes.

Interact Human (ByteDance)

  • Main Topic: AI tool for generating videos of characters speaking dialogue.
  • Key Points:
    • Takes reference images of people or objects, audio clips, and text prompts.
    • Generates videos of characters speaking the audio, with control over who says what.
    • Supports different languages and artistic styles.
  • Examples: Generating videos of two or three characters speaking dialogue, a person interacting with objects, and cartoon scenes.
  • Availability: Only a technical paper has been released, with no indication of open-sourcing.

Polaris

  • Main Topic: A method for improving the reasoning of language models using reinforcement learning.
  • Key Points:
    • Can be applied to any existing model after training.
    • Significantly improves performance, even allowing small models to outperform much larger ones.
    • Uses reinforcement learning to train the model on challenging examples.
  • Technical Details:
    • Filters training data to include only challenging problems.
    • Uses diverse examples and adjusts the sampling temperature.
    • Gradually increases the length and complexity of training examples.
  • Availability: Everything is open-sourced, including the data set, training details, and code on GitHub.

Align Your Flow (Nvidia)

  • Main Topic: A method for distilling image generators to reduce the number of steps required for image generation.
  • Key Points:
    • Generates high-quality images in just one or two steps.
    • Trained using a flow map model.
    • Outperforms other distillation methods like LCM and TCD.
  • Technical Details:
    • Trained on ImageNet or using images generated by other AI models like SDXL or FluxDev.
  • Availability: Code is expected to be released soon.

Conclusion

This week in AI saw significant advancements across various domains, including image and video restoration, 3D modeling, video editing, and language model reasoning. Key highlights include the release of powerful and open-source models like Hunyan 3D 2.1 and Miniax M1, innovative tools like Lauraedit and Part Tracker, and groundbreaking techniques like Polaris and Align Your Flow. These developments showcase the rapid progress and increasing accessibility of AI technology, offering new possibilities for creative expression, problem-solving, and scientific discovery.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.