New AI matches GPT-5, new top image models, AI quests, video to 3D - AI NEWS

AI SearchAbout 8 min readSep 15, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Open-source AI models: Models with publicly available code, allowing for local use, modification, and distribution.
  • Closed-source AI models: Proprietary models with code not publicly available, typically accessed through APIs or platforms.
  • Depth estimation: AI's ability to predict the distance of objects from the camera in an image.
  • Normal estimation: AI's ability to predict the surface orientation of objects in an image.
  • 3D point cloud: A set of data points in three-dimensional space, representing a 3D model.
  • Large Language Models (LLMs): AI models trained on vast amounts of text data, capable of generating human-like text, translating languages, and answering questions.
  • Parameters: Variables that an AI model learns during training, influencing its performance and size.
  • Mixture of Experts (MoE): An AI architecture that combines multiple specialized sub-models (experts) to improve performance and efficiency.
  • Text-to-image generation: AI's ability to create images from textual descriptions.
  • Image editing: AI's ability to modify existing images based on user instructions.
  • Speech-to-text (STT): AI's ability to transcribe spoken language into written text.
  • HDR environment map: A high dynamic range image that captures the lighting information of a scene.
  • Diffusion Transformer: A type of neural network architecture used for generative tasks, such as image and audio generation.

Alibaba FE2E: Depth and Normal Estimation

  • Functionality: Predicts the normal (surface orientation) and depth of an image.
  • Accuracy: Achieves high accuracy in predicting surface orientation and depth, even for complex scenes.
  • Training Data: Requires less training data compared to other depth and normal estimation models.
  • Performance: Outperforms other models in both depth and normal estimation benchmarks.
  • Availability: GitHub repository available with installation instructions for local use.
  • Example: Given an input image of a table with multiple objects, FE2E accurately predicts the depth of each object.
  • Technical Terms: Normal estimation, depth estimation, training data size, benchmark scores.

Winter: 3D Model Creation from Video

  • Functionality: Creates a 3D point cloud model from a video.
  • Process: Generates local point maps from video chunks, then merges them into a global 3D point map.
  • Applications: VR scene generation for real estate or interior design.
  • Accuracy: Captures minute details of the scene, closer to the ground truth compared to competitors.
  • Availability: GitHub repository available with installation instructions for local use.
  • Example: Inputting a video of a room results in a detailed 3D point cloud reconstruction of the room.
  • Technical Terms: 3D point cloud, local point map, global 3D point map, ground truth.

K2 Think: Open-Source Reasoning Model

  • Developer: Developed in the UAE by multiple institutions.
  • Size: 32 billion parameters.
  • Performance: On par with leading reasoning models from OpenAI and DeepSeek, which have hundreds of billions of parameters.
  • Strengths: Especially good at math problems.
  • Weaknesses: Not as strong in coding or humanities.
  • Availability: Free online platform and downloadable model on Hugging Face.
  • Example: Solves Olympiad-level math problems quickly, with a high token output rate.
  • Technical Terms: Parameters, reasoning model, math composite index, hugging face, inference, fine-tuning.

BYU Ernie Models: X1.1 and 4.5 21B Thinking

  • Ernie X1.1:
    • Functionality: Performs well in factual accuracy, instruction following, and agentic capabilities.
    • Performance: On par with GPT-5 and Gemini 2.5 Pro.
    • Hallucination: Hallucinates less often than other leading models.
    • Availability: Free to use online at ernie.bu.com.
    • Features: Web search, document/image upload.
    • Example: Correctly identifies Taylor Swift as American and clarifies the lack of connection between her and a Queen Elizabeth of Japan.
  • Ernie 4.5 21B Thinking:
    • Architecture: Mixture of Experts (MoE) model.
    • Size: 21 billion parameters, with only 3 billion active during use.
    • Efficiency: Achieves performance comparable to larger models with fewer active parameters.
    • Availability: Open-source on Hugging Face under the Apache 2 license (allows commercial usage).
    • Example: Performs on par with Deepseek R1 and Gemini 2.5 Pro with significantly fewer active parameters.
  • Technical Terms: Factual accuracy, instruction following, agentic capabilities, hallucination, Mixture of Experts (MoE), active parameters, Apache 2 license.

Stability AI Stable Audio 2.5: Music Generator

  • Functionality: Generates music from text prompts.
  • Improvements: Improved musical quality and structure compared to previous versions.
  • Features: Audio inpainting, generation of full multi-part compositions.
  • Limitations: Can only generate instrumentals, limited to 3-minute tracks.
  • Quality: Sounds cleaner than previous versions but still slightly behind proprietary models like Suno or Audio in terms of aesthetic quality.
  • Availability: Accessible on the Stable Audio platform via free credits.
  • Example: Generates an ambient house track based on a text prompt.
  • Technical Terms: Audio inpainting, instrumentals, aesthetic quality.

Chat LLM by Abacus AI: All-in-One AI Platform

  • Functionality: Provides access to various AI models, image generators, and video generators in one platform.
  • Features: Artifacts feature for previewing generations, Deep Agent for complex tasks.
  • Pricing: $10 per month.
  • Benefits: Cost-effective compared to paying for each tool separately.
  • Technical Terms: Deep Agent, artifacts.

Tune Out: Anime Background Removal

  • Functionality: Removes the background from anime images.
  • Accuracy: Achieves 99.5% pixel accuracy in background removal for anime-style images.
  • Method: Fine-tuned ByRefnet model using a dataset of high-quality anime images.
  • Performance: Outperforms other background removal tools, especially for complex features like hair.
  • Availability: Instructions for local installation and use are available.
  • Example: Accurately removes the background from an anime image with intricate hair details, where other tools fail.
  • Technical Terms: Image segmentation, fine-tuning, pixel accuracy.

Google AI Quest: AI Literacy Education

  • Target Audience: Students aged 11-14.
  • Objective: Inspire students to use AI for positive impact.
  • Format: Video game where students become AI researchers.
  • Content: Real-world challenges like climate, health, and science.
  • Quests: Flood prediction, detecting diabetic retinopathy, understanding the human brain.
  • Availability: Open and available to all educators and organizations.
  • Technical Terms: AI literacy.

Alibaba Quinn 3 ASR: Speech Recognition Model

  • Functionality: Transcribes speech to text with high accuracy.
  • Languages: Supports English, Chinese, and nine other languages.
  • Features: Autodetects language, handles noisy or low-quality audio, transcribes speech with background music.
  • Performance: Low error rates across multiple languages.
  • Customization: Can be fine-tuned with contextual information for improved accuracy.
  • Availability: Free to use on a Hugging Face space.
  • Example: Accurately transcribes a noisy audio clip with background music and poor audio quality.
  • Technical Terms: Speech recognition, transcription, error rate, fine-tuning.

Dobot Robotics Atom: Humanoid Robot

  • Features: 28 degrees of freedom, fine manipulation, repeatability of 0.05 mm.
  • Processing Power: Chip delivers up to 1,500 trillion operations per second.
  • Capabilities: Autonomous action and teleoperation.
  • Pricing: Starting price around $27,000.
  • Technical Terms: Degrees of freedom, repeatability, teleoperation.

ByteDance Seedream 4.0: Image Generator and Editor

  • Functionality: Generates and edits images.
  • Resolution: Produces images up to 4K resolution.
  • Performance: Number one in text-to-image generation, tied for number one in image editing according to the Artificial Analysis leaderboard.
  • Availability: Free to use on LM Arena.
  • Example: Generates realistic 4K images of cars racing at night.
  • Technical Terms: Text-to-image generation, image editing, ELO score, confidence interval.

Tencent Hunen Image 2.1: Open-Source Image Model

  • Functionality: Generates images up to 2K resolution.
  • Strengths: Good at prompt understanding, incorporating text in images, and generating various styles.
  • Availability: Free to use on a Hugging Face space and downloadable model on GitHub.
  • Hardware Requirements: Requires a GPU with 24 GB of VRAM.
  • Example: Generates an image of two elderly women playing chess in a park.
  • Technical Terms: Inference steps, guidance, quantized FP8 version, VRAM.

Alibaba Quinn 3 Models: Max Preview and Next

  • Quinn 3 Max Preview:
    • Size: Over a trillion parameters.
    • Performance: Outperforms other leading non-thinking models in various benchmarks.
    • Availability: Free to use on chat.quen.ai.
  • Quinn 3 Next:
    • Architecture: Mixture of Experts (MoE) design with 80 billion total parameters, but only 3 billion active.
    • Efficiency: Designed for training and inference efficiency.
    • Performance: Achieves higher performance with lower training costs compared to other models.
    • Availability: Free to use on chat.quen.ai and open-sourced on Hugging Face.
  • Technical Terms: Non-thinking model, Mixture of Experts (MoE), training cost, inference efficiency.

Nvidia Lux DIT: Lighting Estimation AI

  • Functionality: Estimates the lighting of a scene from a single image or video.
  • Output: Generates a lighting map (HDR environment map).
  • Applications: Adding objects that blend seamlessly with the scene.
  • Performance: More similar to the ground truth compared to other HDR environment map generators.
  • Availability: Technical paper released, code coming soon.
  • Example: Generates a lighting map from a low-quality night photo, allowing for the seamless addition of a car into the scene.
  • Technical Terms: HDR environment map, diffusion transformer.

Conclusion

This week in AI saw significant advancements across various domains, including depth and normal estimation, 3D modeling, language models, music generation, image generation and editing, speech recognition, and robotics. Key highlights include the release of several powerful open-source models that rival closed-source alternatives, advancements in AI's ability to understand and generate realistic images and audio, and the development of AI-powered tools for education and creative tasks. The trend towards more efficient and accessible AI models continues, with Mixture of Experts architectures and open-source initiatives playing a crucial role.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.