DeepSeek V4, JIPU GLM4.7 Flash, Emotion AI & NewCoder 14B: A Detailed Overview
Key Concepts:
- DeepSeek V4: Potential next-generation flagship AI model from DeepSeek, focusing on coding ability and architectural redesign.
- GLM4.7 Flash: A 30B parameter Mixture of Experts (MoE) model from ZAI, designed for local deployment with strong coding and reasoning capabilities.
- Constructed Emotion: A theory suggesting emotions are built by the brain combining bodily signals and sensory information.
- NewCoder 14B: An AI model trained via reinforcement learning for competitive programming, achieving high pass rates on challenging benchmarks.
- KV Cache: Memory structure optimizing transformer inference efficiency.
- Sparsity: Technique to reduce compute by selectively activating parts of an AI model.
- FP8 Decoding: Method for increasing performance and reducing memory usage in large models.
- Mixture of Experts (MoE): Architecture where a model utilizes different "expert" networks for different tasks.
- MMLDA (Multi-layered Multimodal Latent Dirichlet Allocation): Probabilistic generative model used for discovering hidden patterns in multimodal data.
- Pass@1: A benchmark metric where the model must provide a correct answer on the first attempt.
1. DeepSeek V4: A Potential Architectural Leap
On January 21st, 2026, evidence emerged suggesting DeepSeek is developing its next flagship model, tentatively named DeepSeek V4, with a potential release around mid-February (Lunar New Year). This isn’t an official announcement, but rather analysis of code updates on GitHub. A significant batch of 114 Flash MLA code files were updated, introducing a new model identifier, “model one,” appearing 28 times alongside “V32” (DeepSeek V3.2). The separation of “model one” from V3.2 suggests a fundamental architectural change, not a minor version update.
Technical differences observed in the code point to substantial modifications:
- KV Cache Layout: Changes indicate a redesign of how the model stores and retrieves information, impacting performance, memory efficiency, and long-context handling.
- Sparsity Handling: Differences suggest compute efficiency optimizations, leveraging sparsity to reduce computational load without sacrificing quality.
- FP8 Decoding Support: Explicit support for FP8 decoding signals a focus on efficiency at scale, optimizing performance on modern hardware.
These changes align with DeepSeek’s recent research into Modified Hierarchical Connections (MHC) and a biologically inspired memory module called Engram, suggesting V4 may integrate these advancements. The company hasn’t confirmed these details.
2. JIPU GLM4.7 Flash: Deployable Reasoning and Coding Power
Zoo AI (JIPU) released GLM4.7 Flash, a 30B parameter model utilizing a Mixture of Experts (MoE) architecture. It’s designed for local deployment, offering strong reasoning and coding capabilities without requiring massive GPU clusters. GLM4.7 Flash supports both English and Chinese and is configured for chat and conversational use.
Key features include:
- Parameter Count: 31 billion parameters, leveraging MoE for efficient computation.
- Context Length: Supports 128,000 tokens of context, facilitating complex reasoning and long-form interactions.
- Interface: Uses a standard interface and chat template for easy integration into existing AI tools.
- Benchmarks: ZAI claims GLM4.7 Flash leads or remains competitive against models like Quen 330BA3B, Thinking 257, and GPT OSS 20B in math reasoning, long-horizon benchmarks, and coding agent tests.
- Inference Support: Compatible with VLLLM, SGLANG, and Transformers-based inference.
- Fine-tuning Ecosystem: A growing ecosystem of fine-tunes and quantizations is available on Hugging Face.
Default settings for generation are temperature 1.0, top P 0.95, and max new tokens 131,072. For tasks requiring precision, like terminal bench and S swb bench verified, settings are adjusted to temperature 0.7, top p 1.0, and max new tokens 16,384. “Preserved thinking mode” is recommended for multi-turn agent tasks to maintain internal reasoning across interactions.
3. Emotion Computing: Modeling Feelings as Computational Processes
Researchers at the Nar Institute of Science and Technology (Japan) and Osaka University developed an AI framework to model emotion formation as a computational process tied to bodily signals. Published in IEEE Transactions on Affective Computing (December 3rd, 2025), the research is based on the “theory of constructed emotion,” which posits that emotions are not innate reactions but concepts built by the brain combining internal (intraception – e.g., heart rate) and external (exterception – e.g., visual stimuli) signals.
The team built a computational model using Multi-layered Multimodal Latent Dirichlet Allocation (MMLDA), a probabilistic generative model that discovers hidden patterns in data without explicit labels. The model was trained on unlabeled data from 29 participants viewing emotion-evoking images and videos, recording physiological responses (heart rate) and verbal descriptions.
Results showed approximately 75% agreement between the model’s discovered emotion categories and participants’ self-reported emotional evaluations, suggesting the model accurately captures emotional experience patterns. Potential applications include emotion-aware robots, improved AI assistants, mental health support, and assistive technologies.
4. NewCoder 14B: Reinforcement Learning for Competitive Programming
News Research released NewCoder 14B, an AI model specifically trained for competitive programming – a challenging coding domain with strict time and memory limits and rigorous testing. Built on top of Quen 314B, NewCoder 14B was improved using reinforcement learning: the model writes code, which is executed in a sandbox, and receives a reward (+1) for passing all tests or a penalty (-1) for failure, exceeding time limits, or exceeding memory limits.
Key findings:
- Benchmark Performance: On Live Codebench V6 (454 problems), NewCoder 14B achieved 67.87% pass@1 (correct answer on the first attempt), a 7.08 percentage point improvement over the base model Quen 314B (60.79%).
- Training Details: Trained on 24,000 verified coding problems using 48 Nvidia B200 GPUs for 4 days, released under the Apache 2.0 license on Hugging Face.
- Training Methodology: Prioritized harder test cases during training and stopped execution immediately upon failure to improve efficiency.
- Reinforcement Learning Variants: Tested GRPO, DAPo, GSPO, and GSPO plus, with the best performance achieved at a context length of 81,920 tokens.
- Context Handling: Handles long prompts in stages (32k, 40k, then extended to 81,920 tokens at evaluation) and ignores solutions exceeding the maximum context window to prevent shortcut learning.
Synthesis/Conclusion:
The developments highlighted demonstrate significant progress across multiple areas of AI. DeepSeek V4 signals a potential architectural shift towards greater efficiency and performance. JIPU’s GLM4.7 Flash offers a powerful, deployable model for reasoning and coding. The Japanese research on emotion computing represents a novel approach to understanding and modeling emotions computationally. Finally, NewCoder 14B showcases the power of reinforcement learning in tackling highly complex coding challenges. These advancements collectively point towards a future where AI is not only more intelligent but also more accessible, efficient, and capable of understanding and responding to human emotions.
AI summaries can miss context or contain errors. Check important details against the original video.





