Inference, Diffusion, World Models, and More | YC Paper Club

Y CombinatorAbout 4 min readMay 29, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • Speculative Decoding: An inference optimization technique using a small "draft" model to predict tokens and a large "target" model to verify them in parallel.
  • SSD (Speculative Speculative Decoding): An advancement that parallelizes the sequential drafting and verification steps to hide latency.
  • Model Predictive Control (MPC): A control strategy using a dynamics model to predict future states and select optimal action sequences.
  • Diffusion Policy/MPC: Using diffusion models to learn multi-step action proposals and dynamics, reducing compounding errors in robotics.
  • World Models: Neural networks that learn the dynamics of an environment to predict future states based on current observations and actions.
  • Representation Collapse: A failure mode in world models where the model maps all inputs to a trivial or identical latent representation.
  • SiGG Regularizer: A technique (Sketching, Isotropic, Gaussian) to ensure latent embeddings remain healthy and prevent representational collapse.
  • Data-Constrained Scaling: A framework for optimizing model performance when data is limited but compute is abundant, utilizing ensembling, regularization, and distillation.

1. Speculative Speculative Decoding (SSD)

Tanishk (Stanford) introduced SSD to address the latency bottlenecks in standard speculative decoding.

  • The Problem: Vanilla speculative decoding is inherently sequential; the draft model must wait for the target model to verify tokens before drafting the next sequence.
  • The Solution: SSD parallelizes drafting and verification. By predicting verification outcomes (how many tokens the target model will accept), the draft model can begin drafting the next round while the target model is still verifying the current one.
  • Key Insight: Inference should be viewed as a capability rather than just a cost. As models scale, the speed of inference (tokens per second) directly correlates to the "peak intelligence" delivered.
  • Performance: SSD achieves significant speedups by hiding drafting latency and increasing the expected tokens per round.

2. Diffusion Model Predictive Control (DMPC)

Stannis (Google DeepMind) presented DMPC, which leverages diffusion models for robotic control.

  • Methodology: DMPC uses a diffusion model to propose multi-step action sequences and a multi-step dynamics model to evolve observations.
  • Advantages:
    • Factorization: Separating action proposals from dynamics allows for runtime adaptation to novel rewards or dynamics (e.g., a robot adapting to a "broken ankle" by updating only the dynamics model).
    • Simplicity: Stronger modeling capabilities allow for the use of simple, sampling-based planners rather than complex optimization algorithms.
  • Real-World Application: Demonstrated on locomotion tasks where the agent adapts to new reward functions (e.g., jumping) or environmental changes without retraining the entire policy.

3. Lay World Model

Isaac Ward presented "Lay World Model," focusing on efficient world modeling.

  • The Challenge: Avoiding "representational collapse," where the model fails to learn meaningful dynamics.
  • The Framework: It uses a Joint Embedding Predictive Architecture (JEPA). The SiGG regularizer is the core innovation, ensuring that latent embeddings are Gaussian-distributed and isotropic, which prevents the model from collapsing into trivial solutions.
  • Key Findings:
    • Efficiency: It is ~50x faster than competitors and runs on a single card with <24GB VRAM.
    • Uncertainty Quantification: The model can detect "surprise" or model error when presented with perturbed inputs (e.g., teleporting objects), providing a built-in safety mechanism for real-world deployment.

4. Deep Learning Generalization & Data-Constrained Scaling

Ashe (Q Labs) and Con Woo presented research on the mechanics of generalization and scaling.

  • Dispelling Mysteries: Andrew Gordon Wilson’s work uses PAC-Bayes theory to explain overparameterization. As models grow, they find "flatter" minima, which are more compressible and generalize better.
  • Data-Constrained Scaling: When data is limited (e.g., 200M tokens), the standard approach of scaling model size leads to overfitting.
  • The "Infinite Compute" Recipe:
    • Aggressive Regularization: Tuning weight decay (up to 30x higher than standard) allows models to scale without overfitting.
    • Ensembling: Training ensembles of smaller models is more data-efficient than training one large model.
    • Distillation: The benefits of large ensembles can be distilled into small, efficient models, retaining ~83% of the performance gains.
  • Key Metric: The "compute asymptote"—the theoretical limit of performance under infinite compute—serves as a new benchmark for evaluating algorithmic efficiency.

Synthesis/Conclusion

The YC Paper Club session highlighted a shift in AI research from "brute force" scaling to algorithmic efficiency. Whether through parallelizing inference (SSD), factorizing robotic control (DMPC), regularizing latent spaces (Lay World Model), or optimizing for data-constrained regimes, the common theme is the pursuit of intelligence per watt and intelligence per sample. The presenters emphasized that as we hit data walls, the focus must shift toward smarter inductive biases, better regularization, and leveraging classical machine learning theories to unlock the next generation of model capabilities.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.