Large Scale AI on Apple Silicon (as mentioned by @AndrejKarpathy ) — Alex Cheema, EXO Labs

AI EngineerAbout 5 min readJun 21, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Quantization of Energy: The idea that energy exists in discrete packets rather than a continuous range.
  • Hardware Lottery: The concept that the success of a research idea is heavily influenced by the available hardware and infrastructure, not solely its intrinsic merit.
  • Causal Consistency: A model where events in a distributed system are ordered based on their causal relationships, allowing for reasoning about data dependencies and system state.
  • Orchestration Layer: A software layer that manages and coordinates the execution of tasks across a distributed system, optimizing resource utilization and ensuring reliability.
  • Pre-fill and Generation Phases (LLMs): The two main stages in generating text with a large language model. Pre-fill is compute-bound, while generation is memory bandwidth-bound.
  • Second-Order Optimization Methods: Optimization algorithms that use information about the curvature of the loss landscape to accelerate training, often requiring more memory.
  • Von Neumann Bottleneck: The limitation in computer architecture where the rate of data transfer between the CPU and memory is slower than the CPU's processing speed.

The Milikan Experiment and Scientific Rigor

The speaker begins by illustrating the challenges of scientific progress with the example of the Milikan experiment.

  • The Problem: Early 20th-century physics faced the problem of infinite energy predicted by existing theories.
  • Planck's Solution: Max Planck proposed that energy is quantized, introducing a constant 'h' to resolve the issue mathematically.
  • Milikan's Experiment (1909): Milikan indirectly measured Planck's constant 'h' by determining the charge of an electron using charged oil droplets in water.
  • The Flaw: Milikan's result, initially celebrated and widely adopted, was later found to be incorrect.
  • Confirmation Bias: Subsequent experiments, influenced by Milikan's established result, were unconsciously adjusted to align with his findings, leading to a 15-year period of widespread acceptance of the wrong value. Scientists were "fudging" experiments to match Milikan's results.
  • Key Takeaway: This highlights the difficulty of maintaining scientific rigor and the potential for confirmation bias to skew results, even with numerous experiments.

Questioning Assumptions: The Rat Maze Experiment

The speaker then presents another example to emphasize the importance of questioning assumptions in scientific research.

  • The Experiment: A scientist attempted to train rats to navigate a maze, consistently exiting a door three positions away from their entry point.
  • The Challenge: Despite controlling for visual cues, smell, and other potential factors, the rats consistently returned to the same exit door.
  • The Discovery: The rats were using auditory cues (sounds) to navigate, a factor the scientist had not initially considered.
  • The Problem: The scientific community disregarded the findings due to established methodologies and assumptions. The experiment wasn't cited and was essentially forgotten.
  • Key Takeaway: Even with rigorous methodology, groundbreaking research can be overlooked if it challenges prevailing assumptions and established practices.

Cargo Cult Science and AI

The speaker connects these historical examples to the field of AI, highlighting similar challenges in adopting new ideas.

  • Fineman's Quote: Richard Fineman's quote, "The first principle is that you must not fool yourself and you are the easiest person to fool," emphasizes the importance of self-awareness and critical thinking in scientific research.
  • Back Propagation Example: Back propagation, introduced in 1963, was repeatedly reinvented but not widely adopted until the late 1980s. Deep convolutional neural networks (CNNs) faced similar skepticism despite their potential.
  • Hardware Limitations: The limitations of CPUs, particularly the von Neumann bottleneck, hindered the adoption of computationally intensive AI techniques.
  • GPU Impact: GPUs, with their improved ratio of computation to memory access, enabled significant advancements in AI.
  • Hardware Lottery: Sarah Hooker's concept of the "hardware lottery" is introduced, emphasizing that the best research ideas don't always win due to factors like hardware availability and existing infrastructure.
  • LLM Inertia: The speaker argues that LLMs are creating inertia by reinforcing the use of languages like Python, as their proficiency in generating Python code leads to increased Python usage, further improving LLM performance in that area.
  • Python Dominance: A study showed Python being the best language for 90-97% of all problems, highlighting the potential for LLMs to further solidify its dominance.

Exo: An Orchestration Layer for AI

The speaker introduces Exo, a project aimed at addressing the challenges of running AI workloads across diverse hardware.

  • The Problem: Lack of a reliable orchestration layer for managing AI tasks across different hardware targets and network configurations.
  • Exo's Solution: Exo is building an orchestration layer that models everything as a causally consistent set of events, enabling reasoning about data dependencies and system state.
  • Causal Graph: Exo uses a causal graph to track the ordering of events across a distributed system, providing guarantees about data location and system behavior.
  • Practical Applications:
    • LLM Generation: Optimizing LLM generation by utilizing different hardware for the pre-fill (compute-bound) and generation (memory bandwidth-bound) phases. For example, using Nvidia's new Blackwell chip for prefill and a system with more memory bandwidth for generation.
    • Training on Apple Silicon: Leveraging the high memory-to-flop ratio of Apple silicon for memory-intensive training methods, such as second-order optimization.
  • New Optimizer: Exo is developing a new optimizer that is twice as efficient per flop as Adam but requires more memory, making it suitable for Apple silicon.
  • Benchmarking and Open Data: Exo is committed to publishing all benchmark data, including suboptimal configurations, to promote transparency and learning.
  • Exo Gym: A tool for running experiments locally to test different distributed algorithms.
  • Upcoming Release: A major release is planned for the end of the week, featuring the orchestration layer, benchmarking website, and tools for testing algorithms on different devices.

Questions and Answers

The speaker answers questions about Exo's current focus and potential future directions.

  • Focus on Orchestration: Exo is primarily focused on the orchestration layer, aiming to provide a reliable and flexible platform for managing AI workloads across diverse hardware.
  • Collaboration with MLX: Exo is collaborating with the MLX team to build out primitives for distributed computing on Apple silicon.
  • AMD vs. Apple Silicon: While AMD GPUs have more flops than Apple silicon, the memory-to-flop ratio is higher on Apple silicon, making it suitable for memory-intensive workloads.
  • Federated Learning: Exo is not currently working on federated learning, but there may be potential synergies with other projects in that area.

Conclusion

The speaker concludes by emphasizing the importance of questioning assumptions, promoting transparency in research, and building robust infrastructure for AI development. Exo aims to contribute to this effort by providing a reliable and flexible orchestration layer for running AI workloads across diverse hardware.

AI summaries can miss context or contain errors. Check important details against the original video.

MAKE IT YOURS

Read. Remember. Reuse.

Free tools

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.