The ONLY way to run your own Deepseek on mobile...

AI JasonAbout 6 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Large Language Models (LLMs) on Edge Devices
  • VRAM Calculation for LLMs
  • Quantization (FP32, FP16, etc.)
  • MLX Framework (Apple Silicon)
  • Termux (Android)
  • Hugging Face Model Repository
  • SwiftUI
  • Cursor (AI Code Editor)
  • Sweetpad (iOS Build Extension)
  • Chat History Implementation
  • App Store Deployment

GPU Needs Calculation for Large Language Models

The video explains how to estimate the VRAM (Video RAM) requirements for running large language models (LLMs). This is crucial for determining if a specific device can handle a particular model.

  • VRAM Consumption: LLMs consume VRAM for two primary purposes: storing model parameters and storing activation memory (intermediate calculation results).
  • Formula: A simplified formula is provided to estimate VRAM needs:
    • VRAM (GB) ≈ (Number of Parameters * Precision) / 8 * 1.2
    • Number of Parameters: Typically found in the model name (e.g., 13B for a 13 billion parameter model) or on Hugging Face.
    • Precision: Refers to the number of bits used to represent each parameter (e.g., FP32, FP16, INT8, INT4). Lower precision reduces VRAM usage but may slightly impact accuracy.
    • Division by 8: Converts bits to bytes.
    • Multiplication by 1.2: Accounts for overhead and activation memory.
  • Quantization: The video explains quantization, which reduces the precision of model parameters (e.g., from 32-bit floating point to 16-bit or lower). This significantly reduces VRAM requirements and computational cost, enabling LLMs to run on devices with limited resources.
  • Example: A 13B parameter model using FP16 precision would require approximately (13 * 16) / 8 * 1.2 = 31.2 GB of VRAM.
  • VRAM Estimator Tools: More accurate estimations can be obtained using specialized tools like VRAM estimators, which consider detailed model parameters found in the config.json file on Hugging Face.

Building Mobile Applications with Local LLMs

The video demonstrates how to build mobile applications that run LLMs locally on both Android and iOS devices.

Android (Termux)

  • Termux: A terminal emulator for Android that allows installing and running Linux packages.
  • Process:
    1. Install Termux from Google Play.
    2. Set up storage access: termux-setup-storage.
    3. Update packages: pkg upgrade.
    4. Install necessary tools: pkg install git clang make python.
    5. Clone the llama.cpp repository: git clone <llama.cpp_repo_url>.
    6. Navigate to the llama.cpp directory: cd llama.cpp.
    7. Build the llama.cpp binary: make.
    8. Run the desired model: ./main -m <path_to_model.gguf> -p "Your prompt".
  • Limitations: Distributing apps built with Termux is difficult because it requires users to install Termux and set up the environment manually.

iOS (MLX Framework)

  • MLX: A machine learning framework developed by Apple, optimized for Apple Silicon (M1, M2, etc.) GPUs. It enables efficient local LLM inference on iOS devices.
  • Process:
    1. Create a new Xcode project.
    2. Add the mlx-examples package dependency: https://github.com/ml-explore/mlx-examples. Specify the main branch.
    3. Select the mlx-llm package.
    4. Import necessary libraries: import MLX and import MLXLLM.
    5. Load a pre-defined model using MLXLLM.ModelRegistry.getModel(named: "DeepSeek-R1.5-1.3B"). The video also shows how to load models from Hugging Face.
    6. Tokenize the input prompt using tokenizer.encode(prompt).
    7. Generate the output using MLXLLM.generate(model: model, tokens: tokens).
    8. Stream the results to the UI.
  • Code Example: The video provides a Swift code example demonstrating how to create a basic chat app with a text field for user input, a button to generate answers, and a display area for the model's output.
  • Cursor Integration: The video demonstrates using Cursor, an AI-powered code editor, to assist in building the iOS app. Cursor can generate code, explain existing code, and debug errors.
  • Sweetpad: An Xcode extension that simplifies building and running iOS apps directly from Cursor.
  • Chat History: The video explains how to implement chat history to provide context to the LLM. It uses the Llama 3 prompt format, which includes system messages, user messages, and assistant messages.
  • Increasing Memory Limit: For larger models, the video shows how to increase the memory limit for the app in Xcode by enabling the "Increase Memory Limit" capability.
  • Adding New Models: The video demonstrates how to add new models from Hugging Face to the model registry by specifying the model path and tokenizer.
  • App Store Deployment: The video briefly covers the process of deploying the app to the App Store, including creating a developer account, adding an app icon, archiving the app, and submitting it for review.

Key Arguments and Perspectives

  • Local LLMs Enable Flexible Pricing: Running LLMs locally eliminates the need for expensive cloud-based inference, allowing for more flexible pricing models (e.g., one-time purchase, usage-based on device resources).
  • MLX Simplifies iOS Development: The MLX framework makes it significantly easier to build iOS apps with local LLMs compared to Android (Termux).
  • Cursor Accelerates Development: AI-powered code editors like Cursor can greatly accelerate the development process by generating code, explaining code, and debugging errors.

Notable Quotes

  • "This is really exciting part about the ability to run local model on your mobile device directly is that your pricing strategy became extremely flexible."
  • "MLX is a machine learning framework that's specifically designed for apple cicon and mpu it lets you run large L model locally on your device."

Technical Terms

  • VRAM: Video Random Access Memory, used to store model parameters and activation memory.
  • FP32, FP16: Floating-point precision formats (32-bit and 16-bit, respectively).
  • Quantization: Reducing the precision of model parameters to reduce VRAM usage and computational cost.
  • Tokenizer: A component that converts text into numerical tokens that the LLM can understand.
  • Inference: The process of using a trained model to generate predictions or outputs.
  • Hugging Face: A platform for sharing and discovering machine learning models and datasets.
  • Apple Silicon: Apple's custom-designed processors (e.g., M1, M2) optimized for performance and efficiency.
  • MLX: Apple's machine learning framework for Apple Silicon.
  • Termux: An Android terminal emulator.
  • SwiftUI: Apple's declarative UI framework for building iOS apps.
  • Cursor: An AI-powered code editor.
  • Sweetpad: An Xcode extension for building iOS apps from Cursor.

Logical Connections

The video logically progresses from explaining the hardware requirements for running LLMs to demonstrating how to build mobile applications that utilize them. It starts with the theoretical aspects of VRAM calculation and quantization, then moves on to the practical steps of setting up the development environment and implementing the code. The video also highlights the advantages of using MLX on iOS compared to Termux on Android.

Synthesis/Conclusion

The video provides a comprehensive guide to running large language models on edge devices, specifically mobile phones. It emphasizes the importance of understanding VRAM requirements and the benefits of quantization. The video then offers practical, step-by-step instructions for building iOS apps with local LLMs using the MLX framework, showcasing the potential for creating innovative and cost-effective AI-powered mobile applications. The use of AI-assisted coding tools like Cursor is also highlighted as a way to accelerate the development process. The key takeaway is that running LLMs locally on mobile devices is becoming increasingly feasible, opening up new possibilities for application development and deployment.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.