The Ultimate Guide to Running ALL Your AI Locally (The Future is Here)

Cole MedinAbout 8 min readJun 12, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Local AI: Running large language models (LLMs) and infrastructure (database, UI) on your own machine, 100% offline.
  • Cloud AI: Using APIs to access LLMs and infrastructure hosted by third-party providers (e.g., OpenAI, Google, Supabase).
  • Open-Source LLMs: LLMs with publicly available code and weights (e.g., Deepseek R1, Quen 3, Mistral).
  • Ollama: An open-source platform for easily downloading and running local LLMs.
  • Quantization: Reducing the precision of LLM parameters to decrease model size and increase speed (e.g., Q4, Q8).
  • Offloading: Splitting LLM layers between GPU (VRAM) and CPU (RAM) to run larger models.
  • OpenAI API Compatibility: A standard API (chat completions) that allows easy swapping between LLM providers (OpenAI, Gemini, Ollama).
  • Local AI Package: A curated set of services (N8N, Supabase, Ollama, Open WebUI, etc.) for building a complete local AI infrastructure.
  • N8N: A no-code/low-code workflow automation platform.
  • Superbase: An open-source database platform.
  • Open Web UI: A chat GPT-like interface for interacting with local LLMs.
  • CRXNG (Seir XNG): An open-source, private web search engine.
  • Caddy: A web server and reverse proxy used for setting up subdomains and HTTPS.
  • Pydantic AI: A Python framework for building AI agents.
  • Digital Ocean: A cloud platform for deploying servers and GPU instances.

What is Local AI?

  • Local AI involves running LLMs and related infrastructure (databases, UIs) entirely on your own machine, ensuring 100% offline operation.
  • This is achieved through open-source LLMs and software, eliminating the need for paid APIs.
  • Examples of open-source LLMs include Deepseek R1, Quen 3, Mistral, and Llama.
  • Open-source software used in local AI includes Ollama, Superbase, N8N, and Open Web UI.
  • A hands-on demo using Ollama is provided, showing how to download and run a Deepseek R1 model directly from the terminal.
    • The process involves downloading the model (1.1 GB in the example) and then interacting with it through a chat interface.
    • The demo illustrates that the LLM runs on the local machine's infrastructure, with the model's parameters loaded onto the graphics card for inference.

Why Local AI? (Pros and Cons)

  • Advantages of Local AI:
    • Privacy and Security: Data remains on your hardware, crucial for businesses in regulated industries (healthcare, finance, real estate).
    • Model Fine-Tuning: Open-source LLMs can be fine-tuned with your own data, creating domain experts.
    • Cost-Effectiveness: No API bills; only electricity or server costs.
    • Speed: Reduced network delays as everything runs on the same infrastructure.
  • Advantages of Cloud AI:
    • Easier Setup: Simple API calls, less initial configuration.
    • Less Maintenance: Providers host and manage the infrastructure.
    • Better Models (Currently): Cloud models like Claude 4 are generally more powerful than local LLMs, although this gap is diminishing.
    • Out-of-the-Box Features: Built-in memory, web search, etc.
  • The video argues that the advantages of local AI will become more prevalent over time as privacy concerns grow and the performance gap between local and cloud LLMs shrinks.

Hardware Requirements

  • LLMs are resource-intensive due to the billions or trillions of parameters they contain.
  • Parameters are interconnected in a web-like structure, with input layers (prompts) and output layers (responses).
  • Fitting an entire LLM into a graphics card requires significant VRAM.
  • LLM Size Ranges and Hardware Recommendations:
    • 7-8 Billion Parameters: 4-5 GB VRAM (e.g., Nvidia 3060 Ti with 8 GB), 25-35 tokens/second.
    • 14 Billion Parameters: 8-10 GB VRAM (e.g., 4070 Ti with 16 GB, 3080 Ti with 12 GB), 15-25 tokens/second, suitable for basic tool calling.
    • 30-34 Billion Parameters: 16-20 GB VRAM (e.g., 3090 with 24 GB, Mac M4 Pro with 24 GB unified memory), 10-20 tokens/second, impressive performance close to cloud AI.
    • 70 Billion Parameters: 35-40 GB VRAM, requires splitting across multiple GPUs (e.g., two 3090s, two 4090s) or an enterprise-grade GPU (H100), 8-12 tokens/second.
  • Recommended Builds:
    • $800: 4060Ti graphics card, 32 GB RAM.
    • $2,000: PC with 3090 and 64 GB RAM, or Mac M4 Pro with 24 GB unified memory.
    • $4,000: Two 3090 graphics cards, 128 GB RAM, or Mac M4 Max with 64 GB unified memory.
  • Specific LLM recommendations based on size range: Deepseek R1 (7B, 14B, 32B, 70B), Quen 3 (8B, 14B, 32B), Mistral Small (22B, 24B).
  • Alternative: Open Router allows testing open-source LLMs without self-hosting.

Tricky Stuff: Quantization, Offloading, and Environment Variables

  • Quantization:
    • Lowers model precision to reduce size and increase speed without significant performance loss.
    • Reduces parameter precision from 16 bits to 8, 4, or 2 bits.
    • Q4 quantization is generally the best balance between size, speed, and quality.
    • Ollama defaults to Q4 quantization.
  • Offloading:
    • Splits LLM layers between GPU (VRAM) and CPU (RAM).
    • Hurts performance but allows running larger models or handling larger contexts.
    • Avoid offloading if possible.
  • Environment Variables (Ollama):
    • Flash Attention: Set to "1" or "true" to improve attention calculation efficiency.
    • Quantize Context: Compress context (system prompt, tool descriptions, conversation history) using Q8 quantization.
    • Context Length: Override the default 2000 token limit (set to 8000 or 32000).
    • Model Limit: Limit the number of models in memory (set to 1 or 2).

Using Local AI Anywhere: OpenAI API Compatibility

  • OpenAI's chat completions API has become a standard for exposing LLMs.
  • Ollama implements this API, allowing easy swapping between OpenAI and local LLMs.
  • The key is to change the base URL to point to the Ollama instance (e.g., http://localhost:11434/v1).
  • An example Python script demonstrates how to switch between OpenAI and Ollama by modifying the base URL and API key.

The Local AI Package

  • A curated set of services for building a complete local AI infrastructure.
  • Includes N8N, Superbase, Ollama, Open Web UI, Flowwise, Quadrant, Neo4j, CRXNG, Caddy, and Langfuse.
  • Requires about 8 GB of RAM to run everything.
  • Services can be removed from the package to reduce resource usage.
  • A beta front-end application is being developed to manage services and environment variables.
  • Installation Steps:
    1. Install Python, Git, and Docker.
    2. Clone the Local AI Package GitHub repository.
    3. Configure environment variables in the .env file.
    4. Start the services using the appropriate command (e.g., python start_services.py --profile gpu-nvidia).
    5. Fix the Superbase pooler issue (if on Windows) by changing the line ending in docker_volumes/pooler/pooler.exs to LF.
  • Updating the Package:
    • Tear down the containers.
    • Pull the latest containers.
    • Start the services again.
  • Access the services in your browser using localhost and the appropriate port (e.g., N8N: localhost:5678, Open Web UI: localhost:8080, Superbase: localhost:8000).
  • Additional LLMs can be pulled into the Ollama container using the docker exec command.

Building Agents with the Local AI Package

  • Open Web UI:
    • A chat GPT-like interface for interacting with local LLMs.
    • Configure the Ollama API connection in the admin panel (settings -> connections).
    • Set the base URL to Ollama (if running Ollama in the container) or host.docker.internal (if running Ollama on the host machine).
  • N8N Agent:
    • Connect an AI Agent node to a chat trigger.
    • Use the Lama Chat Model credentials, setting the base URL to Ollama:11434 or host.docker.internal:11434.
    • Add memory using Postgress credentials, referencing the Superbase database.
    • Create a web search tool using CRXNG.
    • Connect N8N to Open Web UI using the N8N pipe function.
      • Set up a web hook trigger with header authentication.
      • Use a router node to differentiate between main agent requests and metadata requests (title, tags).
      • Configure the N8N pipe function in Open Web UI with the N8N URL, bearer token, input field, and output field.
  • Python Agent:
    • Use Fast API to create an API endpoint for the agent.
    • Use Pydantic AI to define the agent and its tools.
    • Implement a web search tool using CRXNG.
    • Implement header authentication.
    • Fetch and store conversation history in Superbase.
    • Containerize the Python agent using a Docker file.
    • Add the agent to the local AI Docker Compose stack.

Deploying to the Cloud

  • Cloud Platforms:
    • Digital Ocean: Offers both GPU and CPU instances.
    • Tensor Do: Affordable GPU instances.
    • Hostinger: Affordable CPU instances for hosting everything except LLMs.
  • Platforms to Avoid:
    • RunPod, Lambda Labs, vast.ai: These platforms provide access to containers, not the underlying machine, making it difficult to run the local AI package.
  • Deployment Steps (Digital Ocean):
    1. Create a GPU droplet (or CPU droplet for a hybrid setup).
    2. Connect to the droplet via SSH or the web console.
    3. Enable the firewall and open ports 80 and 443.
    4. Clone the Local AI Package GitHub repository.
    5. Configure environment variables in the .env file, including Caddy settings for subdomains.
    6. Set up DNS records for the subdomains.
    7. Start the services using the command python start_services.py --profile gpu-nvidia --environment public.
    8. If Docker Compose is not installed, run the provided commands to install it.
    9. Deploy the Python agent by cloning the repository, configuring environment variables, and adding it to the local AI Docker Compose stack.
  • Access the services in your browser using the subdomains (e.g., n8nyt.dynamis.ai, openwebyty.dynamis.ai, superbaseyt.dynamis.ai).

Synthesis/Conclusion

The video provides a comprehensive guide to setting up and using local AI, covering everything from the basic concepts and hardware requirements to building agents and deploying them to the cloud. It emphasizes the importance of privacy and security, the cost-effectiveness of local AI, and the flexibility of using open-source tools. By following the steps outlined in the video, viewers can create their own private and secure AI infrastructure and build powerful AI agents that run entirely on their own hardware.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.