How to run LLMs locally from an external hard disk #ollama #llm #AI #privatechat #nointernetchat

Geo Joy (Breach Guru)About 4 min readFeb 2, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Ollama: A tool for running large language models (LLMs) locally on your machine.
  • Uncensored LLMs: Language models without the restrictions and policies of models like ChatGPT or Bard.
  • Symbolic Link (Simlink): A shortcut that points from an internal directory to an external directory, allowing applications to access files on an external drive as if they were local.
  • GGUF (GPT Generated Unified Format): A file format becoming the standard for LLMs, supported by Ollama.
  • Hugging Face: A platform for sharing and discovering machine learning models, including those in GGUF format.
  • Quantization Levels (Q3, Q4, Q5, Q6): Different levels of precision for model weights, affecting memory usage and performance.
  • REST API: An interface exposed by Ollama, allowing developers to interact with the LLMs programmatically.

Running Uncensored LLMs Locally with Ollama

Introduction

The video addresses the limitations of mainstream AI assistants like ChatGPT and Bard, which often restrict responses due to alignment with company policies or principles. It proposes a solution: running uncensored LLMs locally using Ollama, providing greater control, privacy, and freedom in generating content. The video also demonstrates how to store these models on an external hard drive for portability.

Ollama: Running LLMs Locally

  • What is Ollama? Ollama is a tool that enables users to run LLMs on their local machines, eliminating reliance on cloud providers. This offers increased control and privacy.
  • Installation: Ollama is available for macOS and Linux, with Windows support coming soon (via WSL). The video demonstrates installation on both macOS and Kali Linux.
  • Models Directory: Ollama provides a directory of available models, frequently updated. Each model has specific requirements, such as RAM. For example, a 7 billion parameter model requires approximately 8GB of RAM, while a 13 billion parameter model needs around 16GB.
  • Model Exploration: Users can explore different versions and tags of models, such as Code Llama (including Python-specific versions) and Llama2 uncensored.
  • Quantization Levels: The video mentions quantization levels (Q3, Q4, Q5, Q6) but notes that a separate video will delve into the details of these levels.
  • Extensions: Ollama offers extensions, including UI-based chatbots and copilot alternatives for offline use.
  • REST API: Ollama exposes a REST API, allowing developers to integrate LLMs into various applications, including mobile apps or local network services.

Using External Hard Drive for Portability

  • Storing Models: The video emphasizes storing models on an external hard drive for portability. This allows users to connect the drive to any computer and instantly access the LLMs.
  • Security: The presenter uses a SanDisk 1TB hard drive with built-in security features, requiring a password for decryption.
  • Directory Structure: The presenter stores models in a specific directory structure on the external drive: ai_models/ollama/models. This directory contains blobs (the model files) and manifests (metadata for the models).
  • Symbolic Link (Simlink): The video explains how to create a symbolic link to map the internal Ollama models directory to the external hard drive. This allows Ollama to access the models stored on the external drive as if they were local.
    • Concept: A simlink creates a shortcut from an internal folder to an external folder, preventing application crashes due to missing directories.
    • Command: The command ln -s <external_path> <internal_path> is used to create the simlink.
    • macOS Example: ln -s /Volumes/SanDisk/ai_models/ollama/models ~/.ollama/models
    • Kali Linux Example: sudo ln -s /media/<username>/SanDisk/ai_models/ollama/models ~/.ollama/models (Note: the trailing slash is removed from the external path in Kali Linux).

Demonstrations and Use Cases

  • Uncensored Responses: The video demonstrates that uncensored models can answer questions that ChatGPT and Bard cannot, such as providing steps to create a virus or generating obscene content.
  • Kali Linux Integration: The video shows how to install and configure Ollama on Kali Linux, highlighting its potential for cybersecurity applications.
  • Short Film Generation: The presenter asks the LLM to create a short film episode with obscene scenes and dialogues similar to Game of Thrones, which the model successfully generates.

Conclusion

The video concludes by emphasizing the portability of the setup, requiring only the Ollama installation script and the external hard drive. It encourages viewers to explore different models from Hugging Face and experiment with local, uncensored LLMs. The main takeaway is that Ollama provides a way to bypass the restrictions of online AI assistants and have complete control over the generated content.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.