Llama.cpp Gets a New Web UI
By Prompt Engineering
Key Concepts
- Llama.cpp: A C++ port of Facebook's LLaMA model, enabling efficient inference of large language models on consumer hardware.
- Web UI (Graphical User Interface): A visual interface for interacting with software, in this case, local language models.
- Ollama: A popular platform for running local LLMs, built on top of Llama.cpp.
- GGUF format: A file format for storing quantized LLM weights, optimized for Llama.cpp.
- In-context retrieval: Providing information directly within the prompt to the model for it to process, as opposed to Retrieval Augmented Generation (RAG).
- Tokens per second: A metric measuring the speed of text generation by a language model.
- Context window: The maximum amount of text (measured in tokens) that a language model can consider at any given time.
- Temperature, Top-K, Top-P: Parameters that control the randomness and creativity of text generation.
- Reasoning effort: A parameter that influences how deeply a model analyzes a prompt, particularly for reasoning tasks.
- Multimodal inputs: The ability of a model to process different types of data, such as text, images, and audio.
- Structured outputs: Generating responses that adhere to a predefined format, such as JSON.
Llama.cpp's New Web UI: A Performance-Focused Interface
The Llama.cpp team has released its own minimalist yet highly efficient and performant web UI for interacting with local models. This development is significant because, while Llama.cpp has been a foundational project for running local models (with platforms like Ollama built upon it), it previously lacked an official, integrated web interface. Users previously relied on third-party solutions like Open Web UI, which, while functional, are described as "extremely bloated," or the non-open-source Llama UI.
Installation and Setup
The installation process is straightforward and varies slightly by operating system. For Windows, the primary option is mentioned, while Linux and Mac users have more choices. The presenter uses Homebrew on an M2 Max with 96GB of VRAM, demonstrating the command brew install llama.cpp.
To run a model, the process involves:
- Starting the Llama.cpp server: This is done via a command-line interface.
- Specifying the Hugging Face repository ID: For example,
gp-oss-20-billioninggufformat. - Defining the port and URL: The server is typically hosted on
localhost.
The presenter notes that the model file (e.g., gp-oss-20-billion.gguf) was pre-downloaded and is approximately 12GB in size. Once the server is running, the web UI becomes accessible via the specified port on localhost.
Feature Comparison: Llama.cpp Web UI vs. Ollama Web UI
The new Llama.cpp web UI presents a simple interface with key information like the model name and total context available. Users can also add files as context.
Llama.cpp Web UI Features:
- Model Name and Context: Clearly displayed.
- File Upload: Ability to add PDF and text files as context.
- Customization: Extensive control over generation parameters:
- Temperature
- Top K
- Top P
- Penalty settings
- Reasoning Effort: While not available for reasoning models in this initial version, the presenter highlights its importance for models like GPT-OSS.
- Multimodal Support: Can attach images and audio files if the model supports them.
- In-context Retrieval: Parses files and includes them in the model's context, distinct from RAG.
- Parallel Conversations: Supports multiple simultaneous chat sessions.
- Structured Outputs: Allows defining custom JSON schemas for model responses.
Ollama Web UI (for comparison):
- Similar Interface: Visually comparable to the Llama.cpp UI.
- System Message and Theme: Options for system message, and light/dark themes.
- Statistics: Some statistics are available.
- Limited Customization: Less granular control over generation parameters compared to Llama.cpp.
- Reasoning Effort: The presenter notes that Ollama does not offer direct control over "reasoning effort" for models like GPT-OSS, unlike Llama.cpp.
- Speed: While comparable in generation speed, Ollama does not directly provide tokens per second or total generated tokens statistics.
Performance and Context Handling
Speed Test (GPT-OSS 20 Billion):
- Llama.cpp: Achieved approximately 84 tokens per second, generating 615 tokens. The UI displays neat statistics.
- Ollama: Generation speed was comparable, but the UI did not directly provide tokens per second or total tokens generated.
Long Context Processing (20,000+ tokens):
- Llama.cpp: When processing a 20-22,000 token PDF file, the model generated around 3,000 tokens at approximately 55 tokens per second. The presenter noted significant CPU stress and fan noise during processing, indicating a heavy load. The generated response was described as "much more comprehensive."
- Ollama: Processing the same file took significantly longer (around 3 minutes 22 seconds for the chain of thought and response). The generated response was described as "really concise" and less comprehensive than the Llama.cpp version. Ollama also caused fans to kick in.
Advanced Features
- Parallel Conversations: The Llama.cpp UI allows running multiple chat sessions concurrently, leveraging Llama.cpp's ability to handle simultaneous users. This applies to both text and multimodal conversations.
- Structured Outputs: Users can define custom JSON schemas in developer settings. The model can then generate responses adhering to this schema, demonstrated with multimodal inputs. This is highlighted as a "extremely powerful" feature.
Conclusion and Recommendation
The presenter highly recommends the new Llama.cpp web UI over third-party solutions for users running models with Llama.cpp. Its performance, extensive customization options (especially for generation parameters and structured outputs), and efficient handling of context make it a superior choice. The UI is described as "a really great option for running local models."
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Diffusion Gemma: Google's First Open Diffusion Model
Prompt Engineering

Unitlab AI: Easiest Way To Label Datasets For Machine Learning
NeuralNine

Top New AI Agent Tools 2025 | Scalable Inference, Bio-Med Agents & More
ManuAGI - AutoGPT Tutorials

Running Google's Gemma LLMs in the browser with MediaPipe Web
Chrome for Developers

Compilers in the Age of LLMs — Yusuf Olokoba, Muna
AI Engineer

4.5 Haiku V/S GPT-5 Mini & GLM-4.6 : Which is the best SMALL & CHEAP CODING MODEL?
AICodeKing