Key Concepts
- Llama CPP: A library for running large language models (LLMs) locally, even on computers without powerful GPUs.
- Fake OpenAI Server: A server spun up using Llama CPP that mimics the OpenAI API, allowing you to use OpenAI's Python library with local LLMs.
- Quantization: A technique to reduce the size and computational requirements of LLMs, making them runnable on less powerful hardware.
- Function Calling: A technique that allows LLMs to call external functions or tools to perform specific tasks.
- Multimodal Models: LLMs that can process multiple types of data, such as text and images.
- Streamlit: A Python library for creating interactive web applications.
- Instructor: A Python library that simplifies function calling by allowing you to define response models that automatically extract data from LLM outputs.
- ChatML & Functionary: Different chat formats used by LLMs. Functionary is specifically designed for function calling.
Setting up Llama CPP
- Cloning the Repository: Use
git clone [Llama CPP repository URL]to download the Llama CPP source code to your local machine. - Building the Library: Navigate to the Llama CPP directory using
cd llama.cppand runmaketo compile the library. Windows users are advised to use Windows Subsystem for Linux (WSL). - Installing Python Libraries: Use
pip install openai llama-cpp-python pydantic instructor streamlitto install the necessary Python packages.
Starting the Fake OpenAI Server
-
Basic Server Startup: Use the command
python -m llama_cpp.server --model [path to model]to start the server. Replace[path to model]with the actual path to your downloaded LLM model (e.g., a Mistral gguf file). -
GPU Acceleration: To utilize a GPU, add the flag
--n_gpu -1to the startup command. This offloads as many layers as possible to the GPU for faster performance. -
Multiple Models with Config File: Create a
config.jsonfile to define multiple models, their aliases, chat formats, and other settings. Start the server usingpython -m llama_cpp.server --config_file config.json.- Example
config.jsonstructure:
{ "models": [ { "model": "[path to model 1]", "alias": "mistral", "chat_format": "chatml", "gpu_layers": -1 }, { "model": "[path to model 2]", "alias": "mixtral", "chat_format": "chatml", "gpu_layers": -1 }, { "model": "[path to model 1]", "alias": "mistral-function-calling", "chat_format": "functionary", "gpu_layers": -1 } ] } - Example
Interacting with the Server using Python
-
Importing OpenAI Library: Use
from openai import OpenAIto import the OpenAI library. -
Creating a Client: Instantiate an OpenAI client, providing a dummy API key and the base URL of your local server (e.g.,
http://localhost:8000/v1).client = OpenAI( api_key="gibberish", # Dummy API key base_url="http://localhost:8000/v1" ) -
Performing Chat Completion: Use
client.chat.completions.create()to send prompts to the server. Specify the model alias and a list of messages.response = client.chat.completions.create( model="mistral", messages=[ {"role": "user", "content": "What is ROI in reference to finance?"} ] ) -
Extracting the Response: Access the LLM's response using
response.choices[0].message.content. -
Streaming Responses: To stream the response, set
stream=Truein thecreate()method. Iterate through theresponseobject and extract the content from each chunk usingchunk.choices[0].delta.content.response = client.chat.completions.create( model="mistral", messages=[ {"role": "user", "content": "What is ROI in reference to finance?"} ], stream=True ) for chunk in response: if chunk.choices[0].delta.content is not None: print(chunk.choices[0].delta.content, end="", flush=True)
Building a Streamlit App
-
Importing Streamlit: Use
import streamlit as stto import the Streamlit library. -
Creating a Title: Use
st.title()to set the title of the app. -
Creating a Chat Input: Use
st.chat_input()to create a chat input field for users to enter prompts. -
Rendering User Messages: Use
st.chat_message("user").markdown()to display user messages in the chat interface. -
Rendering AI Responses: Use
st.chat_message("ai")to create a chat message container for the AI's response. Usest.empty()to create an empty placeholder within the container, and update it with the streamed response chunks.if prompt: st.chat_message("user").markdown(prompt) with st.chat_message("ai"): message_placeholder = st.empty() completed_message = "" for chunk in response: if chunk.choices[0].delta.content is not None: completed_message += chunk.choices[0].delta.content message_placeholder.markdown(completed_message)
Implementing Function Calling
-
Importing Instructor: Use
import instructorto import the Instructor library. -
Patching the Client: Use
client = instructor.patch(client)to patch the OpenAI client, enabling function calling. -
Defining a Response Model: Create a Pydantic model to define the structure of the data you want to extract from the LLM's response.
from pydantic import BaseModel class ResponseModel(BaseModel): ticker: str days: int -
Specifying the Response Model: Set the
response_modelparameter in theclient.chat.completions.create()method to your defined response model. -
Using the Extracted Data: Access the extracted data from the
responseobject using the model's attributes (e.g.,response.ticker,response.days). -
Functionary Chat Format: Start the Llama CPP server with the
--chat_format functionaryflag to enable function calling.
Using Multiple Models Simultaneously
- Configuring Multiple Models: Define multiple models in the
config.jsonfile, specifying their aliases and chat formats. - Selecting the Model: Specify the desired model alias in the
client.chat.completions.create()method. - Chaining LLM Calls: Make multiple calls to the LLM, using different models for different tasks. For example, use one model to extract data and another model to summarize it.
Using Multimodal Models (Llava)
-
Starting the Server with Llava: Start the Llama CPP server with the path to the Llava model, the clip model path, and the
--chat_format lava_1_5flag.python -m llama_cpp.server --model models/llava-1.5-7b-q4/llava-1.5-7b-q4.gguf --clip_model_path models/mmproj-model.gguf --n_gpu -1 --chat_format lava_1_5 -
Formatting the Prompt: When using Llava, the
messagesparameter inclient.chat.completions.create()should be a list of dictionaries, where each dictionary represents a part of the prompt. Usetype: "image_url"to specify an image URL andtype: "text"to specify text.response = client.chat.completions.create( model="llava", messages=[ { "role": "user", "content": [ {"type": "image_url", "image_url": {"url": image_url}}, {"type": "text", "text": prompt} ] } ], stream=True )
Conclusion
The video demonstrates how to build a local AI server using Llama CPP, enabling you to run powerful LLMs on your own computer without relying on external APIs. It covers setting up the server, interacting with it using Python, building a Streamlit app, implementing function calling, using multiple models simultaneously, and working with multimodal models. By following these steps, you can create a versatile AI platform for automating finance tasks and exploring various AI applications.
AI summaries can miss context or contain errors. Check important details against the original video.