How to Build a Fake OpenAI Server (so you can automate finance stuff)

Nicholas RenotteAbout 5 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Llama CPP: A library for running large language models (LLMs) locally, even on computers without powerful GPUs.
  • Fake OpenAI Server: A server spun up using Llama CPP that mimics the OpenAI API, allowing you to use OpenAI's Python library with local LLMs.
  • Quantization: A technique to reduce the size and computational requirements of LLMs, making them runnable on less powerful hardware.
  • Function Calling: A technique that allows LLMs to call external functions or tools to perform specific tasks.
  • Multimodal Models: LLMs that can process multiple types of data, such as text and images.
  • Streamlit: A Python library for creating interactive web applications.
  • Instructor: A Python library that simplifies function calling by allowing you to define response models that automatically extract data from LLM outputs.
  • ChatML & Functionary: Different chat formats used by LLMs. Functionary is specifically designed for function calling.

Setting up Llama CPP

  1. Cloning the Repository: Use git clone [Llama CPP repository URL] to download the Llama CPP source code to your local machine.
  2. Building the Library: Navigate to the Llama CPP directory using cd llama.cpp and run make to compile the library. Windows users are advised to use Windows Subsystem for Linux (WSL).
  3. Installing Python Libraries: Use pip install openai llama-cpp-python pydantic instructor streamlit to install the necessary Python packages.

Starting the Fake OpenAI Server

  1. Basic Server Startup: Use the command python -m llama_cpp.server --model [path to model] to start the server. Replace [path to model] with the actual path to your downloaded LLM model (e.g., a Mistral gguf file).

  2. GPU Acceleration: To utilize a GPU, add the flag --n_gpu -1 to the startup command. This offloads as many layers as possible to the GPU for faster performance.

  3. Multiple Models with Config File: Create a config.json file to define multiple models, their aliases, chat formats, and other settings. Start the server using python -m llama_cpp.server --config_file config.json.

    • Example config.json structure:
    {
      "models": [
        {
          "model": "[path to model 1]",
          "alias": "mistral",
          "chat_format": "chatml",
          "gpu_layers": -1
        },
        {
          "model": "[path to model 2]",
          "alias": "mixtral",
          "chat_format": "chatml",
          "gpu_layers": -1
        },
        {
          "model": "[path to model 1]",
          "alias": "mistral-function-calling",
          "chat_format": "functionary",
          "gpu_layers": -1
        }
      ]
    }
    

Interacting with the Server using Python

  1. Importing OpenAI Library: Use from openai import OpenAI to import the OpenAI library.

  2. Creating a Client: Instantiate an OpenAI client, providing a dummy API key and the base URL of your local server (e.g., http://localhost:8000/v1).

    client = OpenAI(
        api_key="gibberish",  # Dummy API key
        base_url="http://localhost:8000/v1"
    )
    
  3. Performing Chat Completion: Use client.chat.completions.create() to send prompts to the server. Specify the model alias and a list of messages.

    response = client.chat.completions.create(
        model="mistral",
        messages=[
            {"role": "user", "content": "What is ROI in reference to finance?"}
        ]
    )
    
  4. Extracting the Response: Access the LLM's response using response.choices[0].message.content.

  5. Streaming Responses: To stream the response, set stream=True in the create() method. Iterate through the response object and extract the content from each chunk using chunk.choices[0].delta.content.

    response = client.chat.completions.create(
        model="mistral",
        messages=[
            {"role": "user", "content": "What is ROI in reference to finance?"}
        ],
        stream=True
    )
    
    for chunk in response:
        if chunk.choices[0].delta.content is not None:
            print(chunk.choices[0].delta.content, end="", flush=True)
    

Building a Streamlit App

  1. Importing Streamlit: Use import streamlit as st to import the Streamlit library.

  2. Creating a Title: Use st.title() to set the title of the app.

  3. Creating a Chat Input: Use st.chat_input() to create a chat input field for users to enter prompts.

  4. Rendering User Messages: Use st.chat_message("user").markdown() to display user messages in the chat interface.

  5. Rendering AI Responses: Use st.chat_message("ai") to create a chat message container for the AI's response. Use st.empty() to create an empty placeholder within the container, and update it with the streamed response chunks.

    if prompt:
        st.chat_message("user").markdown(prompt)
    
        with st.chat_message("ai"):
            message_placeholder = st.empty()
            completed_message = ""
            for chunk in response:
                if chunk.choices[0].delta.content is not None:
                    completed_message += chunk.choices[0].delta.content
                    message_placeholder.markdown(completed_message)
    

Implementing Function Calling

  1. Importing Instructor: Use import instructor to import the Instructor library.

  2. Patching the Client: Use client = instructor.patch(client) to patch the OpenAI client, enabling function calling.

  3. Defining a Response Model: Create a Pydantic model to define the structure of the data you want to extract from the LLM's response.

    from pydantic import BaseModel
    
    class ResponseModel(BaseModel):
        ticker: str
        days: int
    
  4. Specifying the Response Model: Set the response_model parameter in the client.chat.completions.create() method to your defined response model.

  5. Using the Extracted Data: Access the extracted data from the response object using the model's attributes (e.g., response.ticker, response.days).

  6. Functionary Chat Format: Start the Llama CPP server with the --chat_format functionary flag to enable function calling.

Using Multiple Models Simultaneously

  1. Configuring Multiple Models: Define multiple models in the config.json file, specifying their aliases and chat formats.
  2. Selecting the Model: Specify the desired model alias in the client.chat.completions.create() method.
  3. Chaining LLM Calls: Make multiple calls to the LLM, using different models for different tasks. For example, use one model to extract data and another model to summarize it.

Using Multimodal Models (Llava)

  1. Starting the Server with Llava: Start the Llama CPP server with the path to the Llava model, the clip model path, and the --chat_format lava_1_5 flag.

    python -m llama_cpp.server --model models/llava-1.5-7b-q4/llava-1.5-7b-q4.gguf --clip_model_path models/mmproj-model.gguf --n_gpu -1 --chat_format lava_1_5
    
  2. Formatting the Prompt: When using Llava, the messages parameter in client.chat.completions.create() should be a list of dictionaries, where each dictionary represents a part of the prompt. Use type: "image_url" to specify an image URL and type: "text" to specify text.

    response = client.chat.completions.create(
        model="llava",
        messages=[
            {
                "role": "user",
                "content": [
                    {"type": "image_url", "image_url": {"url": image_url}},
                    {"type": "text", "text": prompt}
                ]
            }
        ],
        stream=True
    )
    

Conclusion

The video demonstrates how to build a local AI server using Llama CPP, enabling you to run powerful LLMs on your own computer without relying on external APIs. It covers setting up the server, interacting with it using Python, building a Streamlit app, implementing function calling, using multiple models simultaneously, and working with multimodal models. By following these steps, you can create a versatile AI platform for automating finance tasks and exploring various AI applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.