Key Concepts
- Open-source AI models: AI models with publicly available code and weights, allowing for local deployment and customization.
- Proprietary AI models: AI models offered as cloud services, requiring data to be sent to external providers and typically charging per token.
- Inference: The process of using a pre-trained AI model to generate outputs based on input prompts.
- Parameters: The weights and biases within a neural network that are adjusted during training and determine the model's behavior.
- Distillation: A technique to train a smaller model (student) to emulate the behavior of a larger model (teacher).
- Quantization: A technique to reduce the memory footprint of a model by using lower precision numbers to represent its parameters.
- Prompt Engineering: The process of designing effective prompts to elicit desired responses from AI models.
- Malicious Prompt Injection: A security vulnerability where crafted prompts can cause AI models to execute unintended or harmful actions.
- Chatbot Arena (LM Arena): A leaderboard that ranks AI models based on blind testing and user voting.
Open-Source AI Models: An Overview
The class focuses on open-source AI models as an alternative to proprietary models, particularly for use cases where data privacy and cost are concerns.
- Customer Problem: A healthcare provider wants to use GenAI to draft responses to patient portal messages but cannot send private health information (PHI) to cloud providers like OpenAI or Google and wants to avoid per-token costs.
- Proprietary Models:
- Run in cloud services outside the user's firewall.
- Charge per token.
- Generally offer the best performance.
- Examples: OpenAI's GPT models, Google's Gemini, Anthropic's Claude.
- Open-Source Models:
- Can be downloaded and run within the user's environment.
- Require hosting infrastructure (hardware or cloud tenant).
- May offer slightly lower performance compared to proprietary models.
- Examples: Llama (Meta), DeepSeek, Quen (Alibaba), Gemma (Google).
- License Types:
- MIT and Apache 2.0 licenses are considered the most open, with no restrictions on use.
- Some licenses, like the Gemma license, may have slight restrictions, such as forbidding dangerous uses.
- Model Size and Inference:
- Models with single-digit billion parameters can run on laptops.
- Models with hundreds of billions or trillions of parameters require more powerful hardware.
- Inference is the process of using a trained model to generate outputs.
Using Open-Source Models: A Practical Demonstration
Marta demonstrates how to use open-source models through LM Studio and a Colab notebook, showcasing their ease of use and potential for solving real-world problems.
- LM Studio: A software application that allows users to run LLMs locally and provides a chat-like UI.
- Hugging Face: A large online repository of open-source LLMs.
- Example: Marta uses Llama 3.23B (3 billion parameters) in LM Studio to answer the prompt "Tell me a joke."
- Colab Notebook: A Jupyter notebook hosted online, providing access to a GPU runtime for running LLMs.
- Patient Portal Assistant:
- The notebook implements a patient portal assistant that drafts responses to patient messages.
- First, the assistant is implemented using OpenAI's API.
- Then, the OpenAI API is swapped with a locally deployed LLM (Llama 3B).
- Multiple LLMs:
- The notebook demonstrates how to combine multiple small LLMs, each with a specialty, to achieve a better solution.
- Tasks: Content moderation (DeepSeek), prioritization (Gemma), and translation (IA23).
- Resource Utilization:
- The notebook shows how to monitor GPU memory utilization when loading and running models.
Developer Perspective: Coding with Open-Source Models
Don demonstrates how easy it is to use open-source models from a coding perspective, comparing it to using proprietary models like OpenAI.
- OpenAI API: Don uses the OpenAI API to generate a bedtime story about a unicorn.
- Ollama: A tool for deploying and running LLMs locally.
- REST Service: Ollama hosts the model as a REST service that can be called using HTTP requests.
- Python Library: Don uses a Python library to call the REST service and get an answer to the question "Why is the sky blue?"
- Ease of Use: From a coding point of view, open-source models are just as easy to use as proprietary models.
Model Comparison and Performance
The class discusses how open-source models compare to proprietary models in terms of performance and provides a framework for evaluating their suitability for different use cases.
- Chatbot Arena (LM Arena): A leaderboard that ranks AI models based on blind testing and user voting.
- Ranking: Proprietary models generally rank higher than open-source models, but open-source models are catching up.
- Blind Testing: Users submit prompts, and the leaderboard randomly picks two models to answer the question. Users then vote on which answer is better.
- Performance Delay: Open-source models may be as good as proprietary models that were released six to eight months prior.
- Use Cases:
- Open-source models may be preferred when data privacy or cost is a concern.
- Open-source models can be used for surrounding tasks like content moderation, prioritization, and translation.
- Parameters: The weights in a neural network that are used in matrix math to take the input and determine the output.
Security Considerations: Malicious Prompt Injection
Marta discusses the security risks associated with LLMs, particularly malicious prompt injection, and provides recommendations for mitigating these risks.
- Code Injection: Malicious prompts can instruct the LLM to write code that opens the door for an attack.
- Denial of Service (DoS): Malicious prompts can cause the LLM to generate a large amount of output, overwhelming the service or resulting in a large bill.
- Prompt Engineering Limitations: Prompt engineering alone cannot guarantee security.
- Security Recommendations:
- Never directly execute code written by an LLM without reading and understanding it first.
- Check inputs and outputs for suspicious content.
- Sanitize outputs to ensure they are in the expected format.
- Manage permissions and prevent LLM-based applications from running code automatically or accessing sensitive information.
Take-Home Exercises
Marta outlines two take-home exercises that allow participants to further explore the concepts discussed in the class.
- Notebook Access: The notebooks used in the class will be linked in the chat.
- Colab Compatibility: The notebooks can be run in Colab without requiring special resources.
- OpenAI API Key: The cells that require an OpenAI API key are labeled and can be skipped.
- Model Comparison Notebook: A notebook that compares the performance of three different LLMs on the patient portal challenge.
Conclusion
Open-source AI models offer a viable alternative to proprietary models, particularly for use cases where data privacy and cost are concerns. While proprietary models generally offer better performance, open-source models are rapidly evolving and can be suitable for many applications. It is important to be aware of the security risks associated with LLMs and to take appropriate measures to mitigate these risks.
AI summaries can miss context or contain errors. Check important details against the original video.