THE SUMMARYAI-generated
Key Concepts
- Language Model (LLM): A machine learning model predicting the next word in a sequence.
- Pre-training & Post-training: Two-stage LLM training: pre-training on large datasets, followed by post-training with instruction following and human feedback.
- Prompting: Crafting input text for LLMs to elicit desired outputs.
- Hallucination: LLMs generating incorrect or nonsensical information.
- Retrieval Augmented Generation (RAG): Enhancing LLMs with external knowledge retrieval to reduce hallucination and improve accuracy.
- Tool Usage/Function Calling: Enabling LLMs to interact with external tools and APIs.
- Agentic Language Model: LLMs that can interact with their environment, reason, and take actions.
- ReAct (Reason and Act): A paradigm for agentic LLMs that combines reasoning and action execution.
- Planning, Reflection, Multi-Agent Collaboration: Design patterns for agentic LLMs.
Language Model Overview
- A language model predicts the next word given input text. For example, given "The students open their," the model might predict "books" or "laptops" with high probability.
- LLMs are trained in two stages:
- Pre-training: Trained on vast amounts of text data from the internet, books, etc., using next-token prediction.
- Post-training: Fine-tuned with instruction following and reinforcement learning with human feedback (RLHF) to improve usability and align with human preferences.
- Instruction following training involves datasets with specific instructions and expected outputs. The model is trained to generate the output based on the given instruction.
- RLHF aligns the model's behavior with human preferences using reward schemes.
- Trained LLMs possess significant world knowledge and are used in applications like AI coding assistants, domain-specific copilots, and conversational interfaces (e.g., ChatGPT).
- LLMs can be accessed via cloud-based APIs or hosted locally, depending on model size and computational constraints.
Using Language Models via API Calls
- Interacting with LLMs typically involves preparing free-form text prompts (instructions or questions) and sending them to the model via an API call.
- The model generates an output, which is then parsed by the surrounding software.
- Prompt engineering is crucial for effective LLM usage.
Prompt Engineering Best Practices
- Clear and Descriptive Instructions: Provide detailed instructions to guide the model.
- Few-Shot Examples: Include example input-output pairs to demonstrate the desired style and format.
- Relevant Context and References: Supply relevant information to ground the model's responses and reduce hallucination.
- Chain of Thought (COT): Encourage the model to reason step-by-step before answering.
- Decompose Complex Tasks: Break down complex requests into smaller, sequential prompts.
- Systematic Tracing and Logging: Maintain logs for debugging and auditing.
- Automated Evaluation: Implement automated evaluation using question-answer pairs to track progress.
Examples of Prompting Techniques
- Clear Instructions: Instead of "Translate to French," use "Translate the following English text to French: [text]".
- Few-Shot Examples: Provide examples like "English: Hello, French: Bonjour" before asking the model to translate a new sentence.
- Contextual Information: "Answer the question based on the following article: [article text]. Question: [question]".
- Chain of Thought: "First, work out your own solution to the problem. Then, compare your solution to the student's solution. Problem: [problem], Student Solution: [solution]".
- Decomposition: Break a complex task into multiple stages, prepending the output of each stage to the next prompt.
Addressing Common Language Model Limitations
- Hallucination: Generating incorrect information.
- Knowledge Cut-off: Lack of up-to-date information.
- Lack of Attribution: Inability to cite sources.
- Data Privacy: Models trained on public data lack access to proprietary data.
- Limited Context Length: Constraints on the amount of input text the model can process.
Retrieval Augmented Generation (RAG)
- RAG addresses limitations by integrating external knowledge retrieval into the LLM process.
- Process:
- Pre-index a knowledge base (e.g., documents, web pages) by chunking the text and converting it into embeddings using an embedding model.
- Store the embeddings in a vector database.
- When a query arrives, convert it into an embedding.
- Perform a nearest neighbor search in the vector database to retrieve relevant text chunks.
- Include the retrieved text chunks in the prompt to the LLM.
- RAG reduces hallucination, provides attribution, enables the use of proprietary data, and optimizes context length.
- Alternative RAG methods include knowledge graph-based approaches (Graph RAG).
Tool Usage (Function Calling)
- Tool usage allows LLMs to interact with external tools and APIs to access real-time information or perform computations.
- Process:
- The LLM is instructed to generate specific output formats (e.g., API calls) for certain types of queries.
- The surrounding software parses the LLM's output and executes the corresponding API call or function.
- The results are fed back to the LLM, which then generates a more human-friendly response.
- Example: A chatbot uses a weather API to provide weather information.
Agentic Language Models
- Agentic LLMs can interact with their environment, reason, and take actions.
- Definition: An agentic LLM can interact with the environment by generating tool usage requests or retrieval requests. The environment provides observations back to the LLM.
- ReAct (Reason and Act): Agentic LLMs reason about tasks and then take actions to accomplish them.
- Reasoning: Breaking down tasks, making plans, and generating actions.
- Action: Using tools, APIs, or code execution to interact with the external world.
- Memory: Storing conversational history and observations to inform future actions.
Example: Customer Support AI Agent
- Customer asks: "Can I get a refund for product X?"
- The agentic system breaks down the request into actions:
- Check the refund policy.
- Check the customer information.
- Check the product details.
- The LLM generates API calls to retrieve the necessary information.
- The LLM draws a conclusion based on the retrieved information and sends a request to a follow-up system or prepares a response draft.
Agentic Language Model Workflow
- Agentic LLMs make iterative calls, reviewing documents or tasks and making external tool calls.
- Examples:
- Research agent: Performs web searches, summarizes information, and prepares reports.
- Software assistant agent: Investigates software bugs, collects relevant code, proposes fixes, and tests them in a sandbox environment.
Benefits of Agentic Language Models
- Agentic patterns enable LLMs to perform more complex tasks than simple input-output interactions.
- They push the boundaries of what AI agents can do in various domains.
Design Patterns for Agentic Language Models
- Planning: Asking the model to break down tasks into simpler steps.
- Reflection: Having the model critique its own output to improve it.
- Tool Usage: Utilizing external tools and APIs to access information or perform actions.
- Multi-Agent Collaboration: Dividing tasks among multiple specialized agents.
Examples of Design Patterns
- Reflection: Ask the model to provide constructive feedback on a code snippet, then use that feedback to refactor the code.
- Tool Usage: Ask the model to generate API calls to perform a specific task.
- Multi-Agent Collaboration: Create separate agents for climate control, lighting control, etc., in a smart home automation system.
Conclusion
Agentic language model usage is a progression of existing language model techniques. By incorporating planning, reflection, tool usage, and multi-agent collaboration, LLMs can perform more complex tasks and interact with the external world in meaningful ways.
AI summaries can miss context or contain errors. Check important details against the original video.





