Key Concepts
- Workflows vs. Agents: Trade-offs between structured, predictable systems and dynamic, AI-driven systems.
- Enterprise-Aware Agents: AI agents that understand and operate within the specific context of a business.
- Golden Workflows: Predefined, optimal sequences of steps for completing specific tasks within an organization.
- Fine-tuning (SFT, RHF): Methods for training LLMs using labeled data (input/output pairs) or reward signals.
- Dynamic Prompting with Search: Using a search engine to retrieve relevant examples and incorporate them into prompts for LLMs.
- Textual Similarity vs. Authoritativeness: Two key components of workflow search, focusing on semantic similarity and the credibility/relevance of the workflow source.
Workflows vs. Agents: A Detailed Comparison
The speaker addresses the core dilemma of choosing between workflows and agents for AI-driven enterprise solutions.
- Workflows: Defined as systems where LLMs and tools are orchestrated through predefined code paths.
- Represented through imperative code or declarative graphs.
- Offer structure and predictability. Running a workflow today will yield similar results tomorrow.
- Analogized to "Toyota" - reliable and predictable.
- Suitable for automating repetitive tasks and encoding existing best practices.
- Lower cost and latency due to the absence of LLM decision-making overhead.
- Easier to debug due to the explicit code or graph structure.
- Humans are in control, allowing for tweaks and engineering to ensure task completion.
- Agents: Defined as systems where LLMs dynamically direct their own processes to achieve a task.
- The core agent loop involves planning, executing actions, and iterating based on environmental feedback.
- Analogized to "Tesla" - innovative but less predictable.
- Suitable for researching unsolved problems and leveraging advanced LLM capabilities.
- Higher cost and latency due to LLM's decision-making process.
- Less logic to maintain, but can be prone to errors ("taking the wrong exit").
The decision to use workflows or agents depends on the current state of LLMs. What doesn't work in an agentic loop now might work in a few months with a new model.
The Synergy Between Workflows and Agents
The speaker proposes a paradigm shift: instead of choosing between workflows and agents, leverage their synergies.
- Agent as Workflow Generator: An agent takes a task and generates a workflow to achieve it. The agent's execution trace becomes a workflow.
- Workflows for Agent Evaluation: Use a library of "golden workflows" to evaluate agents. Assess whether the agent takes the correct steps, not just the final result.
- Workflows for Agent Training: Train agents using golden workflows. This allows agents to execute known tasks precisely while using their reasoning to compose workflows for new tasks.
- Agents for Workflow Building: Use agents to generate workflows from natural language descriptions. Users can then edit and refine the proposed workflow. Glean agents work this way.
- Agents for Workflow Discovery: Deploy agents and save successful execution traces as new workflows. This creates a feedback loop for improving agent performance.
Enterprise-Aware AGI: Onboarding the Super-Intelligent Employee
The speaker discusses the need for enterprise-aware AGI, even in a world with advanced AI.
- AGI is like a super-intelligent new employee who needs onboarding to understand company-specific practices.
- Enterprise-aware AGI is fully onboarded, intelligent, and knows the company's ways of doing things.
- There are many acceptable ways to achieve a task, but a gap exists between acceptable and great outputs.
- Example: Competitor analysis. AGI can do basic research, but it needs to follow company protocols and address key metrics.
Training Agents with Tasks and Golden Workflows
The speaker outlines two main approaches for training agents using task and golden workflow data.
- Fine-tuning:
- Supervised Fine-tuning (SFT): Train the model to mimic the expected output for a given input.
- Reinforcement Learning from Human Feedback (RHF): Optimize the LLM based on ratings or rewards for different workflows.
- Pros: Learns well with large datasets, generalizes across tasks, and combines workflows.
- Cons: Requires forking from frontier LLMs, retraining for any data changes, and limited flexibility for personalization.
- Dynamic Prompting with Search:
- Build a search engine for tasks, using the task-to-golden-workflow data.
- At runtime, find similar tasks in the training data and feed the corresponding workflows to the LLM as examples.
- Offers a spectrum of determinism and creativity.
- When there's no matching workflow, the LLM uses its creativity.
- When there's a high-confidence match, the LLM provides a workflow similar to the training data.
- Example: Competitor analysis. The system retrieves workflows for analyzing competitors and finding customer calls, then the LLM composes a workflow to read customer calls, extract competitors, and run analysis.
Fine-tuning vs. Dynamic Prompting: A Detailed Comparison
The speaker compares fine-tuning and dynamic prompting with search.
- Fine-tuning is strong with large datasets for generalization. Dynamic prompting with search is more flexible and interpretable.
- Fine-tuning is good for generalized behaviors where ground truth labels don't change. Dynamic prompting with search is better for customized behaviors and rapidly changing requirements.
- Analogy: Fine-tuning is like building customized hardware, while dynamic prompting is like writing software.
Building Workflow Search: Textual Similarity and Authoritativeness
The speaker discusses how to build a workflow search engine.
- Two main components:
- Textual Similarity: Find similar-sounding tasks using hybrid search (lexical, vector embeddings, reranking, late interaction).
- Authoritativeness: Determine the credibility and relevance of the workflow source.
- Pure text similarity is not enough in enterprise settings.
- Authoritativeness requires a knowledge graph. Factors include the creator's relationship to the user, success rate, and mentions on platforms like Slack.
- Recommendation system techniques apply to workflow search.
- Authoritativeness signals are hard to encode directly into an LLM.
Conclusion
The speaker concludes by reiterating the key takeaways: workflows offer determinism, agents offer open-endedness, and the synergy between them is powerful. Fine-tuning is suitable for generalized behaviors, while dynamic prompting with search is better for personalized behaviors.
AI summaries can miss context or contain errors. Check important details against the original video.