Key Concepts
- Unstructured content platform
- Generative AI
- Data extraction
- Metadata
- Intelligent Document Processing (IDP)
- Agentic approach
- Directed graph
- LLM as a judge
- Eval sets
- Agent API
Box's AI Journey: From Data Extraction to Agentic Architecture
Background on Box and AI Strategy
Box is an unstructured content platform catering to large enterprises, with a focus on secure AI deployments. Their AI strategy is platform-level, building AI services on top of their existing global infrastructure, which manages over an exabyte of data and hundreds of billions of files.
Initial Foray into AI (2023)
In 2023, Box began exploring generative AI, focusing on features like QA across documents, data extraction, and AI-powered workflows. The initial focus was on data extraction – converting unstructured data into structured data for enterprise use.
The Promise of Early LLMs
Early experiments with GPT-2 and GPT-3 models showed great promise. By pre-processing documents with OCR and using simple prompts, they could extract data with accuracy exceeding specialized, custom-built ML models. This was flexible and performed well across various data types.
Hitting the Limits: Complexity and Accuracy Challenges
As customers demanded more complex data extraction tasks (e.g., 300-page lease documents with 300 fields, risk assessments), the limitations of single-shot AI calls became apparent. Challenges included:
- Complex Documents: Handling large, complex documents with numerous fields.
- OCR Issues: OCR inaccuracies leading to incorrect data extraction.
- Language Diversity: Difficulties in processing multiple languages.
- Attention Span: LLMs struggling to maintain accuracy when extracting a large number of complex fields.
- Accuracy Measurement: Lack of reliable confidence scores from LLMs. Using LLMs as judges provided feedback but didn't guarantee accuracy.
Customers demanded speed, affordability, and high accuracy, pushing Box to seek a more robust solution.
Embracing the Agentic Approach
Box turned to agentic approaches to overcome these limitations. This involved using AI agents with:
- Defined Instructions and Objectives: Clear goals for the agent.
- Model Background and Tools: Access to relevant information and tools.
- Secure Access: Ensuring data security and privacy.
- Memory: Retaining information for context and progress.
- Directed Graph: Orchestrating tasks in a specific sequence.
This approach was initially met with skepticism from engineers who favored traditional methods like improving OCR or fine-tuning ML models.
Agentic Architecture for Data Extraction: A Step-by-Step Process
Box implemented an agentic architecture for data extraction, resembling a LangGraph-style system. The process involves:
- Field Preparation: Preparing the fields to be extracted.
- Field Grouping: Grouping related fields to maintain context (e.g., customer information and addresses).
- Multiple Queries: Performing multiple queries on the document.
- Result Checking: Using tools to verify and double-check the extracted data, including OCR and image analysis.
- Model Voting: Using multiple models from different vendors and having them "vote" on the correct answer.
- LLM as a Judge: Using an LLM to provide feedback and prompt the agent to retry if necessary.
This multi-step process improves accuracy, although it may increase processing time.
Expanding Agentic Capabilities: Deep Research
The agentic foundation enabled Box to develop more advanced features, such as deep research capabilities on user content, similar to what OpenAI or Gemini do on the internet. This involves a complex directed graph with steps for:
- Data Search: Searching for relevant data within the user's Box account.
- Relevance Checking: Verifying the relevance of the found data.
- Outline Creation: Generating an outline of the research topic.
- Plan Preparation: Developing a plan for conducting the research.
- Process Execution: Executing the research plan.
Lessons Learned and Key Takeaways
- Agentic Abstraction is Clean: The agentic architecture provides a clean abstraction layer for building intelligent workflows.
- Easy to Evolve: Agentic systems are easier to evolve and adapt to new challenges by modifying prompts or adding new nodes to the directed graph.
- AI-First Thinking: Encourage teams to adopt an AI-first mindset and explore agentic solutions.
- Build Agentic Architecture Early: If AI models can potentially solve a problem, build an agentic architecture early to leverage their capabilities.
Notable Quotes:
- "If it's plausible that an a set of AI models uh could help you solve that problem, then you should build this AI agentic architecture early."
API Availability
Box provides Agent APIs for customers to call upon these agents and provide arguments for specific tasks.
Evaluation Methods
Box evaluates its agents using:
- Eval Sets: Standard and challenging sets of evaluation data.
- LLM as a Judge: Using LLMs to assess the quality of the agent's output.
- User Feedback: Gathering feedback from users to identify areas for improvement.
Why Agents Over Fine-Tuning?
Box is currently "anti-fine-tuning" due to the challenges of maintaining consistency across multiple models (Gemini, Llama, OpenAI, Anthropic) and the rapid improvements in base model performance. They prefer using prompts, cached prompts, and agentic architectures.
Conclusion
Box's journey highlights the evolution from simple AI-powered data extraction to a sophisticated agentic architecture. By embracing agentic principles, Box has overcome the limitations of early LLMs and built a more robust, flexible, and scalable AI platform for its enterprise customers. The key takeaway is that an agentic approach, with its modularity and adaptability, is crucial for tackling complex AI challenges in real-world applications.
AI summaries can miss context or contain errors. Check important details against the original video.





