Deep Research & AI Decision Frameworks: A Thanksgiving Project & Vision for Copilot
Key Concepts:
- LLM Council: A decision-making framework utilizing multiple Large Language Models (LLMs) with a designated chairman for synthesis.
- DxO (Data eXploration & Optimization): A framework employing specialized LLM roles (Lead Researcher, Critical Reviewer, Data Analyst) for robust analysis, particularly in high-stakes scenarios.
- Ensemble: A method of aggregating responses from multiple LLMs, anonymizing sources, and synthesizing a unified output.
- Metacognition: Thinking about thinking; in this context, tools to enhance and structure decision-making processes.
- Codex-max: A powerful LLM model noted for its speed and efficiency.
- Claude Opus 4.5: A high-performing LLM used for research and critical review.
- GitHub Codespaces: A cloud-based development environment accessible through GitHub.
- Windows 365: A cloud PC service providing access to a Windows environment and applications.
I. Project Overview & Development Environment
The speaker recounts building a personal application over the Thanksgiving weekend to explore the capabilities demonstrated by Karan, focusing on leveraging multiple LLMs for complex decision-making. The development environment consists of a portable Windows 365 instance, GitHub, and GitHub Codespaces – a “turtles all the way down” setup allowing for coding on the go. The speaker emphasizes a workflow of rapidly generating multiple draft branches in GitHub, accepting promising Pull Requests (PRs), and iterating on the code. Currently, the speaker is primarily utilizing Codex-max and Claude Opus 4.5, finding the “auto” model selection option to be efficient for token management. The application is deployed in South Central Canada.
II. Motivation: Joining the Copilot Team & Deep Research Enhancement
The core motivation behind the project was to demonstrate skills relevant to joining the Microsoft Copilot team, specifically in the area of deep research. The speaker aims to build upon existing deep research capabilities by implementing novel decision frameworks powered by multiple LLMs. The goal is to create a system capable of more nuanced and comprehensive analysis than relying on a single frontier model.
III. Implemented Decision Frameworks
The application implements three primary decision frameworks:
- LLM Council: Inspired by Ondrej Karpathy’s concept, this framework allows the user to select from a range of LLMs (GPT, Opus, Gemini, Kimi K2, Grok, etc.), designate a “chairman” model, and submit queries for synthesized responses. This leverages the diverse strengths of different models.
- DxO (Data eXploration & Optimization): Originally developed for healthcare applications, DxO assigns specific roles to different LLMs:
- Lead Researcher (Opus): Performs broad initial research.
- Critical Reviewer (GPT 5.1): Identifies methodological errors, biases (including recency bias), and weaknesses in the research.
- Data Analyst (Gemini/Kimi K2): Provides domain expertise and data-driven insights. The speaker notes that DxO outperformed single frontier models in previous healthcare applications.
- Ensemble: This framework utilizes all available models, anonymizes their responses (labeling them Alpha, Beta, Gamma), and synthesizes a single, unified response. This aims to mitigate individual model biases and provide a more balanced perspective.
The speaker also mentions extending these frameworks to applications like shopping and finance.
IV. Case Study: Selecting the All-Time Best Indian Test Cricket Team
A concrete example of the application’s use is selecting the all-time best Indian Test cricket team. The LLM Council framework was used to analyze player data and generate a team selection. The system identified areas of strong consensus (Gavaskar, Sehwag, Dravid, Tendulkar, Kohli, Kapil Dev, Ashwin, Bumrah) and highlighted key debates (e.g., inclusion of VVS Laxman). The debate process, as captured by the system, revealed the reasoning behind different model recommendations – for example, Phi-1 and Claude emphasized Laxman’s crisis management skills. The system also addressed the captaincy debate, ultimately selecting Kohli. The speaker highlights the streaming nature of the output, presenting the debate as a “chain of debate” rather than a “chain of thought.”
V. Technical Implementation & Results
The application features a streaming output, allowing users to observe the debate process between the LLMs. The DxO framework’s output demonstrates the exhaustive search conducted by the Lead Researcher, followed by the Critical Reviewer’s identification of potential biases (e.g., era bias when comparing players across generations). The speaker emphasizes the ability to account for factors like pitch difficulty and historical context. The entire project was built within a couple of hours and is continuously being refined.
VI. Future Vision & Metacognition
The speaker believes this work will be relevant to the future of Copilot and is actively interviewing for a junior product manager role. The core concept is the next generation of metacognition – providing tools to enhance and structure the decision-making process. The speaker envisions these frameworks and agentic systems being applied to a wide range of domains, including supply chain management, healthcare, and finance. The LLMs act as agents, but the ultimate metacognitive control remains with the user.
Notable Quote:
“This is to me the next generation of metacognition. So if you think about these decision frameworks, you have all these agents, you're working with agents, but the metacognition is still us. And this is tools for metacognition is how I think about it.” – The Speaker.
This project demonstrates a practical application of multiple LLMs in a structured decision-making process, showcasing the potential for enhanced analysis and more informed outcomes. The speaker’s work highlights the importance of combining the strengths of different models and providing tools for users to understand and control the reasoning behind AI-driven recommendations.
AI summaries can miss context or contain errors. Check important details against the original video.





