From Copilot to Colleague: Trustworthy Agents for High-Stakes - Joel Hron, CTO Thomson Reuters

AI EngineerAbout 5 min readJul 24, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • Agentic AI: AI systems that produce output, make judgments, and decisions, moving beyond mere helpfulness.
  • Agency Dials: Levers (autonomy, context, memory, coordination) to tune the level of agency in AI systems based on use case and risk tolerance.
  • Evals (Evaluations): The process of assessing the accuracy and reliability of AI systems, particularly challenging due to variability in human judgment and the complexity of agentic systems.
  • Legacy Applications: Existing software systems with domain-specific logic that can be decomposed and leveraged as tools for AI agents.
  • MVP (Minimum Viable Product): A development approach that focuses on building the smallest, most valuable piece of code, which the speaker argues can be limiting in the context of agentic AI.

Introduction

The speaker, Joel Harren from Thomson Reuters (TR), discusses the evolution of AI assistants from being merely "helpful" to being "productive," capable of making judgments and decisions. He emphasizes the unique challenges of deploying such systems in high-stakes environments like law, tax, and risk management, where accuracy is paramount.

Thomson Reuters Context

  • TR is a long-standing company (over 100 years) with a significant presence in legal, tax, compliance, audit, and risk industries.
  • Customer base includes 97% of the top 100 US law firms, 99% of the Fortune 100, and the top 100 US CPA firms.
  • Key assets:
    • 4,500 domain experts (reportedly the highest employer of lawyers globally).
    • 1.5+ terabytes of proprietary content.
  • Investments:
    • Over $3 billion in acquisitions in recent years.
    • Applied research lab with 200+ scientists and engineers.
    • Over $200 million annually in AI product development.

The Shift to Agentic AI

  • Referencing Y Combinator's call to "build law firms of agents," the speaker highlights the shift from helpful tools to AI systems that produce output and make decisions.
  • Agentic AI Definition: Not a binary state, but a spectrum of agency that can be tuned based on the use case.

Agency Dials Explained

The speaker outlines four key "agency dials" that can be adjusted:

  1. Autonomy:
    • Ranges from simple task execution (e.g., summarizing a document) to self-evolving workflows where the AI plans, executes, and replans its work.
  2. Context:
    • Evolves from using parametric knowledge to incorporating multiple knowledge sources (e.g., controlled knowledge and the web).
    • Advanced systems can permute data sources and update schemas for better future use.
  3. Memory:
    • Moves from stateless retrieval of context to shared and persistent memory across workflow steps and user sessions.
  4. Coordination:
    • Ranges from atomic task execution to delegation to tools and collaboration between multiple agent systems.

Lessons Learned

The speaker shares key lessons learned from two and a half years of building AI systems:

1. The Hard Truth About Evals

  • Challenge: Users expect determinism and certainty, which is difficult to achieve with AI systems.
  • Variability: Even highly trained domain experts show significant swings in accuracy (10%+) when evaluating the same data at different times.
  • Cost: Human evaluation is expensive, especially with frequent iterations.
  • Agentic System Challenges: Referencing source material becomes more difficult, agents can "drift," and building guardrails requires deep expert knowledge.
  • Approach: Focus on rigorous rubrics and using preference as a north star to guide improvement.

2. Legacy Applications: Handicap or Enabler?

  • Initial Approach: Early AI development often involved starting from scratch, leaving behind existing systems.
  • Agentic AI Advantage: Agents allow for the decomposition of legacy applications into reusable tools.
  • Benefit: Leveraging existing domain logic and infrastructure, which are unique assets.

3. Rethinking the MVP

  • Pitfall: Over-indexing on "minimal" in MVP development can lead to chasing rabbit holes.
  • Alternative Approach: Build the entire system first to understand which components require optimization.
  • Mindset Shift: Focus on building the whole product and learning from it, rather than starting with a small component.

Demo Examples

The speaker presents two demo examples:

1. Tax Use Case

  • Functionality: End-to-end generation of tax returns from source documents (W2, 1099, etc.).
  • AI Application:
    • Extracts data from documents.
    • Maps data to fields in a tax engine.
    • Applies tax laws and conditions.
  • Key Takeaways:
    • Leverages existing tax engine as a tool.
    • Uses a validation engine to inspect errors and improve accuracy.
    • Demonstrates how legacy systems can be decomposed and revitalized.

2. Legal Research Use Case

  • Functionality: AI assistant for legal research, preparing for litigation.
  • Data Source: 1.5+ terabytes of proprietary legal content.
  • AI Application:
    • Uses tools from a litigation research product (searching, fetching documents, comparing citations, validating citations).
    • Reasons to an appropriate answer to legal research questions.
  • Key Takeaways:
    • Model writes notes to itself during the research process.
    • Generates a final report summarizing findings.
    • Provides links to hard citations in the product (cases, statutes).
    • Flags the risk associated with citations.

Conclusion

  • Begin with the whole problem in mind when building agentic systems.
  • Treat agency as a lever to be dialed up or down based on risk and use case.
  • Leverage agents to bring life back to old systems by decomposing them into reusable components.
  • Focus on human-in-the-loop evaluation using domain experts.
  • Identify and leverage unique assets (e.g., domain expertise, proprietary content) to create differentiation.

Q&A

  • Question: How would you describe the cyber security postures which are mandated by CISA and government recently such as LLM firewall or LLM guard rails or automated agents for scanning vulnerabilities or any SCM security posture me management? How would you define the cyber security posture for the entire architecture?
  • Answer: TR is heavily focused on compliance with standards like Fed Ramp and conforming to the latest standards like the ISO standard. The space is quickly evolving, and TR is adaptable to it.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.