Spotlight on Databricks: Driving data intelligence with AI

AnthropicAbout 5 min readAug 1, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • General Intelligence vs. Data Intelligence: The distinction between broad AI capabilities and AI specifically tailored to an organization's data.
  • Deterministic vs. Probabilistic Systems: The challenge of building reliable, predictable systems using inherently probabilistic LLMs.
  • Governance: Controlling access to data, models, and tools to ensure security and compliance.
  • Evaluation: Measuring the quality and accuracy of AI systems to ensure they meet required standards.
  • Tool Calling: Using LLMs to select and execute specific functions or tools, creating more complex and deterministic workflows.
  • Multi-Node/Multi-Step Chains (Agents): Decomposing complex tasks into a series of smaller, more manageable steps, each handled by a specialized agent.
  • RAG (Retrieval-Augmented Generation): A technique for improving the accuracy and relevance of LLM outputs by grounding them in external knowledge sources.

1. Introduction and Databricks' Role

  • Craig, Product Management Lead at Databricks, discusses the challenges of deploying AI technology within large organizations.
  • Databricks is a multi-cloud data platform with thousands of customers and billions in revenue.
  • Databricks created popular open-source capabilities like Spark, MLflow, and Delta.
  • The core problem is that large enterprises have fragmented data across various clouds, vendors, and services due to acquisitions and legacy systems.
  • Data silos and a lack of unified expertise hinder the effective use of data for AI initiatives.
  • Databricks aims to help manage data and provide AI capabilities, particularly through Mosaic AI, focusing on data intelligence.

2. General Intelligence vs. Data Intelligence

  • Both general and data intelligence are valuable, but enterprises need to connect AI to their data estate to automate systems and gain insights.
  • Example: FactSet: A financial data company that required users to learn its proprietary query language (FQL).
    • FactSet initially used a one-click RAG approach to translate English into FQL, achieving only 59% accuracy with 15 seconds of latency.
    • Databricks helped FactSet decompose the prompt into individual tasks, creating a multi-step agent chain.
    • This improved accuracy to 85% and reduced latency to six seconds.
    • FactSet then took over the development and improved the accuracy into the 90s, planning to transition to Claude.

3. The Complexity of AI Systems

  • Research from the Berkeley AI Research lab found that most production AI systems are complex, multi-node architectures, not simple single-input/single-output systems.
  • Databricks aims to simplify the creation of these complex systems, especially in areas with financial and reputational risk.
  • Many developers try to build deterministic systems using probabilistic LLMs.

4. Governance and Evaluation

  • To address the challenge of building deterministic systems, Databricks emphasizes governance and evaluation.
  • Governance: Controlling access to data, models, and tools at a granular level.
    • Treating AI agents as principals within the data stack.
    • Databricks governs data, models, tools, and queries.
  • Evaluation: Quantifying and improving the quality of AI systems.
    • Using a golden data set and LLM judges to assess performance.
    • Providing a simplified UI for subject matter experts to correct prompts and answers.
    • The evaluation system includes open-source components in MLflow.

5. Tool Calling and Claude Integration

  • Tool calling involves using LLMs to classify and select from a set of tools (agents, SQL queries, functions).
  • This creates a decision tree that reduces entropy and increases determinism.
  • Before integrating with Anthropic, tool calling was unreliable.
  • Claude significantly improved tool calling capabilities, enabling software engineers to build quasi-deterministic systems.
  • Claude is natively available on Databricks across multiple clouds (Azure, AWS, GCP).
  • Data engineers can use Claude as a principal within their data governance systems.

6. Why Claude and Databricks Together

  • Pairing the strongest model (Claude) with the strongest platform (Databricks).
  • Providing a fully controlled environment for AI development.
  • Enabling high-value use cases and demonstrating the long-term potential of AI.

7. Real-World Examples and Use Cases

  • Banks are prototyping on Claude, but some lack the controls to use it safely.
  • Databricks helps organizations in highly regulated industries gain access to and control over AI technology.
  • Databricks uses Claude to automate the process of filling out analyst questionnaires (e.g., Gartner, Forrester).
    • Previously, this required hundreds of hours from product managers, engineers, and marketing staff.
    • Claude now generates near-final drafts, significantly reducing the workload.
    • Open-source models and non-Anthropic models were not accurate enough for this task.
  • Block (formerly Square): A Databricks customer that built an open-source agentic development environment called Goose.
    • Goose integrates Claude and connects to Block's systems and data.
    • It accelerates developer workflows and improves productivity, resulting in a 40-50% weekly user adoption increase and 8-10 hours saved per week.

8. Conclusion and Call to Action

  • Identify AI use cases and define success metrics.
  • Contact Databricks or Anthropic for assistance.
  • Databricks aims to help organizations gain confidence in AI and address the challenges faced by AI teams.

9. Q&A Highlights

  • The "safe score" in the LLM judges is a guardrail measure, not an adversarial technique.
  • Databricks differentiates itself from point solutions by focusing on the integration of AI and data systems.
  • Composable agentic approaches are encouraged to build deterministic systems that can be tuned at a granular level.
  • While large models like GPT-3.5/4 can perform many tasks, the ability to fine-tune and control individual steps is crucial for high-risk environments.

Main Takeaways

The key to successfully deploying AI in large enterprises lies in addressing data fragmentation, ensuring robust governance and evaluation, and leveraging powerful models like Claude within a well-integrated platform like Databricks. By breaking down complex tasks into manageable steps and focusing on deterministic outcomes, organizations can unlock the full potential of AI while mitigating risks. The integration of AI and data layers is paramount, and tools like Claude and Goose are enabling developers to build sophisticated, high-value applications.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.