Open Source Friday: Building Data Pipelines with dlt and Elvis Kahoro

GitHubAbout 4 min readMay 30, 2026Watch original
THE SUMMARYAI-generated

Key Concepts

  • DLT (Data Load Tool): An open-source Python library designed for Extract, Load, and Transform (ELT) data pipelines.
  • Builder Stack: A modern approach to data engineering that prioritizes composability, serverless execution, and agent-readiness over rigid, monolithic platforms.
  • Schema Contracts: A mechanism to enforce data structure consistency, preventing pipeline failures when source data changes.
  • Agent-Driven Development: Using LLMs (like Claude, Copilot, or Codeex) to generate, maintain, and deploy data pipelines using declarative Python code.
  • Data Observability: Monitoring pipeline health, row counts, and performance metrics through integrated dashboards.

1. Introduction to DLT (Data Load Tool)

DLT is an open-source Python library that simplifies moving data from various sources (e.g., production databases, APIs, CSVs) to destinations (e.g., BigQuery, Snowflake, DuckDB, Hugging Face).

  • Core Philosophy: "Everything should be code." By using Python SDKs, DLT avoids the need for complex, constantly running containers, making it highly accessible for developers and compatible with CI/CD pipelines like GitHub Actions.
  • Scale: The project has seen rapid growth, reaching over 10 million monthly downloads on PyPI and being utilized by over 9,000 companies.

2. The "Builder Stack" vs. Modern Data Stack

The speakers argue that traditional "Modern Data Stack" tools are often too rigid for the current AI-driven landscape.

  • Agent Compatibility: Unlike managed platforms that hide logic behind proprietary abstractions, DLT’s code-first approach allows AI agents to introspect, debug, and generate pipelines effectively.
  • Serverless Elasticity: DLT is designed to handle the "bursty" nature of AI agents, which may run thousands of queries in a short window and then remain idle.
  • Vendor Neutrality: Because DLT is standardized, switching destinations (e.g., from Snowflake to DataBricks) requires changing only one line of code, effectively mitigating vendor lock-in.

3. Step-by-Step: Building Pipelines with AI

The demo showcased a workflow for creating a data pipeline using the DLT CLI:

  1. Initialization: Use uvx dlt-hub --start to scaffold a project with pre-configured skills, transformations, and quality checks.
  2. Context Injection: The agent uses a "context server" to understand API authentication and resource structures (e.g., GitHub issues, OnePassword).
  3. Declarative Definition: The developer (or agent) uses Python decorators to define resources, primary keys, and PII (Personally Identifiable Information) handling.
  4. Validation: Schema contracts (e.g., column_freeze) are applied to ensure data integrity.
  5. Deployment: The same code used locally can be deployed to DLT Pro or cloud environments without modification.

4. Data Quality and Guardrails

A significant portion of the discussion focused on the risks of using AI agents for infrastructure.

  • Human-in-the-loop: The speakers emphasize that agents should be "well-scoped." For example, one agent should be restricted to writing code, while another handles data inspection.
  • Reviewability: Because DLT pipelines are written in readable Python, humans can easily review the code generated by LLMs before execution.
  • SQL-based Checks: Data quality checks are executed as native SQL expressions within the data warehouse, ensuring high performance and avoiding the need to pull data back into the application layer.

5. Notable Quotes

  • “Good APIs are good for both people and agents.” — Elvis Gahoro, on the design philosophy of DLT.
  • “I’m obsessed with developer experience. I want things to be obvious, be smooth, be self-explanatory... If you’re able to do this for humans, like LLMs pick it up very easily.” — Thierry Jean, Lead AI Engineer at DLT Hub.
  • “Everything starts looking like a data pipeline... you will not unsee it.” — The host, reflecting on the ubiquity of data movement.

6. Synthesis and Conclusion

DLT represents a shift toward a "Builder Stack" where data engineering is treated as standard software engineering. By leveraging Python as the primary interface, DLT enables developers and AI agents to collaborate on complex data tasks—from ingestion to transformation and quality assurance—with high transparency and low operational overhead. The project remains committed to its open-source roots while offering a "Pro" layer for enterprises requiring advanced observability and managed deployment.

Actionable Takeaway: For those looking to contribute, the DLT repository on GitHub contains "good first issue" and "quality of life" labels. However, contributors are strongly encouraged to discuss their plans with maintainers in the community Slack before diving into deep code changes.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.