Import EVERYTHING Into Your RAG Agent (Docling & LlamaParse)

The AI AutomatorsAbout 5 min readAug 27, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Agents
  • RAG (Retrieval Augmented Generation) Ingestion
  • Unstructured Data
  • Document Parsing
  • OCR (Optical Character Recognition)
  • Llama Parse
  • Docklin
  • Mistral OCR
  • Vector Database
  • Markdown Format
  • NLQ (Natural Language Query)

Importing Data into AI Agents

The video focuses on importing various file formats into AI agents to enable them to interact with a wide range of data sources. It highlights the challenge that 80-90% of organizational data is unstructured, residing in documents, emails, images, audio, etc. The key is to parse these files into a consistent markdown format for ingestion into a vector database, which then powers the RAG workflow.

Llama Parse: Setup, Usage, and Integration

Setup and Pricing

  • Llama Parse is a service for parsing documents into structured markdown.
  • It offers a free tier with 10,000 credits per month.
  • Pricing is $1 per 1,000 credits. Agentic mode (recommended for PDFs) costs 10 credits per page, allowing for roughly 1,000 pages per month on the free plan.
  • To get started, sign up at llamaindex.ai and access the cloud console.

Usage and Features

  • The Llama Parse console provides a playground interface for testing file parsing.
  • It supports various output formats, including markdown, text, and JSON.
  • Advanced settings allow tweaking the document model and parsing mode (Agentic is recommended for PDFs).
  • It can annotate images with descriptions, which is useful if you don't want to serve the images directly to the user.

N8N Integration: Step-by-Step

  1. Trigger: Use a manual trigger or listen for new files in a Google Drive folder.
  2. HTTP Request (Get PDF): Use an HTTP node to retrieve a PDF file from a public URL (e.g., Superbase bucket).
  3. HTTP Request (Llama Cloud):
    • Create a generic credential in N8N for Llama Cloud using your API key (Bearer {API Key}).
    • Use the "upload file" endpoint from the Llama Cloud API reference.
    • Set the "file" parameter to the binary data from the previous HTTP request.
    • Configure parsing options (model, OCR settings, table extraction) based on the Llama Cloud code snippet for Agentic mode. Example parameters:
      • model: "OpenAI GPT-4-vision-preview"
      • high_res_ocr: true
      • adaptive_long_table: true
      • outline_table_extraction: true
      • output_tables_as_html: true
  4. Switch Node (Status Check):
    • Check the "status" field in the Llama Cloud response (pending, success, error, partial).
    • If "pending," use a "Wait" node (e.g., 10 seconds) and then re-poll the Llama Cloud API using the "get parsing job details" endpoint.
    • If "success," proceed to retrieve the parsed content.
  5. HTTP Request (Get Markdown):
    • Use the "get raw markdown result" endpoint, passing the document ID from the previous response.
    • Authenticate using the Llama Cloud credential.
  6. Superbase Vector Store:
    • Use the Superbase Vector Store node to add the markdown content to your vector database.
    • Configure an embedding model (e.g., text-embedding-3-small).
    • Use a recursive character text splitter with splitAsMarkdown set to true.
  7. AI Agent: Create an AI agent that queries the vector database and responds to user questions.

Docklin: Setup, Usage, and Integration

Overview

  • Docklin is an open-source document parsing and OCR framework by IBM.
  • It supports various file formats (PDF, DOCX, PPTX, XLSX, etc.).
  • It parses files natively and uses OCR when needed.
  • Key advantages: No external APIs, self-hostable, cost-effective, and good for data security/privacy.
  • Drawbacks: Slower processing than Llama Parse, requires deployment as an API.

Setup using Render.com

  1. Deploy Docklin Serve: Use the doclin-serve-cpu-latest image from the Docklin Serve GitHub repository on Render.com.
  2. Resource Requirements: Requires at least the "Standard" package on Render.com due to resource intensity.
  3. Enable UI: Set the environment variable DOCLINE_SERVE_ENABLE_UI to 1.
  4. Deploy: Deploy the web service and wait for it to be set up (approximately 5 minutes).
  5. Access UI: Access the Docklin Serve UI at <your-render-url>/ui.

Usage

  • Upload a file through the UI.
  • Configure OCR options (enable OCR, OCR engine, language).
  • Select output formats (JSON, markdown).
  • Process the file and review the results.

N8N Integration

  1. HTTP Request (Docklin):

    • Use the V1/convert/source endpoint of your Docklin Serve instance.
    • Set the method to POST.
    • Construct a JSON request with options and the URL of the document to be processed. Example:
    {
      "options": {
        "force_ocr": true,
        "ocr_engine": "tesseract",
        "ocr_language": "eng"
      },
      "sources": [
        {
          "url": "<your-superbase-bucket-url>/test.pdf"
        }
      ]
    }
    
  2. Superbase Vector Store: Map the markdown content from the Docklin response to your Superbase vector store.

Security Considerations

  • Public Access Gateway: To secure your Docklin instance, create a separate public app (gateway) that sits in front of a private Docklin Serve app.
  • Authentication: Implement basic username/password authentication and API key requirements in the gateway.
  • Render Setup:
    • Create a private GitHub repository with a Dockerfile, entry point, and Nginx configuration for the gateway.
    • Create a new web service on Render.com, connecting to the private GitHub repository.
    • Set environment variables for the API key, username, password, and the private URL of the upstream Docklin Serve app.
  • N8N Configuration: Configure the N8N HTTP request to use the gateway URL and include the API key in the header (X-API-Key: <your-api-key>).

Mistral OCR

Overview

  • Mistral OCR is a service that supports only PDFs.
  • It is very fast, high quality, easy to use, and cost-effective.
  • Pricing: $1 per 1,000 pages (OCR), $3 per 1,000 pages (annotations for image/diagram extraction).

Usage

  • Refer to the multimodal RAG video for detailed setup instructions.
  • Two workflows are presented:
    • Simple: Ingest PDF, upload to Mistral OCR, retrieve markdown, upload to vector store.
    • Sophisticated: Extract markdown and images, upload images to Superbase bucket, serve images to the AI agent.

Comparison with Llama Parse

  • Llama Parse offers a free tier and more configuration options.
  • Mistral OCR is simpler to set up and can be cheaper at scale.

Conclusion

The video provides a comprehensive guide to importing various file formats into AI agents using different document parsing and OCR solutions. It covers Llama Parse, Docklin, and Mistral OCR, detailing their setup, usage, integration with N8N, and security considerations. The key takeaway is that by effectively parsing unstructured data into a consistent markdown format, organizations can unlock the potential of their untapped data and build more powerful and versatile AI agents.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.