LangExtract: Turn Messy Text into Graph-RAG Insights

Prompt EngineeringAbout 5 min readAug 4, 2025Watch original
THE SUMMARYAI-generated

Lang Extract: Unstructured to Structured Data Conversion

Key Concepts:

  • Lang Extract: An open-source Python package from Google for converting unstructured data into structured data using LLMs.
  • Custom Schema: Defining specific information to extract from text.
  • Entity Extraction: Identifying and categorizing key elements in text (e.g., companies, people, products).
  • Relationship Extraction: Identifying connections between different entities.
  • Knowledge Graphs: Visual representations of entities and their relationships.
  • Few-Shot Examples: Providing a small number of high-quality examples to guide the LLM's extraction process.
  • Retrieval Augmented Generation (RAG): Enhancing LLM performance by retrieving relevant information from a knowledge base.
  • Extraction Passes: Multiple processing iterations to improve extraction accuracy.

Overview of Lang Extract

Lang Extract is a new open-source project from Google designed to convert unstructured data into structured data. A key feature is the ability to define custom schemas, allowing users to extract specific information of interest. The tool also provides visualizations of the extracted information. It's an open-source Python package that uses large language models (LLMs) to extract structured information from unstructured data. Users provide a few short examples along with a schema, and the tool extracts the information accordingly. It can be used with both proprietary models like Gemini and open-source models, although the presenter's examples focus on Gemini. A long-context LLM is recommended for optimal performance.

Installation and Basic Usage

The package is installed using pip. The basic usage involves providing text for information extraction, setting up high-quality examples to guide the extraction, and defining the attributes or entities to extract. The process involves:

  1. Providing an example text.
  2. Identifying attributes or entities to extract.
  3. Providing examples within the text.
  4. Specifying attributes of entities.

The input text and prompt are then passed through Lang Extract using the extract function, along with the few-shot examples. The extracted information can be visualized in an HTML document.

Practical Examples and Capabilities

The video demonstrates several examples to showcase Lang Extract's capabilities:

  • Basic Entity Extraction: Extracting companies, people, products, dates, locations, and prices from a text related to Apple. A Microsoft-related example is used as a few-shot example. The output is a JSONL file that can be visualized as an HTML file, showing the extracted entities and their character boundaries.
  • Entity and Attribute Extraction: Extracting information from a quarterly earnings report, including company tickers, exchanges, and financial metrics with attributes like period, value, and change. Examples also include news articles (funding events, competitive moves) and customer reviews.
  • Relationship Extraction and Knowledge Graph Creation: Extracting medication-related information from clinical notes, including dosage, route, medication frequency, and duration. The example focuses on the relationship between conditions and medications. The prompt identifies medication name, dosage, condition treated, related medications, and interactions. The extraction process uses multiple passes to increase accuracy. The extracted data is used to create a knowledge graph showing the relationships between conditions and medications. For example, hypertension is linked to specific medications used to treat it.

Step-by-Step Example: Medication and Condition Extraction

  1. Input: A clinical note summarizing patient information, problems, and medications.
  2. Prompt: Identify medication name, dosage, condition treated, related medications, and interactions for each medication.
  3. Few-Shot Example: Extract patient information (age, gender), conditions, and medications with their attributes, including a "related condition" attribute linking medications to conditions.
  4. Extraction Passes: Use multiple passes to improve accuracy.
  5. Output: A JSONL file containing the extracted entities and attributes, visualized as an HTML file. The visualization shows the relationships between conditions and medications, enabling the creation of a knowledge graph.

Knowledge Graph Application

The extracted data can be used to create knowledge graphs. The video shows an example where conditions like hypertension, hyperlipidemia, and atrial fibrillation are linked to the medications used to treat them. This demonstrates the ability to visualize relationships between entities extracted from unstructured text.

Key Arguments and Perspectives

The video emphasizes the power of entity extraction and relationship extraction for creating knowledge graphs. These knowledge graphs can be used for downstream applications like retrieval augmented generation (RAG) systems. The presenter suggests that structured data extraction is highly beneficial for metadata usage in information retrieval systems.

Notable Quotes

  • "You don't want to do entity extraction on chunk basis but you want to do it initially on document level and then connect that metadata to chunk level."
  • "...the main power is entity extraction and then figuring out their relationships which will help you create knowledge graphs or relationship graphs."

Technical Terms and Concepts

  • LLM (Large Language Model): A type of artificial intelligence model trained on vast amounts of text data, capable of understanding and generating human-like text.
  • Schema: A structured definition of the data to be extracted, including entities, attributes, and relationships.
  • Entity: A distinct object or concept in the text, such as a company, person, or product.
  • Attribute: A characteristic or property of an entity, such as the ticker symbol of a company or the dosage of a medication.
  • JSONL (JSON Lines): A format where each line is a valid JSON object, often used for storing structured data.
  • RAG (Retrieval Augmented Generation): A framework that combines information retrieval with text generation to improve the accuracy and relevance of LLM outputs.

Logical Connections

The video progresses from a general introduction of Lang Extract to specific examples demonstrating its capabilities. It starts with basic entity extraction, moves to entity and attribute extraction, and culminates in relationship extraction and knowledge graph creation. The examples build upon each other, showcasing the increasing complexity and power of the tool. The connection to RAG is presented as a potential downstream application of the extracted structured data.

Data and Research Findings

The video references a paper on LLM-accelerated annotation for medical information, indicating that the medication and condition extraction example is based on research in that area.

Synthesis/Conclusion

Lang Extract is a powerful tool for converting unstructured data into structured data using LLMs. Its ability to define custom schemas, extract entities and attributes, and identify relationships makes it valuable for creating knowledge graphs and enhancing downstream applications like RAG. The tool's visualization capabilities and open-source nature make it accessible and useful for a wide range of information retrieval and data analysis tasks. While not an officially supported Google product, it represents a promising approach to structured data extraction.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.