Google LangExtract: Watch AI Make Sense of 25,000 Words Instantly!

Mervin PraisonAbout 3 min readAug 1, 2025Watch original
THE SUMMARYAI-generated

Lang Extract: Gemini-Powered Information Extraction Library

Key Concepts:

  • Unstructured to Structured Data Conversion
  • Information Extraction (Character, Emotion, Relationship)
  • Source Grounding
  • Long Context Information Extraction
  • Interactive Visualization
  • Large Language Models (LLMs)
  • Gemini 2.5 Flash
  • JSONL
  • Parallel Processing

Introduction to Lang Extract

Google introduces Lang Extract, a Gemini-powered library designed to convert unstructured data into structured data. This tool facilitates information extraction, offering features like precise source grounding, reliable structured outputs, optimized long context information extraction, interactive visualization, and flexible support for various LM backends. It is adaptable across domains, leveraging LM world knowledge.

Capabilities of Lang Extract

Lang Extract can extract characters, emotions, and relationships from unstructured text. It also allows for defining custom extraction conditions, such as dosage, frequency, and medication, making it suitable for specialized domains like clinical notes and financial documents. The structured output is beneficial for understanding the data and can be used for Retrieval-Augmented Generation (RAG), specifically graph RAG.

Example Applications:

  • Analyzing clinical reports
  • Extracting legal documents
  • Summarizing research papers with source citations

Installation and Setup

Step 1: Installation

  1. Open the terminal on your computer.
  2. Run the command: pip install lang extract
  3. Install libmagic based on your operating system.
  4. Export the Lang Extract API key (Gemini API key) obtained from studio.google.com.

Step 2: Code Creation and Execution

  1. Create a Python file (e.g., app.py).

  2. Import necessary libraries: lang extract and textwrap.

  3. Define the prompt and extraction rules. For example: Extract characters, emotion, and relationships.

  4. Provide high-quality examples to guide the extraction process.

    • Example:
      • Text: "Romeo"
      • Character: Romeo
      • Emotion: Longingly
      • Relationship: Juliet is the sun (metaphor)
  5. Provide the input text to be processed. This can be any unstructured data, including large documents.

  6. Use lx.extract to extract the desired information:

    extraction_results = lx.extract(
        input_text=input_text,
        prompt=prompt,
        examples=examples,
        model_name="gemini-2.5-flash"
    )
    
  7. Visualize the extracted data using lx.visualize_extraction_results, which generates an HTML file.

Step 3: Visualization

  1. Open the generated HTML file in a web browser.
  2. The visualization displays the extracted characters, emotions, and relationships in an interactive format.

Demonstration and Examples

The video demonstrates Lang Extract using radiology reports:

  • Abdominal CT: Unstructured data from an abdominal CT scan is converted into structured data.
  • Chest X-ray: Similarly, unstructured data from a chest X-ray is transformed into a structured format.

Processing Large Documents

The video showcases processing a 25,000+ word document (Romeo and Juliet) to extract characters, emotions, and relationships.

  1. The code is modified to load the text directly from a URL.
  2. Parallel processing is used to speed up the extraction.
  3. The extracted data is saved in a JSONL file and visualized in an HTML file.

Code Modifications:

  • Loading text from URL: input_text = load_text_from_url(url)
  • Parallel processing implementation (details not fully shown in the transcript but implied).

Results:

  • The extraction process took approximately 10 minutes for 147,000 characters.
  • The generated HTML visualization allows interactive exploration of the extracted characters, relationships, and emotions.

Key Arguments and Perspectives

The video argues that Lang Extract offers a significant advancement in converting unstructured data into structured data, making it easier to understand and utilize. The tool's ability to process large documents, provide precise source grounding, and offer interactive visualization makes it valuable for various applications, including clinical analysis, legal document extraction, and research paper summarization.

Conclusion

Lang Extract is a powerful tool for extracting structured information from unstructured data using the Gemini LLM. Its ease of installation, simple code structure, and ability to handle large documents make it a promising solution for various information extraction tasks. The interactive visualization further enhances the usability of the extracted data. The video encourages viewers to try Lang Extract and share their feedback.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.