Summarize Annual Reports in Python (10-K, Financial Documents)

NeuralNineAbout 6 min readJun 3, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • 10K Reports/Annual Reports: Comprehensive documents filed by public companies to inform investors about the company's financial state, risks, and operations.
  • Structured Information Extraction: The process of programmatically extracting specific data points from unstructured documents (like 10Ks) and organizing them into a predefined data model.
  • Large Language Models (LLMs): AI models like Gemini, OpenAI, or Mistral, capable of understanding and generating human-like text, used here to analyze and summarize 10K reports.
  • Data Model (Pydantic): A Python class defined using Pydantic's BaseModel that specifies the structure and data types of the information to be extracted from the 10K report.
  • Gemini API: Google's LLM API used for text analysis and structured output generation.
  • Tokenization: The process of breaking down text into smaller units (tokens) for LLMs to process. Token count is used to estimate API costs.
  • Markdown: A lightweight markup language used to format the extracted information into a readable report.
  • PDF Generation: The process of converting the formatted information into a PDF document for easy sharing and viewing.

1. Introduction and Demonstration

  • The video demonstrates how to extract structured information from 10K reports (annual reports) and summarize them using LLMs in Python.
  • 10K reports are lengthy documents containing financial statements, balance sheets, risk factors, and other information relevant to investors.
  • The goal is to automate the process of extracting key information from these reports, which can be time-consuming and inconsistent when done manually.
  • The Python script extracts information based on a custom data model defined by the user.
  • The extracted information is then used to generate a summarized PDF report containing key financial metrics, business descriptions, risk factors, and management discussions.

2. Setting Up the Environment

  • Obtaining 10K Reports:
    • Search for "[Company Name] Investor Relations" or "[Company Name] Annual Report" to find the company's investor relations page.
    • Download the 10K report as a PDF from the SEC filings section.
  • Obtaining an API Key:
    • The video uses the Google Gemini API.
    • Go to aistudio.google.com/app/apikey and create an API key.
    • A free tier is available for experimentation.
    • Store the API key in a .env file in your project directory. The file should contain the line gemini_api_key = "YOUR_API_KEY".
  • Installing Python Packages:
    • The following packages are required:
      • google-generativeai
      • pydantic
      • python-dotenv
      • ticktoken (optional, for token estimation)
      • markdown2
      • weasyprint
      • PyPDF2
    • Install using pip install [package_name] or uv add [package_name].

3. Token Estimation (Optional)

  • A counter.py script is provided to estimate the number of tokens in a 10K report.
  • This is useful for estimating API costs, as LLM pricing is often based on token usage.
  • The script uses the ticktoken library (designed for OpenAI) to tokenize the text extracted from the PDF.
  • The script loads the PDF using PyPDF2, extracts the text, and then uses ticktoken to count the tokens.
  • Example: The Meta 10K report was estimated to contain approximately 112,000 tokens.

4. Main Script (main.py)

  • Imports:
    • os, json, datetime, typing (List, Optional): Core Python modules for environment variables, JSON handling, date/time manipulation, and type hinting.
    • PyPDF2: For extracting text from PDF files.
    • google.generativeai: For interacting with the Gemini API.
    • dotenv.load_dotenv: For loading environment variables from the .env file.
    • pydantic.BaseModel, pydantic.Field: For defining the data model.
    • markdown2.markdown: For converting Markdown to HTML.
    • weasyprint.HTML: For generating PDF files from HTML.
  • Loading the API Key:
    • load_dotenv() loads the API key from the .env file.
    • os.getenv("gemini_api_key") retrieves the API key from the environment variables.
  • Defining the Data Model (Pydantic):
    • A class AnnualReport(BaseModel) is defined to represent the structure of the extracted information.
    • Each field in the data model corresponds to a specific data point in the 10K report (e.g., company_name, net_income, risk_factors).
    • pydantic.Field is used to provide descriptions for each field, which helps the LLM understand what information to extract.
    • Data types are specified for each field (e.g., str, float, Optional[float], List[str]).
    • Example fields:
      • company_name: str = Field(..., description="The name of the company as reported in the 10K.")
      • cik: str = Field(..., description="The Central Index Key identifier for the company.")
      • fiscal_year_end: datetime = Field(..., description="The fiscal year end date.")
      • total_revenue: Optional[float] = Field(None, description="The total revenue in USD.")
      • risk_factors: List[str] = Field(..., description="A list of risk factors.")
  • Creating the Gemini Client:
    • client = genai.GenerativeModel(model_name="gemini-2.0-flash", api_key=os.getenv("gemini_api_key")) creates a Gemini client using the API key.
  • Crafting the Prompt:
    • The prompt instructs the LLM to analyze the 10K report and fill the data model.
    • The prompt includes the text of the 10K report and the JSON schema of the data model.
    • Example prompt:
      Analyze the following annual report 10k and fill the data model based on it.
      
      [10K Report Text]
      
      The output needs to be in the following data format:
      
      [JSON Schema Definition]
      
      No extra fields allowed.
      
  • Generating Content with Gemini:
    • response = client.generate_content(model="gemini-2.0-flash", contents=prompt, config={"response_mime_type": "application/json", "response_schema": AnnualReport}) sends the prompt to the Gemini API and requests structured output.
    • response_mime_type specifies that the response should be in JSON format.
    • response_schema specifies the Pydantic class to use for validating the response.
  • Validating the Response:
    • ar = AnnualReport.validate_json(response.text) validates the JSON response against the Pydantic data model and creates an AnnualReport object.
  • Generating the PDF Report:
    • A list of Markdown lines (md_lines) is created to format the extracted information.
    • The Markdown lines include headings, bullet points, and formatted text.
    • Conditional statements are used to handle optional fields and lists of data.
    • Example Markdown lines:
      # Annual Report Summary
      
      ## [Company Name]
      
      **CIK:** [CIK]
      
      **Fiscal Year End:** [Fiscal Year End]
      
      - **Total Revenue:** $[Total Revenue]
      
    • The Markdown lines are joined together to create a complete Markdown document.
    • markdown2.markdown(md) converts the Markdown to HTML.
    • weasyprint.HTML(string=html).write_pdf(filename) generates a PDF file from the HTML.
    • The PDF file is named annual_report_[Company Name]_[Year].pdf.

5. Running the Script

  • Run the script using uv run main.py or python main.py.
  • The script will extract the information from the 10K report, generate a PDF report, and save it to the project directory.

6. Customization and Improvements

  • The data model can be customized to extract specific information of interest.
  • The prompt can be adjusted to improve the accuracy and completeness of the extracted information.
  • Different LLMs (e.g., OpenAI, Mistral) can be used.
  • More sophisticated PDF extraction libraries (e.g., Mistral OCR) can be used to improve the quality of the extracted text.
  • A user interface can be built around the script to make it easier to use.
  • The script can be integrated into an agentic system for more automated analysis.

7. Conclusion

  • The video provides a practical guide to extracting structured information from 10K reports using LLMs in Python.
  • The approach is highly customizable and can be adapted to extract a wide range of information.
  • The generated PDF reports can be used to quickly summarize key information and identify potential investment opportunities.
  • The video emphasizes that the extracted information should not be blindly trusted and should be verified before making any investment decisions.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.