Deploy Generative AI Models with Amazon Bedrock & Python

NeuralNineAbout 4 min readMay 27, 2025Watch original
THE SUMMARYAI-generated

AWS Bedrock and Python Integration for Generative AI Applications

Key Concepts:

  • AWS Bedrock: A service for building and scaling generative AI applications using foundation models without managing infrastructure.
  • Foundation Models: Pre-trained AI models like Llama, Mistral, and Claude.
  • Pay-per-use: A pricing model where you only pay for the tokens you use.
  • AWS CLI: Command Line Interface for interacting with AWS services.
  • Boto3: The AWS SDK for Python.
  • Instructor: A library for extracting structured output from language models.
  • Pydantic: A library for data validation and settings management using Python type annotations.
  • UV: A package manager (alternative to pip).
  • Global Interpreter Lock (GIL): A mechanism in Python that allows only one thread to hold control of the Python interpreter.

1. Introduction to AWS Bedrock

  • Main Idea: AWS Bedrock allows developers to use foundation models (like Llama, Mistral, Claude) to build generative AI applications without managing the underlying infrastructure (scaling, load balancing).
  • Value Proposition: Easy to use, pay-per-use pricing, no infrastructure management.
  • Pricing: Based on per input and output token usage (price per thousand tokens).
  • Model Availability: Different models are available in different AWS regions.

2. Setting up AWS Bedrock

  • Prerequisites: An AWS account.
  • Accessing Bedrock: Navigate to the Amazon Bedrock service in the AWS console.
  • Region Selection: Choose a region based on model availability and regulatory requirements (e.g., GDPR in the EU).
    • Example: Mistral Large is available in Paris (EU West 3) but not necessarily in Frankfurt.
  • Model Access: Request access to specific models via the "Model Access" section.
    • Note: Access to Claude models may require providing a reason and can take time to be approved. Llama and Mistral are typically granted immediately.

3. Using the Bedrock Playground

  • Accessing the Playground: Navigate to "Chat and Text" in the Bedrock console.
  • Model Selection: Choose a model to experiment with (e.g., Mistral Large, Llama 3.2).
  • Compare Mode: Compare the outputs of different models side-by-side.
  • Configuration: Adjust parameters like temperature, response length, and system prompt.
  • Example: Asking "Explain to me the global interpreter lock in Python" to compare the responses of Mistral Large and Llama 3.2.

4. Integrating AWS Bedrock with Python

  • Prerequisites:
    • Install the AWS CLI.
    • Configure AWS CLI with access key ID, secret access key, and region.
    • Install Boto3 (AWS SDK for Python).
  • Installation:
    • Using pip: pip install boto3
    • Using UV: uv add boto3
  • Code Example (Basic):
    import json
    import boto3
    
    # Create a Bedrock runtime client
    bedrock_runtime = boto3.client(
        service_name='bedrock-runtime',
        region_name='eu-west-3'  # Replace with your region
    )
    
    # Define the prompt
    prompt = "What is the global interpreter lock in Python?"
    
    # Define the parameters for the Mistral Large model
    parameters = {
        "modelId": "mistral.mistral-large-2402-v1:0",
        "contentType": "application/json",
        "accept": "application/json",
        "body": json.dumps({
            "prompt": f"<s>[INST] {prompt} [/INST]",
            "max_tokens": 200,
            "temperature": 0.2,
            "top_p": 0.9,
            "top_k": 50
        })
    }
    
    # Invoke the model
    response = bedrock_runtime.invoke_model(**parameters)
    
    # Extract and print the response
    response_body = json.loads(response['body'].read().decode('utf-8'))
    print(response_body)
    
  • Explanation:
    • boto3.client('bedrock-runtime', region_name='...'): Creates a client to interact with the Bedrock runtime.
    • parameters: A dictionary containing the model ID, content type, accept type, and the request body (prompt, max tokens, temperature, etc.).
    • bedrock_runtime.invoke_model(**parameters): Invokes the specified model with the given parameters.
    • The response contains the generated text.

5. Structured Output with Instructor

  • Purpose: Extract structured data from the model's output.
  • Prerequisites:
    • Install Instructor and Pydantic: pip install instructor pydantic or uv add instructor pydantic
  • Code Example (Structured Output):
    from enum import Enum
    from pydantic import BaseModel, Field
    import instructor
    import boto3
    import json
    
    # Create a Bedrock runtime client
    bedrock_runtime = boto3.client(
        service_name='bedrock-runtime',
        region_name='eu-west-3'  # Replace with your region
    )
    
    # Apply Instructor to the Bedrock runtime client
    client = instructor.from_bedrock(bedrock_runtime)
    
    # Define an Enum for programming paradigms
    class ProgrammingParadigm(str, Enum):
        FUNCTIONAL = "functional"
        PROCEDURAL = "procedural"
        LOGICAL = "logical"
        CONCURRENT = "concurrent"
        OBJECT_ORIENTED = "object-oriented"
        OTHER = "other"
    
    # Define a Pydantic model for the desired output structure
    class ProgrammingLanguage(BaseModel):
        name: str
        release_year: int
        paradigms: list[ProgrammingParadigm]
        similar_languages: list[str] = Field(..., description="The names of similar programming languages")
    
    # Define the prompt
    prompt = "Give me information about Haskell"
    
    # Create the completion using the structured output
    response = client.completions.create(
        model_id="mistral.mistral-large-2402-v1:0",
        messages=[{"role": "user", "content": prompt}],
        response_model=ProgrammingLanguage
    )
    
    # Print the structured response
    print(response)
    
  • Explanation:
    • instructor.from_bedrock(bedrock_runtime): Applies the Instructor patch to the Bedrock runtime client, enabling structured output.
    • ProgrammingLanguage(BaseModel): Defines a Pydantic model representing the desired output structure (name, release year, paradigms, similar languages).
    • client.completions.create(model_id='...', messages=[...], response_model=ProgrammingLanguage): Creates a completion request, specifying the model ID, the prompt, and the response_model to guide the output structure.
    • The response object is an instance of the ProgrammingLanguage Pydantic model, containing the extracted data.

6. Conclusion

AWS Bedrock provides a convenient way to access and use foundation models for generative AI applications without the burden of infrastructure management. By integrating Bedrock with Python using Boto3 and Instructor, developers can easily build applications that generate both unstructured and structured output, paying only for the tokens they consume. The choice of region and model depends on specific requirements and availability.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.