Azure Top 3 Reliability Actions

THE SUMMARYAI-generated

Key Concepts

  • Availability Zones (AZs)
  • Zone Redundancy
  • Network Gateways (Site-to-Site VPN, ExpressRoute)
  • ExpressRoute Metro
  • ExpressRoute Circuits
  • Peering Locations/Points
  • Azure Reliability Documentation Hub
  • Transient Faults
  • Blast Radius
  • Active-Active Configuration

1. Availability Zones (AZs)

  • Definition: Physically separate locations within an Azure region. Each region has three AZs (AZ1, AZ2, AZ3).
  • Key Feature: Independent and isolated power, cooling, and networking.
  • Benefit: Increased fault tolerance. Workloads can survive data center failures.
  • Implementation: Deploy Azure resources across multiple AZs (zone redundant).
  • Example: Virtual Machine Scale Sets, Azure Kubernetes Service (AKS), App Service Plan.
  • Requirement: Minimum of two instances, ideally three (one per AZ).
  • Storage: Use Zone Redundant Storage (ZRS) for base storage. ZRS replicates data across three AZs.
  • Databases: Enable zone redundancy or high availability options for databases like Cosmos DB, SQL DB (General Purpose), SQL MI, PostgreSQL, SQL.
  • Quote: "When I use multiple availability zones within a region, it means my workload can survive a bigger blast radius."

2. Network Gateways

  • Focus: Ensuring network connectivity to Azure resources is resilient.
  • Types: Site-to-Site VPN, ExpressRoute (primary focus).
  • Recommendation: Use zone redundant SKUs for network gateways.
  • Benefit: Instances of the gateway are instantiated in multiple AZs, eliminating single points of failure.
  • Goal: Ensure every component of the workload is not isolated to a single zone or considered regional.
  • Quote: "What's the point in having all this great zone redundant stuff if I can't actually get to it?"

3. Network Connectivity (ExpressRoute)

  • Context: Resiliency of the connection between on-premises networks and Azure via ExpressRoute.
  • ExpressRoute Circuits: Purchased from a provider with services in peering locations (carrier-neutral facilities).
  • Standard Configuration: Active-active pair of connections to Microsoft edge routers in the same building.
  • Limitation: Standard configuration offers limited resilience; vulnerable to failures affecting the entire collocation facility.
  • ExpressRoute Metro:
    • Definition: Second connection of the active-active pair is located in a different building within the same city.
    • Benefit: Improved resilience at a relatively low cost.
    • Availability: Limited to certain locations and providers.
  • Multiple ExpressRoute Circuits:
    • Benefit: Highest level of resilience.
    • Cost: Higher cost due to purchasing a second circuit.
    • Consideration: Trade-off between latency and resilience when choosing the location of the second circuit's peering point.
    • Options:
      • Second circuit in the same city (like Metro) for lower latency.
      • Second circuit in a different city (hundreds of miles away) for greater resilience, especially when using multiple Azure regions.
  • Gateway Configuration: Configure gateways to route traffic between the multiple ExpressRoute circuits.

4. Azure Reliability Documentation Hub

  • Purpose: Centralized resource for understanding and improving the reliability of Azure workloads.
  • Content:
    • Overview of Azure regions, availability zones, and reliability concepts.
    • Service-specific reliability architectures.
    • Information on handling transient faults (small time duration events).
    • Guidance on implementing retry logic.
    • Availability Zone and regional support details.
    • Configuration instructions for Availability Zones.
    • Capacity planning considerations.
    • Zone down and recovery actions.
    • Multi-region support information.
    • Backup strategies.
    • Service level implications.
  • Process:
    1. Identify the Azure services used in the workload.
    2. Find those services in the Azure Reliability Documentation Hub.
    3. Review the service-specific reliability guidance.
  • Example: App Gateway, Blob Storage.
  • Transient Fault Handling: Use built-in retry policies in Azure storage client libraries.
  • Status: Documentation is a work in progress, with more services being added over time.
  • Quote: "It's designed to help you be successful in your reliability journey."

5. Synthesis/Conclusion

The top three reliability actions, according to Microsoft telemetry, are using Availability Zones for all resources (compute, database, storage), ensuring network gateways are zone redundant, and implementing resilient network connectivity through ExpressRoute Metro or multiple ExpressRoute circuits. The Azure Reliability Documentation Hub is a valuable resource for understanding service-specific reliability considerations and implementing best practices. Prioritizing these actions is crucial for building reliable and resilient cloud services in Azure.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.