Advertise OpenShift AI inference servers from F5 Distributed Cloud

F5 DevCentral CommunityAbout 5 min readJun 19, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

  • AI Inference Models: Machine learning models deployed to make predictions based on new data.
  • OpenShift AI: A platform for developing, deploying, and managing AI/ML models.
  • F5 Distributed Cloud: A platform providing networking, security, and application management services across multiple cloud environments.
  • Customer Edge (CE): A component deployed in the customer's environment (e.g., OpenShift cluster) that connects to F5 Distributed Cloud.
  • Regional Edge (RE): F5 Distributed Cloud's points of presence (PoPs) that provide connectivity and services.
  • IP Anycast VIP: A single IP address advertised from multiple locations, allowing traffic to be routed to the closest available endpoint.
  • Origin Pool: A group of backend servers (in this case, AI inference model endpoints) that F5 Distributed Cloud load balances traffic across.
  • KServe: A Kubernetes-based model serving framework used by OpenShift AI.
  • ExternalName Service: A Kubernetes service that maps a DNS name to an external service, used by KServe to expose models.
  • Performance Operator: An OpenShift operator that enables the use of huge pages for improved performance.
  • Huge Pages: Memory pages larger than the default size, which can improve performance for memory-intensive applications.

Advertising AI Inference Models with OpenShift AI and F5 Distributed Cloud

1. Overview

The video demonstrates how to expose AI inference models deployed in OpenShift AI to the internet using F5 Distributed Cloud. The solution leverages F5 Distributed Cloud's global network and security capabilities to provide a secure and scalable way to access AI models. The core idea is to use a Customer Edge (CE) deployed within the OpenShift cluster to connect to F5 Distributed Cloud's Regional Edges (REs), which then advertise the AI model using an IP Anycast VIP.

2. Architecture and Traffic Flow

The traffic flow is as follows:

  1. Client Request: A client sends a request to the AI inference service (e.g., inference.demos.bd.fi.com).
  2. DNS Resolution: The DNS resolves the domain name to an F5 Distributed Cloud Anycast VIP address.
  3. F5 Distributed Cloud Processing: The request reaches the closest F5 Distributed Cloud point of presence (PoP). F5 Distributed Cloud validates the request and applies security policies.
  4. Load Balancing: F5 Distributed Cloud load balances the traffic to the Customer Edge (CE) within the OpenShift cluster.
  5. Tunneling: The traffic reaches the Customer Edge via a pre-established tunnel (TLS or IPSec).
  6. Service Discovery: The Customer Edge uses DNS to discover the internal endpoint of the AI model within the OpenShift cluster. This is facilitated by KServe, which deploys an ExternalName service for each model.
  7. Internal Load Balancing: The request is load balanced between available instances of the AI model.
  8. Model Inference: Istio ultimately sends the request to the AI model for inference.

3. Configuration Steps

The configuration involves the following steps:

  1. Deploy Inference Service in OpenShift AI: Deploy the AI inference model in OpenShift AI and note the internal endpoint. This endpoint will be used by F5 Distributed Cloud to route traffic to the model.
  2. Deploy Customer Edge in OpenShift: Deploy the F5 Distributed Cloud Customer Edge in the OpenShift cluster. This can be done using a YAML file. The video uses the "deploy as spots" option, which requires the Performance Operator to be installed.
  3. Configure Advertisement in F5 Distributed Cloud:
    • Create Health Check: Define a health check to monitor the AI model's availability. This includes specifying the hostname (inference model's endpoint) and path (/slice/health) to verify the service is running. HTTP/2 is used as the protocol.
    • Create Origin Pool: Configure an origin pool that points to the internal endpoint of the AI model. The origin server is discovered using DNS, and the health check is associated with the origin pool. TLS verification is disabled due to the use of self-signed certificates.
    • Create HTTP Load Balancer (VIP): Create an HTTP load balancer to advertise the AI model to the internet using an Anycast VIP. HTTPS with automatic certificate management is enabled. The origin pool created in the previous step is associated with the load balancer.

4. Detailed Configuration Parameters

  • Health Check:
    • Hostname: The inference model's endpoint (e.g., inference.demos.bd.fi.com).
    • Path: The health check path (e.g., /slice/health).
    • Protocol: HTTP/2.
  • Origin Pool:
    • Origin Server Discovery: DNS.
    • Endpoint: The OpenShift AI internal endpoint.
    • Site: The site where the Customer Edge is deployed.
    • Interface: The Customer Edge's interface in OpenShift (named "outside").
    • TLS Configuration: Disable TLS verification.
  • HTTP Load Balancer:
    • Name: The name of the service.
    • HTTPS with Automatic Certificate: Enabled.
    • HTTP Redirection to HTTPS: Enabled.
    • Origin Pool: The origin pool created in the previous step.
    • Internet Advertisement: Enabled.

5. Security Considerations

F5 Distributed Cloud provides application management and security features, including:

  • API Protection: Protects the AI model's API from unauthorized access.
  • Bot Defense: Mitigates bot traffic.
  • Layer 7 DDoS Protection: Protects against distributed denial-of-service attacks.

6. Demonstration and Verification

The video demonstrates the configuration steps in the F5 Distributed Cloud console. After configuration, the AI model is accessed via the internet using Postman. The response from the AI model is displayed, verifying that the configuration is working correctly.

7. Dashboards and Monitoring

F5 Distributed Cloud provides dashboards for monitoring the health and performance of the AI model. The dashboards provide information about:

  • VIP Status: General status, health, and traffic statistics.
  • Request Sources: Information about the source of requests.
  • Origin Server Status: Health check status and DNS service discovery results.

8. Notable Quotes

  • (Implied) "Exposing an AI inference service is no different to any other application."
  • (Implied) "Although this flow might seem complex the actual configuration in distributed cloud is extremely easy."

9. Conclusion

The video demonstrates a streamlined approach to exposing AI inference models deployed in OpenShift AI to the internet using F5 Distributed Cloud. By leveraging F5 Distributed Cloud's networking, security, and application management capabilities, organizations can securely and scalably deploy AI models and make them accessible to a wider audience. The key takeaways are the ease of configuration, the security benefits, and the comprehensive monitoring capabilities provided by F5 Distributed Cloud.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.