Agentic Observability: Monitoring AI Traffic with F5 NGINX & OpenTelemetry

By F5 DevCentral Community

Share:

Key Concepts

  • Agentic Observability: The practice of monitoring and gaining visibility into the unpredictable, multi-step traffic patterns generated by AI agents.
  • MCP (Model Context Protocol): The communication layer used by AI agents to interact with tools and data sources.
  • Layer 7 Inspection: Analyzing traffic at the application layer to understand the content of requests (e.g., JSON payloads) rather than just network-level metadata.
  • Fan-out: A behavior where a single AI prompt triggers multiple simultaneous background tasks or tool calls.
  • OpenTelemetry (OTel): A framework for collecting and exporting telemetry data (metrics, logs, traces) to observability platforms.
  • Golden Signals: The four key metrics for monitoring system health: Latency, Traffic (Throughput), Errors, and Saturation.

1. The Challenge: AI Traffic vs. Traditional Web Traffic

Traditional web traffic is predictable, following a linear request-response pattern. In contrast, AI agents are inherently unpredictable. A single user prompt can trigger a "fan-out" effect, launching numerous background tasks or tool calls simultaneously.

Standard load balancers fail to address this because they treat AI traffic as generic web requests, creating a "black box" that leads to:

  • Resource Exhaustion: Sudden, silent spikes in backend system load that traditional monitoring cannot correlate to specific AI prompts.
  • Degraded User Experience: Inability to identify which specific AI tool is causing latency, leaving users waiting without clear diagnostic data.
  • Operational Blindness: Teams are forced to guess the cause of performance bottlenecks, scaling issues, or cost overruns.

2. The Solution: Native Agentic Observability with F5 NGINX

The proposed solution involves using NGINX as a native proxy to inspect MCP traffic at Layer 7.

  • No Code Changes: Unlike "bolt-on" AI proxies that introduce friction and points of failure, this approach uses the existing NGINX footprint.
  • Methodology: It utilizes NGINX JavaScript (njs) to parse JSON streams in real-time. It extracts metadata—such as tool names, client identities, and target MCP servers—and maps them to OpenTelemetry span attributes.
  • Flexibility: Because it exports standard OTel span attributes, the data can be piped into any observability platform (e.g., Prometheus/Grafana).

3. Implementation Process

  1. Deployment: Download the mcp.js file from the official Agentic Observability GitHub repository.
  2. Configuration: Place the file in the NGINX configuration folder and load the module within the nginx.conf file.
  3. Dynamic Inspection: Configure the proxy to use the JavaScript module to intercept and parse incoming AI traffic.
  4. Telemetry Export: Map the extracted metadata (tool name, execution status) to OTel attributes, which are then pushed to the monitoring backend.

4. Real-World Applications and Troubleshooting

The demo highlights how this visibility transforms operational decision-making:

  • Pinpointing Latency: Instead of guessing, operators can view P99 latency per specific AI tool to identify which function (e.g., a database lookup vs. a web search) is causing delays.
  • Identifying Rogue Agents: By tracking traffic by "AI Client," teams can identify agents stuck in loops or behaving aggressively and apply targeted rate limiting.
  • Isolating Backend Failures: By monitoring error rates per backend MCP server, teams can isolate "flaky" servers and reroute traffic to stable ones.

5. Notable Statements

  • "Without the ability to see inside the model context protocol or MCP layer, our infrastructure is just a black box."
  • "If you can't see the bottleneck, you can't fix it."
  • "Because the MCP module handles this natively at the proxy layer, there is absolutely no heavy lifting or code changes required by your backend application."

6. Synthesis and Conclusion

The shift toward agentic AI requires a fundamental change in how infrastructure is monitored. By moving observability to the proxy layer (NGINX) and natively inspecting MCP traffic, organizations can eliminate the "black box" effect. This approach provides granular, real-time insights into AI tool execution, client behavior, and backend health without requiring modifications to the underlying AI application code. This enables proactive capacity planning, improved compliance auditing, and rapid troubleshooting of complex AI-driven environments.

Chat with this Video

AI-Powered

Load the transcript when you're ready to chat so the initial page stays lighter.

Ready to summarize another video?

Summarize YouTube Video