Keeping GPUs Ticking Like Clockwork
By The New Stack
Key Concepts
- Agentic AI: AI systems that can act autonomously and make decisions.
- GPU Clusters: Large groups of Graphics Processing Units used for computationally intensive tasks like AI training.
- Fleet IQ: Clockwork's software platform for optimizing large-scale GPU clusters.
- Network Telemetry: The process of collecting and analyzing data about network performance and behavior.
- Fault Tolerance: The ability of a system to continue operating despite failures.
- Congestion Management: Techniques to prevent network slowdowns caused by excessive traffic.
- Quality of Service (QoS): Mechanisms to prioritize certain types of network traffic.
- Software Shim: A layer of software that sits between other software components to enable communication or add functionality.
- eBPF (extended Berkeley Packet Filter): A technology that allows running sandboxed programs in the Linux kernel without changing kernel source code or loading kernel modules.
- RDMA (Remote Direct Memory Access): A technology that allows direct memory access between computers over a network.
- NIC (Network Interface Controller): Hardware that connects a computer to a computer network.
- PCI Bus: A local computer bus used for attaching hardware devices in computers.
- CUDA Kernels: Functions written in CUDA, a parallel computing platform and programming model created by Nvidia.
- Distributed State: The state of a system that is spread across multiple interconnected components.
- Sora: A text-to-video generation model.
Clockwork's Software Platform: Fleet IQ
Clockwork develops a software layer, Fleet IQ, designed to optimize communication within large GPU clusters used for AI workloads. The platform focuses on three core areas to enhance AI efficiency:
- Deep Visibility: Providing comprehensive insights into the GPU fleet, from the network layer up to the application layer, enabling rapid problem detection and resolution.
- Fault Tolerance: Ensuring AI workloads remain resilient to hardware failures, particularly network link disruptions, allowing for uninterrupted operation. This is crucial because interruptions in AI training often necessitate reverting to old checkpoints, wasting significant compute resources.
- Performance Acceleration: Optimizing network communication between GPUs by managing congestion, implementing Quality of Service (QoS), and redirecting flows to prevent contention.
Historical Evolution of Clockwork
Clockwork's journey began in 2018 with a focus on software-based clock synchronization within data centers and cloud environments, achieving nanosecond-level accuracy. This core technology was initially adopted by Fortune 100 financial companies for timestamping financial records and market data.
A significant shift occurred in 2020 with the realization that accurately synchronized clocks could enable precise measurement of packet latency. This led to the development of network telemetry capabilities, allowing for accurate latency monitoring in data center applications. Uber is a notable customer, deploying this technology across hundreds of thousands of VMs.
The evolution continued in 2024 with the application of this technology to GPU clusters. Clockwork achieved highly accurate telemetry for GPU-to-GPU communication, correlating network performance with AI training workloads and libraries like NCCL (Nvidia Collective Communications Library). Alongside monitoring, Clockwork developed dynamic traffic control to actively manage network flows, plugging into communication libraries such as NCCL, TCP, and RDMA.
The progression moved from measurement (clock sync) to control (dynamic traffic control) and then to managing the entire application for both fault tolerance and performance, extending up to PyTorch training workloads.
Addressing Fault Tolerance in GPU Clusters
A major pain point in large-scale GPU clusters is their availability, typically in the high 80s to low 90s percentage, significantly lower than cloud availability (three to four 9s). Network link flaps are a primary cause of downtime.
Network Link Flaps: These occur when connections between GPUs, often optical, are temporarily disrupted. Causes include overheating, dust on optical connectors, or cabling errors. While often transient, these disruptions can halt AI workloads.
Clockwork's approach to fault tolerance involves:
- Rapid Detection: Using their observability tools to quickly identify problematic network links.
- Automatic Failover: Implementing a watchdog timer that triggers an automatic failover when a link issue persists, preventing the problem from propagating to the application layer.
- Seamless Recovery: Monitoring for link recovery and automatically restoring the original connection without the application noticing any interruption.
This process is designed to be non-destructive and ensures that the AI training application continues without interruption.
Clockwork's Technical Stack and Deployment
Clockwork positions itself as a "software shim company." Their deployment involves:
- User-space Agents: Deployed on each GPU node for monitoring, providing deep telemetry without requiring kernel-level access.
- eBPF Modules: Used for controlling TCP flows, enabling dynamic traffic management at the network layer.
- RDMA Library Integration: An agent that communicates with the InfiniBand Verbs (IBS) layer for RDMA network management.
- NCCL/Rippl Integration: Plugging into Nvidia's collective communication libraries for fault tolerance and performance optimization within AI training.
- PyTorch Integration: A recently announced product that instruments PyTorch libraries to monitor and correlate training iteration boundaries with infrastructure performance.
Clockwork's scope spans from communication libraries (like NCCL) down to the transport layer (TCP and packet layers).
Customer Base and Market Position
Clockwork serves a diverse customer base, including:
- Hyperscalers: One of the top five hyperscalers is a customer, utilizing Clockwork's network monitoring as an internal tool.
- Neo Clouds: Multiple neo cloud providers are using GPU clusters with Clockwork's solutions.
- Tenants of Clouds/Hyperscalers: Clockwork's products are also relevant to end-users (tenants) of these clouds, particularly for workload monitoring and resilience in AI training.
Clockwork partners with cloud providers and hyperscalers to reach these end-users.
Common Failures and Remediation
Beyond network link flaps, Clockwork observes other common causes of downtime in GPU clusters:
- Memory Errors: Frequent errors in GPU nodes.
- PCI Bus Failures: GPUs failing to communicate over the PCI bus ("GPUs falling off the bus").
- Firmware Errors (e.g., "excid errors"): Sudden and unpredictable firmware breakdowns.
- Thermal Throttling/Overheating: GPUs being underclocked due to heat, leading to eventual failure.
While Clockwork currently ships remediation for network-related problems and performance bottlenecks, their roadmap includes extending fault tolerance to entire nodes or GPUs failing. They are working on automatically bringing pre-provisioned spare nodes online to maintain uninterrupted AI training.
The Challenge of Distributed State in AI Workloads
A key challenge in achieving resilience for AI workloads is managing distributed state. Unlike single-machine workloads where state can be somewhat tractable, AI workloads involve simultaneous state maintenance across thousands of GPUs. This makes capturing and maintaining a consistent state across the entire cluster during failures extremely difficult. Clockwork's understanding of distributed communication is central to building fault tolerance by comprehending this distributed state.
AI in Clockwork's Own Operations
Clockwork is leveraging AI internally to transform its vast observability data into actionable insights. The goal is to create an AI bot that can interact with customers in natural language, answering questions about their data center performance and issues. This aims to simplify the understanding of complex system data, moving beyond traditional dashboards.
Industry Perspective: The AI Boom and Potential Bubble
Sesh draws parallels between the current AI excitement and the dot-com bubble of the late 1990s, noting the rapid technological advancement and imaginative potential. However, he highlights key differences:
- Revenue Source: In the current AI boom, a significant portion of revenue for major players like Nvidia comes from the operating cash of large hyperscalers, rather than investor funding for nascent companies as was common during the dot-com era.
- Market Correction: While a stock market correction is possible, Sesh believes the underlying revenue streams are more robust, suggesting that revenue might not disappear as dramatically as it did during the dot-com bust.
Future Outlook for AI and the Job Market
- Supply Constraints and Demand: Demand for AI infrastructure in 2026 is largely secured through contracts by hyperscalers, indicating continued real revenue and deployment.
- Software Innovation: Software innovations in AI, including advancements in models (vision, multimodal) and agentic AI use cases, are expected to accelerate.
- Job Market Disruption: Sesh expresses concern about the potential job market dislocation over the next 4-10 years, anticipating seismic shifts in job creation and upheaval due to AI advancements. He hopes for a new equilibrium to be reached within 5-6 years.
Conclusion
Clockwork is at the forefront of optimizing the complex and demanding world of AI workloads running on large GPU clusters. Their platform, Fleet IQ, addresses critical challenges in visibility, fault tolerance, and performance by leveraging deep insights into network and communication layers. The company's evolution from clock synchronization to sophisticated network telemetry and control, coupled with their focus on automated remediation and understanding distributed state, positions them as a key enabler of resilient and efficient AI infrastructure. The broader industry is experiencing unprecedented growth, with potential economic shifts and significant impacts on the job market anticipated in the coming years.
Chat with this Video
AI-PoweredLoad the transcript when you're ready to chat so the initial page stays lighter.
Related Videos

Deterministic Infra for Non-Deterministic AI Agents - Nishant Gupta, Meta Superintelligence Labs
AI Engineer

'No where near normal' but 30-40 oil tankers passing through the Strait 'is better than 0': Mulberry
BNN Bloomberg

'Alphabet has such a dominant position they will be a leader in this space for many years': Clare
BNN Bloomberg

Forget Elon’s Data Centers In Space. This Startup Wants To Float Them At Sea
Forbes

Yahoo Finance Live: Daily Market Coverage - June 29, 2026 9AM-11AM (ET)
Yahoo Finance

Everyone's Buying AI. Smart Investors Are Buying This Instead. - Robert Kiyosaki
The Rich Dad Channel

Mad Money 06/26/26 | Audio Only
CNBC Television