KAYTUS has introduced an OEM AI Managed Service for AI data centers and hyperscale compute clusters, combining onsite hardware maintenance, locally stocked replacement components, diagnostic equipment and certified field engineers. The company said the service is aimed at reducing the time needed to recover failed nodes and preserve available computing capacity for training and inference workloads.
The launch reflects a shift in AI infrastructure from rapid installation toward long-term operation. Large GPU clusters combine dense compute, networking and cooling equipment, and a single node failure can interrupt customer workloads or constrain schedulable capacity. KAYTUS said its service is intended to bring parts, diagnostics and repair capability to a customer site rather than rely on regional warehouses or factory repair cycles.
Under the service model, KAYTUS said it can stock compute nodes, network switches and high-bandwidth network interface cards in a customer data center. It also plans to deploy factory diagnostic and repair tools onsite so engineers can assess and repair individual components or complete nodes without returning systems to a factory. The company said complex hardware repairs can be completed in as little as four hours under the model, subject to the applicable service arrangement.
The offering also includes around-the-clock support in applicable service tiers, with certified onsite engineers backed by Tier 2 specialists. For infrastructure teams, the operational benefit is a more direct route from a hardware fault to diagnosis and restoration. KAYTUS said prolonged recovery can leave billed compute capacity unavailable and increase a provider’s exposure to service-level commitments.
A second component uses AI-assisted failure prediction and scheduled health assessments to identify signs of hardware degradation before an interruption occurs. This is the automation element of the service: preventive inspection and prediction are intended to reduce reliance on reactive repair. The company did not disclose the underlying models, supported telemetry sources or false-positive rates, so customers will need to assess how that prediction capability fits with their existing monitoring and incident-management processes.
KAYTUS cited results from a deployment at an unnamed global cloud service provider supporting more than 100 racks and thousands of accelerators. According to the company, average handling time fell from 48 hours to 12 hours per failed node, while productive compute time increased by 50%, primarily through shorter downtime and faster recovery. Those are vendor-reported results from a customer deployment, not independently verified benchmarks or a guarantee of outcomes for other data centers.
The company said the service is expanding across selected European and Asia-Pacific markets, with service levels and deployment models tailored to a customer’s infrastructure scale. It can also integrate with KAYTUS’s KSManage intelligent operations platform, which the company describes as adding AI-powered operations management to onsite support.
For AI data center operators, the development is less about provisioning new infrastructure than about managing the hardware already in production. A service that combines local spares, onsite diagnostics and predictive inspections could reduce the manual coordination required during failures and make recovery commitments more measurable. The result will depend on each site’s parts inventory, service tier, infrastructure mix and maintenance procedures.
