This article is a technical explanation and implementation example generated using AI. Although the code and procedures presented are based on primary sources, they have not been verified on actual hardware by the author. Behavior may vary depending on the environment and versions.
Before running large-scale AI workloads, it is necessary to safely and practically verify the readiness of the GPU cluster. Based on official information from the NVIDIA Cluster Readiness Engine (NVCRE)—an open-source Kubernetes controller provided by NVIDIA—this article outlines its mechanism, components, and key considerations for use.
- Purpose of NVIDIA Cluster Readiness Engine (NVCRE)
- Three-Tier API Structure
- Catalog and Pass Condition Expression Evaluation
- Test Scale and Adaptive Fault Isolation
- Flexible Execution via the WorkloadRun API
- Positioning within the NVIDIA DSX OS Ecosystem
- Prerequisites and Usage Notes
- Conclusion
- Reference Information
- Update History of This Article
Purpose of NVIDIA Cluster Readiness Engine (NVCRE)
Even when all GPUs, network links, and pods report as healthy, a GPU cluster can experience unexpected performance degradation or failures during large-scale distributed training. Potential causes include a single degraded GPU or a network link that deteriorates under load.
NVCRE is an open-source Kubernetes controller that executes actual distributed workloads on topology-aware node groups prior to production workloads, measures the results, and identifies nodes that fail the tests. It reduces the effort required for operators to manually create manifests or isolate racks individually, proving the 'ready state' of the cluster as an objective fact.
flowchart TD
A[Certification] --> B[Workflow: コミュニケーション / トレーニング]
B --> C[Job: ターゲットノードグループのワークロード実行]
C --> D[結果の集約・特定ノードの障害判定]
D --> E[NVSentinel 連携による隔離・検知]
Three-Tier API Structure
Operations in NVCRE are managed through three hierarchical APIs composed of Custom Resource Definitions (CRDs).
Certification: The top-level resource that specifies the target nodes for testing and the categories to be executed.
Workflow: Manages a single category. It handles catalogs, platform and GPU override applications, iteration count management, orchestration target settings, and the creation of child jobs.
Job: Executes workloads against the target node group, monitors node health, and records metrics and failures.
This hierarchical structure makes it possible to clearly link the cause of a failure to a specific category on a specific node.
Catalog and Pass Condition Expression Evaluation
The built-in catalog covers three domains: NCCL communication variants (all-reduce, all-gather, all-to-all, loopback, and inter-NVSwitch loopback), the NVIDIA DCGM Level 4 diagnostic suite, and NVIDIA NeMo pre-training (8B and 56B parameter Nemotron 5 models).
Common Expression Language (CEL) is used to evaluate pass/fail judgments against the measured metrics. Thresholds are not bundled by default. For example, the following configuration structure is demonstrated.
categories:
- domain: communication
variant: nccl-all-reduce
options:
thresholds:
busBandwidthGBps: "value >= 900"
- domain: training
variant: nemotron5-8b
options:
thresholds:
goodputRatio: "value >= 0.9"
avgTFLOPsPerGPU: "value >= 800"
If target values are not met, NVCRE sets a 'ValidationFailed' condition, recording it separately from the success or failure of the execution itself.
Test Scale and Adaptive Fault Isolation
To handle cases where failures are only visible at specific scales,testScale you can specify the grouping strategy in the field.
Intra-node: Test each node individually.
Intra-rack:
nvidia.com/gpu.clique: Use labels to partition nodes by topology domain.Full-scale: Group all nodes into a single pool.
Diagnose: Perform adaptive failure isolation.
The biggest challenge in multi-node validation is failures such as bandwidth degradation that do not originate from a single node, buttestScale: diagnose when configured, executes topology-aware hierarchical group testing. The engine splits failed groups and repeats execution until the minimum group size (minGroupSize) is reached, isolating a small subset of suspect nodes rather than the entire impact scope.
Flexible Execution via the WorkloadRun API
To reduce the configuration overhead of running multi-node GPU workloads on Kubernetes,WorkloadRun the API is provided.
apiVersion: nvcre.nvidia.com/v1alpha1
kind: WorkloadRun
metadata:
name: nccl-all-reduce
spec:
image: nvcr.io/nvidia/pytorch:26.01-py3
framework:
mpi:
binary: /usr/local/bin/all_reduce_perf_mpi
args: ["-b", "8", "-e", "32G", "-f", "2", "-n", "100"]
mpirunPath: /usr/local/mpi/bin/mpirun
numNodes: 4
bandwidthMeasurement:
logProfileRef: nccl-bandwidth
testType: all_reduce
framework The field allows you to select one oftorch、mpi、exec. Additionally, to prevent deadlocks in busy clusters, using a gang-aware scheduler (such as the KAI Scheduler) is recommended.
Positioning within the NVIDIA DSX OS Ecosystem
NVCRE is a component of the NVIDIA DSX OS operational layer and integrates with other components.
NVIDIA AI Cluster Runtime (AICR): Maintains validated cluster configurations through version-pinned recipes for drivers, operators, kernels, system settings, and more.
NVIDIA Cluster Readiness Engine (NVCRE): Actively generates workloads to proactively validate failures that do not appear in telemetry.
NVSentinel: Passively monitors DCGM metrics and Xid errors for continuous health tracking.
Although NVCRE itself is designed not to cordon or taint nodes, it can translate failed validation results into health events via the NVSentinel NVCRE Certification Monitor, enabling node isolation or workload draining.
Prerequisites and Usage Notes
The prerequisites and notes from the official documentation are as follows.
Environment Requirements: Requires Kubernetes 1.29 or later,
kubectl, Helm 3.x, and the NVIDIA GPU Operator on the target cluster.Specific Hardware Requirements: NVIDIA GB200 NVL72 and GB300 NVL72 catalog entries require the NVIDIA DRA Driver for GPUs to create ComputeDomain resources.
DCGM Requirements: DCGM Level 4 categories require a standalone DCGM service.
Note on Verification with Actual Hardware: The configurations and commands introduced in this article are based on primary sources and represent research information prior to actual hardware verification. Because actual deployment depends on the target environment's version and network configuration, please refer to the official documentation.
Conclusion
This article summarizes the overview of NVCRE, its API hierarchy, CEL-based evaluation expressions, test scaling, the WorkloadRun API, and the roles of the ecosystem, based on the primary source of Validate GPU Cluster Readiness Before AI Workloads Land.
The points to check before execution and the constraints are as follows.
Verify the Kubernetes version of the cluster, the GPU Operator, and the installation status of dedicated drivers as needed.
Consider the impact that active load tests may have on cluster performance and the network fabric prior to production deployment.
Design the integration strategy with tools such as NVSentinel to receive the identification results of failed nodes.
Reference Information
Update History of This Article
The content of this article has been reviewed through an automated review and update workflow leveraging generative AI, and necessary corrections have been applied.
September 26, 2026
- ModificationRevised the unnatural phrasing starting with a comma at the beginning of the summary section to "Based on the primary source of Validate GPU Cluster Readiness Before AI Workloads Land, this article…"

