Decoding 'Validate GPU Cluster Readiness Before AI Workloads Land' from Official Information

AI・機械学習カテゴリを表すパンダのイラスト AI & Machine Learning

This article is a technical explanation and implementation example generated using AI. Although the code and procedures presented are based on primary sources, they have not been verified on actual hardware by the author. Behavior may vary depending on the environment and versions.

Before running large-scale AI workloads, it is necessary to safely and practically verify the readiness of the GPU cluster. Based on official information from the NVIDIA Cluster Readiness Engine (NVCRE)—an open-source Kubernetes controller provided by NVIDIA—this article outlines its mechanism, components, and key considerations for use.


Purpose of NVIDIA Cluster Readiness Engine (NVCRE)

Even when all GPUs, network links, and pods report as healthy, a GPU cluster can experience unexpected performance degradation or failures during large-scale distributed training. Potential causes include a single degraded GPU or a network link that deteriorates under load.

NVCRE is an open-source Kubernetes controller that executes actual distributed workloads on topology-aware node groups prior to production workloads, measures the results, and identifies nodes that fail the tests. It reduces the effort required for operators to manually create manifests or isolate racks individually, proving the 'ready state' of the cluster as an objective fact.

flowchart TD
    A[Certification] --> B[Workflow: コミュニケーション / トレーニング]
    B --> C[Job: ターゲットノードグループのワークロード実行]
    C --> D[結果の集約・特定ノードの障害判定]
    D --> E[NVSentinel 連携による隔離・検知]

Three-Tier API Structure

Operations in NVCRE are managed through three hierarchical APIs composed of Custom Resource Definitions (CRDs).

  • Certification: The top-level resource that specifies the target nodes for testing and the categories to be executed.

  • Workflow: Manages a single category. It handles catalogs, platform and GPU override applications, iteration count management, orchestration target settings, and the creation of child jobs.

  • Job: Executes workloads against the target node group, monitors node health, and records metrics and failures.

This hierarchical structure makes it possible to clearly link the cause of a failure to a specific category on a specific node.


Catalog and Pass Condition Expression Evaluation

The built-in catalog covers three domains: NCCL communication variants (all-reduce, all-gather, all-to-all, loopback, and inter-NVSwitch loopback), the NVIDIA DCGM Level 4 diagnostic suite, and NVIDIA NeMo pre-training (8B and 56B parameter Nemotron 5 models).

Common Expression Language (CEL) is used to evaluate pass/fail judgments against the measured metrics. Thresholds are not bundled by default. For example, the following configuration structure is demonstrated.

categories:

    - domain: communication
      variant: nccl-all-reduce
      options:
        thresholds:
          busBandwidthGBps: "value >= 900"

    - domain: training
      variant: nemotron5-8b
      options:
        thresholds:
          goodputRatio: "value >= 0.9"
          avgTFLOPsPerGPU: "value >= 800"

If target values are not met, NVCRE sets a 'ValidationFailed' condition, recording it separately from the success or failure of the execution itself.


Test Scale and Adaptive Fault Isolation

To handle cases where failures are only visible at specific scales,testScale you can specify the grouping strategy in the field.

  • Intra-node: Test each node individually.

  • Intra-rack: nvidia.com/gpu.clique: Use labels to partition nodes by topology domain.

  • Full-scale: Group all nodes into a single pool.

  • Diagnose: Perform adaptive failure isolation.

The biggest challenge in multi-node validation is failures such as bandwidth degradation that do not originate from a single node, buttestScale: diagnose when configured, executes topology-aware hierarchical group testing. The engine splits failed groups and repeats execution until the minimum group size (minGroupSize) is reached, isolating a small subset of suspect nodes rather than the entire impact scope.


Flexible Execution via the WorkloadRun API

To reduce the configuration overhead of running multi-node GPU workloads on Kubernetes,WorkloadRun the API is provided.

apiVersion: nvcre.nvidia.com/v1alpha1
kind: WorkloadRun
metadata:
  name: nccl-all-reduce
spec:
  image: nvcr.io/nvidia/pytorch:26.01-py3
  framework:
    mpi:
      binary: /usr/local/bin/all_reduce_perf_mpi
      args: ["-b", "8", "-e", "32G", "-f", "2", "-n", "100"]
      mpirunPath: /usr/local/mpi/bin/mpirun
  numNodes: 4
  bandwidthMeasurement:
    logProfileRef: nccl-bandwidth
    testType: all_reduce

framework The field allows you to select one oftorch、mpi、exec. Additionally, to prevent deadlocks in busy clusters, using a gang-aware scheduler (such as the KAI Scheduler) is recommended.


Positioning within the NVIDIA DSX OS Ecosystem

NVCRE is a component of the NVIDIA DSX OS operational layer and integrates with other components.

  • NVIDIA AI Cluster Runtime (AICR): Maintains validated cluster configurations through version-pinned recipes for drivers, operators, kernels, system settings, and more.

  • NVIDIA Cluster Readiness Engine (NVCRE): Actively generates workloads to proactively validate failures that do not appear in telemetry.

  • NVSentinel: Passively monitors DCGM metrics and Xid errors for continuous health tracking.

Although NVCRE itself is designed not to cordon or taint nodes, it can translate failed validation results into health events via the NVSentinel NVCRE Certification Monitor, enabling node isolation or workload draining.


Prerequisites and Usage Notes

The prerequisites and notes from the official documentation are as follows.

  • Environment Requirements: Requires Kubernetes 1.29 or later,kubectl, Helm 3.x, and the NVIDIA GPU Operator on the target cluster.

  • Specific Hardware Requirements: NVIDIA GB200 NVL72 and GB300 NVL72 catalog entries require the NVIDIA DRA Driver for GPUs to create ComputeDomain resources.

  • DCGM Requirements: DCGM Level 4 categories require a standalone DCGM service.

  • Note on Verification with Actual Hardware: The configurations and commands introduced in this article are based on primary sources and represent research information prior to actual hardware verification. Because actual deployment depends on the target environment's version and network configuration, please refer to the official documentation.


Conclusion

This article summarizes the overview of NVCRE, its API hierarchy, CEL-based evaluation expressions, test scaling, the WorkloadRun API, and the roles of the ecosystem, based on the primary source of Validate GPU Cluster Readiness Before AI Workloads Land.

The points to check before execution and the constraints are as follows.

  • Verify the Kubernetes version of the cluster, the GPU Operator, and the installation status of dedicated drivers as needed.

  • Consider the impact that active load tests may have on cluster performance and the network fabric prior to production deployment.

  • Design the integration strategy with tools such as NVSentinel to receive the identification results of failed nodes.


Reference Information

Update History of This Article

The content of this article has been reviewed through an automated review and update workflow leveraging generative AI, and necessary corrections have been applied.

September 26, 2026

  • ModificationRevised the unnatural phrasing starting with a comma at the beginning of the summary section to "Based on the primary source of Validate GPU Cluster Readiness Before AI Workloads Land, this article…"

Document information

Article title
Decoding 'Validate GPU Cluster Readiness Before AI Workloads Land' from Official Information
Published
Updated
Source
https://papanda925.com/?p=17908&lang=en

License: Text and original figures for which this site holds the relevant rights are available under CC BY 4.0 , unless otherwise noted. This article may include content created or edited with generative AI. If code has a separate license notice or a linked GitHub repository license, that license takes precedence for the code. Quotations, third-party materials, images, and trademarks are excluded from this license. Usage policy

Copied title and URL