---
format: "aidr-story-markdown/v1"
id: "094a3a6ba9b159f01f8ac843999936172c5fe12cff7a7aa5363f22dee9d25634"
canonical_url: "https://aidr.today/094a3a6b?lang=en"
title: "Fault tolerant distributed training on Amazon EKS using NVRx"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-16T18:59:25.000Z"
category: "Infra"
topics: ["nvidia","chips","distributed-training","infra","pytorch","eks"]
source_urls: ["https://aws.amazon.com/blogs/machine-learning/fault-tolerant-distributed-training-on-amazon-eks-using-nvrx/"]
summary: "Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery."
---

# Fault tolerant distributed training on Amazon EKS using NVRx

> [Open the canonical story](<https://aidr.today/094a3a6b?lang=en>)

**Published:** 2026-09-16T18:59:25.000Z
**Category:** Infra
**Topics:** nvidia, chips, distributed\-training, infra, pytorch, eks

## Summary

Integrate NVIDIA Resiliency Extension \(NVRx\) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds\. This post covers async checkpointing, in\-process restart, and ft\_launcher in\-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%\+ training efficiency and second\-scale recovery\.

## Sources

- [Story source](<https://aws.amazon.com/blogs/machine-learning/fault-tolerant-distributed-training-on-amazon-eks-using-nvrx/>)

