---
format: "aidr-story-markdown/v1"
id: "b8e2d01df43a03984d3af2635ee11eaee4073bd6695b87dbd3a0188102240120"
canonical_url: "https://aidr.today/b8e2d01d?lang=en"
title: "HeyGen x Google Cloud: Bringing Avatar IV to TPUs"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-15T19:46:17.000Z"
category: "Infra"
topics: []
source_urls: ["https://developers.googleblog.com/heygen-x-google-cloud-bringing-avatar-iv-to-tpus/"]
summary: "HeyGen ported their 18B+ parameter Avatar IV video generation model to Google Cloud's Trillium (v6e) TPUs via torchax and XLA, utilizing FSDP and Ulysses sequence parallelism across an eight-chip mesh. To achieve a 1.86x speedup for real-time streaming, the engineering team pipelined exposed all-to-all collectives, aligned sparse attention block sizes to eliminate mask padding, and bypassed softmax serial dependencies using a precomputed Cauchy-Schwarz upper bound. These custom Pallas kernel and compiler optimizations were deployed only after passing rigorous two-tier quality gates to guarantee byte-identical or mathematically equivalent pixel outputs."
---

# HeyGen x Google Cloud: Bringing Avatar IV to TPUs

> [Open the canonical story](<https://aidr.today/b8e2d01d?lang=en>)

**Published:** 2026-09-15T19:46:17.000Z
**Category:** Infra

## Summary

HeyGen ported their 18B\+ parameter Avatar IV video generation model to Google Cloud's Trillium \(v6e\) TPUs via torchax and XLA, utilizing FSDP and Ulysses sequence parallelism across an eight\-chip mesh\. To achieve a 1\.86x speedup for real\-time streaming, the engineering team pipelined exposed all\-to\-all collectives, aligned sparse attention block sizes to eliminate mask padding, and bypassed softmax serial dependencies using a precomputed Cauchy\-Schwarz upper bound\. These custom Pallas kernel and compiler optimizations were deployed only after passing rigorous two\-tier quality gates to guarantee byte\-identical or mathematically equivalent pixel outputs\.

## Sources

- [Story source](<https://developers.googleblog.com/heygen-x-google-cloud-bringing-avatar-iv-to-tpus/>)

