---
format: "aidr-story-markdown/v1"
id: "4a69b9adf92e074e28a0dbbdaf1a7e3229de45daa17db5c24318b2ab135001cf"
canonical_url: "https://aidr.today/4a69b9ad?lang=en"
title: "How We Made a Text-to-Speech Model Respond in Sub-50 ms"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-21T15:51:10.000Z"
category: "Models"
topics: ["qwen","text-to-speech","inference","chips"]
source_urls: ["https://nari-labs.com/blog/qwen3-tts-speed-cost-frontier/","https://news.ycombinator.com/item?id=49389952"]
summary: "10 RPS with p95 TTFA under 50 ms on a single H100: how we optimized Qwen3-TTS serving."
---

# How We Made a Text\-to\-Speech Model Respond in Sub\-50 ms

> [Open the canonical story](<https://aidr.today/4a69b9ad?lang=en>)

**Published:** 2026-08-21T15:51:10.000Z
**Category:** Models
**Topics:** qwen, text\-to\-speech, inference, chips

## Summary

10 RPS with p95 TTFA under 50 ms on a single H100: how we optimized Qwen3\-TTS serving\.

## Sources

- [Story source](<https://nari-labs.com/blog/qwen3-tts-speed-cost-frontier/>)
- [Discussion](<https://news.ycombinator.com/item?id=49389952>)

