---
format: "aidr-story-markdown/v1"
id: "93a9625d3d92698edb1768ce0502bff01fa0bc12b1142e18ffbeea0547c44cfc"
canonical_url: "https://aidr.today/93a9625d?lang=en"
title: "Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-28T16:15:46.000Z"
category: "Infra"
topics: ["amazon","llm","inference"]
source_urls: ["https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-vllm-omni-on-sagemaker-ai-part-1/","https://aws.amazon.com/blogs/machine-learning/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2/"]
summary: "Deploy a text-to-speech model on Amazon SageMaker AI with the AWS vLLM-Omni Deep Learning Container and stream generated speech over a persistent bidirectional connection. This Part 1 tutorial deploys Qwen3-TTS and streams speech through a Gradio application."
---

# Build real\-time voice applications with vLLM\-Omni on SageMaker AI – Part 1

> [Open the canonical story](<https://aidr.today/93a9625d?lang=en>)

**Published:** 2026-09-28T16:15:46.000Z
**Category:** Infra
**Topics:** amazon, llm, inference

## Summary

Deploy a text\-to\-speech model on Amazon SageMaker AI with the AWS vLLM\-Omni Deep Learning Container and stream generated speech over a persistent bidirectional connection\. This Part 1 tutorial deploys Qwen3\-TTS and streams speech through a Gradio application\.

## Sources

- [Story source](<https://aws.amazon.com/blogs/machine-learning/build-real-time-voice-applications-with-vllm-omni-on-sagemaker-ai-part-1/>)
- [Story source](<https://aws.amazon.com/blogs/machine-learning/generate-images-and-video-with-vllm-omni-on-sagemaker-ai-part-2/>)

