The open model is a full-weights fine-tune of NVIDIA Nemotron Nano 30B trained on telephone and customer support use cases. It delivers GPT 5.6 Terra performance on voice agent tasks with 1/3 the latency and 1/18 the cost, targeting real-time applications where traditional "thinking" compute is disabled to maintain conversation speed.

Deployment on a single NVIDIA B200 GPU supports over 80 concurrent agents with a P95 end-to-end time-to-first-audio-token under 600ms, bringing the cost per minute to approximately $0.0025. The project also introduces PhoneBench, an internal benchmark for real-time AI, and makes weights available on Hugging Face.

Sign in to suggest edits

Key sources

  1. SOURCE@kwindla“TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200”x.com
  2. SUPPORT@kwindla“Ben Shababo at @modal did a bunch of great inference optimization work to achieve the >80 concurrent clients”x.com
  3. SUPPORT@kwindla“We can curate and manage production agent traces, generate synthetic data, create RL environments for training multi-turn models”x.com
Markdown