---
format: "aidr-story-markdown/v1"
id: "40ebeed3ac0fbd10af987688238b2c46adad5a28774cc7b1d27166c1e5b486ab"
canonical_url: "https://aidr.today/40ebeed3?lang=en"
title: "PhoneLLM Matches GPT 5.6 Terra at 1/18 Cost to Top Third Party Latency"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-27T23:51:11.000Z"
category: "Models"
topics: ["nvidia","open-source","agent","inference","benchmark"]
source_urls: ["https://huggingnews.com/ai/phonellm-matches-gpt-56-terra-at-118-cost-to-top-third-party-latency-860728f7","https://x.com/kwindla/status/2093014818647339026","https://x.com/kwindla/status/2093014822145372491","https://x.com/kwindla/status/2093014824687136887"]
summary: "The open model is a full-weights fine-tune of NVIDIA Nemotron Nano 30B trained on telephone and customer support use cases. It delivers GPT 5.6 Terra performance on voice agent tasks with 1/3 the latency and 1/18 the cost, targeting real-time applications where traditional \"thinking\" compute is disabled to maintain conversation speed. Deployment on a single NVIDIA B200 GPU supports over 80 concurrent agents with a P95 end-to-end time-to-first-audio-token under 600ms, bringing the cost per minute to approximately $0.0025. The project also introduces PhoneBench, an internal benchmark for real-time AI, and makes weights available on Hugging Face."
---

# PhoneLLM Matches GPT 5\.6 Terra at 1/18 Cost to Top Third Party Latency

> [Open the canonical story](<https://aidr.today/40ebeed3?lang=en>)

**Published:** 2026-08-27T23:51:11.000Z
**Category:** Models
**Topics:** nvidia, open\-source, agent, inference, benchmark

## Summary

The open model is a full\-weights fine\-tune of NVIDIA Nemotron Nano 30B trained on telephone and customer support use cases\. It delivers GPT 5\.6 Terra performance on voice agent tasks with 1/3 the latency and 1/18 the cost, targeting real\-time applications where traditional "thinking" compute is disabled to maintain conversation speed\. Deployment on a single NVIDIA B200 GPU supports over 80 concurrent agents with a P95 end\-to\-end time\-to\-first\-audio\-token under 600ms, bringing the cost per minute to approximately $0\.0025\. The project also introduces PhoneBench, an internal benchmark for real\-time AI, and makes weights available on Hugging Face\.

## Sources

- [Story source](<https://huggingnews.com/ai/phonellm-matches-gpt-56-terra-at-118-cost-to-top-third-party-latency-860728f7>)
- [Story source](<https://x.com/kwindla/status/2093014818647339026>)
- [Supporting source](<https://x.com/kwindla/status/2093014822145372491>)
- [Supporting source](<https://x.com/kwindla/status/2093014824687136887>)

