---
format: "aidr-story-markdown/v1"
id: "d30ab09b3c36279442e9325d75c07fb2c39686381422c546aaf673bfa17f4110"
canonical_url: "https://aidr.today/d30ab09b?lang=en"
title: "Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-09T08:18:10.000Z"
category: "Infra"
topics: ["tenstorrent","vllm","inference","infra","open-source","hardware"]
source_urls: ["https://vllm.ai/blog/2026-09-07-vllm-tt-plugin","https://lobste.rs/s/twvlv6/serving_llms_on_tenstorrent_hardware"]
summary: "Tenstorrent accelerators join vLLM as an out-of-tree platform plugin, driven by mesh-architecture choices: phase-based scheduling, single-process data paralleli"
---

# Serving LLMs on Tenstorrent Hardware: Inside the vLLM TT Plugin

> [Open the canonical story](<https://aidr.today/d30ab09b?lang=en>)

**Published:** 2026-09-09T08:18:10.000Z
**Category:** Infra
**Topics:** tenstorrent, vllm, inference, infra, open\-source, hardware

## Summary

Tenstorrent accelerators join vLLM as an out\-of\-tree platform plugin, driven by mesh\-architecture choices: phase\-based scheduling, single\-process data paralleli

## Sources

- [Story source](<https://vllm.ai/blog/2026-09-07-vllm-tt-plugin>)
- [Discussion](<https://lobste.rs/s/twvlv6/serving_llms_on_tenstorrent_hardware>)

