---
format: "aidr-story-markdown/v1"
id: "8824a722fbbd4c31d76b2acd736a570a76b2884eb3580f8b8268fa380b4d0ed5"
canonical_url: "https://aidr.today/8824a722?lang=en"
title: "Fast Polynomial Transcendentals for LLMs"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-10-02T04:00:00.000Z"
category: "Research"
topics: ["nvidia","llm"]
source_urls: ["https://arxiv.org/abs/2610.00049"]
summary: "arXiv:2610.00049v1 Announce Type: new Abstract: Graphics processing unit (GPU) generations scale matrix, special-function, and memory pipelines at different rates, so kernel bottlenecks move as hardware evolves. FlashAttention-4 exposed this imbalance inside attention on NVIDIA Blackwell. We test whether short polynomial programs can accelerate other special-function-unit (SFU) operations in large language models (LLMs). We first compare native PyTorch evaluation with packed fused multiply--add (FMA) programs in an isolated IEEE binary16 (FP16) sweep spanning L2-resident and high-bandwidth-memory (HBM)-resident working sets. We then replace native sigmoid, tanh, and sigmoid linear unit (SiLU) with degree-3 or degree-4 bfloat16 (BF16) programs in four GB200 integration tasks: dense SiLU, tanh-softcapped attention, sigmoid attention, and routed-expert Swish-gated linear unit (SwiGLU). The programs combine analytical symmetry, target-format rounding, and packed arithmetic inside consuming kernels. The isolated paths improve by 1.19--2.19x in L2 and 1.00--1.70x in HBM. The dense-SiLU, tanh-softcapped-attention, and routed-expert substitutions improve complete training-step throughput b"
---

# Fast Polynomial Transcendentals for LLMs

> [Open the canonical story](<https://aidr.today/8824a722?lang=en>)

**Published:** 2026-10-02T04:00:00.000Z
**Category:** Research
**Topics:** nvidia, llm

## Summary

arXiv:2610\.00049v1 Announce Type: new Abstract: Graphics processing unit \(GPU\) generations scale matrix, special\-function, and memory pipelines at different rates, so kernel bottlenecks move as hardware evolves\. FlashAttention\-4 exposed this imbalance inside attention on NVIDIA Blackwell\. We test whether short polynomial programs can accelerate other special\-function\-unit \(SFU\) operations in large language models \(LLMs\)\. We first compare native PyTorch evaluation with packed fused multiply\-\-add \(FMA\) programs in an isolated IEEE binary16 \(FP16\) sweep spanning L2\-resident and high\-bandwidth\-memory \(HBM\)\-resident working sets\. We then replace native sigmoid, tanh, and sigmoid linear unit \(SiLU\) with degree\-3 or degree\-4 bfloat16 \(BF16\) programs in four GB200 integration tasks: dense SiLU, tanh\-softcapped attention, sigmoid attention, and routed\-expert Swish\-gated linear unit \(SwiGLU\)\. The programs combine analytical symmetry, target\-format rounding, and packed arithmetic inside consuming kernels\. The isolated paths improve by 1\.19\-\-2\.19x in L2 and 1\.00\-\-1\.70x in HBM\. The dense\-SiLU, tanh\-softcapped\-attention, and routed\-expert substitutions improve complete training\-step throughput b

## Sources

- [Story source](<https://arxiv.org/abs/2610.00049>)

