AMD publicly released a software image for running DeepSeek V4.1 Flash two days after Nvidia’s CUDA version of vLLM supported the model, SemiAnalysis said. The image works out of the box, but Nvidia’s B200 and B300 deliver up to 42 times its performance per dollar, while the H200 delivers up to 14.8 times, the research firm said.

SemiAnalysis had previously reported that the image specified in AMD’s vLLM documentation remained unavailable 23 hours after the model’s release, while Nvidia’s implementation worked across six chip models at launch. The vLLM project, however, announced on Sept. 10 that its launch support was “verified on NVIDIA and AMD GPUs!” SemiAnalysis attributed Nvidia’s advantage to CUDA’s software ecosystem and collaboration with a community of six million developers.

Sign in to suggest edits

Key sources

  1. SUPPORT@semianalysis_“2 days after CUDA vLLM supported DeepSeekv4.1 Flash, AMD finally publicly released its DeepSeek v4.1 Flash image. Functionally, it works out of the box, but performance-wise, it is currently up to 14.8x worse perf per dollar than H200 and up to 42x worse perf per dollar than B200/B300 currently.”x.com
  2. SOURCE@designarena“DeepSeek‑V4.1‑Flash takes 6th overall on Design Arena with an Elo of 1347!”x.com
  3. SUPPORT@semianalysis_“On the Day 0 release of DeepSeekv4.1 Flash, NVIDIA vLLM works out of the box with zero issues across all 6 SKUs: H100, H200, B200, B300, GB200, GB300!”x.com
  4. SUPPORT@semianalysis_“AMD's vLLM documentation points to using vllm/vllm-openai-rocm:deepseekv41-flash-0909, but from hour 0 of the model release to now, hour 23, AMD has still not publicly released the image.”x.com
Markdown