The company released an experimental multimodal model that allows AI agents to process visual inputs with efficiency matching its existing text-based offerings. DeepSeek-V4-Flash-Vision-Exp reaches performance levels close to Opus 4.8 on multimodal agent benchmarks while maintaining the low pricing of the Flash series. Images are tokenized for billing at a rate of up to 384 tokens per image.

Along with the model, DeepSeek launched a free Files API to reduce request bandwidth by allowing users to reference previously uploaded images. Technical specifications for the vision model include a 1M token context window, a 384K token maximum output, and support for JSON output.

Sign in to suggest edits

Key sources

  1. SOURCE@deepseek_ai“multimodal agent performance close to Opus-4.8”x.com
  2. SOURCE@deepseek_ai“Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing”x.com
  3. SOURCE@deepseek_ai“Upload an image once, then reference it by file_id to save request bandwidth”x.com
  4. SUPPORT@maxforai“Vision 版和 V4 Flash 完全一个价”x.com
  5. SUPPORT@teortaxestex“RIP Whale Inference Fleet again”x.com
  6. SUPPORT@kimmonismus“this is the Flash model, the small one!”x.com
  7. SOURCE@stevibe“the official API!”x.com
  8. SUPPORT@omarsar0“DeepSeek-V4-Flash-Vision-Exp advances multimodal agent performance.”x.com
Markdown