← Back to live feed1 story
The company released an experimental multimodal model that allows AI agents to process visual inputs with efficiency matching its existing text-based offerings. DeepSeek-V4-Flash-Vision-Exp reaches performance levels close to Opus 4.8 on multimodal agent benchmarks while maintaining the low pricing of the Flash series. Images are tokenized for billing at a rate of up to 384 tokens per image.
Along with the model, DeepSeek launched a free Files API to reduce request bandwidth by allowing users to reference previously uploaded images. Technical specifications for the vision model include a 1M token context window, a 384K token maximum output, and support for JSON output.
Sign in to suggest edits
Key sources
- SOURCE@deepseek_ai“multimodal agent performance close to Opus-4.8”x.com
- SOURCE@deepseek_ai“Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing”x.com
- SOURCE@deepseek_ai“Upload an image once, then reference it by file_id to save request bandwidth”x.com
- SUPPORT@maxforai“Vision 版和 V4 Flash 完全一个价”x.com
- SUPPORT@teortaxestex“RIP Whale Inference Fleet again”x.com
- SUPPORT@kimmonismus“this is the Flash model, the small one!”x.com
- SOURCE@stevibe“the official API!”x.com
- SUPPORT@omarsar0“DeepSeek-V4-Flash-Vision-Exp advances multimodal agent performance.”x.com