---
format: "aidr-story-markdown/v1"
id: "ea763c119113e0f8edab04cd57a72e2d69f185c3e3b79f2295ecf9ef27563f1e"
canonical_url: "https://aidr.today/ea763c11?lang=en"
title: "DeepSeek V4-Flash-Vision-Exp Hits Opus 4.8 Performance in First Low Cost Vision API"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-21T11:45:45.000Z"
category: "Models"
topics: ["deepseek","vision","llm","inference","api","multi-agent","open-source","agent"]
source_urls: ["https://huggingnews.com/ai/deepseek-v4-flash-vision-exp-hits-opus-48-performance-in-first-low-cost-ad0b1bfc","https://x.com/deepseek_ai/status/2090730032574631962","https://x.com/deepseek_ai/status/2090730039973392531","https://x.com/deepseek_ai/status/2090730042586489333","https://x.com/MaxForAI/status/2090730253782229374","https://x.com/teortaxesTex/status/2090731115061317798","https://x.com/kimmonismus/status/2090738285140173088","https://x.com/stevibe/status/2090774141708415065"]
summary: "The company released an experimental multimodal model that allows AI agents to process visual inputs with efficiency matching its existing text-based offerings. DeepSeek-V4-Flash-Vision-Exp reaches performance levels close to Opus 4.8 on multimodal agent benchmarks while maintaining the low pricing of the Flash series. Images are tokenized for billing at a rate of up to 384 tokens per image. Along with the model, DeepSeek launched a free Files API to reduce request bandwidth by allowing users to reference previously uploaded images. Technical specifications for the vision model include a 1M token context window, a 384K token maximum output, and support for JSON output."
---

# DeepSeek V4\-Flash\-Vision\-Exp Hits Opus 4\.8 Performance in First Low Cost Vision API

> [Open the canonical story](<https://aidr.today/ea763c11?lang=en>)

**Published:** 2026-08-21T11:45:45.000Z
**Category:** Models
**Topics:** deepseek, vision, llm, inference, api, multi\-agent, open\-source, agent

## Summary

The company released an experimental multimodal model that allows AI agents to process visual inputs with efficiency matching its existing text\-based offerings\. DeepSeek\-V4\-Flash\-Vision\-Exp reaches performance levels close to Opus 4\.8 on multimodal agent benchmarks while maintaining the low pricing of the Flash series\. Images are tokenized for billing at a rate of up to 384 tokens per image\. Along with the model, DeepSeek launched a free Files API to reduce request bandwidth by allowing users to reference previously uploaded images\. Technical specifications for the vision model include a 1M token context window, a 384K token maximum output, and support for JSON output\.

## Sources

- [Story source](<https://huggingnews.com/ai/deepseek-v4-flash-vision-exp-hits-opus-48-performance-in-first-low-cost-ad0b1bfc>)
- [Story source](<https://x.com/deepseek_ai/status/2090730032574631962>)
- [Story source](<https://x.com/deepseek_ai/status/2090730039973392531>)
- [Story source](<https://x.com/deepseek_ai/status/2090730042586489333>)
- [Supporting source](<https://x.com/MaxForAI/status/2090730253782229374>)
- [Supporting source](<https://x.com/teortaxesTex/status/2090731115061317798>)
- [Supporting source](<https://x.com/kimmonismus/status/2090738285140173088>)
- [Story source](<https://x.com/stevibe/status/2090774141708415065>)

