Chinese AI developer DeepSeek has unveiled an experimental multimodal version of its V4-Flash model. The new release, named V4-Flash-Vision-Exp, differs fundamentally from the base variant in that it accepts not only text but also images as input. In doing so, the company is expanding the functionality of its flagship V4 line to include the processing of visual data.

What DeepSeek Announced

According to the developer, the experimental V4-Flash-Vision-Exp model can analyze images, screenshots, and other visual inputs alongside text. The company emphasizes that the new version is neither a "trimmed" nor a specialized build: it retains the full set of capabilities of the base model, including text processing, logical reasoning, world-knowledge application, and autonomous AI agent control. In essence, multimodality has been added as an additional layer on top of the already functioning text core.

Model Capabilities and Positioning

DeepSeek places particular emphasis on progress in multimodal agent scenarios. "In multimodal agent benchmarks, V4-Flash-Vision-Exp demonstrates significant progress compared to [the base] V4-Flash, lifting multimodal agent performance to a level close to Opus-4.8," the company stated. This positioning is significant: the developer directly compares its experimental model to one of its competitors' top-tier solutions, claiming that the gap in multimodal agent tasks has been narrowed to a minimum.

Tooling and API Access

Alongside the model, the DeepSeek Harness 0.1.1 tool was released, which, according to the developer, supports the new model "out of the box" — that is, without any additional manual configuration. The updated model is already available via the DeepSeek API, allowing developers and teams to integrate multimodal capabilities into their own products and agents right now, rather than waiting for a general release.

Pricing and Image Handling

The pricing scheme deserves separate attention. For billing purposes, images are converted into tokens — up to 384 tokens per image — and are charged at the rates of the base DeepSeek V4-Flash model. This means that visual requests have not been moved into a separate, more expensive price tier, but are instead folded into the familiar tokenized economics of the existing model, simplifying cost forecasting for API users.

Contradictory Data

When verifying the wording across multiple sources, a discrepancy was found in the strength of the claim regarding the comparison with Opus 4.8. The company's own text uses the cautious phrasing "a level close to Opus-4.8," whereas the headlines of several outlets frame it as "matches Claude Opus 4.8" or "can compete with Anthropic's Opus 4.8." In fact, these are different marketing framings of the same benchmark: "close to" and "matches" are not equivalent claims, and independent third-party verification of these figures in public sources at the time of publication is limited. Moreover, it is important not to conflate this experimental model with the previously announced full multimodal DeepSeek-V4 with a 1 million-token context window, whose release was planned for April: that is a separate element of the V4 line's roadmap, not a product identical to V4-Flash-Vision-Exp.

Context: Where the DeepSeek Line Is Heading

The release of V4-Flash-Vision-Exp fits into DeepSeek's broader strategy of expanding the multimodal and agentic capabilities of the V4 family. If the company had previously outlined plans for a full multimodal V4 model with an extended context window, then the current experimental release looks like an intermediate step: testing the multimodal layer on the basis of the fast Flash version, available via API with transparent tokenized pricing. For the market, this is a signal that competition in the multimodal agent segment, where Western models had previously dominated, continues to intensify.