Build1 distinct publisher2 min readPublished
An independent reconstruction of DeepSeek-VL, VL2 and Janus explains what the new flash vision model was probably built to do. It does not, and says it cannot, say what it was built from.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The number worth carrying out of the paper trail is a ratio. DeepSeek-VL's final pretraining mixture stayed at roughly 70 percent text against 30 percent multimodal data [9], which is about 2.3 parts text for every part of image-bearing data [10]. The schedule around it points the same way. The adaptor trains first while the vision and language components stay frozen; then the language model and adaptor train together on text-only and multimodal data with the main vision encoder still frozen; then supervised fine-tuning runs on multimodal instructions and text conversations [8]. Multimodal data is phased in rather than switched on [11].
What the encoder is for is legible in its parts. A lower-resolution semantic branch based on SigLIP-L runs alongside a higher-resolution branch derived from a SAM-B-style encoder, on the stated grounds that global semantic understanding does not resolve small text, dense documents, OCR or visual grounding [7]. Above it sit a vision-language adaptor and a DeepSeek language model [6]. The instruction data came from a taxonomy of real user scenarios, covering transcription, conversion, multi-image comparison and safety-related prompts among other tasks [5].
VL2 is where the engineering intent stops being ambiguous. Its language component is a Mixture-of-Experts paired with DeepSeek's efficient attention design, so that added multimodal capability does not make every parameter active for every token [16], and it shipped as Tiny, Small and larger variants at the point where the stack is built around high-resolution inputs, structured documents and cheaper inference rather than general image conversation [18].
Janus complicates any reading of the word "vision" in a product name. Its argument is that understanding and generation need different visual granularities, and that a representation that is excellent at recognising a chart is not automatically the right one for generating pixels [20]. An endpoint that accepts images inherits one side of that argument and tells you nothing about the other.
The inspectable portion of all this is narrower than the lineage suggests. DeepSeek released roughly 1.3B and 7B VL variants with project code and weights [13], which is more than a closed endpoint gives you and is also two research generations behind the model now being discussed. The published record explains, in some detail, what the flash vision model was probably designed to be good at. It does not document what it was made from.
Ranked by verification strength, evidence, and original report placement.
DeepSeek released a model identified as deepseek-v4-flash-vision-exp; the immediate framing was that a text-focused model had gained native image input.
The analysis is published on dev.to by zipflow.xyz and states that it is an independent technical analysis of DeepSeek's public research and documentation, that it is not an official DeepSeek statement, and that it does not claim the current Vision-Exp API is available through its upstream channel.
The analysis explicitly separates three things: what DeepSeek's papers disclose, what the current API documentation says, and what still cannot be verified about the newest model's training data.
DeepSeek-VL's 2024 paper, Towards Real-World Vision-Language Understanding, did not frame vision only as captioning; it targeted practical inputs such as web screenshots, PDFs, OCR, charts and knowledge-oriented visual content.
The DeepSeek-VL project described a taxonomy derived from real user scenarios, used to build instruction-tuning data for recognition, transcription, conversion, analysis, commonsense reasoning, logical reasoning, multi-image comparison and safety-related prompts.
The DeepSeek-VL model family combined three major pieces: a hybrid vision encoder, a vision-language adaptor, and a DeepSeek language model.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Well-structured secondary reading, no primary corroboration
Every claim in the cluster traces to one independent, self-declared non-official analysis that paraphrases public DeepSeek papers and documentation without supplied citations or screenshots. The architectural and training-recipe descriptions are internally consistent, specific and falsifiable (SigLIP-L plus SAM-B-style branches, staged freezing, 70/30 mixture, dynamic tiling, MoE), which raises credibility; but nothing is second-sourced, and the one thing the story is titled around — the newest model's training data — is explicitly stated to be unverifiable.
Documented endpoint, zero observed usage
Adoption signals are limited to release-and-documentation facts: an experimental vision endpoint listed in DeepSeek's own docs, prior open-weight VL releases at ~1.3B/7B, and the VL2 Tiny/Small/larger family. There are no deployments, integrations, usage figures or benchmark runs in the cluster, and the only usage disclosure is negative — the author's upstream channel does not expose the model and its examples are not production tests.
Slightly conservative relative to its own evidence
The piece actively deflates the obvious 'DeepSeek gains eyes' framing, foregrounds what cannot be verified, disclaims official status and disclaims access to the model it documents. Its assertions stay at or just below what its evidence base supports; the small negative reflects a well-supported technical reconstruction presented with heavy hedging, offset by the fact that it is nonetheless unverified paraphrase.
Commercial API intermediary, disclosed
The author is zipflow.xyz, which describes an 'upstream channel' for model access, so it has a commercial interest in authority over DeepSeek model coverage and in reader attention to its gateway even while stating that this specific model is not exposed there. The dev.to venue is self-publishing with no editorial gatekeeping. Mitigating this, the conflict and the non-availability are both disclosed up front, and the piece does not attach a purchase path to its claims.
Single publisher, single source, unverified paraphrase
One source item from one publisher supports the whole cluster, the technical content is second-hand summary of papers and documentation that are not supplied, and the body is truncated before its conclusion. Confidence is moderate rather than low because the claims are specific, mutually consistent and framed with explicit epistemic boundaries, but nothing here has been cross-checked.
build
Fail closed, not fluent: separating the jobs an LLM should never have had1 distinct publisher
build
Three manual interventions in a month, and every guard was working as designed1 distinct publisher
build
Six MariaDB versions, one real difference: the only reason to leave 10.6 is the July 2026 clock1 distinct publisher
build
Force the tool call, then hand Lightsail a long-lived key1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
dev.to
1 article · August 24, 2026