Appearance
The problem: MLLMs follow instructions worse than they look
Run any modern multimodal model on a fine-grained instruction-following task and you'll see the same pattern. It can describe an image fluently. It can answer a hard visual question. But ask it to produce an output that satisfies four specific constraints at once, and it quietly drops one. The aggregate score looks fine. The constraint-level breakdown doesn't.
This is a pipeline problem. The standard recipe for training instruction-following MLLMs is one-pass: generate candidate samples, filter with a judge, train, evaluate, ship. Every signal produced along the way, the failed samples, the verifier outcomes, the target model's errors, gets discarded. The model never sees its own failure modes, and the evaluator never checks whether the judge was right in the first place.
Four papers released this month attack that loop from different sides. VISA makes data synthesis self-evolving. VBVR-Pro replaces unreliable LLM judges with deterministic scorers and uses them for reinforcement learning. SciMIF shows where instruction following breaks in scientific domains. And a white paper asks the question underneath all of it: can vision itself be the substrate for reasoning, not just the input to it?
One loop, four fixes
The papers fit together better than their separate abstracts suggest. Each plugs into a different stage of the same feedback cycle.
VISA owns the synthesis stage. VBVR-Pro owns evaluation and the reward path into training. SciMIF makes evaluation meaningful for science. The white paper is the bet that this whole loop, run on visual data instead of text, is a route to something bigger.
VISA: data synthesis that learns from its own failures
The VISA authors start from a simple observation: existing synthesis pipelines generate once and filter. If a sample fails verification, it's gone. If the target model can't handle a constraint type, nobody records that. The pipeline has no memory.
I've run this exact pipeline. The first round of generated data looks great, the second round looks identical, and the model plateaus because every round feeds it the same easy constraints. VISA is the fix. It's an agentic loop where each round analyzes an image to filter out constraints that don't apply and discover new ones that do, samples constraint sets from persistent memory with diversity and difficulty awareness, generates candidates, and verifies them with executable tools and structured judges. Failed samples go through diagnostic-guided recovery. Accepted samples get probed against the target model to estimate difficulty. All of it, verifier signals and failure profiles, writes back to memory.
The practical consequence: later rounds focus on constraints the model actually struggles with instead of re-generating the same easy templates. On MM-IFEval, VISA consistently beats strong baselines on instruction following while preserving general capability across seven public benchmarks. That last part matters. A lot of data-synthesis tricks improve the target benchmark and degrade everything else. VISA's verifier contracts also double as reward signals for RL, so you don't need a separately trained reward model.
Quick Take: The one-pass generate-and-filter pipeline throws away the cheapest training signal you have, feedback from failed samples.
VBVR-Pro: stop asking LLMs to judge things
The most contrarian finding in this cluster is about evaluation. VBVR-Pro's authors ran a systematic study of leading MLLMs as judges and found recurring failure modes in the VLM-as-a-judge approach. When I've used LLM judges for fine-grained visual tasks, I've seen the same things: position bias, self-preference, and a tendency to reward fluent outputs over constraint-satisfying ones. A judge that can't reliably check a constraint is a vibe check.
VBVR-Pro's answer is deterministic, task-specific rule scorers. No learned judge, no prompt engineering, just ground-truth rules per task. These scorers align closely with human judgments and, crucially, work as reward signals for multi-task RL. Post-RL performance across visual reasoning tasks is stronger with these scorers than with the judge-based alternative.
Key Numbers
- 300 procedurally generated tasks in VBVR-Pro, enough to train on without overfitting to a handful of templates.
- 30+ image, video, and interleaved generators for controlled modality studies.
- 7 external benchmarks where VBVR-Pro-trained models transfer, including RISE-Video, MME-CoF-Pro, and BabyVision.
- 22 scientific tasks across 5 disciplines in SciMIF, organized into 10 constraint groups.
The transfer result is the one to sit with. Models trained only on VBVR-Pro's procedural tasks generalize to seven external visual reasoning benchmarks. That's evidence the task space isn't just self-contained. It's capturing something real about visual reasoning.
What each contribution buys you
| Contribution | Stage of the loop | Scale | Key finding | Availability |
|---|---|---|---|---|
| VISA | Data synthesis | 7 general benchmarks preserved | Self-evolving synthesis beats one-pass filtering on MM-IFEval | Not stated in paper |
| VBVR-Pro | Evaluation + RL training | 300 tasks, 30+ generators | Rule-based scorers beat VLM judges; transfer to 7 external benchmarks | Full release: data, models, scorers, code |
| SciMIF | Evaluation (science) | 22 tasks, 5 disciplines, 10 constraint groups | Chemistry is hardest; scale doesn't fix constraint adherence | Code on GitHub |
Notice what's missing from this table: a new model. None of these papers claim a better base MLLM. They claim better data, better evaluation, and better training signals for the models we already have. That's the right bet right now. The field has plenty of capable base models and almost no reliable infrastructure for making them follow instructions precisely.
SciMIF: science is where instruction following goes to die
SciMIF evaluates MLLMs on complex scientific instructions across five disciplines, guided by a taxonomy of 10 constraint groups that covers both general functional requirements and discipline-specific characteristics. The taxonomy is the useful part. It means the benchmark can tell you not just that a model failed, but which kind of constraint it failed on.
The findings are uncomfortable. Chemistry is the hardest discipline for current MLLMs. If you're a chemist hoping a frontier model will follow a multi-step protocol with exact reagent constraints, the evidence says you'll be disappointed. More striking: increasing model scale doesn't improve constraint adherence. Bigger models don't follow fine-grained instructions better. They get more fluent at ignoring them.
The failure modes cluster in two places: fine-grained constraints and instructions requiring deep application of disciplinary knowledge. That second one is the harder problem. A model can parse a constraint about temperature ranges. It cannot know that a particular reaction pathway is thermodynamically implausible unless that knowledge is in its weights, and for niche chemistry it usually isn't.
The modality question: video, interleaved, or image?
VBVR-Pro's mechanism study is the part I want more people to read. The testbed supports controlled comparisons across more than 30 generators, which lets the authors isolate a question that's usually impossible to ask cleanly: does the generation substrate matter for reasoning?
Yes, it does. Video generation is strongest for tasks requiring persistent spatiotemporal state tracking. That makes sense. If a task requires knowing where an object is across ten frames, an image generator has to hallucinate the intermediate states. Video carries them natively. Interleaved generation, meanwhile, is a compute-efficient alternative that covers a surprising fraction of the same ground.
The deeper finding is the existence of vision-native trajectories. Ablations and probing suggest that some reasoning paths only emerge when the model reasons through visual states directly, rather than translating everything to language first. That's the empirical hook for the white paper's thesis.
The white paper: visual general intelligence
The VGI white paper makes the obvious-in-retrospect argument: GPT showed that autoregressive language modeling at scale produces transfer to unseen tasks. Nobody has run the equivalent experiment for vision. What capabilities emerge from autoregressive modeling over images, videos, and geometry, with vision as the core modality and language as a peripheral one?
The paper deliberately doesn't offer a single definition of visual intelligence. It clarifies the principles, input modalities, benchmarks, and learning paradigms that computer vision should pursue in the AGI era. The relationship between vision and language gets reframed. Language becomes one output channel among several.
This is speculative, and the paper knows it. But VBVR-Pro's vision-native trajectories are the kind of evidence that makes the speculation worth funding. If some reasoning paths only exist when vision is the medium, then text-centric training is leaving capability on the table.
Common Pitfalls
Using an LLM judge for fine-grained constraint verification is the most common mistake. The failure modes VBVR-Pro documents are systematic, not occasional. Position bias and self-preference corrupt the signal in ways that aggregate metrics hide. If a constraint is checkable by rules, write the rules.
Running one-pass data synthesis and calling it done is the second. The VISA result is clear: failed samples and verifier outcomes are information. Discarding them means your next training round repeats the same mistakes. A self-evolving loop is the difference between a pipeline that improves and one that plateaus.
Assuming a bigger model follows instructions better is the third. SciMIF's scaling finding contradicts the default instinct. If you're failing on fine-grained constraints at 7B, you'll probably still fail at 70B. The fix is constraint-aware training data and verifiable rewards, not parameter count.
Treating video generation as expensive image generation is the fourth. For spatiotemporal state tracking, video is load-bearing. Swapping in interleaved generation to save compute works for some tasks and silently fails for others. Check whether your task requires persistent state before you optimize the generator budget.
Reporting only aggregate accuracy is the fifth. The interesting failures in all three empirical papers live at the constraint level. Aggregate scores hide the systematic dropping of fine-grained constraints. Break down your evaluation by constraint group before you claim your model follows instructions.
One thing to remember
Every paper in this cluster converges on the same principle: feedback loops beat one-pass pipelines. Whether it's VISA writing verifier signals back to memory, VBVR-Pro using rule-based scorers as RL rewards, or SciMIF exposing which constraints models ignore, the pattern is identical. The signal you need is already being generated by your pipeline. You're just throwing it away.
The Bottom Line
If you're building training data for instruction-following MLLMs, adopt a self-evolving synthesis loop like VISA, because one-pass filtering discards exactly the failure signals that would make your next round better.
If you're evaluating models for scientific or other high-stakes domains, use constraint-level benchmarks like SciMIF and don't expect scale to fix adherence, because fine-grained constraints in chemistry-grade tasks break even the largest current models.
If you're choosing a generation substrate for visual reasoning, use video for tasks with persistent spatiotemporal state and interleaved generation when compute is tight, and watch for vision-native trajectories, because substrate-aware training recipes are likely to become standard within the next year.