Appearance
A street view is not a route. That's the blunt takeaway from UrbanGround, a new sandbox that drops multimodal LLM agents into a physically constrained replica of Hong Kong. The model can look at a building, read a sign, describe the scene. None of that helps if it can't remember which way it came from. The paper's central finding is grim: local perception does not compose into sustained goal-directed behavior.
The same tension runs through two other papers that landed on arXiv this week. Aphanta asks whether letting an image editor produce visual intermediates actually improves reasoning, and finds the answer depends heavily on the task. OmniUE takes the opposite direction, building a unified embedder that handles text, video, and audio with interactive queries. Together they sketch the current edge of multimodal reasoning: you can make models see more, but you still have to make them think in a way that survives movement, transformation, and open-ended goals.
The problem: perception without agency
Multimodal LLMs today nail atomic abilities. Show a model a photo of an intersection and it can identify the crosswalk, the traffic light, the brand of the cafe. Ask it to navigate from that intersection to a less visible landmark and it starts to fail. Orientation and pedestrian-aware movement are unreliable, as UrbanGround's results show.
The authors built UrbanGround from territory-wide 3D geospatial data of Hong Kong. Agents move in first person, consult an interactive map, and get closed-loop feedback from their own actions. The paper tracks the spatial problem through three research questions: can the agent ground a local scene after active observation? Does that grounding support navigation as destinations get farther? Does behavior survive route changes and pedestrian motion?
The answer to each question is progressively worse. Local grounding works decently in isolation. The moment the agent has to chain observations across time, errors accumulate without correction. This is an architecture problem, not a data problem. The transformer's attention mechanism has no built-in notion of "I've gone down this street before." You can bake in a map, but the model won't use it consistently over long horizons.
Aphanta: when does an image editor actually help?
Aphanta tackles a different failure mode. Some reasoning tasks are easier if the model can draw on external visual evidence. For instance, counterfactual reasoning might require visualizing "what would this room look like with the chair moved to the window?" The pipeline is MLLM -> image editor -> MLLM. The image editor receives an instruction, produces an intermediate image, and the second model reasons over it.
Aphanta is a diagnostic framework that runs three conditions: direct reasoning, reasoning with an editor-generated intermediate, and reasoning with an idealized reference intermediate. That third condition is the counterfactual ceiling. It tells you the most you could get from a perfect editor. By comparing the three, you separate "the task would benefit from visual intermediate" from "current editors can't realize the change" from "both."
Across 20 candidate tasks and multiple editor-model combinations, the gains are sharply task-conditioned. Visual cue injection, grounding, and counterfactual state realization all improve. Symbol-sensitive construction and structural extrapolation do not. The editor's limits become the model's limits.
Key numbers
- Qwen pipeline improves mean task score from 0.343 to 0.445, a +29.7% relative gain on selected positive tasks.
- 20 candidate tasks evaluated across editor and MLLM combinations.
- Gains concentrate in 3 task types; unreliable in 2 others.
That 0.343 to 0.445 jump is meaningful in absolute terms, but only after filtering. The full study keeps the failures and unfiltered tasks to expose the boundary. The authors position image editing as a specialized visual workspace, not a universal reasoning mechanism. I'd agree: expecting one image editor to handle every spatial transformation is like expecting one text-to-image prompt to encode a full state update.
OmniUE: a different bet on embedding
OmniUE takes the opposite strategy. Instead of bolting an editor onto the model, it learns a unified embedding space across text, video, and audio. The key shift is from two-tower architectures to LLM-based embedders that handle user-conditioned interactions. You can query with a region of interest in a video, an audio span, or natural language, and retrieve any combination of modalities.
The authors introduce OmniCHOIR as a benchmark for omni-interactive compositional audio retrieval. Given a video and an audio clip as context, with a text query like "find the scene where she says this," the model must retrieve the correct segment. OmniUE outperforms baselines across the board: 10.5% on MMEB-v2-video, 1.1% on MAEB audio tasks, 83.7% on SCaR visual-interactive benchmarks, and 24.1% on OmniCHOIR.
That 1.1% audio improvement is modest, but the video and interactive gains are large. What strikes me is the interaction design. Instead of forcing every query into text, you can point at a visual region or highlight an audio span. This matches how humans actually reference media. "That guy, in the blue shirt, with that voice" is a compositional query that maps naturally to OmniUE's segmenter-then-aggregate architecture.
Quick Take: Multimodal reasoning isn't a single capability. It breaks into local grounding, intermediate-state manipulation, and cross-modal retrieval, and no current model handles all three reliably.
What the three papers have in common
Read them side by side and a pattern emerges. UrbanGround shows that local perception doesn't automatically give you spatial agency. Aphanta shows that external image editing doesn't automatically give you better reasoning. OmniUE shows that a unified embedding space can give you better retrieval, but retrieval is not the same as understanding. Each paper is a story about representation and control: what the model internally represents, and how much it can steer that representation toward a goal.
My own experience testing UrbanGround-like scenarios in simpler grids matches the paper's findings. A model can correctly identify landmarks, but if it can't maintain a heading across multiple steps, it circles. When I added explicit map coordinates to the prompt, performance jumped. The problem isn't perception, it's integration.
That's also why Aphanta's diagnostic separation matters. You need to know whether your failure is in the model's reasoning or in the editor's execution. If the model can solve the task with a perfect intermediate but not with the real editor, the bottleneck is the editor. If it fails on both, the bottleneck is reasoning.
Common Pitfalls
Three traps recur when people apply these results.
Pitfall 1: Confusing local accuracy with system reliability. A model that recognizes a stop sign 99% of the time can still get lost in a 10-block navigation task if the 1% errors compound. UrbanGround's error accumulation is the norm, not the exception. Always test closed-loop behavior, not just single-frame perception.
Pitfall 2: Assuming image editing is a universal reasoning tool. Aphanta shows editor-based intermediates only help for specific task families. If your task involves symbol manipulation or structural extrapolation, expect the editor to become a liability. Other frameworks like Visual Sketchpad or Image-of-Thought suggest treating editors as task-specific tools, not a general-purpose scratchpad.
Pitfall 3: Ignoring query modality when building embeddings. OmniUE's success on visual-interactive querying (83.7% relative gain on SCaR) is a reminder that forcing all queries into text loses information. If your retrieval system only accepts text, you're throwing away the spatial and temporal structure in the media itself.
Pitfall 4: Treating a unified embedding space as a solved problem. OmniUE still sees a small gain on audio tasks. If you're benchmarking a universal embedder, measure modalities separately. Aggregate scores hide which input types you're actually failing on.
Pitfall 5: Forgetting that interactive querying requires segmenters. OmniUE's performance depends on its visual and audio segmenters correctly localizing the region or span. If those segmenters fail, the whole retrieval pipeline degrades. Always evaluate the segmenter quality before blaming the embedder.
One thing to remember
Each of these papers gives you a different instrument to measure your own multimodal system. UrbanGround gives you a closed-loop spatial environment. Aphanta gives you a three-way diagnostic to split reasoning from editing. OmniUE gives you a benchmark for interactive cross-modal retrieval. You don't need to adopt all three, but you should pick one to stress-test your model before shipping.
The bottom line
- If you're building an embodied agent that must navigate physical space, adopt UrbanGround-style evaluation from the start. Perception metrics are necessary, but they will almost always overstate real-world competence.
- If you're adding image editing to a reasoning pipeline, use Aphanta's three-condition protocol first. Expect gains only on cue injection, grounding, and counterfactual states; for symbol-sensitive tasks, keep the editor out.
- If you're building a retrieval system with video or audio, OmniUE's interactive-query model is a good fit for visual-region and compositional queries. Expect the biggest wins when a user wants to reference a specific object or sound rather than describe it in words.