A closer look at visual reasoning

Mixture
of Layers.

Dynamic Layer Routing
for Visual Reasoning

The details are already there.
Let the question decide where to look.

HYBRID ROUTINGPATCH × DEPTH
TEXT QUERY“What does the small sign say?”
ABCD
One image. Different patch needs.

Keep each patch’s position. Choose its visual depth.

Image-wide prior1 − β×Patch evidenceβ→ top-k per patch
Selected features Gated reserve4 patches → 4 tokens
Each patch: (1 − γn) · selected mixture + γn · reserve

Illustrative top-2 routing. Global–local disagreement controls each patch’s reserve gate.

Jeonghwan Kim1*Sofia Stoica1*Jiwan Chung2Ansel Blume1
Hyeonjeong Ha1Zhenhailong Wang1Xin Luna Dong3Heng Ji1

1 University of Illinois Urbana-Champaign2 Yonsei University3 Meta Reality Labs

* Equal contribution · Research manuscript
+18.90 pp

V* accuracy over baseline

Vicuna-13B · DINOv2 w/ Txt · hybrid
90.34%

V* with two vision encoders

Vicuna-13B · CLIP + DINOv2 w/ Txt
0 extra tokens

More depth, same sequence length

Routing preserves the visual token count

01 / Motivation

A single layer
can miss the point.

Recognizing a scene isn’t the same as reading a tiny sign inside it.

Vision encoders represent an image at many depths: from localized visual details to broader semantic context. Yet many multimodal models read only the final or penultimate layer, or combine a fixed set of layers for every question.

Mixture of Layers makes that choice adaptive. A text-conditioned router selects and combines intermediate representations, giving the language model visual features that better match the question.

FIXED REPRESENTATIONS

Every question

The same visual layer

MIXTURE OF LAYERS

A specific question

A relevant mix of layers

02 / The method

Let the question
choose the depth.

One pre-trained encoder. A family of visual representations. Three ways to route between them.

EXPLORE THE ROUTERConceptual illustration · not model outputs
GLOBAL + LOCAL + RESERVE

The detail and the bigger picture.

Combine image-level and patch-level routing with a weighted product of experts. A learned gate uses their disagreement to blend in a stable reserve layer.

routed patch = (1 − γ) · selected layers + γ · reserve

The reserve is the final or penultimate representation used by the original backbone.

01

Encode

Extract intermediate vision features and an aligned text-query embedding.

02

Route

Score candidate layers against the question and select the top-k representations.

03

Fuse & reason

Mix the selected features, project them into the language space, and answer.

Full MoL architecture: the image and text encoders feed layer, patch, or hybrid routers, followed by aggregation, a vision-language connector, and the language model.
THE FULL ARCHITECTURE Three routing variants, one shared idea: query-adaptive access to intermediate visual features. Expand figure ↗

03 / Experiments

Small details.
Measurable gains.

Evaluated across seven fine-grained visual reasoning benchmarks, three language backbones, and multiple vision encoders.

Download all scores ↓
V* BENCHMARK

Accuracy, with a closer look.

Accuracy (%) · higher is better

Fine-grained benchmark accuracy (%)

Scores from the manuscript’s main backbone comparison. Deltas are percentage points, calculated from the displayed scores. Results vary by benchmark and backbone.

Across encoders, too

Route within.
Then combine.

Apply hybrid routing independently to CLIP and DINOv2 w/ Txt, upsample the DINOv2 stream, and fuse the projected features into 576 visual tokens.

Against Interleaved-MoF, this full-system comparison improves V* by 7.99 points. HRBench4K and RealWorldQA decrease; encoder and fusion differences mean this is not a routing-only ablation.

INTERLEAVED-MOF82.35
MoL HYBRID90.34

V* accuracy (%) · Vicuna-13B backbone

04 / Findings & insights

What the layers
tell us.

Routing behavior, receptive fields, and ablations help explain when adaptive visual representations are useful.

01

Different routers use different depths.

In the reported CLIP analysis, image-level routing emphasizes early, detail-preserving layers. Patch-level routing places more mass on later, contextualized features. Hybrid routing balances these preferences.

Layer
Patch
Hybrid

Early ← encoder depth → Late

Qualitative summary
Measured distributions from the paper
Measured CLIP routing probabilities across layers for V*, MMStar, and GQA categories, with receptive-field backgrounds and top-k selection markers.
MEASURED ROUTING BEHAVIOR Curves show average normalized routing probabilities. Lighter backgrounds indicate more localized receptive fields. Expand figure ↗
02

The question is part of the representation.

Text-conditioned hybrid routing scores 84.03% on V*, compared with 77.31% for image-only routing in the text-encoder ablation. Visual evidence becomes more useful when the model knows what to look for.

Image-only77.31
Text-conditioned84.03
03

A stable fallback matters.

Removing the reserve layer reduces accuracy across all five reported reserve-ablation benchmarks. The largest drop is 5.13 points on HRBench4K (43.88 → 38.75).

−3.00 to −5.13 ppWithout the reserve layer · reported ablation
04

More layers isn’t always better.

Small selections can be enough. In the k ablation, hybrid routing achieves 84.03 on V* at k = 1, versus 76.89 at k = 2 and 67.65 at k = 4. The best k depends on the task and router.

k = 184.03
k = 276.89
k = 467.65
V* accuracy (%) · hybrid k ablation

05 / A closer look

Finding the evidence.
Layer by layer.

Selected-layer self-attention maps highlight text, subtle attributes, and spatial relationships in the paper’s qualitative examples.

Click the figure to inspect it
CLIP selected-layer self-attention examples: a clock character, a blue bench relative to a green door, a pink shirt, a leasing-office sign, and relative mask positions.
QUALITATIVE EVIDENCE · CLIP Self-attention maps from selected vision layers. These visualizations illustrate attention patterns; they do not by themselves establish a causal explanation.
THE TAKEAWAY

Better visual reasoning can start with
choosing the right representation.

Adaptive depth is complementary to richer inputs and multiple encoders.

Build on this work

A new perspective
on visual depth.

BibTeX for the current research manuscript. Code will be released soon.

CITATION
@misc{kim2026mixtureoflayers,
  title = {Mixture of Layers: Dynamic Layer
           Routing for Visual Reasoning},
  author = {Kim, Jeonghwan and Stoica, Sofia and
    Chung, Jiwan and Blume, Ansel and
    Ha, Hyeonjeong and Wang, Zhenhailong and
    Dong, Xin Luna and Ji, Heng},
  year = {2026},
  note = {Research manuscript}
}
Paper figure · scroll to inspect
Expanded paper figure