VLMs may recognize anatomy but confuse direction terms (superior ↔ inferior), propagating errors through reasoning.
Video-native VLMs process volumes as ordered slices.
MIS-Ground varies five grounding factors; MIS-SemSam changes only token selection.
1,1603D volumes · 2,320 2D slices
33,864spatial-grounding questions
77anatomical components (67 CT, 10 MRI)
Modality
CTMRI
Slice planes
AxialCoronalSagittal
Coordinates
SRARAS storageViewing
Visual prompts
MaskBoxPoint
MIS-Ground: controlled spatial-grounding tasks
TotalSegmentator torso CT and Osteoarthritis Initiative (OAI) knee MRI. Examples: cross-slice relationships (RQ1), anatomical/colloquial terms (RQ2), and prompts without scans (AB2). CT is 84% of the benchmark, so scores are CT-weighted.
MIS-SemSam: semantic sampling
Rescore top-M/top-p content candidates using probability mass in cosine-kNN neighborhoods.
Precompute content-token neighborhoods, excluding special, control, and modality tokens.
Use default decoding if any candidate is a non-content token.
Select by argmax, or softmax sampling for chain-of-thought.
Training-free; requires logits + embeddings; no extra forward pass; O(|It|·K) lookups/token.
53.4%Qwen3-VL-32B, default decoding
66.5%+ MIS-SemSam (+13.06 points)
Qwen3-VL-32B + MIS-SemSam: 66.5% accuracy
Accuracy by model family and size; MIS-SemSam uses Qwen3-VL. Paired runs change only token selection. M3D (15.9%) and Med3DVLM (0.0%) are omitted; resizing and malformed responses limit interpretation.