Medical Image Spatial Grounding with Semantic Sampling

Andrew Seohwan Yu† Mohsen Hariri† Kunio Nakamura Mingrui Yang Xiaojuan Li Vipin Chaudhary

†Equal contribution Corresponding author: vipin@case.edu

Directional errors in 3D grounding

  • VLMs may recognize anatomy but confuse direction terms (superior ↔ inferior), propagating errors through reasoning.
  • Video-native VLMs process volumes as ordered slices.
  • MIS-Ground varies five grounding factors; MIS-SemSam changes only token selection.

1,1603D volumes · 2,320 2D slices

33,864spatial-grounding questions

77anatomical components (67 CT, 10 MRI)

Modality

CT
MRI

Slice planes

Axial
Coronal
Sagittal

Coordinates

SRA
RAS storage
Viewing

Visual prompts

Mask
Box
Point

MIS-Ground: controlled spatial-grounding tasks

Sample question-answer pairs from MIS-Ground showing CT and MRI slices with labeled anatomical prompts and spatial-relationship questions

TotalSegmentator torso CT and Osteoarthritis Initiative (OAI) knee MRI. Examples: cross-slice relationships (RQ1), anatomical/colloquial terms (RQ2), and prompts without scans (AB2). CT is 84% of the benchmark, so scores are CT-weighted.

MIS-SemSam: semantic sampling

Rescore top-M/top-p content candidates using probability mass in cosine-kNN neighborhoods.

Scoret(c)=∑k∈K(c)wc,k pt(Stid[c,k]),\text{Score}_t(c)=\sum_{k\in\mathcal{K}(c)} w_{c,k}\, p_t\big(S_{\text{tid}}[c,k]\big), wc,k=max⁡(0,Sval[c,k])w_{c,k}=\max(0, S_{\text{val}}[c,k])
  • Precompute content-token neighborhoods, excluding special, control, and modality tokens.
  • Use default decoding if any candidate is a non-content token.
  • Select by arg⁡max⁡\arg\max, or softmax sampling for chain-of-thought.
  • Training-free; requires logits + embeddings; no extra forward pass; O(|It|·K) lookups/token.

53.4%Qwen3-VL-32B, default decoding

66.5%+ MIS-SemSam  (+13.06 points)

Qwen3-VL-32B + MIS-SemSam: 66.5% accuracy

Line plot of overall MIS-Ground accuracy versus model size across model families, with the MIS-SemSam curve highest and crossing the Gemini 3 Flash Preview reference line

Accuracy by model family and size; MIS-SemSam uses Qwen3-VL. Paired runs change only token selection. M3D (15.9%) and Med3DVLM (0.0%) are omitted; resizing and malformed responses limit interpretation.

Grounding results: Qwen3-VL-32B + MIS-SemSam

RQ1: 3D vs. 2D accuracy

Slice direction: 77.6%. Relationships: cross-slice 68.0%, in-plane 71.7%. Overall: 3D 56.9%, 2D 58.8%.

RQ2: Direction terms depend on view

Anatomical vs. colloquial accuracy: standard viewing 69.4% vs. 57.8%; RAS storage 50.2% vs. 59.9%.

RQ3: Prompt effects vary

Standard view: +2–3 percentage points from 57.6%. RAS storage, point prompts: −13.3 points from 75.3% (anatomical); +8.2 from 45.1% (colloquial).

AB1: Text-only anatomical knowledge

Anatomical terms: 69.3% without images, 74.0% with images. Colloquial terms: 40.4% without images.

AB2: Spatial reasoning without scans

Prompts on a blank background: 64.2% accuracy for points, 60.4% for boxes.

Setup: Evaluation protocol

Reasoning; max 8,912 new tokens; T=0.5T{=}0.5. Bayesian credible intervals for accuracy over limited trials.

View, terminology, and decoding affect accuracy

Anatomical terms perform better in standard viewing; colloquial terms perform better in RAS storage.

Visual prompts can raise or lower accuracy, depending on view and terminology.

Qwen3-VL-32B gains 13.06 percentage points by changing token selection, with no retraining or extra forward pass.

Supported in part by NSF awards 2117439 and 2320952. Case Western Reserve University and Cleveland Clinic, Cleveland, OH, USA.