APEX-SAM is a training-free framework for cross-domain few-shot medical image segmentation. It combines quality-aware expert retrieval, anatomy-aware prompt mining, and hybrid multi-modal prompt fusion to segment unseen anatomy without parameter updates.
Abstract
Training-free cross-domain few-shot medical image segmentation aims to segment unseen anatomies without parameter updates, addressing the high cost of dense annotation and domain-specific fine-tuning in clinical practice. Existing support-driven prompting methods face three limitations: support exemplars are randomly selected without quality assurance, geometric alignment is poorly modeled, and multi-modal prompt capabilities remain underexploited.
We present APEX-SAM, a retrieval-augmented framework with three innovations. QAR builds a dual-stream DINO/SigLIP expert bank with diversity-aware selection to ensure support-query compatibility. APM performs style-aligned geometric matching and anatomy-guided point sampling from morphological priors. HMF fuses SAM branches through training-free feature-consensus weighting. Experiments on three cross-domain benchmarks confirm strong performance among training-free methods, with ablations validating each component's contribution.
Method
APEX-SAM improves support selection, prompt construction, and prompt-branch fusion while keeping the full pipeline training-free.
Builds a hierarchical expert bank using structural and semantic descriptors, then selects compatible support exemplars with quality, coverage, and diversity constraints.
Aligns support and query anatomy under cross-modality appearance shifts, then samples positive and negative prompts from morphological priors.
Runs multiple SAM prompt branches and combines them with reliability-weighted feature consensus, avoiding learned fusion layers or fine-tuning.
Benchmarks
The paper evaluates APEX-SAM on Abd-MRI, Abd-CT, and Card-MRI, covering abdominal organs and cardiac structures with substantial modality and anatomy shifts.
Figure 2. Dataset visualization and DINO-feature t-SNE for Abd-MRI, Abd-CT, and Card-MRI.
Results
APEX-SAM achieves consistent gains over training-free and few-shot baselines across abdominal and cardiac segmentation tasks.
Table 1. Dice (%) on Abd-MRI and Abd-CT. Bold indicates the best result and underline indicates the second best result.
| Method | Ref. | Abd-MRI | Abd-CT | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Liver | LK | RK | Spleen | Mean | Liver | LK | RK | Spleen | Mean | ||
| PANet | ICCV'19 | 39.24 | 26.47 | 37.35 | 26.79 | 32.46 | 40.29 | 30.61 | 26.66 | 30.21 | 31.94 |
| SSL-ALP | TMI'22 | 70.74 | 55.49 | 67.43 | 58.39 | 63.01 | 71.38 | 34.48 | 32.32 | 51.67 | 47.46 |
| RPT | MICCAI'23 | 49.22 | 42.45 | 47.14 | 48.84 | 46.91 | 65.87 | 40.07 | 35.97 | 51.22 | 48.28 |
| PATNet | ECCV'22 | 57.01 | 50.23 | 53.01 | 51.63 | 52.97 | 75.94 | 46.62 | 42.68 | 63.94 | 57.29 |
| IFA | CVPR'24 | 50.22 | 35.99 | 34.00 | 42.21 | 40.61 | 46.62 | 25.13 | 26.56 | 24.85 | 30.79 |
| FAMNet | AAAI'25 | 73.01 | 57.28 | 74.68 | 58.21 | 65.79 | 73.57 | 57.79 | 61.89 | 65.78 | 64.75 |
| MAUP | MICCAI'25 | 78.16 | 58.23 | 72.34 | 59.65 | 67.09 | 78.25 | 59.41 | 71.80 | 60.38 | 67.46 |
| APEX-SAM (Ours) | MICCAI'26 | 95.10 | 96.88 | 96.01 | 95.23 | 95.81 | 93.47 | 90.28 | 91.83 | 92.06 | 91.91 |
Table 2. Dice (%) on Card-MRI.
| Method | Ref. | LV-BP | LV-MYO | RV | Mean |
|---|---|---|---|---|---|
| PANet | ICCV'19 | 51.42 | 25.75 | 25.75 | 36.66 |
| SSL-ALP | TMI'22 | 83.47 | 22.73 | 66.21 | 57.47 |
| RPT | MICCAI'23 | 60.84 | 42.28 | 57.30 | 53.47 |
| PATNet | ECCV'22 | 65.35 | 50.63 | 68.34 | 61.44 |
| IFA | CVPR'24 | 50.43 | 31.32 | 30.74 | 37.50 |
| FAMNet | AAAI'25 | 86.64 | 51.82 | 76.26 | 71.58 |
| MAUP | MICCAI'25 | 88.36 | 52.74 | 78.29 | 73.13 |
| APEX-SAM (Ours) | MICCAI'26 | 92.75 | 68.41 | 88.23 | 83.13 |
Qualitative
APEX-SAM produces stable cross-domain masks while remaining transparent about remaining challenges such as low-contrast boundaries, severe shape shifts, retrieval mismatch, and suboptimal prompts.
Figure 3. Successful cross-domain masks and representative failure modes.
Ablation
Each component contributes a measurable improvement, and the thresholded memory update further strengthens the full pipeline.
Table 3. Ablation on core designs (Dice %).
| Configuration | QAR | APM | HMF | Memory | Mean Dice |
|---|---|---|---|---|---|
| Prompt-only baseline | No | No | No | - | 72.4 |
| + QAR | Yes | No | No | Fixed | 80.2 |
| + QAR + APM | Yes | Yes | No | Fixed | 86.3 |
| + QAR + APM + HMF | Yes | Yes | Yes | Fixed | 91.8 |
| Full (Ours) | Yes | Yes | Yes | Thresholded | 95.81 |
Figure 4. Hyperparameter sensitivity on Abd-MRI.
Mean Dice and standard deviation, with selected values highlighted.
Citation
Proceedings metadata such as LNCS volume, page range, DOI, and final paper URL will be added after publication.
Mao, Z., Chen, B., Lei, Q., Tan, J., Sun, K.: APEX-SAM: Anatomy-Aware Prompting with Expert Retrieval for Training-Free Medical Image Segmentation. In: Medical Image Computing and Computer Assisted Intervention - MICCAI 2026. Lecture Notes in Computer Science. Springer, Cham (2026).
@inproceedings{mao2026apexsam,
title = {APEX-SAM: Anatomy-Aware Prompting with Expert Retrieval for Training-Free Medical Image Segmentation},
author = {Mao, Zhihao and Chen, Bangpu and Lei, Qi and Tan, Jiaqi and Sun, Kun},
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
series = {Lecture Notes in Computer Science},
publisher = {Springer},
address = {Cham},
year = {2026}
}