Qualitative results. ReferGaussian results for temporal-state, multi-target, and state-conditioned referring in dynamic 4D scenes.
Benchmark
R4D-Bench-QA contains 12 dynamic scenes and 89 densely annotated sentence-level queries with pixel-level spatiotemporal supervision.
Benchmark overview. R4D-Bench-QA statistics: expression word cloud, query composition, and token-length distribution.
Method
The MLLM-based Refer-Planner performs three semantic stages over a precomputed 4D Gaussian scene.
Grounded-SAM2 produces phrase-conditioned mask tracks, which are iteratively lifted into Gaussian-supported candidate entities.
Candidate entities are summarized with representative observations, temporal support, appearance, state, motion, and relations.
The Refer-Planner resolves identity, temporal, relational, cardinality, exclusion, and zero-target constraints over EntityBank.
Results
Referring segmentation results reported in the paper on R4D-Bench-QA and the public HyperNeRF split.
Table 1. Referring segmentation on R4D-Bench-QA.
| Method | Acc ↑ | vIoU ↑ |
|---|---|---|
| Segment then Splat | 55.6 | 28.4 |
| 4D LangSplat | 58.4 | 32.1 |
| ReferGaussian (Ours) | 76.5 | 34.4 |
Table 2. Per-scene comparison on the four-scene 4D LangSplat HyperNeRF split.
| Method | americano Acc / vIoU |
chickchicken Acc / vIoU |
split-cookie Acc / vIoU |
espresso Acc / vIoU |
Average Acc / vIoU |
|---|---|---|---|---|---|
| LangSplat | 45.19 / 23.16 | 53.26 / 18.20 | 73.58 / 33.08 | 44.03 / 16.15 | 54.01 / 22.65 |
| 4D LangSplat | 89.42 / 66.07 | 96.73 / 90.62 | 95.28 / 83.14 | 81.89 / 49.20 | 90.83 / 72.26 |
| ReferGaussian (Ours) | 99.43 / 68.13 | 82.89 / 60.87 | 90.91 / 55.79 | 98.80 / 67.01 | 93.01 / 62.95 |
Analysis
Additional analyses summarize benchmark construction, MLLM sensitivity, and end-to-end efficiency.
Benchmark construction. R4D-Bench-QA annotates target entity sets, valid temporal windows, and spatiotemporal masks for dynamic referring queries.
MLLM sensitivity. Six-axis radar plots compare referring performance under multiple vision-language backbones.
Efficiency breakdown. Runtime, memory, and MLLM-call statistics show first-start and cached-query costs.
Qualitative
Qualitative grounding on temporal-state and exclusion queries, followed by additional examples.
Qualitative comparison. Query-conditioned grounding on challenging R4D-Bench-QA cases.
Citation
Please cite our ACM Multimedia 2026 paper.
@inproceedings{chen2026r4dgs,
title = {R4DGS: Referring Segmentation in 4D Gaussian Splatting},
author = {Bangpu Chen and Yaxuan Li and Shirui Peng and Xiangtian Si and Liuxin Chu and Xitong Cao and Hongbo Jin and Jiayu Ding},
booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
year = {2026},
doi = {10.1145/3767308.3836021}
}