R4DGS: Referring Segmentation in 4D Gaussian Splatting

Bangpu Chen1, Yaxuan Li2, Shirui Peng3, Xiangtian Si1, Liuxin Chu1, Xitong Cao1, Hongbo Jin4, Jiayu Ding4†
1China University of Geosciences, Wuhan, China  |  2Xinjiang University, Urumqi, China  |  3University of Nottingham Ningbo China, Ningbo, China  |  4Peking University, Shenzhen, China
Corresponding author
chenbangpu@cug.edu.cn
jyding25@stu.pku.edu.cn
ReferGaussian qualitative results

Qualitative results. ReferGaussian results for temporal-state, multi-target, and state-conditioned referring in dynamic 4D scenes.

Benchmark

R4D-Bench-QA

R4D-Bench-QA contains 12 dynamic scenes and 89 densely annotated sentence-level queries with pixel-level spatiotemporal supervision.

R4D-Bench-QA statistics

Benchmark overview. R4D-Bench-QA statistics: expression word cloud, query composition, and token-length distribution.

Method

ReferGaussian Framework

The MLLM-based Refer-Planner performs three semantic stages over a precomputed 4D Gaussian scene.

ReferGaussian framework
Stage 1

Query-Guided Entity Discovery

Grounded-SAM2 produces phrase-conditioned mask tracks, which are iteratively lifted into Gaussian-supported candidate entities.

Stage 2

EntityBank Construction

Candidate entities are summarized with representative observations, temporal support, appearance, state, motion, and relations.

Stage 3

Constraint-Aware Selection

The Refer-Planner resolves identity, temporal, relational, cardinality, exclusion, and zero-target constraints over EntityBank.

Results

Quantitative Comparison

Referring segmentation results reported in the paper on R4D-Bench-QA and the public HyperNeRF split.

Table 1. Referring segmentation on R4D-Bench-QA.

Method Acc ↑ vIoU ↑
Segment then Splat55.628.4
4D LangSplat58.432.1
ReferGaussian (Ours)76.534.4

Table 2. Per-scene comparison on the four-scene 4D LangSplat HyperNeRF split.

Method americano
Acc / vIoU
chickchicken
Acc / vIoU
split-cookie
Acc / vIoU
espresso
Acc / vIoU
Average
Acc / vIoU
LangSplat45.19 / 23.1653.26 / 18.2073.58 / 33.0844.03 / 16.1554.01 / 22.65
4D LangSplat89.42 / 66.0796.73 / 90.6295.28 / 83.1481.89 / 49.2090.83 / 72.26
ReferGaussian (Ours)99.43 / 68.1382.89 / 60.8790.91 / 55.7998.80 / 67.0193.01 / 62.95

Analysis

Benchmark, Model, and Runtime Evidence

Additional analyses summarize benchmark construction, MLLM sensitivity, and end-to-end efficiency.

R4D-Bench-QA dataset construction and annotation overview

Benchmark construction. R4D-Bench-QA annotates target entity sets, valid temporal windows, and spatiotemporal masks for dynamic referring queries.

MLLM sensitivity radar across six evaluation axes

MLLM sensitivity. Six-axis radar plots compare referring performance under multiple vision-language backbones.

Runtime and memory breakdown for the referring pipeline

Efficiency breakdown. Runtime, memory, and MLLM-call statistics show first-start and cached-query costs.

Qualitative

Visual Comparison

Qualitative grounding on temporal-state and exclusion queries, followed by additional examples.

Qualitative results on R4D-Bench-QA

Qualitative comparison. Query-conditioned grounding on challenging R4D-Bench-QA cases.

Citation

BibTeX

Please cite our ACM Multimedia 2026 paper.

@inproceedings{chen2026r4dgs,
  title     = {R4DGS: Referring Segmentation in 4D Gaussian Splatting},
  author    = {Bangpu Chen and Yaxuan Li and Shirui Peng and Xiangtian Si and Liuxin Chu and Xitong Cao and Hongbo Jin and Jiayu Ding},
  booktitle = {Proceedings of the 34th ACM International Conference on Multimedia},
  year      = {2026},
  doi       = {10.1145/3767308.3836021}
}