4D Scaffold Gaussian Splatting

with Dynamic-Aware Anchor Growing

for Efficient and High-Fidelity Dynamic Scene Reconstruction


Woong Oh Cho, In Cho, Seoha Kim, Jeongmin Bae, Youngjung Uh, Seon Joo Kim

Yonsei University, Seoul


AAAI 2026


TL;DR: Our method achieves superior image quality, high rendering speed, and compelling storage with single-frame COLMAP points, surpassing other 4D Gaussian methods (STG, 4DGS, Ex4DGS).

Abstract

Modeling dynamic scenes through 4D Gaussians offers high visual fidelity and fast rendering speeds, but comes with significant storage overhead. Recent approaches mitigate this cost by aggressively reducing the number of Gaussians. However, this inevitably removes Gaussians essential for high-quality rendering, leading to severe degradation in dynamic regions. In this paper, we introduce a novel 4D anchor-based framework that tackles the storage cost in different perspective. Rather than reducing the number of Gaussians, our method retains a sufficient quantity to accurately model dynamic contents, while compressing them into compact, grid-aligned 4D anchor features. Each anchor is processed by an MLP to spawn a set of neural 4D Gaussians, which represent a local spatiotemporal region. We design these neural 4D Gaussians to capture temporal changes with minimal parameters, making them well-suited for the MLP-based spawning. Moreover, we introduce a dynamic-aware anchor growing strategy to effectively assign additional anchors to under-reconstructed dynamic regions. Our method adjusts the accumulated gradients with Gaussians' temporal coverage, significantly improving reconstruction quality in dynamic regions. Experimental results highlight that our method achieves state-of-the-art visual quality in dynamic regions, outperforming all baselines by a large margin with practical storage costs.

Ours spiral


Neural 3D Video dataset





Technicolor dataset


Method


4D Anchor-based framework

Comparison Results


Rendered Results
Our approach achieves high-quality rendering with storage efficiency on the Technicolor dataset.


Neural 3D Video dataset


Please move the slider in the center to compare the videos.
All reported PSNR values are for the dynamic region of the scenes.

STG(25.02dB) cook_spinach Ours(27.39dB)
Ex4DGS(22.09dB) flame_salmon Ours(24.82dB)
4DGS(29.31dB) sear_steak Ours(31.44dB)




Technicolor dataset

STG(24.19dB) Theater Ours(31.12dB)