Geometry-Consistent Memory
Implicit visual-spatial evidence and explicit pose-depth cues are managed through geometry-gated writing, consistent fusion, and hierarchical retrieval.
Geometry as a principle for memory. Consistency as a signal for learning.
A walkthrough of the motivation, memory design, consistency learning, and experiments. The sections below provide a text overview.
Existing multimodal models often retain redundant video observations and produce unstable answers across viewpoints. ConsiSpace uses 3D spatial consistency to compact evidence and turn cross-view agreement into an explicit learning signal.


Implicit visual-spatial evidence and explicit pose-depth cues are managed through geometry-gated writing, consistent fusion, and hierarchical retrieval.
Answer, metric, and topology rewards refine cross-view stability after supervised fine-tuning without requiring additional human annotations.
An average gain of 12.6 points across the four evaluation settings below.
| Benchmark | Setting | Strongest baseline | ConsiSpace | Gain |
|---|---|---|---|---|
| VSI-Bench | Overall | 69.6 | 76.6 | +7.0 |
| OSI-Bench | Overall | 40.3 | 53.0 | +12.7 |
| MMSI-Video-Bench | Sufficient-Coverage | 42.6 | 57.5 | +14.9 |
| MMSI-Video-Bench | Uniform-50 | 43.1 | 58.1 | +15.0 |
Scores use each benchmark’s evaluation protocol. See full experimental results ↗

If ConsiSpace supports your research, please cite our work.
@article{huang2026consispace,
title = {ConsiSpace: Learning Geometric Consistency Matters
for Video Spatial Reasoning},
author = {Huang, Ting and Zhang, Zhenyu and Huang, Wenyuan
and Yang, Jian and Tang, Hao},
journal = {arXiv preprint arXiv:2607.17599},
year = {2026},
doi = {10.48550/arXiv.2607.17599},
url = {https://arxiv.org/abs/2607.17599}
}