Paper Summary
arXiv preprint

Yijun Yuan et al.

IIIS, Tsinghua University

arXiv:2509.16909v1 / 21 Sep 2025

Topics
#Visual SLAM#Dense SLAM#Transformer#Global Refinement

핵심 요약

SLAM-Former는 하나의 transformer를 streaming frontend와 global backend로 함께 쓰고, backend가 정제한 cache를 다음 frontend update에 다시 주입하는 dense SLAM 구조다.

문제모듈형 dense SLAM 해결one transformer 근거tracking / reconstruction
한 문장 요약

SLAM-Former는 하나의 transformer를 streaming frontend와 global backend로 함께 쓰고, backend가 정제한 cache를 다음 frontend update에 다시 주입한다.

Contribution 01

Unified Transformer

frontend tracking과 backend refinement를 하나의 shared backbone으로 처리.

Contribution 02

Cache Sharing

backend가 정제한 map representation을 frontend KV cache로 다시 전달.

Contribution 03

Training Modes

frontend, backend, cache-sharing 상황을 세 가지 mode로 함께 학습.

Contribution 04

Dense SLAM Evidence

tracking, reconstruction, cooperation ablation, runtime을 함께 검증.

핵심 인사이트

흥미로운 점은 global correction을 후처리로 떼어놓지 않고, shared cache를 통해 다음 streaming update의 일부로 만든다는 점이다.

처리 흐름
01RGB Streammonocular frames
02Frontendcausal KV update
03Map Tokenskeyframe memory
04Backendfull attention refinement
05Shared Cacheglobal context returns
06Outputspose / dense map
접근 방식 비교
Modular SLAM

분리된 모듈

해석성은 높지만 모듈 경계에서 오차가 누적될 수 있음.

3D Foundation Models

feed-forward geometry

reconstruction은 강하지만 긴 stream에서 과거 state를 다시 정제하기 어려움.

SLAM-Former

shared token memory

streaming frontend와 global backend가 같은 cache/map-token 표현 위에서 작동.

논문 상세 정리

이하에서는 배경지식, notation, 보조 유도까지 포함해 논문의 세부 내용을 살펴본다.

Comments