☰
SLAM History Book
한국어
/
English
← repos
검색 결과 없음
가이드 렌더링 중...
# Ch.0 — SLAM Solved? 2026년, 핸드폰을 들면 AR 레이어가 벽에 달라붙는다. 실내 배송 로봇은 지도를 받지 않고도 주방과 회의실을 구분한다. [DUSt3R](https://arxiv.org/abs/2312.14132) 계열 모델에 사진 몇 장을 던지면 수 초 안에 3D 구조가 나온다. 이제는 데모라기보다 제품이고, 대체로 배경 기술에 가깝다. 그래서 SLAM을 대체로 풀린 문제로 치는 분위기가 있다. --- 2003년으로 돌아가 보면 풍경이 다르다. Andrew Davison은 Imperial College London의 실험실에서 노트북 한 대와 웹캠 한 대로 실시간 3D 추적을 시연했다. [MonoSLAM](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf)이라 불린 그 시스템은 30Hz의 데스크톱 처리 속도에서 한 프레임당 10여 개의 특징만 주시하며 수십 개 규모의 희박한 지도를 유지했다. 방 하나, 책상 하나였다. 카메라가 책상 밖으로 나가면 지도가 발산했다. 그것이 당시 실시간 단안 SLAM의 대표적인 규모와 한계였다. 당시 시스템이 한 프레임에서 추적한 특징은 오늘날 핸드폰 AR보다 훨씬 적었다. 2003년의 추적 시스템은 *어떤 경로*를 거쳐 2026년의 AR로 이어졌는가. --- SLAM의 역사는 네 가지 서로 다른 전통이 독립적으로 진행되다가 충돌하며 서로를 흡수한 흔적이다. 사진측량학자들은 20세기 전반부터 여러 사진의 광선 다발을 함께 조정했다. 로봇공학자들은 1986년 [Smith-Cheeseman](https://arxiv.org/abs/1304.3111)의 확률적 공간관계 프레임을 거치며 지도를 확률의 언어로 다루기 시작했다. Durrant-Whyte와 Bailey의 [후대 역사 정리](https://www-personal.acfr.usyd.edu.au/tbailey/publications/slamtutorial1.htm)는 "SLAM"이라는 약어가 1995년 ISSRR 워크숍에서 채택됐다고 기록한다. 컴퓨터 비전 연구자들은 실시간 특징점 추적을 발전시켰고, 2020년대의 딥러닝 공동체는 이 구성 요소들을 학습된 모델 안에 다시 배치하고 있다. EKF 기반 SLAM이 graph-based로 교체된 것은 기술의 자연스러운 진화였는가, 아니면 몇 사람의 선택이 가른 우연이었는가. Feature-based와 direct method의 분기는 처음부터 예견된 것이었는가. 딥러닝이 geometry 파이프라인을 대체하는 속도가 이토록 더딘 이유는 무엇인가. Counterfactual은 선택지가 실제로 존재했을 때만 의미가 있다. 그 선택지들은 실제로 존재했다. --- 그 경로를 추적하려면 도구가 필요하다. 연도만 나열하면 연대기가 되고, 기법만 설명하면 교과서가 된다. 이 책은 계보와 예측이라는 두 렌즈로 역사를 읽는다. 어떤 아이디어가 어디서 왔는가. 연구자들이 당시 시점에서 본 미래와 실제로 펼쳐진 미래가 어떻게 갈렸는가. 이 책에는 네 가지 반복 장치가 있다. 각 챕터를 읽을 때 이 장치들을 길잡이로 쓸 수 있다. **계보 도입**은 챕터 첫 한두 단락에 놓인다. 그 챕터의 주인공이 어떤 지적 유산을 물려받았는지를 인물과 연도로 드러낸다. SLAM의 어떤 아이디어도 진공에서 탄생하지 않았다. 계보를 보면 차용의 지형이 보인다. **🔗 차용 박스**는 특정 기법이 어디서 왔는지를 한두 문장으로 명시하는 마진 주석이다. "ORB-SLAM의 이 구조는 Strasdat 2011에서 왔다"처럼. 연구자들은 인용하지만 계보를 명시하지 않는 경우가 많다. 이 박스는 그 계보를 드러낸다. **📜 예언 vs 실제 박스**는 원 논문의 Conclusion·Future Work·Summary 섹션이 짚은 것과 실제로 일어난 일을 대조한다. [Triggs 1999](https://dblp.org/rec/conf/dagstuhl/TriggsMHF99.html)의 BA 종합 논문이 §12 "Summary and Recommendations"에서 대규모 희소 구조 활용을 핵심 지침으로 남긴 자리를, 2010년대 [COLMAP](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf)이 수만 장 규모의 SfM을 오픈소스 실전 도구로 만들며 다른 각도에서 채웠다. 예측의 방향은 맞았지만 경로는 달랐다. 연구자가 당시 시점에서 본 미래와 실제 미래의 간극이 이 장치의 대상이다. **🧭 아직 열린 것**은 챕터 말미에 놓인다. 그 챕터가 다룬 주제에서 2026년 기준 아직 해결되지 않은 항목들이다. SLAM이 풀렸다는 인식 안에 숨어 있는 열린 문제들을 꺼낸다. Ch.19에서 이 항목들을 전 챕터에 걸쳐 수확해 재구성한다. --- 책은 6부로 구성된다. **1부: 선사시대**는 SLAM이 로봇공학에서 태어나기 이전, 사진측량과 고전 컴퓨터 비전이 쌓아 올린 도구들을 추적한다. 왜 bundle adjustment가 여전히 모든 최적화 backend의 뼈대인가? **2부: 고전 SLAM**은 확률적 지도 구축과 EKF의 한계에서 MonoSLAM·PTAM, graph-based SLAM과 최적성 인증까지를 다룬다. 지도와 추정을 함께 유지하는 문제는 어떻게 필터에서 최적화로 확장되었는가? **3부: 성숙기**는 ORB-SLAM과 관성 사전적분·연속시간 추정, direct method, RGB-D, place recognition을 다룬다. 초기 실시간 시연을 넘어 더 넓은 환경과 다양한 센서로 확장한 과정을 추적한다. **4부: 러닝 융합기**는 monocular depth 추정, end-to-end SLAM, 기하와 학습을 결합한 하이브리드 방법을 다룬다. 학습은 기존 파이프라인의 어느 계산을 맡았고, 기하학적 제약은 어디에 남았는가? **5부: 표현의 혁명**은 Neural Radiance Fields와 3D Gaussian Splatting, 동적 장면 표현, 3D foundation model을 다룬다. 지도와 장면을 표현하는 방식이 바뀌면서 추적과 재구성의 관계도 달라졌다. **6부: 막힌 길과 열린 문제**는 SLAM 역사의 실패한 경로들과, 오늘날 "풀렸다"는 인식 뒤에 남아 있는 구조적 미해결 문제들을 꺼낸다. --- 범위를 정해야 지도가 된다. Foundation model과 SLAM의 관계도 현재 연구가 남긴 열린 질문으로 다루되, 미래를 단정하지 않는다. 과거에 무슨 일이 있었고 왜 그랬는가가 재료다. 당시 제약 조건에서 그 선택이 어떤 의미였는지를 드러내는 것이 목표에 가깝다. homogeneous coordinates, epipolar geometry, EKF 공식은 독자가 이미 안다고 가정한다. 계보의 추적이 이 책의 일이고, 어떤 카메라나 LiDAR를 고를 것인가는 다른 책의 주제다. 수식과 정리·증명까지 체계적으로 짚고 싶다면 [SLAM Handbook](https://github.com/SLAM-Handbook-contributors/slam-handbook-public-release)이 있다. Carlone, Kim, Barfoot, Cremers, Dellaert가 편집해 Cambridge University Press에서 2026년에 나온 이 책은 18개 챕터에 SLAM의 현재 이론과 시스템을 총정리한다. 이 역사서는 그 상태에 이르기까지의 경로를 기록한다. 그 Handbook의 Epilogue에서 편집자 5인이 공동으로 남긴 격언 중 하나는 *"If someone tells you 'SLAM is solved,' don't listen to them"*이다. "풀린 문제로 치는 분위기"는 분야 내부의 관찰 대상이지 분야의 합의가 아니다. --- Davison의 2003년 데모 영상은 지금도 인터넷에 남아 있다. 흔들리는 화면, 깜빡이는 랜드마크 점들, 수십 개 규모의 희박한 지도. 거기서 여기까지 오는 사이에는 어떤 일이 있었는가. 그 기록은 MonoSLAM보다 훨씬 앞에서 시작한다. "SLAM"이라는 약어가 1990년대에 정착하기 전, 심지어 Smith-Cheeseman이 확률적 지도를 수식으로 쓰기 전부터 사진측량학자들은 카메라로 3D 구조를 복원하고 있었다. 그 선사(先史)는 사진측량에서 시작한다. --- # Ch.1 — 사진측량과 bundle adjustment: Triggs 이전 100년 오늘날 SLAM 최적화 backend의 뼈대는 독일 측량학에서 태어났다. 20세기 초 Carl Pulfrich가 유리판 위에서 두 시점의 삼각측량을 손으로 계산하던 방법론은, Albrecht Meydenbauer의 사진측량 체계와 결합해 하나의 측량 전통을 형성했다. 그 전통이 1958년 Duane C. Brown의 수치 공식화를 거쳐, 1999년 Bill Triggs, Philip McLauchlan, Richard Hartley, Andrew Fitzgibbon의 손에서 컴퓨터 비전 언어로 번역되었다. Bundle adjustment의 1999년 종합은 100년 된 측량 유산을 컴퓨터 비전 커뮤니티가 쓸 수 있는 언어로 옮긴 작업이었다. Triggs et al.(1999)은 Pulfrich의 기하학에서 시차 원리를, Brown(1958)의 군사 측량에서 reprojection 정식화를 물려받았다. solver 골격은 Levenberg-Marquardt가 제공했다. --- ## 1. 20세기 초 유리판과 Stereophotogrammetry 1901년 [Carl Pulfrich](https://en.wikipedia.org/wiki/Carl_Pulfrich)는 함부르크 자연과학자 회의에서 Zeiss 광학연구소가 제작한 **입체 측량기(stereocomparator)**를 발표했다(1899년 뮌헨에서 입체 거리계 시제품을 먼저 공개한 뒤의 정식 공개). 두 카메라 시점에서 같은 점을 찍고, 유리판 위의 좌표 차이를 읽어 거리를 산출하는 장치였다. 원리는 단순했다: 두 시점의 시차(parallax)가 깊이와 역비례한다. 수학은 그리스 시대의 삼각법이었고, 새로운 것은 광학 기기의 정밀도였다. 한 세대 앞선 흐름으로, [Albrecht Meydenbauer](https://www.uni-marburg.de/de/fotomarburg/histfoto/gliederung/messbilder/messbildverfahren)는 건축물 보존을 위한 **건축 사진측량(architectural photogrammetry)**을 체계화했다. 1858년 그는 베츨라 대성당 외벽을 측량하다 추락할 뻔한 뒤, 사진으로 직접 측량을 대신할 수 있다는 생각을 품었다. 1885년 그는 프로이센 왕립 사진측량국(Königlich Preussische Messbild-Anstalt)을 설립했다. 이 두 흐름이 합쳐진 전통이 20세기 항공 측량으로 이어졌다. 비행기 위에서 지형을 찍고, 두 시점 사진으로 3차원 지도를 만드는 aerotriangulation이다. 수동 계산기의 시대였다. > 🔗 **차용.** 현대 SLAM의 스테레오 깊이 추정은 Pulfrich의 stereocomparator와 같은 원리다. 두 카메라 간 baseline과 시차로 깊이를 구한다. 125년 전 유리판이 픽셀 배열로 바뀌었을 뿐이다. --- ## 2. 1958년 Brown과 수치 bundle adjustment Pulfrich와 Meydenbauer가 광학 기기로 해결한 문제를, Brown은 수식으로 옮겼다. [Duane C. Brown](https://digital.hagley.org/08206139_solution)은 미국 공군 탄도미사일 개발 체계의 측량 엔지니어였다. 위성 궤도와 지상 좌표를 함께 추정하는 문제, 즉 다수의 카메라 시점과 다수의 지상 제어점을 동시에 최적화하는 문제를 다루었다. 1958년 보고서 "A Solution to the General Problem of Multiple Station Analytical Stereotriangulation"(RCA-MTP Data Reduction Technical Report No. 43, AFMTC-TR-58-8)에서 Brown은 **bundle adjustment**를 수치적으로 공식화한 초기 문헌 중 하나를 남겼다(같은 시기 Helmut Schmid도 공동 발명자로 함께 거론된다). **Reprojection error**는 카메라 $i$에서 관측된 2D 이미지 좌표 $x_{ij} \in \mathbb{R}^2$와, 3D 점 $X_j \in \mathbb{R}^3$을 내부 행렬 $K_i$·외부 행렬 $[R_i | t_i]$로 투영한 예측 좌표 $\pi(K_i, R_i, t_i, X_j)$의 차이다. 이 차이를 최소화한다: $$E = \sum_{i,j} \| x_{ij} - \pi(K_i, R_i, t_i, X_j) \|^2$$ "Bundle"이라는 이름은 각 카메라 중심에서 관측된 3D 점들로 뻗어 나가는 광선 다발(bundle of rays)에서 왔다. 그 광선들이 3차원 점에서 교차하도록 카메라 자세와 점 위치를 동시에 조정한다. Brown의 작업은 군사 측량을 배경으로 했지만, bundle adjustment는 이후 사진측량 문헌에서 계속 발전했다. > 🔗 **번역.** Triggs et al.(1999)은 논문을 "컴퓨터 비전 공동체의 잠재적 구현자"를 위한 사진측량 bundle adjustment 종합으로 규정했다. 두 분야가 같은 최적화 문제를 서로 다른 용어와 관행으로 다루던 상황에서, 이 논문은 사진측량의 축적을 컴퓨터 비전 쪽으로 옮기는 연결점이 됐다. 군사 보안 때문에 컴퓨터 비전이 이를 독립적으로 재발견했다는 직접 근거는 확인되지 않는다. --- ## 3. Levenberg와 Marquardt — 비선형 최적화의 선구자 Brown이 최소화해야 할 목적함수를 손에 쥐었다면, 그것을 실제로 푸는 도구는 전혀 다른 곳에서 왔다. Reprojection error 최소화는 비선형 최소제곱 문제다. 해석적 해가 없으므로 반복 수치 최적화가 필요하다. 1944년 [Kenneth Levenberg](https://cs.uwaterloo.ca/~y328yu/classics/levenberg.pdf)는 Gauss-Newton과 steepest descent를 댐핑 파라미터 $\lambda$로 보간하는 방법을 발표했다. $\lambda$가 클수록 steepest descent에 가까워져 안전하게 수렴하고, 작을수록 Gauss-Newton의 빠른 수렴을 활용한다. 기본 형태에서는 선형화한 정규방정식의 행렬 $\mathbf{J}^{\top}\mathbf{J}$에 $\lambda \mathbf{I}$를 더해 수치 안정성을 높인다. 컴퓨터 비전보다 20년 앞선 시점이었다. 1963년 [Donald Marquardt](https://epubs.siam.org/doi/10.1137/0111030)는 같은 아이디어를 독립적으로 재발견해 더 명시적으로 공식화했다. **Levenberg-Marquardt(LM) 알고리즘**이라는 이름으로 굳어졌다. LM 알고리즘이 컴퓨터 비전에서 BA의 표준 solver가 되기까지 약 35년이 더 걸렸다. --- ## 4. 1999년 Triggs et al. — 100년 유산 통합 이 수치 도구와 BA의 계산 구조를 컴퓨터 비전 독자에게 체계적으로 정리한 문헌이 Triggs와 동료들의 종합 논문이었다. 1999년 Vision Algorithms Workshop에서 Bill Triggs, Philip McLauchlan, Richard Hartley, Andrew Fitzgibbon은 ["Bundle Adjustment — A Modern Synthesis"](https://link.springer.com/chapter/10.1007/3-540-44480-7_21)를 발표했다. 이 논문은 20세기 측량학과 항공 사진측량에 흩어져 있던 BA 이론을 컴퓨터 비전 커뮤니티의 언어로 번역해 종합했다. Triggs et al.은 sparse BA의 구조적 성질을 명시했다. Hessian 행렬의 희소 블록 구조(Schur complement trick)를 이용하면 카메라-점 결합 최적화를 훨씬 효율적으로 수행할 수 있다. 또한 gauge freedom(기준틀의 임의성)을 명시적으로 다루었다. 이 논문이 나오고 7년 후, Noah Snavely의 [Photo Tourism(2006)](https://phototour.cs.washington.edu/Photo_Tourism.pdf)은 인터넷에 흩어진 사진 수백 장에서 노트르담·트레비 분수 같은 유명 랜드마크를 자동 재구성했다. 그로부터 10년 후 Johannes Schönberger의 [COLMAP(2016)](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf)은 수만~수십만 장 규모의 robust incremental SfM을 오픈소스로 공개하며, 이미 백만 장대까지 가 있던 연구 흐름을 누구나 재현할 수 있는 도구로 가져왔다. Triggs의 언어가 없었다면 그 경로는 훨씬 느렸을 것이다. --- ## 5. Reprojection Error — 개념의 형성 Triggs et al.이 최소화할 오차 함수를 정리했다면, 그 함수가 현재 형태로 정착한 과정은 별도의 계보를 이룬다. 이 오차 함수가 지금 형태로 자리 잡기까지 두 번의 전환이 있었다. 20세기 초 항공 삼각측량사들은 오차를 "지상 좌표계에서의 거리 차이"로 쟀다. 3D 공간에서 직접 비교하는 방식이어서, 카메라 렌즈가 틀어졌거나 캘리브레이션이 나빠도 그 오차는 지상 좌표 잔차에 녹아들어 보이지 않았다. Brown은 1958년 보고서에서 비교 대상을 이미지 면으로 옮겼다. "3D 점을 이미지로 투영한 위치"와 "실제 이미지 관측"을 픽셀 단위로 맞추는 방식이다. 이렇게 하면 캘리브레이션 오차, 렌즈 왜곡, 외부 파라미터 오차가 하나의 잔차에 함께 드러난다. 통계적으로도 더 깔끔하다. 카메라 이미지 노이즈는 픽셀 단위의 등방성 가우시안으로 모델링할 수 있고, 그러면 reprojection error 최소화는 최대우도 추정과 같아진다. Triggs et al.(1999)은 그 공식을 컴퓨터 비전 교과서 언어로 다듬어 표준화했다. 이 reprojection error minimization이 2026년 기준 factor graph 기반 SLAM backend의 시각 관측을 다루는 핵심 최적화 문제다. > 🔗 **차용.** SLAM에서 visual landmark의 관측 모델 $z = \pi(K, T, p) + \epsilon$은 Brown(1958)의 reprojection 공식을 직접 계승한다. Gauss-Newton으로 이를 최소화하는 SLAM backend는 1958년 항공 삼각측량 solver와 수학적으로 동일한 구조를 가진다. --- ## 6. SLAM Backend의 뼈대 — 2026년까지 "SLAM"이라는 약어는 1995년 ISSRR 워크숍과 뒤이은 1990년대 문헌을 거치며 널리 쓰이기 시작했지만, 그 backend의 수학은 사진측량의 bundle adjustment와 같은 비선형 최소제곱 구조를 물려받는다. 오늘날 ORB-SLAM3는 g2o를 통해 SE(3) 자세와 3D landmark 위치를 동시 최적화한다. LIO-SAM은 GTSAM의 factor graph 위에서 비선형 최적화를 수행한다. DROID-SLAM은 GRU-based optical flow로 업데이트 방향을 구하지만, differentiable dense bundle adjustment 레이어를 계산 그래프 안에 남겨 둔다. Lie group과 factor graph가 1999년의 행렬 표기를 대체했고, 신경망이 사람이 설계하던 계산 일부를 넘겨받았지만, 관측 잔차를 최소화하는 공통 구조는 이어진다. 시각 BA는 다수 시점의 reprojection error로 카메라 자세와 맵을 추정하고, LiDAR 시스템은 거리나 점–평면 잔차 등 센서에 맞는 제약을 쓴다. Pulfrich의 유리판이 픽셀 배열로 바뀌고, 손 계산이 GPU로 바뀌었을 뿐이다. 이 연속성은 분야의 강점이자 취약점이다. 오랫동안 축적된 최적화 이론과 실용 검증을 그대로 활용할 수 있다. 그러나 BA의 전제(static world, point feature, Gaussian noise)가 현실 환경과 어긋나면 그 검증의 범위도 함께 벗어난다. --- > 📜 **예언 vs 실제.** Triggs et al.(1999)은 수천 대 카메라와 수백만 점을 다루는 대규모 BA로의 확장을 주요 도전으로 꼽은 것으로 널리 읽힌다. 그 방향성은 이후 20년에 걸쳐 달성되었다. 2006년 Snavely의 Photo Tourism이 인터넷 사진 수백 장으로 랜드마크를 재구성했고, 2016년 COLMAP은 그 흐름의 robust incremental SfM 구현체를 표준화했다. 이후의 확장은 최적화 문제를 그대로 키우는 일만으로 이루어지지는 않았다. incremental BA와 visibility graph pruning 위에 vocabulary tree 루프 클로저가 얹힌, 엔지니어링 층의 결과였다. --- ## 🧭 아직 열린 것 **비선형 BA의 global optimum 보장.** LM 알고리즘은 국소 최솟값(local minimum)에 수렴한다. 초기값이 나쁘면 틀린 구조에 수렴한다. 초기화를 위한 방법들, 즉 5-point algorithm, PnP, epipolar geometry 추정이 차례로 등장했지만 기하 solver와 이를 감싸는 강건 추정 절차는 구분해야 한다. RANSAC은 이상치를 가르는 외부 절차로 쓰이며, 얻은 해에 반복 정제를 적용할 수 있다. 대규모 환경에서 전역 최적을 보장하는 convex relaxation 기반 접근들이 연구되고 있으나, 실시간 SLAM 수준의 속도와 규모에서는 아직 실용화되지 않았다. **사진측량 수준 정밀도와 Visual SLAM의 간극.** 항공 사진측량의 정확도는 촬영 조건과 산출물의 요구 규격에 따라 평가한다. 교정된 카메라와 고품질 GCP(지상 기준점)가 있고, 최적화는 오프라인에서 수행한다. 실시간 Visual SLAM은 같은 수식 구조를 쓰면서도 GPS 없는 환경과 저해상도 카메라, 그리고 즉각적 추정이라는 제약 아래서 동작한다. 두 분야의 정확도를 비교하려면 영상 좌표 오차와 지상 좌표 오차를 구분하고, 거리·해상도·기준점 조건과 평가 성분을 함께 명시해야 한다. --- BA의 전제(static world, point feature, Gaussian noise)가 무너지기 시작하는 것은 카메라가 이동하는 물체를 만났을 때다. 측량사는 다리를 측량하지 로봇 축구 경기장을 측량하지 않았다. 그에 앞서 최적화에 넣을 대응점부터 찾아야 한다. Ch.2는 Harris corner와 optical flow를 통해 이미지에서 그 입력을 확보해 온 경로를 따른다. --- # Ch.2 — Classical CV 도구상자: Harris에서 SIFT까지, 그리고 ORB까지 Ch.1의 번들조정은 카메라 자세와 3D 점을 동시에 최적화하는 backend 문제를 다루었다. 그러나 그 최적화가 작동하려면 먼저 이미지에서 "대응하는 점"을 찾아야 한다. 측량사는 야지에서 직접 타깃을 세웠고, 컴퓨터 비전은 그 역할을 알고리즘에 맡겨야 했다. feature detection·description은 그렇게 시작된 문제다. 1970년대 후반 Hans Moravec은 Stanford Cart 프로젝트에서 카메라로 환경의 두드러진 점을 찾으려 했다. 그 작업은 1980년 Stanford 박사논문 ["Obstacle Avoidance and Navigation in the Real World by a Seeing Robot Rover"](https://frc.ri.cmu.edu/~hpm/project.archive/robot.papers/1975.cart/1980.html.thesis/index.html)로 정리됐다. Moravec은 주변 패치의 밝기 변화로 추적할 점을 고르는 정량적 기준을 제시했다. 1988년 Chris Harris와 Mike Stephens가 그 직관을 autocorrelation matrix의 eigenvalue로 공식화했다. Lucas와 Kanade는 그보다 7년 앞서 픽셀 추적의 틀을 세웠다. Lowe는 두 개념을 흡수해 scale과 rotation에 불변인 서술자를 만들었다. Rublee는 특허 없이 더 빠르게 같은 일을 했다. SLAM의 front-end는 이 계보 위에서 돌아간다. --- ## 2.1 코너라는 개념: Moravec에서 Harris까지 카메라가 조금 움직였을 때 영상 패치가 크게 변하는 점을 "코너"라 부른다. Moravec(1977)의 기준은 단순했다. 인접 픽셀과의 Sum of Squared Differences(SSD)가 상하좌우 모든 방향에서 크면 코너로 간주한다. Harris와 Stephens는 1988년 Alvey Vision Conference에서 ["A Combined Corner and Edge Detector"](https://www.bmva.org/bmvc/1988/avc-88-023.html)를 발표하며 이를 연속 미분으로 대체했다. 이미지 $I$에서 점 $(x,y)$ 주변 창 $W$를 이동량 $(\Delta x, \Delta y)$로 움직일 때 강도 변화를 근사하면: $$M = \sum_{(x,y) \in W} \begin{pmatrix} I_x^2 & I_x I_y \\ I_x I_y & I_y^2 \end{pmatrix}$$ $M$의 두 eigenvalue $\lambda_1, \lambda_2$로 점의 성격을 구분한다. 둘 다 크면 코너, 하나만 크면 엣지, 둘 다 작으면 평탄한 영역. Harris는 고유값을 직접 구하지 않고 $R = \det(M) - k \cdot \text{tr}(M)^2$ 점수를 사용해 eigenvalue 분해를 피했다. $k$는 보통 0.04–0.06. > 🔗 **차용.** Harris(1988)의 autocorrelation matrix 아이디어는 Moravec(1977)의 SSD 기반 코너 탐색을 연속 미분으로 정제한 것이다. 개념의 원형은 Stanford Cart 보고서에 있었다. 1994년 Jianbo Shi와 Carlo Tomasi는 ["Good Features to Track"](https://cecas.clemson.edu/~stb/klt/shi-tomasi-good-features-cvpr1994.pdf)(CVPR 1994)에서 Harris 점수 대신 $\min(\lambda_1, \lambda_2)$를 직접 사용하는 것이 optical flow 추적에 더 안정적임을 보였다. 이 기준이 Shi-Tomasi 코너 검출이다. OpenCV는 `goodFeaturesToTrack` 함수로 이를 구현했다. 30년이 지난 오늘도 그 함수는 그대로다. --- ## 2.2 추적의 원형: Lucas-Kanade와 KLT Harris의 행렬 $M$은 점을 찾는다. 찾은 점을 다음 프레임에서 다시 찾는 것은 별개 문제다. Bruce Lucas와 Takeo Kanade는 1981년 ["An Iterative Image Registration Technique"](https://www.ijcai.org/Proceedings/81-2/Papers/017.pdf)에서 프레임 간 픽셀 이동을 밝기 불변 가정(brightness constancy assumption) 아래 최소화 문제로 정식화했다. 밝기 불변 가정: 픽셀 $(x,y)$의 강도는 움직임 전후로 같다. $$I(x, y, t) = I(x + u, y + v, t + 1)$$ 테일러 전개 후 선형화하면: $$I_x u + I_y v + I_t = 0$$ 이 방정식 하나에 미지수가 둘이다. Lucas-Kanade는 $3\times3$ 또는 $5\times5$ 창 안의 픽셀들이 같은 $(u,v)$로 움직인다는 가정을 추가해 overdetermined 시스템을 만들고 최소자승으로 푼다. $$\begin{pmatrix} \sum I_x^2 & \sum I_x I_y \\ \sum I_x I_y & \sum I_y^2 \end{pmatrix} \begin{pmatrix} u \\ v \end{pmatrix} = -\begin{pmatrix} \sum I_x I_t \\ \sum I_y I_t \end{pmatrix}$$ 왼쪽 행렬은 Harris의 구조 행렬 $M$과 동일하다. 코너 검출과 optical flow는 같은 수학을 쓴다. Tomasi와 Kanade는 1991년 tech report ["Detection and Tracking of Point Features"](https://cecas.clemson.edu/~stb/klt/tomasi-kanade-techreport-1991.pdf)에서 추적 창의 품질을 eigenvalue 기준으로 선택하고 Newton-Raphson 반복으로 displacement를 정제하는 구체적 구현을 제시했다. 이후 Bouguet(Intel, 2000)이 이미지 피라미드 기반 coarse-to-fine 전략을 더해 큰 이동에서도 수렴하도록 확장했고, 이 조합이 KLT(Kanade-Lucas-Tomasi) 추적기로 정착했다. [VINS-Mono](https://arxiv.org/abs/1708.03852)(2018) 같은 실시간 VIO가 여전히 이 계보의 front-end를 돌린다. 1981년의 최소자승 추적기가 40여 년 뒤 스마트폰 드론의 VIO에서 돌아가는 셈이다. > 🔗 **차용.** Lucas-Kanade(1981) → KLT tracker → Qin et al. VINS-Mono(2018): 37년 전 제안된 optical flow가 실시간 VIO의 feature tracking backbone으로 그대로 살아있다. --- ## 2.3 SIFT — 불변성과 특허 KLT는 같은 카메라가 조금씩 이동하는 상황에 맞다. 다른 카메라로, 다른 날 찍은 이미지에서 같은 점을 연결하는 문제는 다른 차원이다. 시점이 달라지면 같은 점이 패치 모양과 크기, 방향까지 달라져 단순 픽셀 비교가 통하지 않는다. **서술자(descriptor)**가 필요한 이유다. David Lowe(UBC)는 1999년 ICCV에서 아이디어를 발표했다. 당시 발표 제목은 "Object Recognition from Local Scale-Invariant Features"였고, 128차원 벡터를 비교해 서로 다른 사진에서 같은 물체를 찾아내는 시연을 보였다. 5년 뒤인 2004년 IJCV에 완성판 ["Distinctive Image Features from Scale-Invariant Keypoints"](https://www.cs.ubc.ca/~lowe/papers/ijcv04.pdf)가 실렸고, 이것이 오늘날 SIFT로 인용되는 논문이다. SIFT(Scale-Invariant Feature Transform)는 두 단계로 구성된다. **검출 단계**: DoG(Difference of Gaussians)를 여러 scale에서 계산해 local extremum을 keypoint로 선택한다. DoG는 Laplacian of Gaussian의 근사다. $L(x,y,\sigma) = G(x,y,\sigma) * I(x,y)$를 Gaussian 스무딩 이미지라 하면: $$D(x, y, \sigma) = L(x, y, k\sigma) - L(x, y, \sigma)$$ 여기서 $k$는 인접 scale 간 비율이며 보통 $2^{1/s}$이다($s$는 octave당 스케일 수). 여러 octave에 걸쳐 극값을 찾으면 scale 변화에도 같은 점을 탐지할 수 있다. **서술자 단계**: keypoint 주변 $16\times16$ 창을 $4\times4$ 블록으로 나누고 각 블록의 gradient 방향 히스토그램(8빈)을 연결해 128차원 벡터를 만든다. keypoint의 dominant gradient 방향을 기준으로 회전시키므로 회전 불변성도 확보한다. SIFT는 크기와 회전, 일부 affine 변형에 강인한 128차원 서술자였다. KITTI 이전 시대, SLAM 벤치마크가 없던 시절에도 시점과 크기가 다른 영상의 대응점을 찾는 데 SIFT가 유용했던 이유다. Lowe는 2000년 3월 SIFT를 특허 출원했고, 2004년 3월 등록됐다(US6711293B1, 우선권 1999년 3월). 이 특허는 상업용 사용에 비용을 부과했고, 2020년 3월 만료 전까지 SIFT를 대체하려는 시도의 동기 중 하나가 되었다. > 📜 **예언 vs 실제.** Lowe는 2004년 SIFT 논문의 "9 Conclusions"에서 서술자의 확장 가능성을 "view matching for 3D reconstruction, motion tracking and segmentation, robot localization, image panorama assembly, epipolar calibration"로 나열했다. 방향 자체는 대부분 맞았다. SfM·SLAM·파노라마·초기 영상 기반 로봇 위치추정이 2000년대 후반 SIFT에 기댔다. 다만 장기 대응 문제에서 SIFT의 위치는 CNN 이후 흔들렸다. 2012년 AlexNet 이후 물체 인식 쪽 수요는 CNN으로 이동했고, SLAM용 local descriptor 자리도 SuperPoint·R2D2 같은 학습 서술자가 점차 가져갔다. 응용 영역 예측은 적중, 서술자 형태는 기술변화. --- ## 2.4 SURF — 속도와 정확도 절충 SIFT의 128차원 서술자는 정확했지만, 당시 데스크톱 CPU에서 이미지 한 장을 처리하는 데 수백 밀리초가 걸려 실시간 SLAM에 쓰기 어려웠다. Herbert Bay(ETH Zurich)는 2006년 ECCV에 ["SURF: Speeded-Up Robust Features"](https://people.ee.ethz.ch/~surf/eccv06.pdf)를 발표하고 탐지와 서술 양쪽의 계산량을 줄였다. DoG 대신 *Hessian 행렬의 행렬식*으로 keypoint를 탐지한다. integral image를 이용한 box filter로 Gaussian 이차 미분을 근사해 계산 속도를 높인다. 서술자는 64차원으로 SIFT의 절반. keypoint 주변을 $4\times4$ 하위 영역으로 나누고, 각 영역에서 Haar wavelet 응답 $d_x, d_y$의 합 $(\sum d_x,\, \sum d_y,\, \sum|d_x|,\, \sum|d_y|)$ 4값을 연결해 $4\times4\times4=64$차원을 구성한다. 128차원 확장(SURF-128)도 존재하나 기본값은 64차원이다. SURF는 SIFT보다 3–7배 빨랐다. 그러나 정확도 비교는 차원 수뿐 아니라 검출기와 서술 방식, 평가 조건에 따라 달랐고, Bay도 특허를 피하지 못했다(ETH Zurich 특허). SIFT는 속도 때문에 밀렸고, SURF는 정확도와 특허 두 가지 때문에 밀렸다. 두 문제를 동시에 푼 것이 ORB였다. > 🔗 **차용.** Lowe(1999/2004)의 DoG scale-space → Bay(2006)의 Hessian integral image: scale-invariance를 얻는 두 가지 답. DoG는 이론적으로 우아하고, Hessian 근사는 공학적으로 빠르다. --- ## 2.5 ORB — binary descriptor와 특허 해방 2011년 Ethan Rublee(Willow Garage), Vincent Rabaud, Kurt Konolige, Gary Bradski는 ICCV에 ["ORB: An Efficient Alternative to SIFT or SURF"](https://www.gwylab.com/download/ORB_2012.pdf)를 발표했다. 제목이 직접적이다. Willow Garage는 ROS의 산실이기도 했다. 로보틱스 연구자가 쓸 수 있는 feature를 만들겠다는 동기가 제목에 그대로 담겼다. ORB는 두 기존 기법을 조합하고 개선했다. **검출**: [FAST](https://www.edwardrosten.com/work/rosten_2006_machine.pdf)(Features from Accelerated Segment Test, Rosten & Drummond 2006). 픽셀 주변 16개 점을 순환하며 충분히 밝거나 어두운 연속 호가 있으면 코너로 판정한다. SIFT의 DoG보다 10배 이상 빠르다. ORB는 FAST에 Harris 점수를 추가해 응답이 강한 것만 남긴다. **서술자**: [BRIEF](https://www.cs.ubc.ca/~lowe/525/papers/calonder_eccv10.pdf)(Binary Robust Independent Elementary Features, Calonder et al. 2010). keypoint 주변 패치에서 무작위로 선택한 점 쌍의 밝기를 비교해 비트열을 만든다. 256비트가 기본이다. 유클리드 거리 대신 Hamming 거리로 매칭하므로 두 비트열에 XOR를 적용한 뒤 1인 비트 수를 세는 popcount로 거리를 구한다. BRIEF의 약점은 회전 불변성 부재였다. Rublee는 FAST 코너의 intensity centroid 방향으로 패치를 회전 보정해 **rBRIEF(rotated BRIEF)**를 만들었다. 방향 추정이 들어오면서 회전이 있는 영상 쌍에서도 BRIEF를 활용할 수 있게 됐다. $$\theta = \text{atan2}(m_{01},\, m_{10}), \quad m_{pq} = \sum_{x,y} x^p y^q I(x,y)$$ 계산 속도는 SIFT의 100배였고, 특허가 없었으며, OpenCV에 즉시 통합됐다. [ORB-SLAM](https://arxiv.org/abs/1502.00956)(Mur-Artal et al. 2015)은 이름 그대로 ORB를 기반으로 했고, 이후 삼부작까지 이어졌다. ORB-SLAM3는 2021년에도 front-end를 바꾸지 않았다. > 🔗 **차용.** Calonder et al.(2010)의 BRIEF → Rublee et al.(2011)의 ORB: binary descriptor에 intensity centroid 기반 방향 추정을 추가해 회전 불변성을 확보했다. --- ## 2.6 학습 기반 descriptor ORB가 실용적 정점이라면, 그 뒤의 질문은 자연스럽다. 손으로 설계한 규칙이 아닌 학습된 규칙이 더 나은가. 2016년 Yi et al.의 [LIFT](https://arxiv.org/abs/1603.09114)(Learned Invariant Feature Transform, ECCV 2016)는 검출·방향 추정·서술자 세 단계를 CNN으로 대체하려 했다. 단계별로 따로 학습한 세 네트워크를 파이프라인으로 연결하는 구조였다. 2018년 DeTone et al.의 [SuperPoint](https://arxiv.org/abs/1712.07629)(CVPRW 2018)는 homographic adaptation이라는 자기지도 학습법으로 keypoint 검출과 256차원 서술자를 동시에 학습했다. 합성 데이터로 사전 학습한 뒤 실제 이미지에 적응했다. 이후 SLAM 연구에서 널리 시험된 learned local feature 가운데 하나가 됐다. 그러나 2026년 기준으로도 전통 descriptor가 사라지지 않았다. ORB는 임베디드 장치에서 SuperPoint보다 빠르고, 도메인 밖 이미지에서 일반화가 불안정한 learned descriptor보다 예측 가능한 동작을 보인다. [AnyLoc](https://arxiv.org/abs/2308.00688)(Keetha et al. 2023)처럼 DINOv2 기반 feature가 장소 인식에 도입되었지만, ORB-SLAM3는 2021년 발표 이후 여전히 ORB를 쓴다. 1977년 Moravec의 직관이 2020년대 로봇 위에서 돌아가고 있다. --- ## 2.7 🧭 아직 열린 것 **학습 기반 descriptor의 일반화 한계.** SuperPoint, R2D2, DISK 등 learned descriptor는 학습 도메인에서 전통 방법을 능가하지만 새로운 환경(underwater, thermal, low-light)에서는 일관성이 없다. 어느 쪽이 낫다는 합의는 2026년에도 없다. **Wide-baseline 매칭의 실패 모드.** Harris나 ORB 기반 매칭은 큰 시점 변화에서 성능이 떨어질 수 있으며, 그 정도는 장면과 회전축·매칭 조건에 따라 다르다. Affine-covariant detector(ASIFT, MSER)가 일부 보완했지만, 완전한 해법은 없다. [DUSt3R](https://arxiv.org/abs/2312.14132)(Wang et al. 2023)는 matching 자체를 회피했지만, 이것이 descriptor 문제의 종말인지 우회인지는 아직 판단하기 이르다. --- Harris의 직관과 Lowe의 불변성이 기반을 놓았고, Rublee의 속도 최적화가 그것을 현장으로 끌어냈다. 이 기법들은 각자 이미지 한 장 혹은 두 장 사이에서 일하도록 설계됐다. 수십, 수백 장의 이미지를 동시에 기하적으로 일관되게 연결하려면 다른 층이 필요했다. --- *참고 문헌* - Harris, C. & Stephens, M. (1988). A Combined Corner and Edge Detector. *Proc. Alvey Vision Conference*. - Lucas, B. D. & Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. *IJCAI*. - Shi, J. & Tomasi, C. (1994). Good Features to Track. *CVPR*. - [Lowe, D. G. (2004). Distinctive Image Features from Scale-Invariant Keypoints.](https://doi.org/10.1023/B:VISI.0000029664.99615.94) *IJCV 60(2)*. - Bay, H., Tuytelaars, T. & Van Gool, L. (2006). SURF: Speeded-Up Robust Features. *ECCV*. - Calonder, M. et al. (2010). BRIEF: Binary Robust Independent Elementary Features. *ECCV*. - [Rublee, E. et al. (2011). ORB: An Efficient Alternative to SIFT or SURF.](https://doi.org/10.1109/ICCV.2011.6126544) *ICCV*. - DeTone, D., Malisiewicz, T. & Rabinovich, A. (2018). SuperPoint: Self-Supervised Interest Point Detection and Description. *CVPRW*. [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) --- # Ch.3 — Structure from Motion: Longuet-Higgins에서 COLMAP까지 Harris와 Lowe가 이미지 안에서 "볼 만한 점"을 골라내는 방법을 다듬는 동안, 다른 계보는 그 점들이 두 장의 사진에 동시에 찍혔을 때 무엇을 알 수 있는가를 물었다. 특징을 *검출*하는 문제와 특징으로부터 *공간을 재구성*하는 문제는 같은 시기에 각자 발전했고, 2000년대 중반에야 하나의 파이프라인으로 합쳐졌다. 1981년 케임브리지 이론심리학자 H.C. Longuet-Higgins는 *Nature*에 세 페이지짜리 논문을 실었다. 제목은 "[A Computer Algorithm for Reconstructing a Scene from Two Projections](https://cseweb.ucsd.edu/classes/fa01/cse291/hclh/SceneReconstruction.pdf)". 그는 두 장의 사진에 찍힌 같은 점들의 좌표 여덟 쌍으로 카메라의 상대 운동과 장면 구조를 복원하는 선형 알고리즘을 제시했다. 로봇공학자도 컴퓨터 비전 연구자도 아니었다. 사진측량과 motion perception에 선행 연구가 있었으므로 이 논문을 SfM 전체의 시작으로 볼 수는 없지만, 현대 컴퓨터 비전의 two-view geometry를 여는 영향력 있는 전환점이었다. 이후 self-calibration, incremental SfM, Bundler와 대규모 인터넷 사진 재구성이 이어졌고, Johannes Schönberger의 2016년 COLMAP은 그 축적을 강건한 범용 오픈소스 파이프라인으로 정리했다. --- ## 3.1 Essential Matrix와 8-point Algorithm Longuet-Higgins의 출발점은 단순했다. 두 카메라로 같은 점을 찍으면, 그 점의 이미지 좌표 쌍 사이에 대수적 제약이 존재한다. 좌표계를 정규화하면 이 제약은 행렬 하나로 집약된다. 그는 이것을 **essential matrix** $\mathbf{E}$로 정의했다. 두 카메라의 중심을 각각 $\mathbf{O}_1$, $\mathbf{O}_2$, 대응점을 정규화 좌표 $\mathbf{x}_1$, $\mathbf{x}_2$라 하면 제약은: $$\mathbf{x}_2^\top \mathbf{E} \mathbf{x}_1 = 0$$ $\mathbf{E}$는 카메라 사이의 회전 $\mathbf{R}$과 이동 $\mathbf{t}$로부터 $\mathbf{E} = [\mathbf{t}]_\times \mathbf{R}$로 인수분해된다. 여기서 $[\mathbf{t}]_\times$는 $\mathbf{t}$의 skew-symmetric 행렬이다. Essential matrix는 스케일 모호성을 제거하면 자유도가 5이다. 그러나 5개 대응점으로 푸는 non-linear 5-point algorithm([Nistér 2004](http://www.cad.zju.edu.cn/home/gfzhang/training/SFM/2004-PAMI-David%20Nister-An%20Efficient%20Solution%20to%20the%20Five-Point%20Relative%20Pose%20Problem.pdf))이 등장하기 전까지, 표준 접근은 rank-2 제약과 단위 스케일 제약을 강제하기 전 단계에서 행렬을 9개 원소 중 스케일 1개를 고정해 8개의 미지수로 보고 8개 대응점으로 선형 시스템을 푸는 것이었다. 이것이 **8-point algorithm**이다. Longuet-Higgins 자신은 정확히 8개 점으로 유일해를 구하는 절차를 제시했다. 구현은 간단했고, 계산량도 작았다. 문제는 수치 안정성이었다. 이미지 좌표가 수백~수천 픽셀 단위이면 계수행렬의 원소 크기가 크게 달라져 SVD가 불안정해진다. > 🔗 **차용.** Hartley는 1997년 정규화된 8-point algorithm([In Defense of the Eight-Point Algorithm](https://www.cse.unr.edu/~bebis/CS485/Handouts/hartley.pdf))에서 이미지 좌표를 평균 0, 평균 거리 $\sqrt{2}$로 선형 변환한 뒤 fundamental matrix를 추정하고 원래 좌표계로 되돌리는 방식을 내놓았다. 카메라 내부 파라미터로 좌표를 보정하는 정규화와는 목적이 다르다. Longuet-Higgins의 기하학은 그대로 두고, 수치 조건만 고쳤다. 이 정규화 절차는 이후 다중 뷰 기하 교재와 구현에서 널리 채택됐다. Fundamental matrix $\mathbf{F}$는 essential matrix의 일반화다. 카메라 내부 파라미터 $\mathbf{K}$를 알지 못해도 $\mathbf{x}_2^\top \mathbf{F} \mathbf{x}_1 = 0$이 성립한다. 두 카메라의 내부 파라미터를 각각 $\mathbf{K}_1$, $\mathbf{K}_2$라 하면 관계는 $\mathbf{F} = \mathbf{K}_2^{-\top} \mathbf{E} \mathbf{K}_1^{-1}$이다. 같은 카메라로 찍은 경우($\mathbf{K}_1 = \mathbf{K}_2 = \mathbf{K}$)에는 $\mathbf{F} = \mathbf{K}^{-\top} \mathbf{E} \mathbf{K}^{-1}$로 단순화된다. SfM 파이프라인에서 $\mathbf{K}$를 모를 때는 $\mathbf{F}$를 먼저 추정하고, $\mathbf{K}$를 알 때는 $\mathbf{E}$를 직접 푼다. --- ## 3.2 Tomasi-Kanade Factorization 1981년 이후 십 년간 SfM은 주로 두 장 사진 사이의 기하학으로 연구되었다. 여러 장 사진을 동시에 처리하는 방법은 별도의 문제였다. 1992년 Carlo Tomasi와 Takeo Kanade가 CMU에서 **[factorization method](https://people.eecs.berkeley.edu/~yang/courses/cs294-6/papers/TomasiC_Shape%20and%20motion%20from%20image%20streams%20under%20orthography.pdf)**를 발표하면서 이 문제의 윤곽이 드러났다. $F$장의 프레임에서 $P$개의 포인트를 관측한다면, 이미지 좌표를 $2F \times P$ 행렬 $\mathbf{W}$로 쌓을 수 있다. 각 행에서 그 프레임의 점 좌표 중심을 빼 병진 성분을 제거한 뒤, orthographic(scaled orthographic) 카메라 모델 아래의 $\mathbf{W}$는 rank 3 이하가 된다. 원 논문(Tomasi & Kanade 1992)은 이 가정에서 출발했다. 그러면: $$\mathbf{W} = \mathbf{M} \mathbf{S}$$ 여기서 $\mathbf{M}$은 $2F \times 3$ 모션 행렬, $\mathbf{S}$는 $3 \times P$ 구조 행렬이다. SVD로 $\mathbf{W}$의 상위 3개 특이값만 유지하면 rank-3 factorization을 얻는다. 다만 $\mathbf{M}\mathbf{A}$와 $\mathbf{A}^{-1}\mathbf{S}$도 같은 $\mathbf{W}$를 만들므로, 카메라 행들의 직교·동일 노름 제약으로 가역행렬 $\mathbf{A}$를 정하는 metric upgrade가 뒤따른다. 이 절차는 한 번의 저랭크 SVD와 metric upgrade로 모든 프레임의 모션과 모든 포인트의 3D 위치를 함께 추정했다. 비용은 $2F \times P$ 행렬의 SVD에 지배되며, 반복 비선형 bundle adjustment보다 구조가 단순하고 구현하기 쉬웠다. > 🔗 **차용.** Nistér, Naroditsky, Bergen의 2004년 CVPR 논문 "Visual Odometry"는 실시간 에고모션 추정을 이 계보의 응용 문제로 돌려놓은 것으로 후속 문헌에 널리 인용된다. Tomasi-Kanade의 batch factorization을 그대로 쓰는 대신 짧은 윈도우 안에서 프레임 간 상대 포즈를 풀어나가는 쪽으로 방향이 옮겨갔고, 이는 batch 정확도 대신 latency를 택하는 흐름의 초기 지점으로 남았다. Orthographic/affine 가정이 한계였다. Affine 카메라는 원근 왜곡(perspective distortion)을 무시한다. 이 모델은 장면의 깊이 변화가 카메라까지의 거리에 비해 충분히 작을 때(예를 들어 원거리의 작은 물체를 촬영할 때)에만 유효하다. 카메라와 가까운 장면, 시야각이 넓은 렌즈, 혹은 전경·배경 깊이 차이가 큰 환경에서는 오차가 컸다. 1990년대 후반부터 perspective camera로의 확장이 여러 방향에서 시도되었고, 이는 bundle adjustment의 재발견으로 이어졌다. --- ## 3.3 Hartley & Zisserman과 정전(canon)화 Tomasi-Kanade의 factorization이 다중 시점 문제의 틀을 잡았다면, 남은 과제는 원근 카메라로의 확장과 흩어진 수학을 하나의 언어로 묶는 일이었다. 2000년 Richard Hartley와 Andrew Zisserman의 680쪽 교과서 *[Multiple View Geometry in Computer Vision](https://www.robots.ox.ac.uk/~vgg/hzbook/)*이 나왔다. 1981년부터 1990년대까지 여기저기 흩어진 SfM 수학을 사영기하(projective geometry)의 언어로 통합했다. Hartley & Zisserman은 essential matrix, fundamental matrix, homography, camera calibration, bundle adjustment를 사영기하의 단일 프레임워크로 묶었다. 각자 따로 다뤄지던 개념들의 공통 뿌리가 한 교과서 안에서 드러났다. Hartley & Zisserman은 bundle adjustment를 특히 무게 있게 다뤘다. Ch.1에서 Triggs et al.(1999)의 종합을 통해 살펴본 reprojection error 최소화 문제를 사영기하 프레임워크 안에 놓고 *robust cost function* $\rho$를 명시적으로 얹었다. outlier가 섞인 실제 데이터에서 최적화가 무너지지 않도록 Huber나 Cauchy 함수로 오차를 눌렀다. Levenberg-Marquardt로 풀되, Jacobian의 희소 구조를 써서 계산량을 줄였다. 2000년대 초반 SLAM·VO 논문 대부분이 이 교과서를 표준 참조로 달았다. 개념 정의가 이 책 하나로 통일되면서, Photo Tourism 같은 대규모 응용은 개념 재정의 없이 구현에 집중할 수 있었다. --- ## 3.4 Photo Tourism과 Bundler — 인터넷 규모 SfM 2006년 Noah Snavely, Steven Seitz, Richard Szeliski는 SIGGRAPH 논문 "[Photo Tourism](https://doi.org/10.1145/1179352.1141964)"을 발표했다. 인터넷에 업로드된 관광지 사진들(피렌체 두오모, 로마 트레비 분수)을 모아서 3D 재구성을 시도했다. 카메라도 날씨도 구도도 제각각이었고, 일부 사진은 관계없는 실내 컷이 섞여 있었다. 체계적으로 촬영한 데이터셋이 아니라, 수천 명이 아무 순서 없이 올린 이미지들이었다. Snavely의 파이프라인은 다음 순서로 작동했다. SIFT 특징 검출과 매칭으로 이미지 쌍 사이의 대응점을 찾는다. Fundamental matrix로 기하적으로 불일치하는 매칭을 RANSAC으로 제거한다. 연결성이 높은 이미지 쌍부터 시작해 카메라를 하나씩 추가하는 incremental SfM을 수행한다. 카메라를 추가할 때마다 bundle adjustment로 전체 포즈와 포인트를 재최적화한다. 논문이 보고한 데이터셋은 Notre Dame 대성당(2,635장 후보 중 597장 등록)·Trevi 분수(Rome, 466장 중 360장)·Yosemite Half Dome(1,882장 중 325장)·Great Wall(120장 중 82장)·Trafalgar Square(1,893장 중 278장) 등이었고, 평균 reprojection error는 1,611×1,128 픽셀 이미지에서 약 1.5 픽셀이었다. 이 결과는 통제되지 않은 인터넷 사진을 수백 장 규모의 일관된 재구성으로 묶는 전환점이 됐다. 이 파이프라인의 구현체가 Bundler였다. Snavely가 오픈소스로 풀었고, SfM 연구자들의 기본 출발점이 되었다. --- ## 3.5 COLMAP — 공학적 성숙 > 📜 **예언 vs 실제.** Snavely et al. 2006 "Discussion and future work" 섹션은 "Ultimately, we wish to scale up our reconstruction algorithm to handle millions of photographs"라고 명시하며, 더 나은 이미지 등록 순서, 렌즈 왜곡 모델링, 반복 구조 처리, 비연결 구조 재구성을 남은 과제로 꼽았다. 규모 확장은 COLMAP(Schönberger 2016)과 OpenSfM이 수만~수십만 장 규모로 이어받았고, 실시간·온라인 처리는 SfM이 아니라 SLAM 계보가 별도로 답했다(incremental refinement 대신 fixed-lag smoother와 loop closure로). 이는 규모 확장 방향의 진전이지만, 여기 든 규모만으로 수백만 장이라는 목표가 달성됐다고 할 수는 없다. 2016년 Johannes Schönberger와 Jan-Michael Frahm은 CVPR 논문 "[Structure-from-Motion Revisited](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf)"를 발표했다. Bundler 이후 십 년간 쌓인 개선들을 체계적으로 묶은 재설계였다. COLMAP이 Bundler와 가장 크게 달라진 점은 세 곳이다. 첫째, 카메라 추가 순서. Bundler는 연결성이 높은 쌍부터 시작했지만, 어떤 쌍을 먼저 확장할지에 대한 체계적 기준이 없었다. COLMAP은 초기 이미지 쌍 선택과 카메라 등록 순서를 triangulation angle, feature track 길이, visibility score 기반으로 자동화했다. 재구성의 안정성이 크게 높아졌다. 둘째, bundle adjustment 주기. 매 카메라 추가 후 full bundle adjustment는 비용이 크다. COLMAP은 local bundle adjustment(최근 추가된 카메라와 공유 포인트가 많은 카메라들만 묶어서 최적화)와 주기적 global bundle adjustment를 교대하는 방식을 도입했다. 셋째, 기하적 검증. 매칭된 특징점 쌍에 대해 fundamental matrix와 homography 두 모델로 각각 RANSAC을 돌린다. Fundamental matrix는 일반적인 비평면 장면, homography는 평면 장면이나 순수 회전을 모델링한다. COLMAP은 두 모델의 inlier 수를 비교해 장면 유형을 판별하고, 어느 쪽에도 들어오지 않는 매칭을 걸러낸다. 불량 매칭과 평면-퇴화(planar degeneracy) 상황에서 Bundler보다 버텼다. > 🔗 **차용.** COLMAP의 incremental bundle adjustment 전략은 Snavely의 Bundler 파이프라인을 모듈화하고 각 단계의 품질 관리를 추가한 것이다. 알고리즘의 핵심 수학(essential matrix 추정, triangulation, Levenberg-Marquardt)은 Hartley & Zisserman 교과서의 것이다. COLMAP은 엔지니어링 판단을 체계화했다. COLMAP이 널리 쓰인 것은 성능 때문만이 아니었다. 코드베이스가 정돈되어 있었고, 문서도 충분했으며, CUDA 가속으로 대규모 이미지 집합을 처리할 수 있었다. 2020년 NeRF가 나온 뒤 많은 NeRF 학습 코드가 COLMAP 출력(카메라 포즈 + sparse point cloud)을 입력으로 받았고, 3D Gaussian Splatting의 대표 구현도 같은 전처리를 썼다. COLMAP은 SfM 도구이자 여러 3D 재구성 파이프라인의 공통 입구가 되었다. --- ## 3.6 SfM과 SLAM의 분화 SfM과 SLAM은 같은 수학을 쓰면서도 처리 조건이 다른 문제를 푼다. 이 구분이 뚜렷해진 것은 2000년대 초였다. SfM은 *오프라인*이다. 모든 이미지를 수집한 뒤 처리하므로 시간 제약이 없고, 전체 데이터를 반복 참조하면서 global bundle adjustment를 여러 번 돌릴 수 있다. 카메라 포즈가 틀렸으면 되돌아가 다시 계산하면 된다. SLAM은 *온라인*이다. 센서 데이터가 실시간으로 유입되고, 현재 시점의 로봇 위치를 그 자리에서 내놓아야 한다. 과거 데이터를 무한정 참조할 수 없으며, 지도가 자라면서 계산량이 커지고, 루프를 완주해 처음 방문한 장소로 돌아왔을 때 accumulated drift를 교정해야 한다. 두 분야는 온라인 처리의 제약에서 차이가 난다. SLAM은 이동 중 재방문을 탐지하고 누적된 drift를 교정해야 하며, 그 교정은 루프가 연결하는 여러 포즈에 걸칠 수 있다. SfM과 SLAM은 bundle adjustment와 영상 검색 같은 도구를 공유하지만, 언제 어떤 상태를 갱신할지 정하는 실행 조건이 다르다. 불확실성 전파도 달랐다. SLAM은 현재 포즈의 불확실성을 실시간으로 추적하고 새 관측마다 갱신한다. EKF나 factor graph 형태의 probabilistic 표현이 필요하다. SfM에서는 최적화가 끝난 뒤 covariance를 사후에 계산하면 되고, 실시간 추적은 필수가 아니다. Davison의 [MonoSLAM(2003)](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf)은 스스로를 "real-time SfM"으로 불렀다. 그러나 EKF 상태벡터에 카메라 포즈와 landmark를 함께 유지하는 구조는 SfM의 global batch와 달랐다. 2000년대를 거치며 두 분야는 각자의 문제 설정을 가진 독립된 계보로 갈라졌다. --- ## 3.7 🧭 아직 열린 것 **동적 물체 포함 SfM.** COLMAP을 비롯한 주류 범용 SfM 파이프라인은 정적인 세계를 가정한다. 장면의 포인트가 움직이지 않는다는 전제로 bundle adjustment를 풀기 때문에, 자동차나 보행자가 많은 장면에서는 오염된 매칭이 최적화를 왜곡한다. RANSAC이 일부를 걸러내고 Dynamic SfM 연구는 segmentation이나 물체별 motion을 모델링하지만, 이 장에서 검토한 공개 구현 가운데 COLMAP과 같은 범용 도구로 정착한 사례는 확인되지 않았다. **SfM과 SLAM의 경계 흐려짐.** 2023년 [DUSt3R](https://arxiv.org/abs/2312.14132)(Wang et al.)는 사전 훈련된 네트워크 하나로 이미지 두 장을 받아 dense point map과 카메라 포즈를 동시에 냈다. 특징점 매칭도 RANSAC도 bundle adjustment 초기화도 거치지 않았다. [MASt3R](https://arxiv.org/abs/2406.09756)(2024)로 확장되면서 수십 장 재구성도 됐다. 전통적인 SfM 파이프라인의 각 모듈이 하나씩 대체되고 있다. COLMAP이 NeRF·3DGS의 입구였다면, DUSt3R 류는 그 입구마저 바꾸려 한다. 이 패러다임이 COLMAP을 실질적으로 밀어낼지, 특정 도메인에서만 이길지는 아직 모른다. --- SfM이 정밀한 오프라인 재구성을 다듬는 동안, 움직이는 로봇은 이미지가 모두 수집되기 전에 포즈 추정을 내놓아야 했다. Randall Smith와 Peter Cheeseman이 [1986년에 던진 질문](https://people.csail.mit.edu/brooks/idocs/Smith_Cheeseman.pdf)(불확실한 공간관계를 어떻게 전파하는가)이 그 압박 아래서 SLAM이라는 별개의 분야를 키웠다. --- # Ch.4 — Smith-Cheeseman과 EKF-SLAM의 흥망 1부에서 다룬 photogrammetry·SfM·bundle adjustment는 카메라가 정지해 있거나, 촬영 후 오프라인으로 모든 이미지를 한꺼번에 처리할 여유가 있다고 전제했다. Hartley-Zisserman의 기하학, RANSAC의 강건 추정, Levenberg-Marquardt의 반복 최적화는 세상을 측정하는 법을 알았지만, 움직이는 로봇이 *지금 이 순간* 어디에 있는지는 묻지 않았다. 지도를 만들면서 동시에 자신의 위치를 알고, 불확실성이 쌓이는 와중에도 추정을 포기하지 않는 문제는 SRI International의 작은 메모에서 열렸다. Randall Smith와 Peter Cheeseman은 1986년 로봇이 공간 속에서 무언가를 측정할 때 그 측정값이 얼마나 불확실한지를 수학적으로 다루려 했다. SRI International에서 나온 그들의 아이디어는 Kalman(1960)의 필터 수학을 이어받되, 단일 상태 추정이 아닌 *공간관계의 네트워크* 전체에 불확실성을 전파하는 방향으로 확장했다. 그로부터 수년 뒤 Sydney에서 Hugh Durrant-Whyte가, MIT에서 John Leonard가 이 수학에 "로봇이 지도를 만들면서 동시에 자신의 위치를 추정한다"는 문제 정식을 결합했다. "SLAM"이라는 약어는 그 접합의 산물이다. --- ## 4.1 불확실 공간관계의 수학 — Smith, Self, Cheeseman (1988) 1986년 SRI International의 Randall Smith, Matthew Self, Peter Cheeseman은 로봇이 여러 장소를 거쳐 측정값을 누적할 때 오차가 어떻게 전파되는지를 수식으로 잡으려 했다. 그 작업은 UAI 1986에서 발표되고 1988년 논문 ["Estimating Uncertain Spatial Relationships in Robotics"](https://arxiv.org/abs/1304.3111)으로 나왔다. 질문 자체는 명료했다. 로봇이 A에서 B를 측정하고 B에서 C를 측정했을 때, A에서 C까지의 불확실성은 어떻게 계산되는가? [Kalman 필터](https://www.cs.unc.edu/~welch/kalman/kalmanPaper.html)는 이미 있었다. 레이더 추적, 탄도 계산, 위성 궤도 보정에 1960년부터 쓰였다. Smith, Self, Cheeseman이 한 일은 Kalman의 공분산 전파 방정식을 공간 변환의 합성(composition)에 맞게 재공식화한 것이다. 로봇 pose $\mathbf{x}_r$과 landmark 위치 $\mathbf{m}_i$를 하나의 state vector에 담고, 그 전체의 joint covariance $\mathbf{P}$를 유지한다. $$\mathbf{x} = [\mathbf{x}_r^\top,\ \mathbf{m}_1^\top,\ \ldots,\ \mathbf{m}_N^\top]^\top$$ $$\mathbf{P} = \begin{bmatrix} \mathbf{P}_{rr} & \mathbf{P}_{rm} \\ \mathbf{P}_{mr} & \mathbf{P}_{mm} \end{bmatrix}$$ off-diagonal 블록 $\mathbf{P}_{rm}$은 로봇 위치 불확실성과 landmark 위치 불확실성의 상관관계를 담는다. 이 상관관계를 추적해야 일관된 추정이 가능하다. 논문은 이를 명시적으로 증명했고, SLAM 분야 전체가 이 출발점에 섰다. > 🔗 **차용.** Smith-Self-Cheeseman(1988)의 공간관계 수학은 Kalman(1960)의 공분산 전파를 직접 계승한다. 단일 이동 물체를 추적하던 기법이 로봇과 지도 요소 전체를 동시에 추적하는 틀로 바뀌었다. --- ## 4.2 "SLAM"이라는 이름의 정착 Smith-Self-Cheeseman의 1988년 논문에는 "SLAM"이라는 단어가 없다. Oxford에서 Sydney로 옮긴 Hugh Durrant-Whyte와 MIT의 John Leonard가 1990년대 초 각자의 연구실에서 같은 문제를 다른 이름으로 부르고 있었다. 두 그룹이 서로를 인용하기 시작하면서 공통 용어가 필요해졌고, "SLAM"은 그렇게 수렴해 굳었다. 정확히 어느 문서에서 처음 쓰였는지는 연구자마다 기억이 다르다. 공식 선점 논문은 없다. Leonard와 Durrant-Whyte의 1991년 논문 ["Simultaneous Map Building and Localization for an Autonomous Mobile Robot"](https://doi.org/10.1109/IROS.1991.174711)이 이 문제를 로봇공학 메인스트림에서 제목으로 명시한 초기 사례로 자주 인용된다. "Mapping"과 "Localization"이 분리 불가능하게 얽혀 있다는 것, 그것을 동시에(simultaneously) 해야 한다는 것, 이 직관이 약어 이전에 있었다. "Simultaneous Localization and Mapping", 줄여서 SLAM이다. 이후 10년간 이 이름은 분야 전체를 모으는 구심이 됐다. > 🔗 **병행 계보.** [Bar-Shalom의 다중 표적 추적](https://archive.org/details/trackingdataasso0000bars)(multi-target tracking, 1988년 단행본으로 정리됨)과 초기 SLAM은 여러 미지 상태의 불확실성과 data association을 함께 다룬다는 수학적 접점이 있다. 표적·추적기와 landmark·robot pose의 대응은 두 문제를 이해하기 위한 비유로는 유용하지만, 현재 확인되는 인용만으로 Leonard와 Durrant-Whyte가 이 치환을 직접 차용했다고 단정할 수는 없다. --- ## 4.3 EKF-SLAM의 공식 Extended Kalman Filter(EKF)가 SLAM에 적용된 것은 자연스러운 수렴이었다. 1988년 이전부터 비선형 시스템 추정에 사용되던 EKF는 predict-update 두 단계로 작동한다. predict 단계에서는 로봇이 움직이면 모션 모델 $f(\cdot)$로 state를 예측하고, Jacobian $\mathbf{F}$로 공분산을 전파한다. $$\hat{\mathbf{x}}^- = f(\hat{\mathbf{x}}, \mathbf{u})$$ $$\mathbf{P}^- = \mathbf{F}\mathbf{P}\mathbf{F}^\top + \mathbf{Q}$$ update 단계에서는 센서 측정값 $\mathbf{z}$가 오면 관측 모델 $h(\cdot)$의 Jacobian $\mathbf{H}$로 Kalman gain $\mathbf{K}$를 계산해 state와 공분산을 갱신한다. $$\mathbf{K} = \mathbf{P}^-\mathbf{H}^\top(\mathbf{H}\mathbf{P}^-\mathbf{H}^\top + \mathbf{R})^{-1}$$ $$\hat{\mathbf{x}} = \hat{\mathbf{x}}^- + \mathbf{K}(\mathbf{z} - h(\hat{\mathbf{x}}^-))$$ $$\mathbf{P} = (\mathbf{I} - \mathbf{K}\mathbf{H})\mathbf{P}^-$$ EKF-SLAM은 이 두 단계를 반복한다. 구조는 단순하지만 처음부터 확장성의 천장을 안고 있었다. 문제는 state 차원이다. 6DOF pose에 3D landmark $N$개를 담으면 state vector 차원은 $6 + 3N$이고, 공분산 행렬은 $(6+3N)^2$ 원소의 $O(N^2)$ 구조를 이룬다. 관측 차원이 고정된 한 번의 update에서도 전체 공분산 갱신에는 $O(N^2)$ 비용이 든다. Kalman gain은 $\mathbf{S} = \mathbf{H}\mathbf{P}^-\mathbf{H}^\top + \mathbf{R}$의 역행렬뿐 아니라 상태–관측 공분산과의 곱으로 계산하며, $\mathbf{S}$의 크기는 관측 차원에 달려 있다. landmark 100개면 $306 \times 306 \approx 9.4$만 원소, 1,000개면 $3006 \times 3006 \approx 900$만 원소다. 2000년대 초 일반 PC로 실시간을 유지할 수 있는 landmark 수는 수십에서 백 단위가 한계였다. [Andrew Davison의 MonoSLAM(2003)](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf)이 실시간 시연에서 landmark 수십 개 수준에 갇힌 것은 우연이 아니었다. EKF-SLAM의 $O(N^2)$ 벽이 그 숫자를 결정했다. --- ## 4.4 확장성의 벽 2003년 ICCV에서 Davison이 웹캠 하나로 실시간 3D 추적을 시연했을 때, 수십 개 수준의 feature로 책상 하나 크기의 공간을 매핑했다. 당시 상업용 SLAM 시스템이 없던 환경에서 실시간 단안 추적은 드문 시연이었다. 그 한계는 공분산 행렬의 크기에서 왔다. 100 landmarks에서 covariance 행렬은 $306 \times 306$ (6DOF pose + 3D landmark 100개 기준, state 차원 $6 + 3 \times 100 = 306$). 1,000개면 $3006 \times 3006$. 매 시간 단계마다 이것을 역행렬 연산과 함께 갱신해야 한다. 더불어 EKF는 joint 분포 전체를 한 덩어리로 유지하기 때문에, 새 landmark가 추가되면 기존 모든 landmark와의 cross-correlation이 즉시 생성된다. 지도가 커질수록 update 비용은 landmark 수의 제곱에 비례해 증가한다. 2000년대 중반까지 시도된 해법은 submap이었다. 전체 지도를 겹쳐지는 소영역으로 나누고, 각 submap에서만 EKF를 돌린 뒤 submap 사이를 별도 연결 구조로 잇는다. [Chong과 Kleeman(1999)](http://www.cs.cmu.edu/afs/cs/Web/People/motionplanning/papers/sbp_papers/integrated1/chong_feature_map.pdf)이 초기 형태를 제안했다. 그러나 submap 경계에서의 정보 손실과 루프 클로저의 어려움, 그리고 구현 복잡도가 submap 접근을 실용화하는 데 마찰을 일으켰다. > 🔗 **차용.** Chong-Kleeman(1999)의 submap 분할과 현대 SLAM의 local window 최적화는 계산 범위를 제한한다는 목적을 공유한다. ORB-SLAM의 local map과 VINS-Mono의 sliding window에서도 그 관심사를 볼 수 있다. 다만 지도 분할과 상태 주변화의 방식이 다르므로 같은 원리의 직접 계승으로 단정하지는 않는다. --- ## 4.5 Consistency 문제: Julier-Uhlmann의 반례 EKF-SLAM의 더 깊은 결함은 2001년 ICRA에서 터졌다. Simon Julier와 Jeffrey Uhlmann이 EKF 기반 SLAM의 거동을 수치 실험으로 분석하며 필터가 자기 자신을 너무 믿는다는 것을 보였다. 그들이 IEEE ICRA에 낸 논문 제목은 ["A Counter Example to the Theory of Simultaneous Localization and Map Building"](https://doi.org/10.1109/ROBOT.2001.933257)이었다. 도발적이었고, 내용도 그랬다. 2차 문헌들이 이 논문을 인용하며 요약하는 핵심은, EKF-SLAM이 asymptotically *overconfident*하다는 것이다. 즉, 실제 추정 오류는 커지는데 필터가 계산하는 공분산(불확실성)은 실제보다 작게 수렴한다. 이것이 inconsistency다. 원인은 linearization error에 있다. EKF는 비선형 모션 모델과 관측 모델을 일차 Taylor 전개로 근사한다. 이 근사 오류가 매 단계 누적되면 공분산이 실제 오류를 과소 평가하기 시작한다. 로봇이 "나는 여기 있다"고 과도하게 확신하면, 이후 측정값을 필터가 덜 신뢰하게 되어 오류가 교정되지 않고 쌓인다. 2007년 [Shoudong Huang과 Gamini Dissanayake](https://doi.org/10.1109/TRO.2007.903811)는 이 inconsistency의 원인을 더 정밀하게 해부했다. 논문의 핵심 진단은 두 가지였다. 현재 상태 추정치에서 평가된 Jacobian들 사이의 기본 제약(constraint)이 무너지는 것이 EKF-SLAM 비일관성의 주된 원인이고, 그 결과 로봇 방향각(yaw)의 분산이 실제로는 유지되어야 하는데도 잘못 0으로 수렴할 수 있다는 것이었다. 선형화 시점에 따라 시스템의 관측 가능한 자유도가 달라지고, 관측 불가능한 방향에 필터가 임의의 정보를 주입하게 된다는 이후 observability 기반 계열의 해석은 이 논문에서 출발한다. > 📜 **예언 vs 실제.** Julier와 Uhlmann의 2001년 반례 이후, consistent estimation을 달성하려는 필터 설계 시도가 이어졌다. Unscented Kalman Filter(UKF), Invariant EKF, robust covariance 등 필터 계열의 변형들이 10년 가까이 제안됐다. 그러나 2026년 시점에서 되돌아보면 최적화 기반 추정도 이 문제를 다루는 주요 경로가 되었다. [iSAM](https://www.cs.cmu.edu/~kaess/pub/Kaess08tro.pdf)(Kaess et al., 2008), [g2o](http://ais.informatik.uni-freiburg.de/publications/papers/kuemmerle11icra.pdf)(Kümmerle et al., 2011), GTSAM이 널리 쓰였고, 필터 기반 VIO도 함께 발전했다. 반복 최적화는 보존한 상태를 재선형화할 수 있지만, 그것만으로 consistency가 보장되지는 않는다. 관측 불가능한 방향과 주변화 과정의 선형화도 함께 관리해야 한다. --- ## 4.6 FastSLAM — 분할통치 EKF-SLAM의 $O(N^2)$ 벽을 다른 방식으로 공격한 것이 [FastSLAM](https://cdn.aaai.org/AAAI/2002/AAAI02-089.pdf)이다. Michael Montemerlo, Sebastian Thrun(Stanford), Daphne Koller, Ben Wegbreit가 2002년 AAAI에서 발표했다. FastSLAM은 Rao-Blackwellization을 이용했다. 로봇 경로 $x_{0:t}$가 주어지면 각 landmark의 위치 추정이 *서로 독립*이 된다. 따라서 경로를 particle filter로 표현하고(각 particle이 하나의 가능한 경로를 대표), 각 particle마다 별도의 landmark EKF를 독립적으로 운용할 수 있다. particle $K$개, landmark $N$개면 per-step 복잡도는 $O(K \log N)$으로, EKF-SLAM의 $O(N^2)$와 달리 $K$가 고정되어 있을 때 $N$에 대해 로그 규모로 증가한다(KD-tree 기반 landmark 탐색 사용 시). landmark 수가 많아져도 per-particle EKF는 서로 독립이라 $N \times N$ 전체 공분산을 유지할 필요가 없다. $K$는 수십~수백 수준으로 고정되므로 실질적 이득이 컸다. FastSLAM은 작동했다. 실내 환경에서 수백 개 landmark까지 실시간을 유지했고, 기술 이전도 빨랐다. 그러나 particle depletion 문제가 쌓였다. 지도가 커지면 대부분의 particle이 불량 경로를 대표하게 되고, effective sample 수가 급감한다. 루프 클로저 상황에서 경로 가중치 재조정이 어렵다. 무엇보다 particle 수를 늘려도 large-scale 환경에서 드리프트가 축적되는 문제는 해결되지 않았다. [FastSLAM 2.0](https://www.ijcai.org/Proceedings/03/Papers/165.pdf)(Montemerlo et al. 2003)이 proposal distribution을 개선했지만, 방법론이 필터 패러다임 안에 갇혀 있는 한 확장성의 천장이 있었다. 그 천장을 결국 피해 간 방법은 필터 계열이 아니었다. --- ## 4.7 EKF의 퇴장 그래프 기반 접근이 2005년 이후 빠르게 현실화되면서 EKF-SLAM은 주력에서 물러났다. [Feng Lu와 Evangelos Milios의 1997년 그래프 아이디어](https://doi.org/10.1023/A:1008854305733)가 [Olson-Leonard-Teller(2006)](https://april.eecs.umich.edu/pdfs/olson2006icra.pdf)의 efficient solver와, 이후 g2o·GTSAM·iSAM2의 실시간 인수분해 기법과 결합하자, EKF의 장점이었던 "incremental update"는 더 이상 차별점이 아니었다. 루프 클로저에서는 로봇이 출발점으로 돌아왔을 때 지도 오류를 수정한다. EKF는 이 순간 전체 공분산 행렬을 업데이트해야 하며, 비용은 $O(N^2)$다. 그래프 최적화는 pose 그래프에 새 엣지 하나를 추가하고 sparse 행렬을 재분해한다. Sparse 구조를 쓰면 비용이 훨씬 낮다. 2010년 무렵부터 새로운 SLAM 시스템에서 backend로 EKF를 선택하는 경우는 드물어졌다. 매우 제한된 연산 자원이나 실시간 필터 요구처럼 특수한 제약이 있는 경우에만 잔존했다. > 📜 **예언 vs 실제.** Durrant-Whyte와 Bailey의 [2006년 IEEE Robotics & Automation Magazine 튜토리얼](https://people.eecs.berkeley.edu/~pabbeel/cs287-fa09/readings/Durrant-Whyte_Bailey_SLAM-tutorial-I.pdf)은 SLAM의 확장성 문제를 논하며 submap 분해와 information filter가 대규모 환경에서의 해법이 될 것으로 전망했다. Information filter(EKF의 역공분산 형태)는 sparse information matrix를 이용해 landmark가 늘어도 연산이 느려지지 않을 것으로 기대됐다. 실제 전개는 달랐다. Information filter 계열(SEIF 등)은 sparsity를 강제로 유지하는 과정에서 marginalization error가 생겼다. Submap은 일부 시스템에 흡수되었으나 주류 해법이 되지 못했다. 2010년대를 지배한 것은 factor graph + iterative 최적화였다. --- ## 4.8 🧭 아직 열린 것 **Filter vs Optimization의 공존.** EKF가 backend 주력에서 물러났다고 해서 사라진 것은 아니다. 2026년 기준으로 자율주행 일부 구현은 여전히 필터 기반을 선호한다. 최적화 기반 SLAM은 반복 수렴이 필요하고, 실시간 보장이 어려운 경우가 있다. 저비용 임베디드 시스템에서 sparse EKF나 UKF가 재등장하는 사례가 있다. "필터는 죽었다"는 선언은 정확하지 않다. 용도와 제약에 따라 공존한다. **비가우시안 불확실성.** EKF의 가장 근본적인 가정은 불확실성이 가우시안 분포를 따른다는 것이다. 현실의 센서 오류는 다중 모드(멀티모달)이거나 heavy-tail 분포를 갖는 경우가 많다. 특히 여러 위치 가설이 생기는 perceptual aliasing(서로 다른 장소가 같아 보이는 것) 상황에서 단일 가우시안은 실제 불확실성을 단순화한다. Particle filter는 이론상 비가우시안을 표현하지만 고차원 state에서는 비실용적이다. Stein particle, normalizing flow, 학습 기반 uncertainty estimation이 시도되고 있으나, 2026년 기준으로 이것이 실시간 SLAM에서 검증된 형태는 제한적이다. --- EKF-SLAM이 100 landmark 수준의 실시간 천장에 부딪힐 무렵, Imperial College의 Andrew Davison은 별도 센서 없이 카메라 한 대로 실시간 단안 SLAM을 시연했다. 숫자의 한계는 그대로였지만 그것을 다루는 방식이 달라졌다. --- # Ch.5 — MonoSLAM → PTAM: 실시간의 몽상과 분리 혁명 EKF-SLAM은 상태와 불확실성을 함께 추정하는 지도 구축 방법을 제시했지만, 공분산 행렬이 landmark 수 $N$에 대해 $O(N^2)$로 커지는 구조적 벽에 부딪혔다. Davison과 Klein은 여기서부터 각자 다른 방향으로 걸었다. 2003년, Davison은 Imperial College 실험실에서 웹캠 한 대를 노트북에 꽂았다. 1988년 Smith와 Cheeseman이 세운 확률 공간관계 수학, 그 위에 Leonard와 Durrant-Whyte가 얹은 EKF-SLAM 틀을 그대로 가져왔지만 센서는 카메라 하나뿐이었다. IMU도 스테레오도 레이저도 없는 상태에서 Shi-Tomasi 1994 코너 검출기와 Kalman 예측-갱신 루프만 붙여 실시간으로 돌렸다. 당시 기준으로 무모한 조합이었다. 4년 뒤 2007년, Oxford의 Klein과 Murray가 다른 답을 냈다. tracking과 mapping을 두 스레드로 쪼갠 것이다. 그 분리가 이후 10년 Visual SLAM의 뼈대가 되었다. --- ## 1. 2003년의 데모 2003년 ICCV에서 Davison은 [Real-Time Simultaneous Localisation and Mapping with a Single Camera](https://doi.org/10.1109/ICCV.2003.1238654)를 공개했다. 단안 카메라만으로 실시간 SLAM을 시연한 보기 드문 결과였다. 당시 SLAM 분야의 주류는 레이저 센서였다. LiDAR는 2D 거리를 직접 제공했고, 스테레오 카메라는 픽셀 수준에서 깊이를 복원했다. 단안 카메라는 깊이 정보 자체가 없었다. 단안으로 3D 구조를 추정하려면 최소 두 프레임이 필요했고, 초기 깊이 추정의 불확실성이 EKF 상태 벡터 전체로 전파되었다. 이론적으로 가능했지만 실시간으로 돌린다는 것은 별개의 문제였다. MonoSLAM은 IMU나 스테레오 카메라 없이 단안 카메라만 썼다. 이 구성은 추가 하드웨어와 스테레오 캘리브레이션을 요구하지 않았다. 단안 실시간 SLAM의 가능성은 증명됐지만, EKF 구조는 센서와 지도를 확장하는 데 맞지 않았다. --- ## 2. EKF의 아름다움과 벽 2007년 IEEE PAMI에 실린 [MonoSLAM](https://doi.org/10.1109/TPAMI.2007.1049)는 Davison, Ian Reid, Nicholas Molton, Olivier Stasse가 공동 집필했으며, ICCV 2003 데모의 완성된 논문 형태였다. MonoSLAM의 상태 벡터는 [Smith-Cheeseman(1988)](https://arxiv.org/abs/1304.3111)과 [Leonard-Durrant-Whyte(1991)](https://ieeexplore.ieee.org/document/174711/)의 정식(Ch.4)을 단안 카메라에 직접 이식했다. 카메라 상태 $\mathbf{x}_v \in \mathbb{R}^{13}$(위치 3, 사원수 방향 4, 속도 3, 각속도 3)과 landmark 집합 $\mathbf{y}_i \in \mathbb{R}^3$를 하나의 벡터 $\mathbf{x} = (\mathbf{x}_v^\top, \mathbf{y}_1^\top, \ldots, \mathbf{y}_N^\top)^\top \in \mathbb{R}^{13+3N}$에 담고, 그 전체 공분산 $(13+3N)\times(13+3N)$ 행렬 $\mathbf{P}$를 매 프레임 predict-update 루프로 유지했다. predict 단계에서는 카메라 운동 모델 $f$의 자코비안 $\mathbf{F}$로 공분산을 전파했고($\mathbf{P}^- = \mathbf{F}\mathbf{P}\mathbf{F}^\top + \mathbf{Q}$), update 단계에서는 투영 함수의 자코비안 $\mathbf{H}_i$로 칼만 이득을 계산해 상태와 공분산을 갱신했다. EKF predict-update 수식 자체는 Ch.4 §4.3의 것과 동일하다. 달라진 것은 상태 벡터 안에 카메라 속도·각속도가 함께 들어간 점이었다(이동 물체인 카메라의 동역학 모델이 필요했기 때문이다). 공분산 갱신 $(\mathbf{I} - \mathbf{K}_i\mathbf{H}_i)\mathbf{P}^-$의 지배 비용은 $(13+3N)^2$ 행렬 곱셈으로, landmark 수 $N$에 대해 $O(N^2)$였다. 논문 §III은 30 Hz 실시간 처리에서 유지 가능한 feature 수의 상한이 "약 100개" 수준이라고 명시한다. > 🔗 **차용.** MonoSLAM의 EKF 상태 벡터 구조는 Smith-Cheeseman-Durrant-Whyte(1988-1991)의 확률적 공간관계 표현을 단안 카메라에 직접 이식한 것이다. Kalman 필터 자체는 1960년부터 있었지만, 로봇 pose와 landmark를 같은 벡터에 넣는 "augmented state vector" 관행이 확립된 것은 Leonard-Durrant-Whyte 1991의 스타일이었다. 이 숫자는 시스템 한계를 드러냈다. Davison은 논문에서 sub-mapping 전략으로의 확장을 향후 방향으로 제시했다. 그러나 EKF 내부에서 계층적 구조를 만드는 것은 근본적으로 어려웠다. 공분산 행렬이 모든 landmark 간 상관관계를 빠짐없이 담고 있었기 때문이다. [Shi-Tomasi(1994)](https://doi.org/10.1109/CVPR.1994.323794) 코너가 MonoSLAM의 시각 특징으로 선택된 것도 이 맥락에서 읽힌다. "Good Features to Track"의 선택 기준은 추적하기 좋은 점을 고르는 것이었다. 애초에 추적이 실패할 가능성이 낮은 코너만 상태 벡터에 넣으면 EKF의 갱신이 더 안정적이었다. PAMI 논문은 광각 렌즈에서 매 프레임 약 12개의 특징이 안정적으로 보이도록 map management를 구성한다고 명시한다. 이 한정된 수의 특징이 모두 잘 추적되는 한, EKF는 돌아갔다. > 🔗 **차용.** PTAM이 아니라 MonoSLAM에서 이미 Shi-Tomasi 1994의 코너 검출기가 쓰였다. "좋은 특징을 선택해서 추적한다"는 설계 철학은 Shi-Tomasi → MonoSLAM → PTAM의 직접적인 계보다. --- ## 3. 2007년, 같은 해 그 한정된 숫자가 EKF의 천장을 드러냈다. Oxford에서 그 천장을 올려다보던 사람이 Klein이었다. 2007년 ISMAR에 Klein과 Murray가 [Parallel Tracking and Mapping for Small AR Workspaces](https://doi.org/10.1109/ISMAR.2007.4538852)를 올렸다. 같은 해 PAMI에는 Davison의 MonoSLAM 정식판이 실렸고, 두 논문의 연구 계보도 맞닿아 있었다. Klein은 당시 Murray 그룹 박사과정이었다. Murray 그룹은 Oxford Active Vision Laboratory의 직계였고, 몇 년 전까지 Davison이 박사과정 학생으로 있던 바로 그 방이었다. Murray는 Davison의 지도교수였다. 같은 연구실 계보에서 나온 PTAM은 MonoSLAM이 보여 준 단안 실시간 가능성을 이어받되, EKF 대신 다른 구조를 택했다. 단안 실시간의 가능성은 확인됐다. PTAM은 확장을 위해 EKF 대신 tracking과 mapping을 분리한 구조를 택했다. --- ## 4. 분리 PTAM은 Tracking(카메라 pose 추적)과 Mapping(3D 지도 구축)을 분리해 두 개의 병렬 스레드에서 실행했다. EKF에서 이 둘은 같은 루프 안에 섞여 있었다. 매 프레임마다 예측-갱신 한 사이클을 돌리면서, 카메라가 움직이면 상태를 예측하고 이미지에서 landmark를 찾으면 다시 갱신했다. PTAM은 Tracking과 Mapping을 나눴다. Tracking 스레드는 매 프레임 카메라 pose를 추정하는 일만 한다. 현재 keyframe 집합에서 보이는 3D 점들의 2D 투영과 실제 관측을 매칭해서 pose를 실시간으로 계산한다. Mapping 스레드는 새 keyframe이 추가될 때마다 bundle adjustment를 실행한다. Tracking 스레드가 독립적으로 돌아가기 때문에 Mapping이 느려져도 무방했다. Mapping 스레드의 bundle adjustment는 keyframe 집합 $\mathcal{K}$와 3D 점 집합 $\mathcal{P}$에 대해 재투영 오차의 합을 최소화했다: $$\min_{\{\mathbf{T}_k\}, \{\mathbf{p}_j\}} \sum_{k \in \mathcal{K}} \sum_{j \in \mathcal{P}_k} \rho\!\left(\left\|\mathbf{z}_{kj} - \pi(\mathbf{T}_k,\, \mathbf{p}_j)\right\|^2_{\mathbf{\Sigma}_{kj}}\right)$$ 여기서 $\mathbf{T}_k \in SE(3)$는 keyframe $k$의 pose, $\mathbf{p}_j \in \mathbb{R}^3$는 3D 점, $\pi$는 카메라 투영 함수, $\mathbf{z}_{kj}$는 keyframe $k$에서 점 $j$의 관측 픽셀 좌표, $\mathbf{\Sigma}_{kj}$는 측정 공분산, $\rho$는 Huber 함수 등의 robust kernel이다. Mapping 스레드는 이 최적화를 Levenberg–Marquardt로 반복해서 풀었다. 비동기로 돌기 때문에 Tracking 스레드의 실시간성에 영향을 주지 않았다. > 🔗 **차용.** PTAM의 Mapping 스레드에서 실행되는 bundle adjustment는 [Triggs et al. 1999 "Bundle Adjustment — A Modern Synthesis"](https://doi.org/10.1007/3-540-44480-7_21)의 직접 적용이다. 1부에서 다룬 사진측량의 100년 전통이 실시간 SLAM backend의 중심으로 들어온 지점이다. EKF-SLAM에서는 공분산 행렬의 크기 때문에 대규모 joint update가 부담이었고, 스레드 분리는 keyframe BA를 tracking과 비동기로 실행하게 했다. Mapping 스레드가 비동기로 bundle adjustment를 실행하면서, 지도에 들어갈 수 있는 landmark 수가 EKF의 $O(N^2)$ 제약을 벗어났다. PTAM이 사용한 keyframe의 수는 수백 개였다. 각 keyframe에는 수백 개의 patch feature가 있었다. MonoSLAM의 수십 landmark 규모와는 다른 세계였다. 초기 맵 구축 방법도 달랐다. PTAM은 사용자가 카메라를 천천히 움직이는 초기화 단계에서 [Nistér 2004](https://doi.org/10.1109/TPAMI.2004.17)의 5-point 알고리즘 계열(PTAM 논문은 그 후속인 Stewénius·Engels·Nistér 2006을 인용)로 essential matrix를 추정하고, 첫 keyframe 쌍에서 초기 3D 구조를 복원했다. Essential matrix $\mathbf{E}$는 두 카메라 좌표계 사이의 순수 기하 관계를 담는 $3\times 3$ 행렬로, 대응점 쌍 $(\mathbf{p}, \mathbf{p}')$에 대해 ${\mathbf{p}'}^\top \mathbf{E}\, \mathbf{p} = 0$을 만족한다. $\mathbf{E}$는 내부적으로 $\mathbf{E} = \mathbf{t}_\times \mathbf{R}$ ($\mathbf{t}_\times$는 병진의 반대칭 행렬, $\mathbf{R}$은 회전)으로 분해되므로 자유도가 5이다. 최소 5쌍의 대응점으로 문제를 정할 수 있지만 해가 유일한 것은 아니며, 일반적으로 복소수 범위에서 최대 10개의 후보가 나온다. Nistér의 기여는 이 5-point 연립방정식을 효율적으로 풀어 RANSAC 루프 안에서 실시간으로 돌릴 수 있게 한 것이다. PTAM은 이 solver를 초기화 단계에서 RANSAC과 함께 사용해 첫 두 keyframe 사이의 상대 pose를 추정하고 초기 3D 점군을 삼각측량으로 복원했다. > 🔗 **차용.** PTAM의 5-point essential matrix 초기화는 David Nistér 2004 "An Efficient Solution to the Five-Point Relative Pose Problem"이 열어 놓은 minimal-solver 계보를 따른다(PTAM 논문은 그 후속 Stewénius·Engels·Nistér 2006 ISPRS를 직접 인용). 5-point solver는 단안 카메라의 초기 맵 구축에 필요한 최소 대응쌍을 사용하는 minimal solver였고, PTAM은 이 솔버를 RANSAC 루프에 태워 초기 두 keyframe의 상대 pose를 실시간에 가깝게 추정했다. > 🔀 **구조적 평행.** Leonard-Durrant-Whyte의 submap과 PTAM의 keyframe map은 모두 계산을 관리하기 위해 전역 지도를 선택된 지역 단위로 다룬다. 그러나 PTAM 원 논문에서 submap 계열을 직접 설계 근거로 삼았다는 인용 관계는 확인되지 않는다. PTAM의 keyframe 선택은 병렬 tracking·mapping과 local bundle adjustment의 필요에서 설명하는 편이 정확하다. 후속 ORB-SLAM은 여기에 covisibility graph를 더했다. --- ## 5. 새 아키텍처의 확산 PTAM은 AR(증강현실) 워크스페이스를 대상으로 설계되었다. 논문 제목에도 "Small AR Workspaces"가 명시되어 있다. Tracking 스레드의 재현성이 좋았고, 실시간성이 확실했기 때문에 AR 응용에 바로 쓸 수 있었다. 2010년대 초 Metaio(독일 AR 스타트업, 2015년 Apple에 인수)와 Qualcomm의 Vuforia SDK는 PTAM과 유사한 tracking/mapping 분리 구조를 채용했다. 이런 상용 SDK는 안정적인 planar AR을 소비자 스마트폰으로 확산시켰다. 2015년 Raul Mur-Artal, J.M.M. Montiel, Juan D. Tardós가 발표한 [ORB-SLAM](https://arxiv.org/abs/1502.00956)은 PTAM의 구조를 계승했다. 특징점은 patch에서 ORB 디스크립터로 바꾸고, keyframe 관리는 covisibility graph로 정교화했으며, loop closure를 새로 얹었다. 2018년 Qin, Li, Shen의 [VINS-Mono](https://arxiv.org/abs/1708.03852) 역시 sliding window 최적화 + loop closure의 이중 스레드 구조를 갖는다. tracking/mapping 분리의 계보가 VIO로 확장된 사례다. --- ## 6. Davison vs Klein & Murray — 관점 비교 2007년에 두 논문이 나왔다. MonoSLAM PAMI는 2003년 데모의 완성판이었다. PTAM은 같은 해에 MonoSLAM의 한계를 돌파하는 새 구조로 나왔다. MonoSLAM이 EKF를 붙들고 있었던 이유 중 하나는 상태와 불확실성을 함께 추적하는 확률적 표현에 있었다. EKF는 상태의 불확실성을 공분산 행렬로 명시적으로 관리했다. 지도의 각 landmark가 얼마나 불확실한지, landmark 간 공분산이 어떻게 연결되는지를 수학이 추적했다. 이 관점에서 bundle adjustment는 최소자승 최적화였고, 불확실성 표현을 줄이는 대신 확장성을 얻는 거래로 읽혔다. PTAM은 그 대가를 감수하는 설계를 택했다. AR 응용에서 중요한 것은 카메라 pose의 실시간 추적이었다. 지도의 불확실성을 센티미터 단위로 추적할 필요는 없었다. Bundle adjustment로 지도를 주기적으로 refine하면 충분했다. 이후 SLAM 분야의 방향은 이 거래 쪽으로 기울었다. 2010년대 이후 graph-based 최적화와 bundle adjustment가 주류가 되었고, EKF-SLAM은 계산 자원이 극도로 제한된 응용 외에서는 대부분 전면에서 물러났다. 다만 MonoSLAM이 붙들었던 확률론적 관심사가 사라진 것은 아니었다. Davison의 lab은 PTAM 계보로 건너뛰는 대신 이후 몇 단계에 걸쳐 factor graph 기반 추정, 그리고 Gaussian Belief Propagation(GBP)·Robot Web 쪽으로 옮겨갔다. 23년 뒤 SLAM Handbook Ch.18에서 Davison은 같은 흐름을 EKF→BA→factor graph→GBP로 이어지는 representation 변경의 연속으로 해석한다. 본인이 MonoSLAM을 직접 호명해 평가하는 대목은 없고, 각 표현 교체가 시스템 전체의 재설계를 유발한다는 일반 원리로 치환해 서술한다. --- ## 📜 예언 vs 실제 > **Davison 2007 PAMI MonoSLAM**: Davison은 Conclusion에서 더 큰 실내·실외 환경, 더 빠른 움직임, 가림·조명 변화가 있는 복잡한 장면을 다음 과제로 꼽았다. 구체 수단으로 sub-map 전략과 100 Hz 이상의 고프레임률 CMOS 카메라를 거론했고, sparse map을 "higher-order entities"(표면 등)의 dense 표현으로 확장할 여지도 함께 언급했다. > > 이 예측들의 운명은 각기 달랐다. Sub-map과 PTAM의 keyframe 구조, ORB-SLAM의 covisibility graph는 계산 범위를 나눈다는 관심사를 공유하지만, 이 비교만으로 직접 계승을 확인할 수는 없다. 그러나 EKF를 유지하면서 계층적 확장을 달성한 시스템은 나오지 않았다. 계층화는 BA 기반 아키텍처 전환과 함께 왔다. 고프레임률 카메라는 2010년대 이벤트 카메라 연구에서 다른 경로로 구체화되었다. 동적 장면 강건성은 2026년 기준 여전히 열려 있다. DynaSLAM, FlowSLAM 등 여러 시도가 있었지만 "기본 파이프라인에 포함된 해법"은 아직 없다. IMU 통합은 Davison이 Future Work에서 직접 지목하진 않았지만(관련 연구는 논문 본문에서 참조) 2010년대 Visual-Inertial Odometry(VIO) 연구 붐이 맡은 방향이다. 확률론적 일관성이라는 관심사 자체는 폐기되지 않고 factor graph·GBP 쪽으로 옮겨갔다. Davison은 23년 뒤 Handbook Ch.18에서 표현 변화에 따른 시스템 재설계를 설명한다. MonoSLAM을 그 계보의 한 단계로 놓는 것은 이 역사서의 해석이다. > **Klein & Murray 2007 PTAM**: Klein과 Murray는 §8(Failure modes / Mapping inadequacies)에서 시스템의 한계로 corner-기반 추적의 모션 블러 취약성, point cloud 중심의 지도가 가진 기하 이해 부족, 그리고 "not designed to close large loops in the SLAM sense"를 열거했다. 즉, 큰 루프의 전역 일관성 확보가 PTAM의 설계 범위 밖임을 분명히 했다. > > 2015년 ORB-SLAM은 이 한계들을 정면으로 겨냥했다. [DBoW2](http://doriangalvez.com/papers/GalvezTRO12.pdf) 기반 appearance loop closure와 covisibility graph 기반 keyframe 관리가 얹혔고, 특징은 patch 대신 ORB descriptor로 교체됐다. PTAM이 "우리 문제가 아니다"라고 선을 그은 곳에서 ORB-SLAM이 지도 확장을 이어받은 구도다. Klein & Murray 자신이 명시적으로 "appearance-based loop closure가 답"이라고 적은 것은 아니지만, 한계 지점의 지적이 후속 계보의 출발점으로 정확히 맞았다. --- ## 🧭 아직 열린 것 **Monocular scale 복원.** MonoSLAM부터 PTAM까지, 단안 카메라 시스템은 모두 scale ambiguity를 안고 있다. 이미지 한 장에서 절대 거리를 알 수 없다는 것은 기하학적 사실이다. IMU를 추가하면 중력 방향과 가속도계 판독값으로 scale이 observability를 갖는다. 그러나 IMU 없는 순수 monocular 시스템에서 scale 복원은 2026년에도 근본적으로 해결되지 않았다. 학습 기반 monocular depth estimation([MiDaS](https://arxiv.org/abs/1907.01341), [Depth Anything](https://arxiv.org/abs/2401.10891))이 단일 이미지에서 상대적 깊이를 추정하지만, 이것을 metric scale로 변환하려면 여전히 외부 참조(지면 가정, 사전 알려진 물체 크기 등)가 필요하다. **단일 VO 시스템의 환경 범용성.** MonoSLAM은 책상 주변의 작은 실내 환경만 다루었다. PTAM은 "Small AR Workspaces"라고 스스로 범위를 제한했다. 이후 ORB-SLAM2가 실내·실외·RGB-D를 아우르려 했지만, 조명 변화가 극단적인 환경이나 low-texture 공간에서는 여전히 tracking failure가 발생한다. 단일 파이프라인이 실내 복도, 야외 도심, 야간 환경, 텍스처 없는 흰 벽 전부를 견고하게 처리하는 시스템은 2026년 기준 아직 없다. Multi-modal fusion(카메라 + LiDAR + IMU)이 일부 커버하지만, 카메라 단독 시스템의 범용성은 여전히 미결이다. **저조도·동적 환경에서의 특징점 추적.** MonoSLAM이 요구했던 것은 충분한 조명과 정적인 장면이었다. 2007년의 PTAM도 마찬가지였다. 2026년 현재 이 두 가정은 여전히 대부분의 feature-based SLAM 시스템에서 암묵적으로 유지된다. 저조도에서 ORB feature는 검출 자체가 실패하고, 움직이는 사람이 많은 장면에서는 dynamic point가 static point로 잘못 분류된다. 이 문제를 학습 기반 optical flow나 semantic segmentation으로 우회하는 시도가 있지만, 실시간 범용 해법으로 자리잡은 시스템은 아직 없다. --- PTAM이 확립한 tracking/mapping 분리는 한 가지를 해결하지 못했다. keyframe이 쌓일수록 누적 오차가 loop에서 폭발했다. 그 답은 PTAM과 같은 해에 나온 것이 아니었다. 1997년, CMU 지하 복도에서 Feng Lu와 Evangelos Milios가 레이저 스캔 문제를 붙들고 있던 바로 그 시점에 이미 형태를 갖추고 있었다. --- # Ch.6 — Graph SLAM 혁명 1997년 Feng Lu와 Evangelos Milios는 레이저 스캔을 하나씩 누적 지도에 붙이는 방식이 등록 오차 때문에 일관되지 않은 지도를 만들 수 있다고 지적했다. 대신 각 스캔의 local frame과 frame 사이의 상대 공간 관계를 모두 유지하고, 그 제약을 동시에 풀어 전체 pose를 맞췄다. 다만 Lu-Milios가 이 방향의 유일한 시조는 아니다. 그보다 10여 년 앞서 LAAS의 [Chatila와 Laumond(1985)](https://www.semanticscholar.org/paper/Position-referencing-and-consistent-world-modeling-Chatila-Laumond/c34a678e40a7d80cb3683f07fc837179fd9bf3ee)가 이동 로봇의 참조 좌표계와 일관된 월드 모델을 논의했고, 1999년 [Gutmann과 Konolige](https://www.semanticscholar.org/paper/Incremental-mapping-of-large-cyclic-environments-Gutmann-Konolige/3c1bda51b8ca59f1836ed1b96c485d905804989a)가 대형 순환 환경의 증분 지도 작성에 포즈 정합을 적용했으며, 2000년대 초 Thrun 그룹이 *full SLAM* 문제로 이 접근을 정식화했다. [Folkesson과 Christensen(2004)](http://www.hichristensen.net/hic-papers/folkesson-icra2004.pdf), Konolige, Dellaert도 뒤이어 각자의 정식화를 내놓았다. Lu-Milios 1997의 분명한 기여는 "레이저 스캔 정합 + 배치 최대우도 추정"이라는 구체적인 파이프라인을 완결된 형태로 제시한 데 있다. Smith-Cheeseman이 확률 지도의 수학적 토대를 놓고 Davison이 실시간 단안 SLAM의 가능성을 보인 사이, 이 병렬 기여자들은 SLAM을 전체 궤적의 동시 추정 문제로 바꾸고 있었다. 2000년대의 EKF-SLAM은 landmark 수가 늘수록 $O(N^2)$ 공분산 갱신에 막혔고, Klein과 Murray의 PTAM(2007)은 별도의 BA 기반 keyframe 구조로 tracking과 mapping을 나눠 실시간 최적화의 가능성을 보였다. 필터와 나란히 발전한 graph-smoothing 해법은 여러 연구 집단의 작업을 거쳐 SLAM backend의 한 축이 되었다. --- ## 6.1 레이저 스캔에서 포즈 그래프로: Lu-Milios 1997 [Lu & Milios 1997. "Globally Consistent Range Scan Alignment"](https://doi.org/10.1023/A:1008854305733)이 등장하기 전까지, 연속 레이저 스캔의 정합(alignment)은 ICP(Iterative Closest Point) 계열의 국소 정합으로 이어 붙이는 경우가 많았다. ICP는 두 스캔을 국소적으로 잘 맞추지만, 드리프트가 누적되면 수십 미터 이후 지도가 뒤틀린다. 루프를 다시 돌아왔을 때 출발점과 지도가 맞지 않는다. Lu와 Milios의 아이디어는 단순했다. 로봇의 포즈 시퀀스 $x_1, x_2, \ldots, x_n$을 노드로, 각 포즈 쌍 사이의 상대 측정값을 엣지로 표현하면, 지도 구성 문제는 그래프 위의 에너지 최소화 문제가 된다. 엣지 하나하나는 두 포즈 사이의 상대변환 $\hat{z}_{ij}$와 그 불확실성 $\Omega_{ij}$를 담는다. 전체 비용 함수는 $$F = \sum_{(i,j) \in \mathcal{E}} e_{ij}^T \Omega_{ij} e_{ij}, \quad e_{ij} = z_{ij} - h(x_i, x_j)$$ 여기서 $h(x_i, x_j)$는 두 포즈로부터 기대 상대변환을 계산하는 함수이며, $z_{ij}$는 실제 측정된 상대변환, $\Omega_{ij} = \Sigma_{ij}^{-1}$는 측정 불확실성의 역행렬인 정보 행렬이다. 루프 클로저도 이 공식에 자연스럽게 포함된다. 나중에 같은 장소를 다시 방문했을 때 얻은 상대 측정값을 그래프에 엣지로 추가하면, 전체 최적화가 그 제약을 반영하여 모든 포즈를 조정한다. EKF에서 루프 클로저는 covariance를 $O(N^2)$ 단위로 갱신하는 무거운 작업이었다. 포즈 그래프에서는 엣지 하나로 제약을 표현하지만, 이를 추정에 반영하려면 그래프를 다시 최적화해야 한다. > 🔗 **차용.** Lu-Milios의 포즈 그래프 최적화 정식화는 [Levenberg(1944)](https://www.ams.org/qam/1944-02-02/S0033-569X-1944-10666-0/)와 [Marquardt(1963)](https://www.stat.cmu.edu/technometrics/70-79/VOL-14-03/v1403757.pdf)의 비선형 최소자승 알고리즘을 기반으로 한다. 수십 년 앞서 비선형 파라미터 추정을 위해 개발된 수치 최적화 기법이 실내 레이저 맵핑의 백엔드에 도착했다. 당시 Lu-Milios의 해법은 모든 포즈를 동시에 푸는 배치(batch) 선형 시스템이었다. 스캔 수가 늘어나면 선형 시스템의 크기도 함께 커진다. 그래서 개념 증명의 성격이 강했다. 그러나 전역 일관성을 달성할 수 있으며, 그 도구가 필터가 아닌 최적화임을 보여주었다. 같은 시기 Gutmann-Konolige는 증분성에, Folkesson-Christensen은 데이터 연관 강건성에, Thrun 그룹은 대규모 실환경 적용에 방점을 찍으며 같은 결론을 서로 다른 문제에서 구체화했다. --- ## 6.2 희소성의 발견: 정보 행렬과 포즈 그래프의 확장 Lu-Milios의 아이디어가 발표된 후 5년간, 여러 그룹이 같은 방향에서 확장을 시도했다. 공통된 발견은 정보 행렬(information matrix, $\Omega = \Sigma^{-1}$)의 **희소성(sparsity)**이었다. EKF-SLAM의 covariance 행렬 $\Sigma$는 조밀(dense)하다. 로봇이 새 landmark를 관측할 때마다 기존 모든 landmark와의 상관관계가 갱신된다. 로봇 포즈를 marginalize한 상태에서 $n$개의 2D landmark가 있으면 $\Sigma$는 $2n \times 2n$ 행렬이고, 갱신 비용은 $O(n^2)$다. 100개 landmark 정도에서 실시간성이 무너지는 이유다. 반면 포즈 그래프의 정보 행렬은 다르다. 로봇의 포즈 $x_i$와 $x_j$가 직접 측정 관계에 있을 때만 $\Omega$의 $(i,j)$ 블록에 비영(non-zero) 항이 생긴다. 연속 이동 시 인근 포즈들만 엣지로 연결되고, 먼 포즈들은 직접 연결되지 않는다. $\Omega$는 그래프 토폴로지를 반영한 띠형(banded) 희소 구조를 가진다. 루프 클로저가 없는 순수 주행 시나리오에서 이 구조는 정확히 tridiagonal에 가깝다. Sebastian Thrun 그룹의 [Sparse Extended Information Filter(SEIF)](http://www.cs.cmu.edu/~thrun/papers/thrun.tr-seif02.pdf), Edwin Olson의 연구는 이 희소성을 명시적으로 활용하기 시작했다. 희소 선형 대수 풀이기(sparse solver)를 쓰면 계산 비용이 $O(n^2)$에서 크게 줄어들 수 있었다. 실제 복잡도는 그래프 구조에 의존하지만, 로봇이 제한된 지역 내에서 움직이는 현실 시나리오에서는 $O(n \log n)$ 수준이 가능했다. > 🔗 **차용.** Thrun 그룹의 sparse information filter(SEIF)와 [Eustice의 exactly sparse delayed-state filter](https://web.mit.edu/2.166/www/handouts/eustice_et_al_ieeetro_2006.pdf)는 정보 행렬의 희소성이 필터 기반에서도 활용 가능하다는 것을 보였다. 이 희소성 통찰은 Dellaert의 factor graph 공식화와 Bayes tree 자료구조로 이어지는 맥락을 형성한다. 2006년 ICRA에서 [Olson, Leonard, Teller](https://april.eecs.umich.edu/pdfs/olson2006icra.pdf)는 stochastic gradient descent로 포즈 그래프를 최적화하는 방법을 발표했다. 수렴 보장은 없었다. 그래도 수백 노드 규모에서 충분히 빠르게 돌았고, Olson의 구현 코드는 이후 커뮤니티 전반에 퍼졌다. --- ## 6.3 Factor Graph와 Square Root SAM 2006년 Dellaert와 당시 박사과정이던 Kaess가 발표한 [Square Root SAM](https://doi.org/10.1177/0278364906072768)은 SLAM 백엔드를 factor graph로 정식화했다. Dellaert는 Georgia Tech에서 확률론적 그래픽 모델(probabilistic graphical model)을 연구해 왔다. Square Root SAM은 SLAM을 베이지안 추론 문제로 표현하고 factor graph 위에서 그 추론을 수행했다. **Factor graph**(변수 노드와 factor 노드를 엣지로 연결한 이분 그래프)에서 변수 노드는 로봇 포즈와 landmark의 위치, factor 노드는 관측값 또는 사전 확률(prior)이다. Factor $f_k(x_{i_1}, x_{i_2}, \ldots)$는 연결된 변수들 사이의 확률적 제약을 나타낸다. 전체 결합 확률은 $$p(X) \propto \prod_k f_k(X_k)$$ 이며, MAP 추정은 이 확률을 최대화하는 $X^*$를 찾는 것이다. Gaussian factor 하에서 이것은 비선형 최소자승 문제가 된다. Jacobian 행렬 $J$에 QR 분해를 적용하면 상삼각(upper triangular) 행렬 $R$이 남는다. $R^T R = J^T J = \Omega$이며, $R$이 바로 "square root information matrix"다. 이 $R$의 희소 구조는 Jacobian 자체가 아니라 변수 제거(variable elimination) 순서와 factor graph 토폴로지가 결정한다. 적절한 ordering(예: AMD, COLAMD)을 선택하면 fill-in을 최소화하여 희소한 $R$을 얻을 수 있다. 이 공식화는 EKF의 covariance 갱신보다 수치적으로 안정하다. 지도 전체의 랜드마크와 포즈를 일관된 방식으로 함께 최적화할 수 있으며, 루프 클로저는 새 factor를 추가하는 것으로 표현된다. --- ## 6.4 iSAM과 iSAM2: 온라인 증분 추론 Square Root SAM은 배치(batch) 방법이었다. 새 관측이 들어올 때마다 전체 $J^T J$를 다시 분해해야 한다. 밀집 행렬의 비용은 $O(n^3)$이며, 희소 행렬에서는 연결 구조와 소거 순서에 따라 달라진다. 온라인 로봇 시스템에서는 실용적이지 않았다. 2008년 [Kaess, Ranganathan, Dellaert가 발표한 **iSAM**(incremental Smoothing and Mapping)](https://www.cs.cmu.edu/~kaess/pub/Kaess08tro.pdf)은 이 문제를 Givens rotation으로 접근했다. 새 변수와 factor가 추가될 때, 기존 QR 분해를 처음부터 다시 수행하는 대신 새 행만 추가하여 Givens rotation으로 $R$을 갱신한다. iSAM1의 본질적 한계는 재선형화 스케줄이었다. 비선형 factor를 선형화한 결과로 만든 $R$은 현재 추정값 근처의 1차 근사일 뿐이다. 로봇이 이동하면서 추정값이 선형화 지점에서 멀어지면 근사 오차가 누적된다. iSAM1의 대응은 **주기적 전면 재선형화(periodic full relinearization)**였다. 몇십 스텝마다 전체 factor graph를 처음부터 다시 선형화하고 QR 분해를 처음부터 다시 수행했다. 루프 클로저로 $R$에 채움(fill-in)이 발생해 희소 구조가 손상되는 것은 이 스케줄이 촉발되는 가시적 증상이었지만, 비용의 근본은 "전체를 주기적으로 다시 푼다"는 스케줄 자체에 있었다. 증분적으로 보이던 알고리즘이 주기마다 배치 알고리즘으로 되돌아가는 구조였다. 2012년 [iSAM2](https://doi.org/10.1177/0278364911430419)는 Bayes tree라는 자료구조로 이 문제를 해결했다. Bayes tree는 factor graph에 variable elimination을 적용하여 얻는 chordal Bayes net으로부터 구성되는 트리 구조다. Bayes net의 클리크(clique)를 노드로, 클리크 간 공유 변수(separator)를 엣지로 가진다. 새 factor가 추가될 때 Bayes tree에서 영향받는 클리크를 특정하고, 해당 서브트리만 factor graph로 되돌려 재선형화·재최적화한다. **Fluid relinearization**은 선형화 오차가 임계값을 넘는 factor만 선택적으로 골라 다시 선형화하고, 그 영향이 Bayes tree의 separator를 타고 필요한 만큼만 전파한다. iSAM1의 "주기마다 전체" 스케줄이 "필요한 factor만, 영향받는 clique만"으로 대체된 셈이다. 루프 클로저가 발생해도 연결 clique 집합이 국소적으로 한정되는 경우가 많아 전체 재계산을 피할 수 있었다. > 🔗 **차용.** Bayes tree의 자료구조적 아이디어는 확률론적 그래픽 모델 문헌의 junction tree(join tree) 알고리즘 계보를 잇는다. Koller-Friedman의 [*Probabilistic Graphical Models*](https://mitpress.mit.edu/9780262013192/probabilistic-graphical-models/) 같은 표준 교과서가 다루는 제거 순서·chordal 그래프 기반 추론 기법이 대표적이다. 인공지능 추론 커뮤니티의 기법 계열이 실시간 로봇 SLAM에 이식된 것이다. iSAM2는 [GTSAM(Georgia Tech Smoothing and Mapping)](https://gtsam.org) 라이브러리로 패키징됐다. C++ 코어에 Python 바인딩을 얹은 형태다. Dellaert가 Georgia Tech 재직 중 Google과도 일하던 시기에도 GTSAM 개발은 이어졌다. GTSAM의 공개 문서와 사례에는 자율주행, 드론, 로봇팔 보정 등 여러 응용이 포함된다. --- ## 6.5 g2o: ROS 생태계의 범용 그래프 최적화기 뮌헨 공대(TUM)·프라이부르크의 Rainer Kümmerle, Giorgio Grisetti, Hauke Strasdat, Kurt Konolige, Wolfram Burgard는 2011년 ICRA에서 [g2o](https://doi.org/10.1109/ICRA.2011.5979949)(general graph optimization)를 발표했다. g2o는 "어떤 종류의 그래프 최적화든 플러그인 방식으로 처리한다"는 원칙으로 설계된 실용적인 오픈소스 구현이었다. g2o의 설계는 세 개념을 분리한다. vertex(변수 노드)와 edge(factor/제약)가 그래프를 구성하고, solver가 희소 선형 시스템을 푼다. 사용자는 vertex 타입과 edge의 오차 함수·Jacobian을 정의하면, g2o가 Gauss-Newton 또는 Levenberg-Marquardt로 전체 최적화를 수행한다. 희소 풀이기는 Cholmod, CSparse, Eigen 중 선택하거나 외부 라이브러리로 교체할 수 있다. ROS(Robot Operating System)가 2010년대 초 모바일 로봇 연구에 널리 퍼지면서 g2o도 그래프 기반 SLAM의 대표 구현 가운데 하나가 됐다. ORB-SLAM과 LSD-SLAM이 g2o를 채택했지만, gmapping은 particle-filter 계열이고 Cartographer는 Ceres를 쓰므로 ROS SLAM 전체가 g2o로 통일된 것은 아니다. 영향력은 컸지만 범용 표준 하나가 모든 backend를 대체한 역사는 아니었다. --- ## 6.6 왜 분야가 여기로 수렴했나 Chatila-Laumond(1985), Lu-Milios(1997), Gutmann-Konolige(1999), Folkesson-Christensen(2004), Thrun 그룹, Dellaert(2006), Kaess(2012)까지 여러 그룹이 서로 다른 문제에서 출발해 그래프 제약을 유지하고 푸는 도구를 발전시켰다. 문제 모델링 방식이 바뀌었다. EKF-SLAM은 현재 상태의 최적 추정값과 불확실성을 유지하면서 과거를 marginalize한다. 이 필터 패러다임에서 과거 포즈는 사라지고, 누적 오차는 현재 추정값 속에 잠복한다. 루프 클로저를 닫으려면 현재 covariance에 무거운 갱신이 필요하다. 그래프 SLAM은 과거 포즈를 버리지 않는다. 포즈·landmark·관측값 모두 그래프에 살아 있고, 루프 클로저는 새 엣지를 추가하는 것으로 표현된다. 재최적화가 전체 궤적을 일관성 있게 조정한다(이산 keyframe 대신 시간 연속적인 궤적으로 그래프를 재정식화하는 계열은 Ch.7c Continuous-Time SLAM 참조). 이미 지나간 포즈도 수정 대상이 된다는 점이 필터와의 본질적 차이다. 계산 비용도 달랐다. EKF의 갱신 비용은 $O(N^2)$ (landmark 수 $N$에 대해), 정보 저장은 $O(N^2)$다. 그래프 방법은 희소 Cholesky(또는 QR) 분해를 활용하면 복잡도가 크게 줄어든다. 다만 갱신 비용은 그래프 연결, 분해 과정의 fill-in, 소거 순서와 다시 계산할 영역에 달려 있다. 공간적으로 좁은 곳을 움직인다는 조건만으로 $O(N \log N)$을 보장할 수는 없다. > 📜 **예언 vs 실제.** Dellaert의 Square Root SAM(2006)이 제시한 배치 방식의 한계는 같은 그룹에서 곧바로 증분화 방향으로 이어졌다. 2008년 iSAM이 Givens rotation 기반 증분 갱신으로 이를 다뤘고, 2012년 iSAM2는 Bayes tree로 루프 클로저 상황의 효율성까지 끌어올렸다. GTSAM·Ceres·g2o는 비선형 최소제곱 문제를 다루지만, 사용 가능한 solver와 증분 자료구조는 서로 다르다. 세 논문은 동일한 문제 의식을 단계적으로 해소했으며, 이 계보는 거의 예고한 대로 실현됐다. 마지널리제이션(marginalization)의 유연성도 한몫했다. 그래프에서 오래된 포즈를 marginalize할 때 그 정보가 남은 변수들에 연결 factor로 보존된다. 필터도 과거 상태를 제거하며 그 정보를 현재 추정에 전달한다. 양쪽 모두 압축 과정의 근사와 재선형화 제약을 살펴야 한다. 슬라이딩 윈도우 최적화나 keyframe 선택 같은 공학적 트레이드오프가 여기서 등장한다. --- ## 6.7 비선형성과 강건성: 실무 엔지니어링의 층위 그래프 최적화의 이론적 우아함과 실제 구현 사이에는 간격이 있다. 그 간격을 메우는 작업이 2010년대 SLAM 엔지니어링의 상당 부분을 차지했다. 초기값 의존성이 한 문제다. 가우스-뉴턴이나 LM 최적화는 초기 포즈 추정이 참값에서 크게 벗어나 있으면 지역 최솟값(local minimum)에 수렴한다. 루프 클로저에서 잘못된 대응 관계가 섞이면 초기값이 훼손된다. 그래서 루프 클로저 검증과 아웃라이어 rejection이 백엔드 이전 단계의 핵심 작업이 됐다. 이 지역 최솟값 문제 자체를 볼록 완화(SDP)로 우회하여 전역 최적성을 증명 가능한 형태로 푸는 계열은 Ch.6b(Certifiable SLAM)에서 별도로 다룬다. 표준 최소자승은 아웃라이어에 취약하다는 것도 실무에서 금방 드러났다. Huber 비용이나 Cauchy 비용 같은 robust kernel을 쓰면 잘못된 매칭의 영향을 줄일 수 있다. g2o와 GTSAM 모두 robust kernel을 선택 가능하게 한다. 어느 kernel을 쓸지는 환경과 센서 특성에 따라 달라지며, 2026년에도 이 선택은 여전히 엔지니어의 경험에 의존한다. Marginalization 근사도 문제다. iSAM2의 Bayes tree는 정확한 증분 추론을 제공하지만, 변수 수가 계속 증가하면 트리가 커진다. 실제 시스템에서는 오래된 포즈를 marginalize하여 트리 크기를 관리한다. 이 marginalization 과정에서 발생하는 fill-in이 information matrix를 조밀하게 만들 수 있다. 어떻게 truncate할지, Prior factor로 어떻게 근사할지가 구현 품질을 가른다. > 📜 **예언 vs 실제.** g2o가 표방한 범용성은 특정 센서 목록을 모두 기본 제공한다는 뜻이 아니라, 사용자가 상태를 vertex로, 관측 제약을 edge로 정의해 같은 최적화 뼈대에 얹을 수 있다는 뜻이었다. 이후 연구들은 line·plane·관성·객체 제약을 각자의 시스템에 맞는 사용자 정의 edge로 구현했다. 어느 시스템을 사례로 들 때에는 그 시스템이 실제로 쓰는 추정기와 edge 구현을 확인해야 한다. 2026년에도 g2o의 핵심 유산은 고정된 factor 목록보다 이 확장 인터페이스에 있다. --- ## 🧭 아직 열린 것 어느 robust kernel을 선택해야 하는가. Huber, Cauchy, Geman-McClure, DCS 등 여러 선택지가 있지만, 주어진 환경과 센서에 어느 kernel이 최적인지를 사전에 결정하는 원칙적인 방법이 없다. 이 선택은 여전히 엔지니어의 직관과 경험에 의존한다. 학습 기반으로 cost function 자체를 최적화하는 연구가 있으나, 온라인 증분 시스템에 통합하는 것은 풀리지 않은 문제다. 비가우시안 상황을 factor graph 안에서 표현하는 것은 아직 열려 있다. GTSAM·g2o의 기본적인 연속 최적화는 가우시안 residual 모델에서 출발하지만, robust kernel·max-mixture·hybrid factor 같은 확장도 있다. 그래도 루프 클로저의 오매칭 확률이나 다중 가설 포즈를 정확하고 실시간으로 표현하는 범용 해법은 없다. Bayes tree의 증분 효율은 새 factor가 건드리는 clique가 국소적일 때 가장 크다. 대규모 지도에서 loop-closure 제약이 조밀해지면 영향을 받는 subtree와 fill-in이 커져 계산량과 메모리가 늘 수 있다. 계층적 관리와 submap 분할은 이 확장을 제어하는 접근이다. --- 2010년대 들어 백엔드 논쟁은 잦아들었다. g2o와 GTSAM 같은 대표 도구가 널리 쓰이면서, 연구자들의 관심은 백엔드 위에 무엇을 얹느냐로 옮겨갔다. 어떤 feature로, 얼마나 멀리서 루프를 인식하는가가 새 물음이 되었다. 프론트엔드가 새 경쟁 무대였다. 한 가지 질문은 남았다. g2o와 GTSAM이 내놓은 해가 실제 전역 최솟값인가. Ch.6b는 프론트엔드 계보를 Ch.7에서 이어가기 전에 이 certifiability 문제를 다룬다. --- # Ch.6b — Certifiable SLAM: 지역 최솟값을 넘어서 Lu-Milios에서 g2o·GTSAM에 이르는 계보는 한 가지 문제를 남겨두었다. 포즈 그래프 최적화는 비볼록 문제이고, Gauss-Newton·LM이 내놓는 해는 지역 최솟값일 수 있다. 실무자들은 "odometry 초기값이 있으면 대체로 잘 풀린다"는 경험칙에 의존했지만, 어느 현장에서는 백엔드가 엉뚱한 지점에서 수렴하고도 실패 신호를 내지 않았다. 2015년 MIT의 Luca Carlone이 그 경험칙을 수학으로 대체하기 시작했다. Carlone의 Lagrangian duality 시도는 2019년 Rosen의 SE-Sync, Briales-Gonzalez-Jimenez의 Cartan-Sync, Yang-Carlone의 TEASER, Papalia의 CORA로 이어졌다. 이때 쓰인 도구는 모두 SLAM 바깥에서 왔다. 오퍼레이션스 리서치의 Shor relaxation, 수학 최적화의 Burer-Monteiro factorization, 미분기하의 Riemannian optimization, 그래프 이론의 Kirchhoff Matrix-Tree였다. --- ## 6b.1 지역 최솟값이라는 오래된 불안 Ch.6 §6.7은 그래프 SLAM 백엔드의 첫 문제로 초기값 의존성을 꼽았다. 비용 함수가 회전 변수 $\boldsymbol{R}_i \in \mathrm{SO}(3)$ 위에서 비볼록이기 때문에, 초기 추정이 참값에서 멀면 Gauss-Newton은 잘못된 해의 분지에 수렴한다. Handbook §6.1의 parking garage 예시에서는 같은 입력의 무작위 초기화 네 번 중 하나만 SE-Sync가 도달한 전역 최솟값에 붙고, 나머지 셋은 육안으로도 바닥이 접힌 지역 최솟값에 안착한다. 2000년대 후반까지 커뮤니티는 두 갈래로 대응했다. odometry를 신뢰해 초기값 품질을 확보했고, 루프 클로저 검증과 아웃라이어 제거를 전단에서 철저히 했다. 둘 다 유효했지만, 수렴한 값이 진짜 최솟값인지 판정하는 도구는 아니었다. Huang과 Dissanayake가 2010년 무렵 짚은 문제는 단순했다. 초기값이 아무리 좋아도 데이터 자체가 모호하면 최적화기는 틀린 답에 가서 멈출 수 있다. PGO가 NP-hard라는 것도 그 무렵 정식화됐다. 그런데도 현장에서는 g2o가 대체로 잘 풀렸다. 이론은 최악을 말하고 실무는 평균을 보는 간극을 2010년대 중반 백엔드 이론 연구자들이 파고들었다. Gauss-Newton이 수렴했다고 해서 전역 최적이라는 뜻은 아니다. 이상적인 2차 조건 아래의 국소 최솟값은 기울기가 0이고 헤시안이 양의 준정부호지만, 정지 기준이나 수치 문제 때문에 그 조건을 만족하기 전에 멈출 수도 있다. 백엔드가 "수렴했다"고 신호를 보내면 실패가 가장 눈에 띄지 않는다. > 🔗 **차용.** Ch.6의 robust kernel(Huber, Cauchy)과 GNC는 [Black & Rangarajan (1996)](https://cs.brown.edu/people/mjblack/Papers/ijcv1996.pdf)의 robust statistics·이중성 정리를 공유한 뿌리에서 갈라졌다. 한쪽은 비용 가중으로 아웃라이어 영향을 줄였고, 반대쪽은 같은 원리를 비볼록성 회피에 전용했다. --- ## 6b.2 Shor relaxation — 바깥에서 들어온 무기 PGO의 비볼록성은 회전 제약 $\boldsymbol{R}_i \in \mathrm{SO}(d)$에서 온다. 직교성 조건 $\boldsymbol{R}^\top \boldsymbol{R} = \boldsymbol{I}$은 이차식이다. 3차원에서 $\det(\boldsymbol{R})=+1$ 자체는 삼차식이지만, 열벡터의 오른손 좌표계 조건을 외적 관계로 쓰면 이차 제약으로 표현할 수 있다. 목적함수도 이차이므로 PGO는 **QCQP**(Quadratically Constrained Quadratic Program)로 정식화할 수 있다. 그리고 QCQP에는 1987년 이후 오퍼레이션스 리서치 분야에서 검증된 convex relaxation 도구가 있었다. [Naum Shor의 1987 relaxation](https://link.springer.com/article/10.1007/BF01582220)이다. Shor의 아이디어는 $\boldsymbol{x}^\top \boldsymbol{M}\boldsymbol{x} = \mathrm{tr}(\boldsymbol{M}\boldsymbol{x}\boldsymbol{x}^\top)$ 항등식으로 리프팅 변수 $\boldsymbol{X} \triangleq \boldsymbol{x}\boldsymbol{x}^\top$을 도입해 원 QCQP를 "$\boldsymbol{X} \succeq 0$이면서 rank-1" 위의 선형 목적 문제로 바꾸는 것이다. rank-1 제약을 버리면 볼록한 **SDP**를 얻는다. 탐색 공간이 $n$에서 $n(n+1)/2$로 늘지만 볼록성을 얻는다. $$d^* = \min_{\boldsymbol{X}\in\mathbb{S}^n} \mathrm{tr}(\boldsymbol{C}\boldsymbol{X}) \;\; \text{s.t.} \;\; \mathrm{tr}(\boldsymbol{A}_i\boldsymbol{X})=b_i,\; \boldsymbol{X}\succeq 0.$$ 쓸모는 이중성 부등식 $d^* \le p^*$에 있다. SDP 최솟값은 원 QCQP 최솟값의 아래쪽 경계다. 후보해 $\hat{\boldsymbol{x}}$가 있을 때 $f(\hat{\boldsymbol{x}}) - d^*$가 그 후보의 최적성 간극의 상한이 된다. 여기서 "certifiable"이라는 이름이 나온다. 전역적으로 못 풀어도, 가진 해가 얼마나 나쁜지의 상한은 풀 수 있다. SDP 해 $\boldsymbol{X}^*$가 rank-1로 떨어지면 $\boldsymbol{X}^* = \boldsymbol{x}^*\boldsymbol{x}^{*\top}$에서 $\boldsymbol{x}^*$가 원 QCQP의 전역 최적해다. 이 "favorable situation"이 SLAM에서 얼마나 자주 일어나는지가 이후 논문들의 주제가 된다. 이 계보의 출발점은 Carlone이 2015년 IROS와 ICRA에서 발표한 두 편의 논문, [Carlone et al. 2015 "Lagrangian duality in 3D SLAM"](https://arxiv.org/abs/1506.00746)과 [Carlone & Dellaert 2015 "Planar pose graph optimization"](https://doi.org/10.1109/ICRA.2015.7139264)이다. 2D PGO에서 duality gap이 대개 0임을 경험적으로 보였고, 3D로 확장 가능함을 시사했다. Carlone은 2014년 TRO 서베이에서 g2o·GTSAM 초기화 기법을 정리한 직후였고, odometry와 루프 클로저가 충돌할 때 최적화가 자주 틀린 지점에서 멈추는 것을 본 뒤였다. 2015년 논문은 "duality gap이 보통 0"임을 보고할 뿐, 언제 성립하는지의 닫힌 조건은 주지 못했다. 같은 시기 [Briales & Gonzalez-Jimenez (2017)](https://arxiv.org/abs/1702.03235)의 Cartan-Sync가 SO(3) synchronization으로 같은 프로그램을 밀었다. 수학 쪽에서는 Boumal·Absil·Sepulchre가 Riemannian optimization을, 최적화 쪽에서는 Burer-Monteiro의 low-rank SDP factorization이 2003년부터 자리 잡고 있었다. 흩어진 재료들이 2019년 한 편의 논문에서 조립된다. --- ## 6b.3 SE-Sync — Rosen 2019가 조립한 것 [Rosen, Carlone, Bandeira, Leonard의 SE-Sync (IJRR 2019)](https://arxiv.org/abs/1612.07386)는 certifiable SLAM의 캐논이다. Rosen은 MIT에서 John Leonard의 박사과정을 마쳤고, Leonard는 Ch.4의 Durrant-Whyte와 함께 1990년대 초 "SLAM"이라는 이름을 자리잡게 한 MIT 연구자였다. 공저자 Afonso Bandeira는 SDP·synchronization 수학 쪽 전문가로 rank-deficient 2차 임계점의 전역성 증명을 맡았다. Shor relaxation, translation elimination, Burer-Monteiro low-rank parameterization, Boumal의 Riemannian staircase를 PGO라는 한 문제 위에서 맞물리게 했다. 조립의 순서는 세 단계다. 첫째, 회전 고정 시 translation이 선형 최소자승이 된다는 관찰에서 $\boldsymbol{t}$를 닫힌 형태로 소거한다(Problem 6.2). Ch.6의 graph SLAM 계보가 오래전부터 알던 사실을 Carlone이 2014년 TRO 서베이에서 명시했고, Rosen이 convex relaxation의 첫 단계로 집어넣었다. 둘째, 남은 rotation-only 문제 $\min_{\boldsymbol{R}\in\mathrm{SO}(d)^n} \mathrm{tr}(\tilde{\boldsymbol{Q}}\boldsymbol{R}^\top\boldsymbol{R})$에 Shor relaxation을 적용해 SDP로 리프팅한다(Problem 6.3). 셋째, $dn \times dn$ 차원 SDP는 그대로 풀면 interior-point method가 수천 포즈에서 무너지므로 Burer-Monteiro 재파라미터화 $\boldsymbol{Z} = \boldsymbol{Y}^\top \boldsymbol{Y}$로 Stiefel manifold 위의 저차원 비제약 문제로 바꾼다(Problem 6.4). 두 정리가 이 조립을 정당화한다. Theorem 6.1 **exact recovery**: 측정 노이즈가 어떤 상수 $\beta$보다 작으면 SDP relaxation의 유일 해 $\boldsymbol{Z}^*$가 rank $d$이고, 원 MLE의 전역 최솟값을 정확히 복원한다. 일반적인 벡터 QCQP의 rank-1 조건과 달리, 여기서는 $\boldsymbol{Z}=\boldsymbol{R}^\top\boldsymbol{R}$이고 $\boldsymbol{R}\in\mathbb{R}^{d\times dn}$이므로 exact solution의 rank가 $d$다. 다만 $\beta$는 ground-truth에 의존해 사전에는 모른다. Theorem 6.2는 Boumal et al.의 결과로, Stiefel manifold 위에서 찾은 2차 임계점이 rank-deficient하면 곧 전역 최솟값임을 보장한다. 이 두 정리가 Riemannian Staircase를 가능케 한다. rank를 작게 두고 시작해 2차 임계점을 찾고 rank-deficiency를 검사하고, 안 맞으면 rank를 하나 올린다. rank가 $dn + 1$에 닿으면 모든 $\boldsymbol{Y}$가 row rank-deficient가 되므로 유한 단계 내 반드시 멈춘다. 실무 데이터셋에서는 보통 한 계단이면 끝난다. sphere·torus·garage 벤치마크에서 SE-Sync는 g2o·GTSAM 수준 속도로 수렴하며 a posteriori certificate를 함께 냈다. g2o·GTSAM은 빨랐지만 답을 언제 믿을지 알려 주지 않았고, Rosen의 알고리즘은 suboptimality bound까지 계산한다. 이 bound가 0이면 해는 증명 가능하게 전역 최적이다. Lu-Milios 이후 20년 만에 백엔드가 해의 전역 최적성을 인증할 수 있게 됐다. 간극이 남았다는 사실만으로 그 해가 비최적이라고 판정하는 것은 아니다. > 📜 **예언 vs 실제.** Rosen은 IJRR 2019 논문 §8.2에서 "우리가 보인 algebraic simplification은 anisotropic noise·outlier·다양한 센서 모달리티로 확장될 수 있을 것"이라 적었다. 그 예언은 부분적으로 적중했다. 2023년 Holmes-Barfoot의 landmark-SLAM 확장, 2024년 Papalia의 CORA 범위 측정 확장, Yang-Carlone의 TEASER 계열이 실제 뒤따랐다. 그러나 "visual SLAM의 perspective projection까지 SE-Sync가 덮는다"는 가장 야심찬 확장은 2026년에도 오지 않았다. Projection이 rational function이라 polynomial optimization으로 편입되기 어렵다는 구조적 장벽이 드러났다. > 🔗 **차용.** SE-Sync의 심장에 있는 Burer-Monteiro factorization은 [Burer & Monteiro (2003)](https://link.springer.com/article/10.1007/s10107-002-0352-8)의 low-rank SDP 해법이다. 그 위에 [Boumal-Voroninski-Bandeira (2016)](https://arxiv.org/abs/1605.08101)이 Riemannian 언어로 2차 임계점의 전역성을 보였고, Rosen이 SLAM 맥락에 가져왔다. 순수 수학에서 로봇 백엔드까지 16년이다. --- ## 6b.4 Graph Laplacian과 Fisher Information의 뜻밖의 등가 전역 최적해를 찾았더라도, 그 추정이 참값과 얼마나 가까운지는 별개의 질문이다. Cramér-Rao Lower Bound와 Fisher Information Matrix가 이 질문을 다룬다. 회전을 고정한 단순 PGO 모델에서 Rosen-Khosoussi-Barfoot의 결과에 따르면 FIM은 그래프의 weighted reduced Laplacian의 Kronecker product로 정확히 떨어진다. $$\mathcal{I} = \boldsymbol{J}^\top \boldsymbol{\Sigma}^{-1} \boldsymbol{J} = \boldsymbol{L}_w \otimes \boldsymbol{I}_3.$$ 그래프 구조만 알면 실제 측정 없이도 추정 정확도의 근사를 얻는다. Kirchhoff의 Matrix-Tree Theorem에 따라 reduced Laplacian의 determinant는 가중 spanning tree 수와 같고, 이것이 D-optimality(정보 행렬의 행렬식)에 대응한다. 알제브라 연결성(Fiedler value)은 E-optimality(최악 분산)에 대응한다. 1847년 Kirchhoff가 전기 회로망을 위해 증명한 정리는 180년 뒤 측정 선택·active SLAM의 이론 기반이 되었다. active SLAM에서 "FIM 최대화"는 Laplacian 스펙트럼 조작으로 환산된다. [Kasra Khosoussi와 Timothy Barfoot의 2014년 이후 작업](https://arxiv.org/abs/1709.08601)이 이 연결을 정립했다. Khosoussi는 Sydney에서 Dissanayake·Huang 지도로 박사과정을 밟았고, 이후 MIT와 Toronto를 거쳤다. 3D PGO로 일반화된 형태에서는 Laplacian과 SE(3) adjoint representation의 Kronecker 결합이 등장해 위상·기하 정보를 분리해 다루게 한다. "측정 선택 기준"을 FIM 전체 대신 6배 작은 Laplacian으로 근사 가능하다는 것이 Ch.6에서 간단히 언급한 "루프 클로저 선택"의 수학적 근거가 된다. Ch.4 §4.8이 짚은 EKF-SLAM의 consistency 문제도 같은 질문으로 이어진다. Julier-Uhlmann이 2001년 지적한 EKF의 over-confidence는 CRLB로 재해석하면 근사 선형화가 Fisher information을 과대 추정한다. Handbook §6.2는 FIM을 convex relaxation과 나란히 다룬다. 전역 최솟값과 그 정확도는 쌍으로 다뤄야 한다. > 🔗 **차용.** [Kirchhoff의 Matrix-Tree Theorem(1847)](https://en.wikipedia.org/wiki/Kirchhoff%27s_theorem)은 전기 회로망 분석 도구로 태어나 조합론을 거쳐 측정 설계 문헌으로 이식됐고, 2010년대 Khosoussi를 통해 SLAM active perception의 언어가 됐다. 한 정리의 180년 이주 경로다. --- ## 6b.5 확장과 한계 — TEASER, CORA, 그리고 Lasserre의 벽 SE-Sync가 나온 뒤 연구 범위는 아웃라이어에 강건한 certifiable estimator와 range·landmark·anisotropic noise 같은 확장된 측정 모델로 넓어졌다. 아웃라이어 쪽이 먼저였다. Ch.6 §6.7이 짚었듯 루프 클로저 검증이 완벽하지 않으면 오매칭이 섞이고, Huber·Cauchy 커널로도 일정 비율 이상의 아웃라이어 앞에서는 최적화가 무너진다. 2017년 무렵 certifiable 계보가 이에 답해야 했다. 대표는 [Yang, Shi, Carlone의 TEASER (TRO 2020)](https://arxiv.org/abs/2001.07715)다. 3D 점군 등록에서 99% 아웃라이어에서도 전역 최적 해를 찾는다. truncated least squares 비용을 GNC 래퍼에서 풀되 회전 부분 문제에 SDP relaxation을 붙여 certificate를 함께 낸다. 비결은 스케일·translation·rotation을 각각 certifiable subproblem으로 쪼개 단계마다 전역 최적 보장과 함께 넘기는 데 있었다. 이어진 [Yang & Carlone (2022)](https://arxiv.org/abs/2109.03349)는 이를 Lasserre moment relaxation으로 일반화해 "certifiably robust estimation"이라 명명했다. Range-aided SLAM은 [Papalia et al. CORA (2024)](https://arxiv.org/abs/2302.11614)의 자리다. 범위 측정 $(\|\boldsymbol{t}_j - \boldsymbol{t}_i\| - \tilde r_{ij})^2$는 노름의 제곱근을 포함해 그대로는 이차식이 아닌데, Papalia는 보조 단위벡터 $\boldsymbol{b}_{ij} \in S^{d-1}$로 bearing lifting해 QCQP에 다시 집어넣었다. CORA는 단일 로봇에서는 tight한 relaxation이 멀티로봇에서는 일반적으로 exact하지 않음을 보여 "언제 Shor가 통하는가"의 범위를 좁혔다. Landmark 쪽에서는 [Holmes & Barfoot (2023)](https://arxiv.org/abs/2308.05631)이 Schur complement로 landmark를 미리 소거해 SE-Sync에 바로 넣을 수 있는 형태로 만들었다. Holmes·Khosoussi·Rosen은 Handbook Ch.6을 공저했다. 이 계보가 2025년 한 테이블에 모였다. 그러나 벽도 드러났다. anisotropic noise와 truncated-quadratic outlier를 POP(Polynomial Optimization Problem)로 일반화하면 Lasserre moment relaxation이 필요한데, 유도된 SDP가 **degenerate**해 constraint qualification이 실패하고 Riemannian Staircase의 수렴 조건이 깨진다. Yang의 2022년 sparse monomial basis 같은 우회가 있지만 전용 solver는 일반 local solver보다 느리다. 속도와 증명 가능성을 동시에 확보하는 알고리즘은 아직 없다. Visual SLAM·VIO에는 perspective projection과 IMU preintegration의 구조적 비호환이라는 더 깊은 장벽이 있다. > 📜 **예언 vs 실제.** Carlone이 2015년 ICRA에서 "Lagrangian dual이 tight한 인스턴스가 왜 대부분인지 이론적 해명이 필요하다"고 적었다. 10년이 지났고, 답은 부분적으로만 나왔다. Rosen-Carlone-Bandeira-Leonard의 exact recovery 정리가 "노이즈가 $\beta$ 이하"라는 충분조건을 주었지만, 실제 SLAM 인스턴스에서 $\beta$를 사전에 계산하는 방법은 없다. tightness가 언제 깨지는지에 대한 **사전**(a priori) 조건은 2026년 기준 여전히 per-instance certificate로 대체되어 있다. --- ## 🧭 아직 열린 것 **Tightness가 깨지는 경계.** SE-Sync의 exact recovery 정리는 "노이즈 $\beta$ 이하"라는 충분조건만 주고, 실제 인스턴스에서 $\beta$ 계산법은 없다. 사전에 tight 여부를 판정할 수 있어야 알고리즘 설계가 진전된다. 무거운 아웃라이어나 극단적 희소 그래프에서 relaxation의 실패 양상에 대한 계통적 연구는 아직 초기 단계다. **Visual SLAM·VIO와의 통합.** perspective projection $\pi(\boldsymbol{X}) = [X/Z, Y/Z]$는 rational function이라 polynomial optimization에 그대로 들어오지 않는다. 분모를 곱해 polynomial로 바꿔도 feature마다 새 변수·보조 제약이 추가되어 ORB-SLAM3의 수천 맵포인트에서 SDP 크기는 실시간 영역 밖이다. Forster의 2015년 IMU preintegration도 exponential map과 bias drift가 얽혀 POP 편입이 어렵다. 2026년 기준 Ch.7·Ch.8·Ch.13의 visual/VIO 주류는 certifiable guarantee 바깥에 있다. **Online certification과 스케일.** SE-Sync는 배치다. 새 측정마다 SDP를 다시 풀어 certificate를 갱신하는 증분 certifiable SLAM은 아직 성숙하지 않았다. iSAM2가 배치 SAM에 풀어낸 증분화를 certifiable 쪽에서 반복해야 하는 셈이다. warm-start, rank 증분, 부분 certificate 합성 모두 열린 연구고, 도시 규모 그래프에서 moment relaxation solver의 속도도 여전히 문제다. **Outlier-majority.** TEASER처럼 다수 아웃라이어가 섞인 정합에서 성과를 보인 추정기도 있다. 다만 그 결과가 임의의 SLAM 문제와 오염 조건에서 식별 가능성이나 인증을 보장하는 것은 아니다. 식별 가능한 해가 여러 개인 경우에는 list-decodable regression 같은 다중 가설 접근도 검토한다. 2024년 Cheng·Shi·Carlone의 후속 작업이 있었지만 TEASER 같은 표준 도구는 없다. --- 경험칙은 10년의 이론 프로그램으로 대체됐다. Carlone-Khosoussi-Rosen-Holmes-Barfoot-Dissanayake가 공저한 *The SLAM Handbook* Ch.6은 이 주제를 34페이지로 다룬다. 같은 10년 동안 Ch.12·Ch.13·Ch.16의 학습 기반 SLAM은 다른 경로로 나아갔다. 한쪽은 해의 전역성을 증명하는 쪽, 다른 쪽은 신경망이 해를 직접 예측하는 쪽이다. 두 계보가 만날지, 분야를 둘로 나누어 지낼지는 2026년에도 답이 없다. Ch.19는 이 장의 열린 항목들을 "백엔드 이론의 공백" 아래 다시 묶는다. 본류는 Ch.7로 돌아가, 백엔드를 주어진 것으로 놓고 그 위에서 작동할 프론트엔드를 묻는다. --- # Ch.7 — Feature-based 계보: ORB-SLAM 삼부작 Ch.6의 graph SLAM 계보는 pose graph optimization을 SLAM의 공통 언어로 굳혔다. Kümmerle(2011)의 g²o와 Kaess(2012)의 iSAM2는 대규모 지도에서 반복 최적화를 실현했고, loop closure의 비용을 현실적인 수준으로 낮췄다. 3부는 그렇게 자리를 잡은 백엔드 위에서 프론트엔드로 시선을 옮긴다. 남은 과제는 어떤 특징을 어떻게 뽑아 추적할 것인가였다. Klein과 Murray가 2007년 PTAM으로 tracking과 mapping을 두 스레드로 분리했을 때, 그것은 실험실 데모였다. 가능성을 입증한 아이디어였지만, 소규모 실내 장면을 넘어서면 무너졌다. Raúl Mur-Artal이 2015년 Zaragoza 대학에서 그 구조를 가져올 때, 그는 세 가지를 함께 들고 왔다. Rublee(2011)의 ORB 디스크립터, Gálvez-López(2012)의 DBoW2 visual vocabulary, 그리고 Strasdat(2011)의 Essential graph 아이디어. PTAM이 빠른 프로토타입이었다면 ORB-SLAM은 10년짜리 표준이었다. --- ## 7.1 ORB-SLAM (2015): 설계의 삼각대 [Mur-Artal, Montiel & Tardós 2015. ORB-SLAM](https://doi.org/10.1109/TRO.2015.2463671)은 IEEE Transactions on Robotics에 실린 논문이다. 제목은 단순하다. ORB feature를 쓰는 SLAM이지만, 논문의 선택 하나하나가 설계 판단이다. 시스템의 뼈대는 Tracking, Local Mapping, Loop Closing 세 스레드다. PTAM도 두 스레드(Tracking과 Mapping)였다. Mur-Artal은 Loop Closing이라는 세 번째 스레드를 추가했다. Loop Closing은 DBoW2로 장소를 인식하고, Essential graph를 통해 포즈 그래프를 최적화하며, 마지막으로 전역 Bundle Adjustment를 실행한다. 이 분리 덕분에 Tracking은 지도 수정을 기다리지 않고 실시간을 유지한다. > 🔗 **차용.** PTAM(Klein & Murray, 2007)의 Tracking–Mapping 분리가 ORB-SLAM의 Tracking–LocalMapping으로 직접 이어졌다. Mur-Artal은 논문 §3에서 이 부채를 명시했다. ORB-SLAM은 두 스레드 구조를 세 스레드로 확장하며 루프 클로저를 독립 모듈로 격리했다. Mur-Artal이 front-end에서 ORB(Oriented FAST and Rotated BRIEF) descriptor를 고른 데는 이유가 있었다. SIFT와 SURF는 특허 문제가 있었고, BRIEF는 빠르지만 회전에 취약했다. ORB는 FAST 키포인트에 회전 불변성을 덧붙인 것으로, Rublee et al.이 2011 ICCV에서 발표했다. 원 논문 실험에서 계산 속도가 SIFT보다 약 두 자릿수 차이, 즉 100배 수준으로 빠르다고 보고됐고 binary 형태라 해밍 거리로 매칭한다. CPU에서도 실시간으로 동작한다. ORB가 scale invariance를 얻는 방식은 image pyramid다. 원본 이미지를 스케일 팩터 s(ORB-SLAM에서는 1.2)로 8단계 축소해 피라미드를 만들고, 각 레벨에서 독립적으로 FAST 키포인트를 검출한다. 키포인트의 방향(orientation)은 intensity centroid로 정의한다: 패치 내 픽셀 intensity의 1차 모멘트로 중심을 구하고, 이 방향각 θ를 BRIEF 비트 비교 쌍에 적용해 회전 불변 descriptor를 만든다. 결과는 256-bit binary vector다. 두 descriptor 사이의 유사도는 XOR 후 popcount, 즉 해밍 거리로 계산한다. > 🔗 **차용.** [Rublee et al. 2011. ORB](https://doi.org/10.1109/ICCV.2011.6126544)의 descriptor가 시스템 이름 자체가 되었다. ORB는 Zaragoza 팀이 설계한 것이 아니다. Mur-Artal은 있는 도구를 가져와 파이프라인을 조립했다. front-end의 선택이 10년간 시스템의 이름으로 불린 경우다. 키프레임 선택 정책이 PTAM과 다르다. PTAM은 키프레임을 공격적으로 추가했다. ORB-SLAM은 covisibility graph 기반으로 중복을 제거한다. **Covisibility graph**는 키프레임 사이의 공유 landmark 수를 엣지 가중치로 삼는 그래프다. 공유 landmark가 15개 이상인 키프레임 쌍이 연결된다. Local Mapping은 이 그래프를 이용해 local window를 선택하고, 그 안에서만 Bundle Adjustment를 수행한다. KITTI 시퀀스 00(전체 4.5 km 루프)에서 ORB-SLAM은 1.2% translation drift를 기록했다. 당시 비교 대상이었던 PTAM은 큰 루프를 닫지 못했다. 절대 scale의 모호성은 두 단안 시스템에 공통으로 남았다. ORB-SLAM이 같은 시퀀스에서 루프를 닫고 drift를 흡수한 것은 Essential graph와 DBoW2 덕분이다. **Essential graph**는 covisibility graph의 부분 그래프다. 공유 landmark가 100개 이상인 엣지, spanning tree, 루프 클로저 엣지만 남긴다. 루프가 감지되면 이 그래프 전체를 포즈 그래프로 최적화한다. 수천 개의 키프레임이 있어도 Essential graph의 엣지는 sparse하다. 최적화가 수 초 안에 끝난다. > 🔗 **차용.** Essential graph의 아이디어는 [Strasdat et al. 2011. Double Window Optimisation](https://doi.org/10.1109/ICCV.2011.6126517)의 계층적 최적화 구조에서 왔다. Strasdat는 local window와 global window를 분리해 최적화 비용을 낮췄다. Mur-Artal은 이를 Essential graph라는 sparse 포즈 그래프로 일반화했다. 루프 클로저의 장소 인식은 DBoW2가 담당한다. [Gálvez-López & Tardós 2012. DBoW2](https://doi.org/10.1109/TRO.2012.2197158)는 binary descriptor용 vocabulary tree다. ORB descriptor를 k-medians(k-means++ seeding)로 계층적 클러스터링해 트리 구조의 vocabulary를 만든다. 트리 분기 수 $k_w$와 깊이 $L_w$가 고정되면 leaf 노드(word) 수는 $k_w^{L_w}$가 된다. DBoW2 논문은 $k_w=10$, $L_w=6$로 1백만 단어 규모의 vocabulary를 학습한 예를 보고하며, ORB-SLAM 공개 구현도 비슷한 수준의 vocabulary를 사용한다. 각 word에는 TF-IDF(Term Frequency–Inverse Document Frequency) 가중치가 붙는다: 특정 word가 전체 키프레임 데이터베이스에 자주 등장할수록 낮은 IDF 가중치를 받아 discriminative한 word가 더 큰 영향력을 갖는다. 키프레임은 이 가중 BoW 벡터로 표현되고, inverted index에 저장된다. 새 프레임이 들어오면 vocabulary tree를 내려가 word를 결정하는 데 O(log(k^L))=O(L)이 걸리고, inverted index로 후보 키프레임을 바로 조회한다. 전체 지도를 순회하지 않는다. Tracking 스레드는 매 프레임마다 현재 포즈를 추정한다. 이전 프레임과의 feature matching 뒤 motion-only bundle adjustment로 포즈 $\mathbf{T}_{cw} \in SE(3)$를 정제한다. 다음은 3D–2D correspondence $\{(\mathbf{X}_i, \mathbf{u}_i)\}$에 대한 기본 reprojection 목적함수다: $$\mathbf{T}^* = \arg\min_{\mathbf{T}} \sum_i \left\| \mathbf{u}_i - \pi(\mathbf{T}\mathbf{X}_i) \right\|^2$$ 여기서 $\pi$는 카메라 투영 함수, $\mathbf{X}_i$는 맵 포인트의 월드 좌표, $\mathbf{u}_i$는 이미지 좌표다. 실제 포즈 최적화는 강건 손실과 관측 가중치를 사용하며, 맵 포인트를 고정한 채 현재 카메라 포즈만 바꾼다. EPnP와 RANSAC은 재위치 추정에서 초기 포즈 후보를 얻는 데 쓰인다. 이웃 키프레임과 맵 포인트를 함께 바꾸는 local BA는 별도의 Local Mapping 스레드가 맡는다. --- ## 7.2 ORB-SLAM2 (2017) — stereo/RGB-D ORB-SLAM(2015)은 mono-only였다. 카메라 하나만으로는 scale을 알 수 없다. "이 복도가 10m인가 100m인가"를 이미지 픽셀에서 읽어낼 방법이 없다. Mur-Artal과 Tardós는 2016년에 stereo와 RGB-D 확장 작업을 시작했다. [Mur-Artal & Tardós 2017. ORB-SLAM2](https://doi.org/10.1109/TRO.2017.2705103)는 stereo와 RGB-D를 추가해 이 문제를 해결한다. stereo는 기선(baseline)을 알므로 depth를 직접 삼각측량한다. RGB-D는 depth 센서가 측정값을 준다. 두 경우 모두 scale을 알 수 있다. 구조는 mono와 동일한 세 스레드다. front-end만 센서 종류에 따라 달라진다. stereo는 rectified 이미지 쌍에서 ORB를 추출하고 좌우 매칭으로 depth를 구한다. 좌우 대응이 있는 특징점은 **stereo 관측**으로, 한쪽에서만 검출된 특징점은 **monocular 관측**으로 쓴다. 깊이를 구한 점은 기선 길이에 비례한 임계값으로 가까운 점과 먼 점을 다시 구분한다. **Stereo 초기화**는 mono와 달리 첫 프레임부터 즉각 수행된다. mono 초기화는 두 프레임 사이의 Essential Matrix나 Homography를 통해 맵을 구성하고 scale 모호성이 남는다. Stereo는 첫 키프레임에서 좌우 이미지 간 수평 시차(disparity) $d$와 기선 $b$, 초점 거리 $f$로 depth를 계산한다: $$Z = \frac{b \cdot f}{d}$$ depth $Z$가 임계값 $Z_{\max}=40b$ 이하인 특징점은 즉시 3D 맵 포인트로 등록된다. RGB-D 초기화도 동일한 원리다. depth 이미지에서 픽셀 $(u, v)$의 depth 값 $Z$를 읽고, 역투영(back-projection)으로 3D 좌표를 얻는다. 두 경우 모두 scale이 고정되므로 첫 프레임 직후 Local BA를 바로 실행할 수 있다. EuRoC MAV(Micro Aerial Vehicle) 데이터셋 Machine Hall 01 시퀀스에서 ORB-SLAM2(stereo)는 Table II에서 절대 translation 오차 0.035 m를 기록했다. 같은 표는 Stereo LSD-SLAM을 비교 대상으로 삼고 있어, 해당 조건에서 ORB-SLAM2의 오차가 더 작았음을 확인할 수 있다. KITTI 오도메트리에서도 ORB-SLAM2가 당시 발표된 방법 중 상위권이었다. 2017년 5월 논문이 IEEE TRO에 실리던 날, Mur-Artal과 Tardós는 GitHub에 소스를 함께 올렸다. Zaragoza 팀의 두 사람이 mono·stereo·RGB-D 세 모드를 단일 코드베이스로 공개한 것이다. 이후 GitHub star는 수천을 넘었고, ROS 래퍼가 커뮤니티에서 만들어졌다. --- ## 7.3 ORB-SLAM3 (2021): Atlas와 Visual-Inertial 2021년 IEEE Transactions on Robotics에 실린 [Campos et al. 2021. ORB-SLAM3](https://doi.org/10.1109/TRO.2021.3075644)는 저자 목록이 달라진다. Mur-Artal이 아니라 Carlos Campos가 1저자다. Mur-Artal은 Tardós와 함께 공저자로 이름을 올렸다. Campos는 Zaragoza 대학에서 Tardós 지도 아래 박사 과정을 밟았다. 계보가 한 세대 내려온 것이다. ORB-SLAM3는 **Atlas**(멀티맵)와 **Visual-Inertial** 모드를 추가했다. Atlas는 여러 개의 분리된 지도를 동시에 유지하는 구조다. 추적이 실패하면 기존 지도를 닫고 새 지도를 시작하며, 나중에 같은 장소를 재방문했을 때 두 지도를 병합한다. ORB-SLAM과 ORB-SLAM2도 기존 지도에서 재위치 추정을 시도했지만, 복구하지 못한 뒤 새 지도를 시작하면 이전 지도와 함께 관리하고 병합하는 데 한계가 있었다. ORB-SLAM3 논문은 Atlas를 이 실패 양상에 대한 해법으로 제시했다. ORB-SLAM3는 실패 후 재초기화하고 이전 지도를 기억한다. Visual-Inertial(VI) 모드는 IMU 데이터를 통합한다. Campos는 Forster et al.이 RSS 2015에서 "IMU Preintegration on Manifold" 제목으로 제안하고 2016년 IEEE TRO에서 [On-Manifold Preintegration for Real-Time Visual-Inertial Odometry](https://doi.org/10.1109/TRO.2016.2597321)로 확장한 방식을 그대로 가져왔다. IMU는 빠른 모션에서 Visual SLAM이 잃기 쉬운 추적을 보완한다. VI-SLAM은 단안 카메라의 scale ambiguity도 해결한다. IMU의 가속도계 측정이 중력 방향과 함께 절대 scale을 제공한다. 키프레임 $i$와 $j$ 사이의 IMU 측정을 한 번만 적분해 놓는다. 가속도계·자이로스코프 측정값을 $\tilde{\mathbf{a}}_t = \mathbf{a}_t + \mathbf{b}_a + \mathbf{n}_a$, $\tilde{\boldsymbol{\omega}}_t = \boldsymbol{\omega}_t + \mathbf{b}_g + \mathbf{n}_g$로 모델링하면(bias $\mathbf{b}$, noise $\mathbf{n}$), 두 키프레임 사이의 상대 회전·속도·위치 변화량을 다음과 같이 preintegration한다: $$\Delta\mathbf{R}_{ij} = \prod_{k=i}^{j-1} \mathrm{Exp}\bigl((\tilde{\boldsymbol{\omega}}_k - \mathbf{b}_g)\Delta t\bigr)$$ $$\Delta\mathbf{v}_{ij} = \sum_{k=i}^{j-1} \Delta\mathbf{R}_{ik}\,(\tilde{\mathbf{a}}_k - \mathbf{b}_a)\Delta t$$ $$\Delta\mathbf{p}_{ij} = \sum_{k=i}^{j-1}\!\left[\Delta\mathbf{v}_{ik}\Delta t + \tfrac{1}{2}\Delta\mathbf{R}_{ik}\,(\tilde{\mathbf{a}}_k - \mathbf{b}_a)\Delta t^2\right]$$ 여기서 $\mathrm{Exp}(\cdot)$는 $\mathfrak{so}(3)$의 지수 사상이다. bias가 BA 중 갱신되면 전체 재적분 없이 1차 선형 근사로 보정한다. ORB-SLAM3는 이 preintegrated 항($\Delta\mathbf{R}$, $\Delta\mathbf{v}$, $\Delta\mathbf{p}$)을 factor graph의 inertial edge로 추가해 visual reprojection residual과 함께 최적화한다. > 🔗 **차용.** Campos는 Forster et al.의 On-Manifold Preintegration(TRO 2016, 원형은 RSS 2015) 공식을 ORB-SLAM3 Inertial 통합의 핵심으로 가져왔다. Forster의 수식은 연속 IMU 측정을 bias 보정과 함께 SO(3) 매니폴드 위에서 적분하는 방법을 제공한다. ORB-SLAM3는 이 공식을 factor graph 최적화에 연결했다. EuRoC MAV 전체 11개 시퀀스 평균 RMSE ATE(절대 궤적 오차)에서 ORB-SLAM3(mono-inertial)는 Table II에서 0.043 m로 보고된다. 같은 표에서 VINS-Mono는 0.110 m로 집계되며, Kimera(stereo-inertial)는 0.119 m였다. VI 모드와 Atlas가 결합하면 무인기나 핸드헬드 장치가 조명이 달라지거나 추적에 실패한 뒤에도 이전 지도로 돌아올 수 있다. --- ## 7.4 왜 2020년대에도 Baseline인가 2023년에도 학회 논문들은 ORB-SLAM3를 비교 대상으로 표에 넣었다. 새 방법이 발표될 때 "ORB-SLAM3보다 얼마나 낫냐"가 기준선이었다. ORB feature는 조명 변화에 어느 정도 내성이 있고, binary descriptor라 계산이 빠르며, 많은 수를 실시간으로 뽑아 추적 실패를 줄인다. learned feature가 특정 데이터셋에서는 더 정확하지만, 새로운 환경에서 무너지는 경우가 있다. ORB의 동작은 예측 가능하다. 코드가 공개되어 있고, ROS 통합이 잘 되어 있으며, 수천 개의 실사용 사례가 문서화되어 있다. 실험실에서 새 시스템을 평가할 때 ORB-SLAM3를 돌려보는 것이 첫 번째 단계가 된 지 오래다. mono·stereo·RGB-D·IMU를 단일 코드베이스가 지원하기 때문에 "우리 방법 vs ORB-SLAM3(stereo)" 혹은 "우리 방법 vs ORB-SLAM3(mono-inertial)"을 나란히 비교할 수 있다. 하나의 baseline이 여러 설정을 커버한다. Learned alternative도 ORB-SLAM3를 일관되게 능가하지 못한다. DROID-SLAM(Teed & Deng, 2021)은 여러 시퀀스에서 ORB-SLAM3를 이긴다. 그러나 논문 자체가 보고하듯 EuRoC·TartanAir 같은 대용량 시퀀스에서는 24 GB급 GPU가 필요하고 TartanAir에서는 8 fps로 실시간이 아니다. 반면 ORB-SLAM3는 CPU-only로 돌고, 커뮤니티 보고로는 ARM/임베디드 플랫폼에서도 기본 동작이 확인된다. --- ## 📜 예언 vs 실제 > 📜 **예언 vs 실제.** Mur-Artal은 2015년 ORB-SLAM 논문 Section IX-C에서 두 가지 Future Work를 제시했다. 하나는 "Points at Infinity"로, 시차가 부족해 일반 맵 포인트로 편입할 수 없는 먼 점들을 회전 추정에 활용하자는 것이었다. 다른 하나는 "Dense Map Reconstruction"으로, compact한 키프레임 선택이 dense reconstruction의 좋은 skeleton을 제공한다는 제안이었다. 10년 뒤 시점에서 보면 첫 번째 방향은 VI-SLAM 및 후속 연구에서 부분적으로 흡수됐고, 두 번째 방향은 2020년대 NeRF-SLAM·Gaussian Splatting 계열이 "sparse skeleton + dense overlay" 구도를 다른 재료로 실현하는 쪽으로 귀결됐다. 저자가 지목한 RGB-D/stereo/IMU 같은 후속 모달리티 확장은 이 Section이 아니라 ORB-SLAM2(2017)·ORB-SLAM3(2021)에서 별도의 문제의식으로 덧붙었다. > 📜 **예언 vs 실제.** Campos et al.은 2021년 ORB-SLAM3 논문 Conclusions에서 ORB-SLAM3의 주된 실패 모드가 저텍스처 환경임을 인정하며, 네 가지 data association 문제에 적합한 photometric 기법의 개발을 다음 방향으로 제시했다(내시경 영상 응용을 예로 들었다). 2023-2025년 사이 그 방향보다 먼저 두드러진 흐름은 SuperPoint·LightGlue 같은 learned front-end를 ORB-SLAM3에 이식하는 연구였고, photometric 계열의 통합은 DSO·LDSO 쪽 맥락에서 별도로 이어졌다. ORB-SLAM3 공식 저장소 main branch는 2026년 현재도 전통 ORB descriptor를 유지한다. 저자 예언의 중심축(photometric)과 실제 학계 관심(learned feature)은 어긋난 채로 굴러갔다. --- ## 🧭 아직 열린 것 **Long-term map reuse.** Atlas가 멀티맵 유지를 가능하게 했지만, 조명이 크게 달라진 환경에서 지도 병합은 여전히 실패한다. 아침에 만든 지도와 저녁에 재방문할 때의 장소를 같은 곳으로 인식하는 것이 목표인데, 외관 변화가 크면 DBoW2의 place recognition이 놓친다. seasonal change가 있는 outdoor 환경에서 이 문제는 장기 자율주행 연구의 과제로 남아 있다. 2024년 기준 완전한 해답은 없다. **Pure vision baseline의 자리.** learned feature 기반 시스템들이 표준 benchmark에서 ORB-SLAM3를 이기기 시작했다. SuperPoint + SuperGlue 조합, LightGlue, 그리고 DINOv2 기반 feature들이 특정 시퀀스에서 더 낮은 오차를 보인다. 그러나 일반화 가능성은 다른 문제다. training distribution 밖의 환경에서 learned feature가 전통 ORB보다 나쁜 결과를 내는 경우가 보고된다. "일관되게 능가한다"는 주장을 하려면 아직 더 넓은 실험이 필요하다. **대규모 outdoor에서의 drift.** ORB-SLAM3는 도심 주행이나 수 km 이상의 경로에서 LiDAR SLAM 대비 여전히 열세다. GPS-denied 환경에서 urban-scale localization을 순수 카메라로 달성하는 것은 2026년 기준 미해결이다. 시각 조건의 변화, 동적 객체, 텍스처 없는 구간이 복합되면 drift가 누적된다. LiDAR 측량 정밀도와의 격차는 좁혀지고 있으나 닫히지는 않았다. --- ORB-SLAM 삼부작이 feature-based 계보의 표준을 세운 같은 시기, Newcombe와 Engel은 정반대의 선택을 하고 있었다. 특징점을 뽑지 않고 이미지 전체의 밝기 정보를 직접 쓰겠다는 것이었다. 두 계보는 2010년대 내내 나란히 발전했고, 서로를 비교 대상으로 삼으면서 각자의 한계를 드러냈다. ORB-SLAM3가 2021년 EuRoC 벤치마크를 주도하는 동안, DSO는 TUM 복도에서 ORB-SLAM2를 눌렀다. 같은 시간표, 다른 출발점이었다. --- # Ch.7b — 흔들리는 센서에서 제약식으로: IMU Preintegration의 발명 2009년 시드니의 ACFR(Australian Centre for Field Robotics)에서 박사과정생 Todd Lupton과 지도교수 Salah Sukkarieh는 빠른 기동 중 IMU 측정을 factor graph에 묶는 문제를 다뤘다. 드론이 빠르게 기동할 때 IMU는 200Hz로 측정값을 내놓지만, factor graph에 이를 전부 넣기는 어려웠다. 키프레임은 초당 몇 개뿐인데, 그 사이 수십·수백 개 IMU 측정을 어떻게 한 묶음으로 만들 것인가. Lupton이 IROS에 낸 답이 preintegration의 씨앗이었다. 6년 뒤 2015년 RSS, Christian Forster가 Davide Scaramuzza·Luca Carlone·Frank Dellaert와 함께 그 씨앗을 SO(3) 매니폴드 위로 옮겼을 때 IMU는 factor graph의 일등시민이 되었다. Ch.7의 ORB-SLAM3, Ch.8의 VI-DSO, Ch.17의 LIO-SAM은 이 수식을 "Forster 2016을 썼다"로 처리했다. FAST-LIO는 raw IMU 전파와 iterated filter update를 쓰는 다른 경로다. --- ## 7b.1 MEMS와 센서의 민주화 Preintegration이 필요해진 이유는 IMU가 싸졌기 때문이다. 스트랩다운 관성항법의 뿌리는 1950년대 항공우주에 있다. 잠수함·미사일의 ring laser gyro는 수만 달러 장비였고 로봇공학 커뮤니티가 쓸 일은 없었다. 흐름을 바꾼 것은 MEMS(Micro-Electro-Mechanical Systems)였다. Analog Devices의 ADXL, InvenSense의 MPU 시리즈가 6축 IMU를 수 달러로 끌어내렸다. iPhone에 IMU가 들어간 것이 2007년, 2010년대 초반에는 연구용 드론·핸드헬드 장비가 당연히 MEMS IMU를 달았다. 스마트폰 수십억 대가 단가를 떨어뜨리는 시점과 Visual SLAM이 monocular scale ambiguity(Ch.5 §🧭)를 진지하게 고민하는 시점이 겹쳤다. 측정 모델은 단순하다. 가속도계는 중력이 섞인 specific force $\tilde{\mathbf{a}} = \mathbf{R}_w^b(\mathbf{a}^w - \mathbf{g}^w) + \mathbf{b}^a + \boldsymbol{\eta}^a$를, 자이로스코프는 angular velocity $\tilde{\boldsymbol{\omega}} = \boldsymbol{\omega}_b^b + \mathbf{b}^g + \boldsymbol{\eta}^g$를 준다. 여기서 $\mathbf{b}$는 bias, $\boldsymbol{\eta}$는 white noise다. 중력이 항상 섞이고, bias는 시간에 따라 천천히 떠다니며(random walk), MEMS 노이즈는 고주파다. IMU를 쓸 때에는 중력 방향을 추정하고 world frame을 그 방향에 정렬하는 경우가 많다. IMU는 온도·전원 상태마다 bias가 조금씩 달라지는 까다로운 동반자였다. --- ## 7b.2 첫 시도 — Lupton & Sukkarieh (2009 / 2012) 문제는 factor graph의 시간 축이었다. Ch.6가 정리한 Kaess의 iSAM2는 키프레임 단위의 pose를 노드로 삼는다. 그런데 IMU는 키프레임 사이에 수십 번 측정을 던진다. 이 측정을 전부 노드로 만들면 그래프가 폭발하고, 버리면 정보가 사라진다. Lupton과 Sukkarieh의 [Visual-Inertial-Aided Navigation for High-Dynamic Motion (IROS 2009, TRO 2012)](https://doi.org/10.1109/TRO.2011.2170332)이 내놓은 답은 우회였다. 키프레임 $i$에서 $j$ 사이의 IMU 측정을 *한 번만* 수치 적분해 상대 증분을 만들어 놓자. 그 증분을 하나의 factor로 삼으면 IMU 원측정은 그래프에 들어갈 필요가 없다. "pre-integration"이라는 이름이 여기서 나왔다. 아이디어는 맞았지만 구현에 두 장애물이 있었다. 회전 표현이 Euler angle이었다(gimbal lock이 있고 매니폴드가 아니다). 더 큰 문제는 bias였다. BA 한 번 돌 때마다 bias 추정치가 바뀌고, bias가 바뀌면 증분도 달라진다. 증분을 매번 재적분하면 키프레임당 수백 개 측정을 다시 처리해야 한다. Lupton도 이를 줄이기 위한 1차 편향 보정을 제안했다. 6년 뒤 Forster는 회전의 매니폴드 구조를 다루면서 그 보정을 SO(3) 위에서 정식화했다. --- ## 7b.3 변곡점 — Forster-Carlone (2015 / 2017) 2015년 RSS, ETH Zürich의 박사과정생 Christian Forster가 Scaramuzza(UZH), Carlone(Georgia Tech, 이후 MIT), Dellaert(Georgia Tech, GTSAM의 창시자)와 함께 [IMU Preintegration on Manifold for Efficient Visual-Inertial Maximum-a-Posteriori Estimation](https://www.roboticsproceedings.org/rss11/p06.pdf)을 냈다. 2017년 IEEE TRO에 확장판 [On-Manifold Preintegration for Real-Time Visual-Inertial Odometry](https://doi.org/10.1109/TRO.2016.2597321)가 실렸다. UZH의 민첩한 드론 실험, Georgia Tech의 GTSAM factor graph 언어, Carlone의 최적화 이론이 한 논문에서 만났다. 재정의가 셋이었다. $\Delta\mathbf{R}_{ij}$를 SO(3) 매니폴드 위 상대 회전으로 엄밀히 정의하고, $\Delta\mathbf{v}_{ij}, \Delta\mathbf{p}_{ij}$를 *중력과 초기 상태에 독립*이 되도록 재정의한 것이 첫째였다. 이 양들은 물리적 증분이 아니라 수학적으로 state-independent하게 만들어진 양이다. 사전적분량을 다시 계산하지 않고도 IMU factor를 평가할 수 있지만, 잔차에는 양 끝 pose와 velocity뿐 아니라 중력과 편향 보정도 들어간다. 둘째, noise를 지수사상 끝으로 밀어내는 right Jacobian trick으로 공분산 $\boldsymbol{\Sigma}_{ij}$를 해석적으로 propagate했다. 셋째는 **Bias 1차 Jacobian 선형 보정**이었다. BA 반복 중 bias가 조금 바뀌었을 때 증분 전체를 재적분하지 말고 precomputed 편미분으로 1차 근사 보정하자는 것이다. Lupton의 Euclidean 선형화와 같은 아이디어지만 SO(3) 위에서 작동한다. 키프레임 사이를 처음 적분할 때 편미분을 함께 계산해 두면, 선형화점이나 bias가 조금 바뀔 때 측정값 전체를 다시 적분하지 않고 1차 보정으로 factor를 갱신할 수 있다. 이 계산 절감 덕분에 IMU factor를 실시간 최적화 안에서 다루기 쉬워졌다. GTSAM에 Forster의 구현이 레퍼런스로 올라갔다. 후속 시스템들은 수식을 다시 쓰지 않고 `ImuFactor`를 `#include`했다. > 🔗 **공통 언어.** Forster의 manifold preintegration은 SO(3)의 지수사상과 right Jacobian으로 회전 불확실성을 다룬다. [Barfoot 2017. *State Estimation for Robotics*](https://doi.org/10.1017/9781316671528)는 뒤이어 이 Lie group 계산을 상태 추정의 공통 체계로 정리했다. 2017년 책이 2015년 논문의 원전인 것은 아니지만, 두 작업은 같은 수학적 언어를 보여 준다. Forster의 기여는 Lupton의 preintegration 아이디어를 Euler angle이 아닌 SO(3) 위에서 전개한 데 있다. > 🔗 **차용.** Bias 1차 Jacobian 아이디어 자체는 [Lupton & Sukkarieh 2012](https://doi.org/10.1109/TRO.2011.2170332)가 먼저 제시했다. Forster et al. TRO 2016 §VIII-B는 이 부채를 명시하며 "we follow [Lupton-Sukkarieh] but operate directly on SO(3)"라 적었다. Euclidean 근사를 매니폴드 위로 옮기자 같은 수학이 실시간이 되었다. --- ## 7b.4 실무 VIO 3파의 정립 Forster 공식이 자리 잡자 2017-2022년 사이 VIO 시스템이 세 갈래로 뻗었다. 첫 갈래는 필터 계열이고 뿌리는 Forster보다 앞서 있다. 2007년 UC Riverside의 Anastasios Mourikis와 Stergios Roumeliotis가 ICRA에 낸 [MSCKF(Multi-State Constraint Kalman Filter)](https://doi.org/10.1109/ROBOT.2007.364024)가 출발점이었다. stochastic cloning으로 과거 여러 카메라 포즈를 필터 상태에 유지한다. 관측한 3D 점은 상태에 추가하지 않고, 선형화한 관측식을 점 Jacobian의 left null space로 투영해 점 위치 오차에 대한 의존성을 제거한다. preintegration 없이 EKF 뼈대로 Visual-Inertial을 실시간으로 돌린 초기에 큰 영향을 준 사례였다. 2021년 화성에서 NASA JPL의 Mars helicopter Ingenuity가 돌린 추정기가 MSCKF 계열이었다. University of Delaware의 Guoquan Huang 그룹이 2020년 [OpenVINS](https://doi.org/10.1109/ICRA40945.2020.9196524)로 오픈소스화했다. 두 번째 갈래는 최적화 계열이다. HKUST의 Shaojie Shen 그룹과 박사과정생 Tong Qin이 2018년 TRO에 낸 [VINS-Mono](https://doi.org/10.1109/TRO.2018.2853729)가 대표작이다. Forster 공식을 그대로 받아 sliding-window tightly-coupled BA 안에 IMU factor로 심고, 초기화 단계에서 scale과 gravity 방향을 분리 추정하는 절차를 정리했다. 코드가 공개되어 2019-2022년 학회의 VIO baseline이 됐다. Ch.7의 ORB-SLAM3가 EuRoC 11개 시퀀스 평균 ATE 0.043m로 보고될 때 같은 표에서 VINS-Mono는 0.110m였다. 세 번째 갈래는 direct 계열이다. Ch.8에서 다룬 VI-DSO(2018), Basalt(2019), [DM-VIO(2022)](https://doi.org/10.1109/LRA.2021.3140129)가 여기 속한다. TUM Cremers 그룹이 DSO의 photometric BA 위에 Forster의 inertial factor를 얹은 적층 구조였다. DM-VIO는 *delayed marginalization*을 더했다. IMU 초기화가 수렴하기 전에 섣불리 marginalize하면 잘못된 prior가 고정돼 장기 drift를 유발하는데, 두 marginalization prior를 병렬로 유지하다 gravity와 scale이 관측된 뒤 최종 prior로 합치는 방식이다. --- ## 7b.5 Observability — 무엇을 못 보는가 Huang 그룹이 2010년대 초부터 정리한 분석에 따르면, 단서 없는 visual-inertial 시스템의 null space는 **4차원**, 즉 3차원 global position과 1차원 yaw-around-gravity다. 절대 좌표와 중력 축 회전은 IMU와 카메라만으로는 영원히 알 수 없다. GPS를 더하면 position이, 자기장이나 외부 anchor를 더하면 yaw가 복원된다. 순수 VIO는 이 4차원 부분공간을 구조적으로 볼 수 없다. *roll과 pitch는 보인다*. 가속도계가 중력으로 수평을 읽기 때문이다. IMU를 붙이면 Ch.5가 지적한 monocular scale ambiguity도 해결된다. 더 까다로운 쪽은 degenerate motion이다. 전역 yaw의 비관측성은 운동에 관계없이 남으며, 시차나 관성 운동의 변화가 부족하면 깊이·scale·편향을 추가로 구별하기 어려워질 수 있다. 드론의 hover나 자동차의 일정 속도 직진에서는 이런 운동 자극 부족을 살펴야 한다. 이륙·제동·선회가 정보를 더할 수 있지만, 특정 동작 하나가 scale 추정을 보장하지는 않는다. > 📜 **예언 vs 실제.** Forster et al.은 2017년 TRO §IX에서 세 방향을 꼽았다. time-synchronization과 online extrinsic calibration의 통합, long-term operation에서 bias random walk 가정 검증, event camera·rolling shutter 같은 비동기 센서로의 확장. 2026년 시점에서 첫 번째는 VINS-Mono·Kalibr·OpenVINS가 시간 offset을 상태 변수로 올리며 표준화됐고, 두 번째는 navigation-grade IMU에서는 맞지만 consumer MEMS에서는 온도·전원 변동이 여전히 남았으며, 세 번째는 Le Gentil의 GP 연속시간 preintegration이 답의 한 갈래가 되었다. 예언은 대체로 적중했으나 저자들이 그린 단일 확장이 아니라 세 갈래로 분화했다. --- ## 7b.6 Continuous-time 분기 2021년 RSS, 시드니의 UTS(University of Technology Sydney)에서 Cédric Le Gentil과 지도교수 Teresa Vidal-Calleja가 [Continuous Integration over SO(3) for IMU Preintegration](https://roboticsproceedings.org/rss17/p075.pdf)을 냈다. 같은 시드니였다. Lupton의 ACFR에서 몇 km 떨어진 곳에서 같은 문제를 다른 각도로 다시 본 셈이다. Forster의 preintegration은 discrete하다. IMU 측정 사이를 piecewise-constant로 가정하고 Euler integration한다. LiDAR·event camera처럼 비동기 센서가 섞이면 IMU 표본 사이 시각의 상태를 평가해야 한다. 비동기성 자체가 구간 상수 근사를 깨뜨리는 것은 아니지만, 시간 보간의 정확도가 중요해진다. Le Gentil의 답은 IMU를 **Gaussian Process**로 모델링해 angular velocity를 연속 함수로 본 것이다. 임의의 시간 $\tau$에서 상태를 평가할 수 있으니 비동기 측정이 자연스럽게 들어온다. 이 방향은 B-spline·STEAM·GPMP 계보와 만난다. --- ## 7b.7 차용의 지형 > 🔗 **차용.** Factor graph 위에서 IMU factor를 평가·최적화하는 골격은 Ch.6에서 정리한 [Dellaert의 GTSAM](https://gtsam.org/) 전통 그대로다. Forster의 `ImuFactor`는 GTSAM의 `NoiseModelFactor` 인터페이스에 꽂혀 visual reprojection factor와 나란히 하나의 `Values` 객체로 최적화되었다. 소프트웨어 구조의 상속이었다. > 🔗 **차용.** Bias를 random walk로 다루는 방식은 Ch.4의 Kalman filter state propagation 관습에서 왔다. Lupton 이전부터 항법 커뮤니티가 "bias를 상태에 포함하고 process noise를 작게 주는" 모델을 썼고, preintegration 시대에는 이것이 bias random walk factor로 재해석됐다. --- ## 🧭 아직 열린 것 **Visual-inertial observability의 실시간 감지.** 4차원 null space와 degenerate motion 표는 이론적으로 정리됐지만, 실제 시스템이 "지금 내가 degenerate 구간에 있다"를 판단하는 메커니즘은 미완성이다. Hesch·Li·Huang 계열의 FEJ(First-Estimate Jacobian)가 선형화 시점의 null space를 보존하지만, 런타임에 degenerate 조건의 시작·종료를 포착해 제어 루프에 피드백하는 널리 합의된 방법은 2026년 기준 없다. 드론 제어와 VIO 추정이 계산 자원을 공유할 때 이 공백은 추정 실패를 제어기가 늦게 알아차리는 위험으로 이어진다. **Preintegration과 continuous-time의 통합.** Forster의 이산 증분과 Le Gentil의 GP 연속표현은 같은 문제를 다른 수학 언어로 푼다. LiDAR·event·frame 카메라를 섞어 쓸 때 어떤 표현을 밑바닥에 깔 것인가는 아직 엔지니어링 선택의 문제다. B-spline 연속시간 BA가 부분적 답을 내놓았지만 배포 시스템 다수는 여전히 Forster의 이산 factor를 쓴다. **학습 기반 IMU bias 모델.** Bias random walk 가정은 navigation-grade IMU에서는 맞지만, consumer MEMS에서는 온도 hysteresis와 전원 과도 현상 때문에 어긋난다. TLIO·RoNIN 계열이 LSTM·Transformer로 IMU-only odometry의 bias를 학습했고, 최근에는 conditional diffusion으로 bias 분포 자체를 모델링하는 시도가 나왔다. 이 접근을 Forster factor 안에 어떻게 넣을지, 학습 모델을 결합한 뒤에도 preintegration의 수학적 우아함이 어디까지 유지될지는 다음 질문이다. --- Lupton이 시드니에서 시작한 아이디어가 6년 동안 Euler angle의 벽에 갇혀 있었고, Forster가 SO(3)로 옮겨 bias Jacobian의 자물쇠를 풀었고, Le Gentil이 다시 시드니에서 연속시간으로 가지를 쳤다. 세 세대의 작업이 ORB-SLAM3, VI-DSO, LIO-SAM의 한 줄 뒤에 쌓여 있다. --- # Ch.7c — 시간이 매끈하게 흘러야 할 때: Continuous-Time Trajectory Ch.7b가 정리한 preintegration은 IMU 측정을 이산 키프레임 사이의 relative factor로 압축하는 공학이었다. 그 압축은 "키프레임"이라는 단위를 전제로 성립한다. 두 키프레임 사이에 100번 들어온 관성 샘플이 하나의 factor로 접히려면, factor 양 끝점의 시각이 명확해야 한다. 카메라 셔터가 글로벌하게 한 번 열리고 닫히는 시스템에서는 그 가정이 무해하다. 문제가 생기는 자리는 따로 있었다. 2012년 [Paul Furgale·Timothy Barfoot·Gabe Sibley](https://asrl.utias.utoronto.ca/~tdb/bib/furgale_iros12.pdf)는 ICRA 논문에서 이 질문을 확률적 batch estimation 문제로 정식화했다. rolling shutter로 찍은 한 장의 이미지에서 각 행은 다른 시각의 자세로 투영된다. spinning LiDAR가 한 바퀴를 도는 사이에도 차량은 움직이고, IMU와 카메라는 서로 다른 속도로 측정값을 낸다. 이 센서들을 하나의 최적화로 묶는 방법은 자세를 "프레임"이 아니라 "시간 t의 함수"로 두는 것이었다. 세 저자는 유한한 temporal basis function의 계수를 상태로 추정했고, 이 논문은 continuous-time trajectory estimation을 로봇 센서 융합에 정식화한 초기의 영향력 있는 작업 가운데 하나가 됐다. 이후 continuous-time trajectory는 rolling-shutter 카메라, spinning LiDAR, event camera처럼 한 프레임이나 한 scan 안에서도 측정 시각이 달라지는 센서를 다루는 도구로 자리 잡았다. 앞 장의 visual-inertial 문법은 대부분 discrete keyframe 위에서 완성됐지만, 여기서는 그 틀로 다루기 어려운 센서의 시간을 잇는다. --- ## 7c.1 Discrete-time의 한계 Ch.7b의 preintegration이 해결한 것은 "IMU가 카메라보다 빠르다"는 단일 축이었다. 해결하지 못하는 축은 네 개가 더 있다. 첫째, rolling shutter. consumer CMOS 카메라는 한 프레임을 위에서 아래로 일정 시간에 걸쳐 읽는다. 빠르게 움직이는 카메라에서 첫 행과 마지막 행은 서로 다른 자세에서 찍힌다. Ch.8의 DSO·LSD-SLAM이 한 프레임에 하나의 자세를 가정할 때 이 왜곡은 모델 밖에 있었다. Continuous-time 궤적을 쓰면 각 행의 실제 노출 시각에서 자세를 질의할 수 있다. 둘째, spinning LiDAR motion distortion. Ch.17에서 보았듯 Velodyne HDL-64E는 10 Hz로 한 바퀴를 돈다. 그 100 ms 사이에 차량이 10 m/s로 달리면 한 scan을 얻는 동안 차량은 총 1 m를 이동하고, 각 점은 그 구간의 서로 다른 자세에서 찍힌다. LOAM은 이 왜곡을 odometry 루프 안에서 간접 보정했지만, 원리적 해법은 "각 점이 찍힌 순간의 자세"를 질의할 수 있는 궤적 표현이었다. 셋째, event camera. Ch.18이 기록한 DVS는 픽셀마다 μs 단위로 비동기 이벤트를 쏟는다. 이벤트에는 "프레임"이 없다. [Mueggler et al. 2015](https://arxiv.org/abs/1502.00796)는 SE(3) B-spline 궤적 위에서 event SLAM을 정식화했다. 넷째, high-rate IMU를 이질 주파수 센서 여럿과 동시 결합하는 일. 한 시스템에 200 Hz IMU, 20 Hz 카메라, 10 Hz LiDAR가 들어오면, 이산 상태 노드를 모든 측정 시각마다 두는 것은 현실적이지 않다. 상태의 수가 측정의 수를 따라가는 순간 factor graph는 부풀어 오른다. 네 문제의 구조는 같다. 측정 시각 $t_i$가 제어되지 않는다. 관측은 아무 때나 들어오고, 추정기는 그 시각의 자세를 알아야 한다. "측정 시각 / 추정 시각 / 질의 시각"의 분리가 continuous-time representation의 이점이다. --- ## 7c.2 Parametric spline: Furgale 계보 Furgale·Barfoot·Sibley가 2012년에 고른 도구는 B-spline이었다. 궤적을 basis function의 합 $\mathbf{p}(t) = \sum_k \Psi_k(t)\,\mathbf{c}_k$로 쓰고, 계수 $\mathbf{c}_k$를 최적화 변수로 둔다. B-spline의 핵심은 local support다. 한 시점 t에서 0이 아닌 basis는 소수(보통 4개)뿐이고, 나머지는 정확히 0이다. 임의 시각 $t_i$의 자세를 질의하는 비용이 상수이고, factor graph sparsity가 그대로 유지된다. > 🔗 **차용.** B-spline의 수학적 뼈대는 [de Boor (1978) *A Practical Guide to Splines*](https://link.springer.com/book/10.1007/978-1-4612-6333-3)의 고전이다. Furgale이 한 일은 그 뼈대를 SE(3) 위로 끌어올리고, 계수를 factor graph의 변수 노드로 배치한 것이다. 수치해석 교과서의 도구가 SLAM 최적화로 이식된 경로다. 계수 간격을 좁게 잡으면 과적합하고, 넓게 잡으면 빠른 움직임을 놓친다. 간격 선택은 경험에 의존했다. linear B-spline을 SE(3)에 그대로 얹으면 보간 결과가 매니폴드를 벗어난다. 2013년 Oxford의 [Steven Lovegrove et al.](https://www.roboticsproceedings.org/rss09/p11.html)가 cumulative B-spline을 제안했다. basis를 합이 아니라 누적 곱 형태 $T(t) = \prod_k \exp\bigl(\tilde\Psi_k(t) \log(T_k T_{k-1}^{-1})\bigr) \cdot T_0$로 재배치하면 각 인자가 Lie group에 닫혀 있다. 이 형식은 이후 rolling-shutter·event camera·VIO 논문에서 널리 쓰이는 표현이 됐다. Basalt, [Mueggler event SLAM](https://arxiv.org/abs/1502.00796), [Kerl et al. 2015 dense rolling shutter VO](https://doi.org/10.1109/ICCV.2015.172)가 cumulative B-spline을 사용했다. Parametric spline은 계산이 가볍고 코드가 단순하다는 이점 때문에 실시간 VIO와 event 시스템에서 꾸준히 쓰였다. 별도 motion prior가 없는 spline에서는 관측이 드문 구간의 매끄러움이 basis와 제어점에 의존한다. 계수나 미분에 prior를 추가할 수도 있지만, GP 기반의 다른 갈래는 궤적의 사전 분포를 구성의 출발점으로 삼는다. --- ## 7c.3 SDE 기반 GP: Barfoot 계보와 STEAM 2014년, 토론토의 Barfoot 그룹이 두 번째 갈래를 열었다. [Barfoot, Tong, Särkkä 2014 "Batch Continuous-Time Trajectory Estimation as Exactly Sparse Gaussian Process Regression"](https://www.roboticsproceedings.org/rss10/p01.pdf)은 제목이 그대로 주장이었다. 궤적을 basis 합이 아니라 Gaussian process로 두겠다. 궤적의 사전 분포는 kernel $\mathcal{K}(t, t')$로 주어지고, 관측이 들어오면 posterior를 조건부 Gaussian으로 닫는다. GP의 순수 형태는 문제가 하나 있다. 관측 수 $N$이 크면 kernel matrix $K$의 역행렬 비용이 $O(N^3)$이다. Barfoot·Tong·Särkkä는 이 비용을 회피하는 kernel의 가족을 보였다. 궤적이 linear time-invariant stochastic differential equation $\dot{\mathbf{x}}(t) = A\mathbf{x}(t) + L\mathbf{w}(t)$의 해로 정의될 때, 그 kernel $K$의 역행렬 $K^{-1}$은 block-tridiagonal 구조를 가진다. factor graph로 읽으면 연속한 상태 노드 사이에만 binary factor가 있고, 멀리 떨어진 노드 사이에는 factor가 없다. > 🔗 **차용.** "SDE에서 유도한 GP motion prior를 factor graph로 표현한다"는 틀은 [Särkkä 2013 *Bayesian Filtering and Smoothing*](https://users.aalto.fi/~ssarkka/pub/cup_book_online_20131111.pdf)이 정리한 SDE-GP 연결을 Barfoot 그룹이 SLAM으로 끌어온 것이다. Rasmussen-Williams의 GP 교과서는 kernel을 닫힌 형식으로 쓰지만, 실시간 SLAM은 sparse inverse를 원한다. Särkkä의 SDE 표현이 그 다리였다. 실무 귀결이 **STEAM** (Simultaneous Trajectory Estimation and Mapping)이다. 2015년 RSS에서 [Sean Anderson·Barfoot 2015 "Full STEAM Ahead"](https://www.roboticsproceedings.org/rss11/p45.pdf)가 constant-velocity prior 기반 STEAM을 공식화했다. 상태를 자세 $\mathbf{p}(t)$와 속도 $\mathbf{v}(t)$로 augment하고, 속도의 white noise 적분으로 자세가 따라가는 구조다. Anderson은 같은 해 sparsity 증명을 더 엄밀한 형태로 완성했고, 그 증명은 이후 Barfoot 그룹의 모든 continuous-time 논문을 떠받치는 토대가 됐다. STEAM의 두 번째 이점은 GP interpolation이었다. 제어점(control pose)을 소수만 두고, 제어점 사이 임의 시각의 자세를 posterior mean으로 질의할 수 있다. spinning LiDAR의 한 scan 안에서 10,000개의 점이 각자 다른 시각에 찍혀도, 제어점은 scan당 하나만 둔다. 추정할 상태 수를 관측 수보다 훨씬 작게 유지할 수 있다. 다만 각 관측의 잔차를 평가하고 누적하는 계산은 여전히 필요하다. 2019년 Tang·Barfoot의 [STEAM 오픈소스](https://github.com/utiasASRL/steam)가 공개되면서 학계·산업계에서 직접 쓸 수 있는 라이브러리가 됐다. 같은 해 Dellaert 그룹의 GTSAM에도 GP continuous-time factor가 contrib로 들어갔다. --- ## 7c.4 Lie group 위의 continuous-time Parametric이든 nonparametric이든 SLAM은 SE(3) 위의 궤적을 원한다. Euclidean 상의 spline·GP를 SE(3)로 끌어올리는 일은 기술적으로 간단하지 않다. tangent space에서 선형 보간을 한 뒤 exponential map으로 매니폴드에 얹는 방식이 통용된다. B-spline 쪽에서는 [Sommer, Demmel et al. 2020 "Efficient Derivative Computation for Cumulative B-Splines on Lie Groups"](https://arxiv.org/abs/1911.08860)가 SE(3) cumulative spline의 Jacobian을 닫힌 형식으로 정리했다. CVPR에 실린 이 논문은 rolling-shutter VIO·event camera·visual-inertial 시스템에서 실시간 미분이 가능한 B-spline 궤적의 표준 공식을 제공했다. Basalt와 Cremers 그룹 후속 작업이 이 정리 위에 섰다. GP 쪽에서는 Anderson·Barfoot이 "local variable" 구도를 제안했다. 각 제어 자세 $T_k$ 근처에서 local perturbation $\xi_k(t) = \log(T(t)\,T_k^{-1})$를 정의하고, 그 위에서 GP를 운용한다. 전역 매니폴드 위에서 직접 GP를 정의하는 것은 어렵지만, 각 제어점 근방의 tangent space에서는 Euclidean GP가 성립한다. 제어점 사이를 건너뛸 때 adjoint가 등장하는데, 그 수학적 근거는 Ch.7b preintegration의 on-manifold 논의와 같다. 두 도구가 같은 Lie group 문법을 공유한다는 사실이 2015년 이후 분명해졌다. > 🔗 **차용.** [Anderson-Barfoot 2015 ICRA](https://doi.org/10.1109/ICRA.2015.7138984)는 GP를 Lie group local variable에 적용하는 구도를 체계적으로 정리했다. 이들이 쓴 트릭("연속한 두 제어점 사이에서만 GP를 돌리고, 제어점 사이를 건너뛸 때 adjoint로 보정")은 이후 여러 continuous-time LiDAR·VIO 논문으로 이어졌다. 여기서 비교한 spline과 GP 구현의 차이 중 하나는 motion prior를 두는 방식이다. spline의 계수를 직접 추정할 수도 있고, 계수나 미분에 별도 사전 분포를 줄 수도 있다. GP는 SDE에서 유도된 사전 분포가 constant-velocity 혹은 white-jerk 등으로 내장돼 있다. 관측이 드문 구간에서 GP는 prior가 채우고, spline은 인접 관측이 채운다. 둘을 결합하려는 시도(Johnson et al. 2020)도 있었지만, 실무에선 응용에 따라 한쪽을 고른다. --- ## 7c.5 응용으로 내려온 계보: LiDAR와 VIO 초기 calibration 응용을 거쳐 활용 범위가 넓어졌다. 2022년 전후로 continuous-time은 세 현장에서 자주 쓰이는 해법이 됐다. 첫째, LiDAR motion distortion. Paris의 [Pierre Dellenbach et al. 2022 "CT-ICP"](https://arxiv.org/abs/2109.12979)는 각 scan을 "시작 자세"와 "끝 자세" 두 개로 파라미터화하고 그 사이를 선형 보간했다. 간단한 continuous-time 모델이지만, KITTI·NCLT·Newer College 벤치마크에서 기존 LOAM·FAST-LIO의 정확도를 상회했다. 같은 해 Toronto의 [Keenan Burnett et al. 2022 "Are We Ready for Radar to Replace Lidar?"](https://arxiv.org/abs/2206.05432)와 [STEAM-ICP](https://github.com/utiasASRL/steam_icp)가 GP 기반 continuous-time을 Aeva FMCW LiDAR에 적용했다. Aeva 센서가 각 점마다 도플러 속도를 함께 출력하는데, 이 속도는 STEAM의 속도 상태와 직접 대응한다. 둘째, rolling-shutter VIO. Basalt·[Cremers 그룹 rolling-shutter VO](https://doi.org/10.1109/CVPR.2016.71)·[OKVIS](https://doi.org/10.1177/0278364914554813) 후속작들이 이미지 각 행이 찍힌 시각을 B-spline 궤적에 질의한다. 글로벌 셔터를 가정하고 우회하는 기존 VIO와 달리 rolling shutter 자체를 모델 안에서 처리한다. 셋째, event camera. Ch.18이 기록한 2010년대의 좌절 이후, 2020년대에는 continuous-time 궤적을 쓰는 event SLAM이 여러 갈래로 나왔다. 각 이벤트의 μs 타임스탬프를 B-spline 혹은 GP에 질의해 그 순간의 자세를 얻고, event-image consistency로 residual을 계산한다. event가 "프레임이 없는 센서"라는 사실과 continuous-time이 "프레임 가정이 필요 없는 표현"이라는 사실이 자연스럽게 맞물렸다. > 🔗 **차용.** CT-ICP는 [Besl·McKay 1992 ICP](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf)에서 이어진 ICP 계열의 point-to-plane 목적함수에 scan 내부 continuous-time linear 보간을 얹은 조합이다. 고전 registration과 Furgale의 continuous-time 정신이 30년의 간격을 두고 한 시스템에서 만났다. --- ## 📜 예언 vs 실제 > 📜 **예언 vs 실제.** Furgale·Barfoot·Sibley의 2012년 ICRA 논문은 고속 IMU와 sweeping laser가 discrete-time state를 비대하게 만드는 문제를 출발점으로 삼았고, 기존 rolling-shutter 보간 연구도 관련 사례로 언급했다. 논문이 직접 검증한 것은 camera–IMU self-calibration이었다. 이후 Barfoot·Tong·Särkkä(2014)가 SDE 기반 sparse GP를 정식화했고, rolling-shutter VIO와 event SLAM에서도 continuous-time trajectory가 쓰였다. 이는 후속 계보이지, 2012년 논문의 Future Work가 두 방향을 직접 예고했다는 뜻은 아니다. 이후 discrete keyframe 기반 ORB-SLAM이 주류를 이루고 continuous-time이 특수 센서에 주로 쓰이는 분업이 굳어졌지만, Burnett의 STEAM-ICP처럼 FMCW LiDAR의 도플러 속도를 continuous-time state와 결합한 전개도 나왔다. --- ## 🔗 차용 요약 이 장의 밑에는 세 계보가 더 놓여 있다. Särkkä의 SDE-GP 교과서가 Barfoot·Tong·Särkkä 2014의 방정식을 받쳤고, de Boor의 1978년 spline 교과서가 Furgale 2012의 basis function을 제공했다. Anderson·Barfoot의 2015년 local-variable 기법은 GP를 Lie group 위로 옮겼다. Continuous-time trajectory estimation은 수치해석·확률론·Lie group 미분기하가 SLAM 안에서 만난 결과다. --- ## 🧭 아직 열린 것 **Learning-based continuous-time prior.** SDE가 주는 motion prior는 constant-velocity나 white-jerk 같은 물리 가정을 내장한다. 실제 주행·보행·UAV 궤적은 이 가정을 어기는 경우가 많다. 2023-2024년 neural SDE나 neural ODE로 데이터 기반 prior를 학습해 continuous-time factor graph에 꽂으려는 시도들이 나왔다. 아직 실시간 sparse 구조를 유지한 채 learned prior를 얹은 시스템은 검증 단계다. **VIO와 continuous-time의 통합.** Ch.7b의 preintegration은 keyframe 기반 VIO의 사실상 표준으로 남아 있다. continuous-time 궤적이 preintegration을 대체할 수 있는지, 혹은 두 도구가 공존하는 하이브리드가 더 나은지는 2026년 기준 결론이 없다. Le Gentil의 [GP-augmented preintegration 계보](https://arxiv.org/abs/2007.04144)가 한 다리를 놓으려 시도 중이지만, ORB-SLAM3·VINS-Fusion 수준의 배포 시스템에서는 여전히 discrete-time preintegration이 주력이다. **Edge deployment를 위한 online sliding window.** STEAM과 B-spline 기반 시스템은 제어점 수가 누적되면 최적화가 느려진다. marginalization으로 과거 제어점을 제거하면서 continuous-time posterior의 일관성을 유지하는 문제는 기술적으로 까다롭다. 자동차·드론 같은 임베디드 플랫폼에서는 이 공백이 먼저 메워져야 한다. --- Ch.7b가 discrete-time preintegration의 효율을 밀어붙였다면, 여기서는 continuous-time이라는 다른 시간 표현의 계보를 따라왔다. 두 도구는 경쟁하지 않는다. 하나의 SLAM 시스템에 IMU preintegration factor와 continuous-time LiDAR factor가 나란히 들어가는 구성이 2024년 이후 꾸준히 보고되고 있다. Ch.8은 다시 시각 계보로 돌아가, DSO와 VI-DSO가 Forster factor를 쓰는 한편 direct photometric 접근이 별도의 결론에 이르는 과정을 다룬다. --- # Ch.8 — Direct 계보: DTAM에서 DSO까지 Richard Newcombe는 Andrew Davison의 박사과정 학생이었다. MonoSLAM의 30-landmark 한계가 드러난 Imperial College에서, 그의 2011년 선택은 정반대였다. 모든 픽셀을 쓰기로 한 것이다. Davison이 "몇 개의 점만 추적하면 충분하다"는 EKF의 논리에 기대어 실시간을 증명했다면, Newcombe는 GPU 한 장을 얹고 화면 전체를 써도 실시간이 가능하다는 것을 보여줬다. DTAM은 MonoSLAM의 직계 후손이지만, 그 방법론적 DNA는 완전히 뒤집혀 있다. ORB-SLAM 계보는 feature를 먼저 뽑고 그 feature만 추적하는 방식이었다. Harris 코너와 ORB 디스크립터가 걸러낸 수백 개의 점만 남고 나머지 픽셀은 버려진다. Direct 계보는 대응점의 기하 오차 대신 픽셀 밝기의 차이를 직접 측정값으로 삼았다. 모든 픽셀을 쓰는 dense 방법부터 일부만 선택하는 sparse 방법까지 여기에 속한다. 같은 해 뮌헨에서는 Daniel Cremers가 다른 경로를 걷고 있었다. Computer vision의 variational 방법론(Gauss-Newton image alignment, 광학 흐름의 수식 언어)을 SLAM 전체에 이식하는 작업이었다. Cremers의 제자 Jakob Engel은 2014년 LSD-SLAM을, 2016년 DSO를 내놓았다. 두 논문은 서로 다른 밀도에서 같은 질문을 던졌다. feature를 추출하는 대신 픽셀의 밝기를 직접 비교하면 어떤 일이 생기는가. --- ## 1. 모든 픽셀: DTAM 2011년 ICCV에서 Newcombe와 공동저자 Lovegrove, Davison이 발표한 [Newcombe, Lovegrove & Davison 2011. DTAM](https://doi.org/10.1109/ICCV.2011.6126513)은 "Dense Tracking and Mapping in Real-Time"의 약자다. 이름 그대로 모든 픽셀을 사용해 추적과 지도 구축을 실시간으로 수행한다. 시스템은 두 부분으로 구성된다. 추적 단계에서는 현재 프레임 전체를 cost volume과 비교하는 photometric alignment를 수행한다. 특징점 추출이나 디스크립터 매칭 없이 픽셀 intensity의 차이만 최소화한다. 지도 구축 단계에서는 multi-baseline stereo 방식으로 depth map을 추정하고, total variation regularization으로 smooth한 dense 3D 모델을 유지한다. $$E(\mathbf{u}) = \sum_{i} \rho\left( I_i\bigl(\pi(KT_i\mathbf{p}(\mathbf{u}))\bigr) - I_r\bigl(\pi(\mathbf{p}(\mathbf{u}))\bigr) \right) + \lambda \,\text{TV}(\mathbf{u})$$ 여기서 $\mathbf{u}$는 역 깊이(inverse depth) 맵, $\mathbf{p}(\mathbf{u})$는 $\mathbf{u}$로 역투영한 3D 점, $K$는 카메라 내부 행렬, $T_i$는 참조 프레임 기준 $i$번 프레임의 rigid body 변환, $\pi$는 원근 투영, $\rho$는 Huber loss, $\text{TV}(\mathbf{u}) = \|\nabla \mathbf{u}\|_1$은 total variation regularizer이다. 이 최적화를 실시간으로 돌리려면 GPU가 필요하다. DTAM은 그 전제를 숨기지 않았다. 당시 Nvidia GTX 480 한 장(논문 §3의 commodity 시스템 설정)에서 실행되었다. > 🔗 **차용.** DTAM의 dense volumetric 접근은 depth camera 기반 연구, 특히 [Curless & Levoy 1996](https://doi.org/10.1145/237170.237269)의 TSDF 아이디어에서 부분 영감을 받았으나, 단안(monocular) 카메라에 적용했다는 점이 핵심 차이다. 이후 Newcombe 자신이 주도한 [KinectFusion](https://doi.org/10.1109/ISMAR.2011.6092378)(2011, ISMAR)이 오히려 depth sensor 버전으로 이 아이디어를 완성시키는 역방향 흐름이 나타난다. 실내 scene 전체가 실시간으로 복원되는 영상은 2011년 ICCV 발표 직후 YouTube에 공개되어 수만 회 조회를 기록했다. 그러나 GPU 없이는 돌아가지 않았고, 조명 변화에 취약했으며, 실외 대규모 환경으로는 확장되지 않았다. --- ## 2. 엣지의 추적: LSD-SLAM [Engel, Schöps & Cremers 2014. LSD-SLAM](https://doi.org/10.1007/978-3-319-10605-2_54)은 DTAM의 dense를 포기하는 대신 GPU 의존성도 함께 버렸다. "Large-Scale Direct Monocular SLAM"은 semi-dense 방식으로, 이미지에서 gradient magnitude가 임계값 이상인 픽셀만 추적한다. 벽의 평탄한 영역은 무시하고, gradient가 충분한 엣지 근방 픽셀만 살린다. 코너 detector는 쓰지 않으며, 오직 intensity gradient의 세기가 픽셀 선택 기준이다. 추적 단계는 SE(3)에서의 direct image alignment다. 현재 프레임을 키프레임에 direct로 warping하여 photometric residual을 Gauss-Newton으로 최소화한다. 지도는 키프레임 기반이며 각 키프레임마다 semi-dense depth map을 유지한다. 키프레임 간 연결은 pose graph로 관리하고, loop closure는 appearance-based relocalization으로 후보를 찾은 뒤 depth consistency check로 검증한다. > 🔗 **차용.** Gauss-Newton photometric registration은 이미지 정렬 분야의 고전이다. [Lucas & Kanade 1981](https://www.ijcai.org/Proceedings/81-2/Papers/017.pdf) tracker와 그 역방향 합성([Baker & Matthews 2004](https://doi.org/10.1023/B:VISI.0000011205.11775.fd))이 LSD-SLAM frontend의 직접 조상이다. Cremers 그룹은 variational image processing 커뮤니티의 언어를 SLAM 파이프라인 전체로 이식했다. CPU에서 실시간으로 동작한다는 점이 LSD-SLAM의 실용적 의미였다. 키프레임만 들고 pose graph를 최적화하는 구조는 PTAM의 tracking/mapping 분리와 표면적으로 닮았지만, 내부는 달랐다. ORB나 BRIEF 같은 binary descriptor가 없고, 픽셀 강도가 유일한 측정값이다. LSD-SLAM은 실외 대규모 환경에서도 동작하는 장면을 공개했다. 자전거를 타고 수십 미터를 이동하는 동안 semi-dense map이 구축되는 데모는 direct 방식의 확장 가능성을 보여줬다. KITTI 벤치마크에서 당시 상위권 feature-based 방법과 비교 가능한 수준이었다. 그러나 조명 변화가 문제였다. 터널 진입, 창문 역광, 갑작스러운 플래시에서는 photometric consistency 가정이 깨져 시스템이 즉시 불안정해졌다. --- ## 3. Sparse Direct의 완성: DSO [Engel, Koltun & Cremers 2018. DSO (PAMI)](https://doi.org/10.1109/TPAMI.2017.2658577)는 2016년 arXiv에 먼저 공개되었다. "Direct Sparse Odometry"는 LSD-SLAM보다 sparse하고 DTAM보다 훨씬 적은 픽셀을 쓰되, photometric calibration을 철저히 했다. 시스템은 각 키프레임에서 gradient가 높은 픽셀 약 2,000개를 선택한다. ORB-SLAM2의 기본 설정(nFeatures=1000)에 비해 많고, LSD-SLAM의 semi-dense(gradient 있는 픽셀 전체)보다 훨씬 적다. 이 픽셀들에 대해 sliding window bundle adjustment를 수행하는데, 최적화 변수가 camera pose뿐 아니라 inverse depth, affine brightness 파라미터 $(a_i, b_i)$까지 포함한다. 윈도우를 벗어난 프레임은 marginalization으로 제거되며, Schur complement로 제거할 변수를 소거해 남은 창의 최적화 크기를 제한한다. 전체 계산량은 창 크기와 픽셀 수, 반복 횟수에 달려 있다. DSO는 카메라 photometric 모델을 세 층으로 분리한다. 첫째, vignetting(렌즈 주변부로 갈수록 밝기가 감소하는 효과)은 사전 캘리브레이션으로 보정한다. 둘째, camera response function(gamma curve, 센서가 빛을 비선형으로 기록하는 특성)도 사전에 역함수를 추정해 linear intensity domain으로 변환한다. 셋째, 프레임의 노출 시간 $t_i$는 입력 메타데이터로 사용하고, 남은 affine brightness 변화 $(a_i, b_i)$를 실시간으로 추정한다: $$E_{pj} = \sum_{\mathbf{p} \in \mathcal{N}_p} w_{\mathbf{p}} \left\| \left( I_j\!\left[\mathbf{p}'\right] - \frac{t_j e^{a_j}}{t_i e^{a_i}} I_i[\mathbf{p}] - \left(b_j - \frac{t_j e^{a_j}}{t_i e^{a_i}} b_i\right) \right) \right\|_\gamma$$ 여기서 $t_i, t_j$는 노출 시간, $(a_i, b_i)$와 $(a_j, b_j)$는 각 프레임의 affine brightness 파라미터(gain과 bias), $\|\cdot\|_\gamma$는 Huber loss이다. Vignetting은 전처리 단계에서 photometric calibration으로 보정되며, 위 잔차는 보정된 intensity에 적용된다. 카메라의 노출 변화·vignetting·response curve를 별도 캘리브레이션 단계와 실시간 최적화 변수로 분리해 처리한 것은 direct SLAM에서 DSO가 처음이었다. > 🔗 **차용.** Photometric camera calibration의 형식적 기반은 [Debevec & Malik 1997](https://doi.org/10.1145/258734.258884)의 HDR 복원 작업에서 비롯된다. DSO도 camera response function을 보정한 밝기 값을 쓰지만, response function 자체는 사전 보정한다. 온라인으로 바꾸는 것은 프레임별 affine brightness 파라미터다. TUM monocular dataset에서 DSO는 ORB-SLAM2를 여러 시퀀스에서 능가한다고 보고했다. 특히 feature가 희박한 환경(평탄한 벽이 많은 실내 복도)에서 DSO가 ORB-SLAM2보다 낮은 ATE를 기록했다. photometric 정보를 직접 쓰면 원론적으로 더 많은 정보를 활용한다는 주장의 경험적 근거였다. > 📜 **예언 vs 실제.** DSO는 사전 photometric calibration을 요구했고, 그 의존성은 곧 후속 연구의 표적이 되었다. 2018년 Bergmann, Wang, Cremers의 [online photometric calibration](https://doi.org/10.1109/LRA.2017.2777002)은 SLAM 실행 중 노출 시간·response function·vignetting attenuation을 함께 추정했다. 논문은 auto-exposure video를 실시간으로 보정해, 평가 데이터에서 사전에 보정한 영상과 대등한 VO 정확도를 보고했다. DSO가 드러낸 사전 photometric calibration 의존성에는 이 후속 연구가 직접 답했다. > 📜 **예언 vs 실제.** DTAM은 GPU 한 장에 의존한 실시간 dense SLAM이었고, dense 재구성의 접근성을 넓히는 문제는 이후에도 남았다. 그 실현 경로는 직진이 아니었다. 순수 mono dense는 NeRF와 3DGS가 등장하는 2020년대까지 실시간 배포 가능한 형태로 나오지 않았다. 대신 Newcombe 자신이 주도한 KinectFusion이 RGB-D depth sensor를 사용해 GPU dense 재구성을 2011년에 바로 완성했다. sensor 교체로 문제를 우회한 것이다. --- ## 4. VI-DSO와 계보의 확장 2018년 von Stumberg, Usenko, Cremers는 DSO에 IMU를 결합한 [VI-DSO](https://doi.org/10.1109/ICRA.2018.8462905)를 ICRA 2018에서 발표했다. photometric direct method의 실패 모드인 조명 급변 상황에서 IMU의 관성 측정이 pose 추적을 보조할 수 있고, mono 카메라의 scale ambiguity도 IMU로 해소할 수 있다. VI-DSO는 DSO의 windowed photometric bundle adjustment에 IMU preintegration factor를 추가한다. IMU preintegration 방식은 [Forster et al.의 2017년 논문](https://doi.org/10.1109/TRO.2016.2597321)에서 차용했다. 결과적으로 scale이 복원되고 극단적 조명 조건에서 robustness가 향상되었다. Cremers 그룹의 후속 작업들, [Basalt](https://arxiv.org/abs/1904.06504)(2019)와 [DM-VIO](https://doi.org/10.1109/LRA.2021.3140129)(2022)도 같은 방향을 이었다. direct photometric frontend에 tightly coupled inertial backend를 붙이는 구조다. 이 계보는 feature-based VIO(VINS-Mono, OpenVINS)와 병렬로 진행되면서 각자의 생태계를 형성했다. > 🔗 **차용.** VI-DSO의 IMU preintegration은 [Forster et al. 2017. On-Manifold Preintegration (IEEE TRO)](https://doi.org/10.1109/TRO.2016.2597321)의 manifold preintegration 공식을 그대로 사용한다. DSO의 photometric layer 위에 Forster의 inertial layer가 올라간 적층 구조다. --- ## 5. direct method의 한계 Direct method는 descriptor로 요약한 특징점 대신 영상 밝기 자체를 잔차에 넣는다. 그렇다고 모든 픽셀이나 gradient가 거의 없는 픽셀까지 유효한 것은 아니다. DTAM은 dense한 광도 정보를 사용했고 LSD-SLAM은 semi-dense 영역을, DSO는 영상 전역에 고르게 퍼진 **충분한 intensity gradient**의 픽셀을 골랐다. 차이는 낮은 gradient를 무조건 쓰는 데 있지 않고, 별도의 descriptor 없이 밝기 변화가 주는 미분 가능한 잔차를 직접 최적화하는 데 있다. 그럼에도 2026년 기준 대다수 배포 시스템은 feature-based다. 첫째, photometric calibration 의존성이다. DSO가 가정하는 vignetting 보정, response curve 보정, 노출 제어는 consumer camera에서 그냥 얻어지지 않는다. 스마트폰 카메라는 HDR 합성, auto-exposure, 실시간 화이트 밸런스를 자체적으로 적용하며, 그 파이프라인은 사용자에게 공개되지 않는다. DSO의 photometric 모델은 이런 카메라에서 기본 가정이 깨진다. 둘째, 조명 변화다. 자동 노출, 역광, 플리커처럼 프레임 간 밝기가 급변하는 상황에서는 direct 방식의 핵심 가정인 photometric consistency가 바로 깨진다. DSO의 affine brightness 모델은 완만한 밝기 변동만 흡수할 수 있어, 실외에서 구름이 지나가거나 실내에서 형광등이 깜빡이는 장면은 여전히 추적 실패의 주 원인으로 남았다. 셋째, DSO가 controlled dataset에서 ORB-SLAM2를 이기는 시퀀스가 있어도, 실제 배포에서는 ORB-SLAM 계열이 더 널리 쓰였다. ORB-SLAM은 여러 카메라 모델에서 별도 photometric calibration 없이 동작한다. 카메라를 교체해도 바로 돌아간다. DSO는 카메라마다 vignetting·response curve를 따로 캘리브레이션해야 했다. 넷째, [SuperPoint](https://arxiv.org/abs/1712.07629)(2018)·[LightGlue](https://arxiv.org/abs/2306.13643)(2023) 같은 학습 기반 feature가 "feature는 정보를 버린다"는 direct method의 핵심 비판을 약화시켰다. 기존 handcrafted descriptor보다 훨씬 많은 정보를 보존하면서도 descriptor matching의 실용적 장점을 유지한다. direct method가 feature-based를 공격하던 그 지점에, learned feature가 자리를 메운 것이다. --- ## 🧭 아직 열린 것 **조명 급변 환경에서의 direct tracking.** Direct method의 근본 전제, 장면의 밝기 분포가 프레임 간 보존된다는 가정은 자동 노출 카메라, 강한 역광, 터널-야외 전환 상황에서 즉각 붕괴한다. VI-DSO의 IMU 보조가 부분적으로 완화하지만, 조명 모델 자체를 동적으로 추정하는 완전한 해법은 아직 없다. 학습 기반 photometric 보정이 대안으로 탐색 중이지만, 실시간 배포 가능한 형태로 나오지 않았다. **Textureless + direct의 이중 약점.** Feature-based는 코너가 없는 벽 앞에서 실패한다. Direct는 gradient가 없는 면에서 밝기 잔차가 자세 변화에 거의 민감하지 않다. 두 방식 모두 실내 복도, 대규모 창고, 균질한 실외 지형 같은 환경에서 약하다. Semi-dense LSD-SLAM은 gradient 있는 픽셀을 선택적으로 쓰는 방식으로 절충했지만, 그 픽셀이 충분히 분포하지 않는 상황의 degeneracy는 해결하지 못했다. **Learned photometric model로의 이행 가능성.** 현재 direct SLAM의 photometric 모델은 단순 affine brightness 보정이나 고정 camera response function으로 표현된다. Neural radiance field 계열의 연구들은 장면의 appearance를 neural network로 모델링하는 방식을 탐색하고 있다. 이것이 실시간 direct SLAM의 photometric layer로 들어올 수 있는지, 들어온다면 direct와 learned의 경계가 어디에 그어지는지는 2026년 현재 열린 질문이다. 한편 direct 계보와 나란히, 다른 방향의 탈출구가 이미 2011년에 열려 있었다. Newcombe 자신이 KinectFusion을 통해 보여줬다. 단안 카메라의 photometric 가정을 지키는 대신, 센서 자체를 바꾸면 된다. depth 정보를 직접 측정하는 RGB-D 카메라는 밝기 변화에 무관하게 dense 재구성을 가능하게 했다. direct method가 photometric consistency를 수식으로 지키려 했다면, RGB-D는 그 가정 자체를 질문 목록에서 지웠다. --- # Ch.9 — Dense/RGB-D: KinectFusion부터 BundleFusion까지 2011년 11월, Richard Newcombe(Imperial College London)는 ISMAR에서 KinectFusion을 발표했다. 함께 공개된 데모 영상에는 손에 든 Kinect 센서 하나가 실시간으로 방 전체를 3D 메시로 채워가는 장면이 담겼다. 그 장면은 Newcombe 자신이 같은 해 발표한 DTAM이 단안 카메라로 꿈꾸던 것을 RGB-D 센서로 실제로 해낸 것이었다. 계보는 선명하다: 1996년 Curless와 Levoy가 그래픽스 커뮤니티를 위해 고안한 TSDF 표현, 1992년 Besl과 McKay가 로봇공학에 제공한 ICP 추적, 그리고 2010년 Microsoft가 $150에 출시한 Kinect 센서. 이 세 줄기가 교차한 지점에서 dense SLAM의 짧고 강렬한 시대가 열렸다. Davison의 MonoSLAM(Ch.5)이 단안 카메라로 sparse landmark를 추적하던 실시간 추적이라는 목표는 이어졌지만, KinectFusion은 깊이 스트림과 GPU 병렬 계산으로 다른 구조를 택했다. Newcombe의 DTAM(Ch.8)이 직접 광도 최적화로 dense 재구성을 시도하면서 GPU의 가능성을 열었고, KinectFusion은 그 가능성을 RGB-D 센서로 닫았다. --- ## 9.1 Kinect 이전의 dense 재구성 2011년 이전에도 dense 3D 재구성은 가능했다. *실시간*만 빠져 있었다. 오프라인 파이프라인들은 스테레오 혹은 structured light 스캐너로 취득한 포인트 클라우드를 시간을 들여 병합했다. 실내 스캔 장비는 수십만 달러였다. 연구실 바깥에서 이 기술을 쓰기는 어려웠다. SLAM 커뮤니티는 이미 sparse landmark로 충분히 실용적인 결과를 얻고 있었고, dense 재구성은 그래픽스 쪽 문제로 분류해 두고 있었다. [Curless와 Levoy의 1996년 SIGGRAPH 논문 "A Volumetric Method for Building Complex Models from Range Images"](https://graphics.stanford.edu/papers/volrange/volrange.pdf)는 이 시기의 그래픽스 쪽 접근을 대표한다. 이 방법은 **TSDF(Truncated Signed Distance Function)**를 사용한다. 3D 공간을 균일한 복셀 그리드로 나누고, 각 복셀에 가장 가까운 표면까지의 부호 있는 거리를 누적한다. 부호 관행은 센서에서 표면 방향으로 진행할 때 표면 앞(free space)이 양수, 표면 뒤(solid 내부)가 음수다. Truncated란 이 값을 절댓값 기준 일정 한계 $t$ 이내로 잘라낸다는 뜻으로, $\text{TSDF}(x) = \text{clip}(d(x), -t, +t)$ 형태가 된다. 새 깊이 프레임이 들어올 때마다 이 값을 가중 평균으로 갱신하면, 노이즈가 점진적으로 평균화되면서 표면이 점점 선명해진다. 표면 추출은 TSDF의 zero-crossing에 marching cubes를 적용하면 된다. 이 방법은 정확했다. 그러나 복셀 그리드는 메모리를 많이 먹었고, 실시간 갱신은 당시 하드웨어로는 불가능했다. Curless-Levoy의 논문은 이후 15년 동안 그래픽스 교과서에 머물렀다. 그 15년 사이에 두 가지가 바뀌었다. GPU가 GPGPU 시대로 진입했고, Kinect가 등장했다. --- ## 9.2 KinectFusion과 TSDF Microsoft는 2010년 Xbox 360용 Kinect를 약 $150에 출시했다. 구조광(structured light) 방식으로 깊이를 측정하는 이 센서는 VGA 해상도의 깊이 맵을 30Hz로 스트리밍했다. 정밀도는 연구용 ToF 카메라보다 낮았지만 가격도 훨씬 낮았다. 출시 몇 주 만에 오픈소스 드라이버가 공개됐고, 연구 활용이 뒤따랐다. Newcombe는 그 무렵 Microsoft Research Cambridge로 자리를 옮겼고, Shahram Izadi 팀과 함께 GPU 기반 dense SLAM을 준비하고 있었다. Kinect가 공급한 깊이 스트림이 이 파이프라인과 결합됐다. 결과가 2011년 ISMAR에서 발표된 [Newcombe et al. 2011. KinectFusion](https://doi.org/10.1109/ISMAR.2011.6092378)이다. > 🔗 **차용.** KinectFusion의 핵심 표현인 TSDF는 Curless & Levoy(1996)가 오프라인 3D 스캐닝을 위해 고안한 것이다. Newcombe 팀은 이를 GPU의 병렬 복셀 갱신으로 실시간화했다. 파이프라인은 네 단계로 구성된다. 깊이 전처리 단계에서는 원시 깊이 맵에서 bilateral filter로 노이즈를 줄이고 표면 법선을 계산한다. ICP 추적 단계에서는 현재 프레임의 포인트 클라우드를 이전 TSDF에서 ray-cast한 가상 표면에 정렬한다. [Besl & McKay(1992)](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf)의 **ICP(Iterative Closest Point)**를 point-to-plane 형태로 바꾸고, 3단계 영상 피라미드에서 coarse-to-fine으로 각각 최대 4·5·10회 반복한다. 각 반복에서는 수많은 vertex-normal 대응의 정규방정식 항을 GPU가 병렬로 누적한다. 결과는 카메라의 6-DoF 포즈다. point-to-plane ICP의 목적함수는 다음과 같다. 현재 프레임의 포인트 $\mathbf{p}_i$를 변환 $T = (R, \mathbf{t})$로 움직인 뒤 대응 점 $\hat{\mathbf{p}}_i$(ray-cast 표면)과 법선 $\hat{\mathbf{n}}_i$에 대해 $$E(R, \mathbf{t}) = \sum_i \bigl(\hat{\mathbf{n}}_i^\top (R\,\mathbf{p}_i + \mathbf{t} - \hat{\mathbf{p}}_i)\bigr)^2$$ 을 최소화한다. 원래 Besl-McKay의 point-to-point($\|R\mathbf{p}_i + \mathbf{t} - \hat{\mathbf{p}}_i\|^2$)와 달리 법선 방향 오차만 측정하므로, 표면에 접하는 방향의 미끄러짐에 덜 민감하다. 소회전 근사 $R \approx I + [\boldsymbol{\omega}]_\times$를 적용하면 $E$는 6-DoF 벡터 $(\boldsymbol{\omega}, \mathbf{t})$에 대한 선형 최소제곱 문제로 바뀌어 GPU에서 병렬 감소(parallel reduction)로 한 번에 풀린다. > 🔗 **차용.** KinectFusion의 추적 단계는 Besl & McKay(1992) ICP를 직접 계승한다. 고전 로봇공학 문헌의 기법을 GPU 밀도로 다시 꺼낸 것이다. TSDF 통합 단계에서는 추정된 포즈로 깊이 맵을 복셀 그리드에 투영해 TSDF 값을 갱신한다. 논문의 대표 실험 설정은 512³ 복셀로 약 3m 한 변 크기의 방 규모 볼륨을 덮는다(§4.2, Fig. 13). 표면 렌더링 단계에서는 TSDF의 zero-crossing을 ray marching으로 찾아 표면의 vertex map과 normal map을 만든다. 이 결과가 다음 ICP 추적의 참조 표면이 된다. Newcombe는 같은 해 DTAM을 단안 카메라 dense SLAM으로 발표했다. KinectFusion은 그 자매 연구다. DTAM이 GPU를 써서 단안의 광도 일관성을 최적화했다면, KinectFusion은 같은 GPU를 깊이 통합에 투입했다. 두 논문의 저자 목록이 겹치는 이유다. > 🔗 **차용.** KinectFusion과 DTAM은 같은 해 같은 연구자가 발표한 두 dense 시스템이다. 두 시스템은 GPU를 이용한 dense 처리라는 방향을 공유하지만, DTAM의 광도 최적화와 KinectFusion의 깊이 정합·TSDF 통합은 지도 표현과 목적함수도 다르다. 512³ TSDF가 30Hz로 갱신됐고, 실내 방 한 칸을 몇 분 안에 dense mesh로 복원했다. 고정된 방 규모 볼륨에서는 dense model-to-frame ICP가 많은 표면 측정을 함께 사용해 낮은 추적 drift를 보였다. 다만 그 모델도 과거의 추정 포즈로 쌓이므로 절대 기준이 되는 표면은 아니었다. 한계도 명확했다. 512³ 복셀 그리드는 고정된 공간 범위만 다룰 수 있었다. 방을 벗어나면 복셀이 포화되거나 기존 복셀을 덮어써야 했다. loop closure가 없었다. 그리고 Kinect의 IR 구조광은 햇빛 아래에서 작동하지 않았다. 실외는 처음부터 범위 밖이었다. --- ## 9.3 Kintinuous — rolling volume KinectFusion이 발표된 직후 Whelan은 Imperial College에서 이 한계를 다뤘다. 고정 크기 TSDF 볼륨이 문제라면, 카메라를 따라 이동하면 된다. 2012년 7월 RSS 워크숍(RGB-D: Advanced Reasoning with Depth Cameras, Sydney)에서 [Whelan 등이 발표한 Kintinuous](https://www.cs.cmu.edu/~kaess/pub/Whelan12rssw.pdf)는 "rolling TSDF volume"을 도입했다. 카메라가 볼륨 경계에 가까워지면 반대쪽 슬라이스를 메시로 출력하고 해제한 뒤, 새 슬라이스를 앞에 붙인다. 메모리는 일정하게 유지되면서 카메라는 무한히 이동할 수 있다. 실내 복도 전체를 걷는 데모는 KinectFusion이 보여주지 못한 것이었다. 그러나 loop closure는 여전히 없었다. 긴 복도를 걸어서 원점으로 돌아왔을 때 두 끝이 맞지 않는 문제는 해결되지 않았다. 이 오정합은 표면의 세밀함과 별도로 전역 지도의 일관성을 제한했다. --- ## 9.4 ElasticFusion: Surfel과 비강체 변형 Whelan은 Kintinuous 이후 방향을 바꿨다. TSDF 복셀 대신 surfel을 선택했다. **surfel(surface element)**은 위치, 법선, 반경, 색상을 가진 점이다. 컴퓨터 그래픽스에서 [Pfister 등(2000)](https://www.merl.com/publications/docs/TR2000-10.pdf)이 렌더링 표현으로 제안한 개념이었다. 복셀 그리드에 비해 불규칙하고 표면에 밀착하는 구조다. > 🔗 **차용.** ElasticFusion의 surfel 표현은 Pfister 등(2000)의 그래픽스 렌더링 기법을 SLAM의 맵 표현으로 이식한 것이다. [Whelan et al. 2016. ElasticFusion](https://doi.org/10.1177/0278364916669237)의 핵심 기여는 두 가지다. 첫째, surfel 기반 dense map을 채용했다. 둘째, *non-rigid deformation*을 이용한 loop closure를 구현했다. 기존 dense SLAM의 loop closure는 어려웠다. 전역 메시나 복셀 그리드를 loop closure 정보에 맞춰 수정하려면 비용이 컸다. ElasticFusion은 surfel 집합을 변형 그래프(deformation graph)와 연결하고, loop closure가 감지되면 그래프를 변형해 전체 맵에 오차를 분산시켰다. surfel 맵 수준에서의 비강체 변형이었다. 구체적으로, deformation graph의 각 노드 $g_k$는 위치 $\mathbf{v}_k$와 회전 $R_k$, 이동 $\mathbf{t}_k$를 가진다. surfel $s$는 가장 가까운 $K$개 노드의 영향권 안에 놓이고, surfel의 변형 후 위치는 $$\tilde{\mathbf{p}}_s = \sum_{k \in \mathcal{N}(s)} w_k \bigl(R_k (\mathbf{p}_s - \mathbf{v}_k) + \mathbf{v}_k + \mathbf{t}_k\bigr)$$ 로 계산된다(가중치 $w_k$는 거리 기반 감쇠). loop closure 제약이 추가되면 그래프 노드들의 $(R_k, \mathbf{t}_k)$를 Gauss-Newton으로 최적화해 오차를 전역에 분산한다. TSDF를 통째로 다시 쌓지 않고도 dense 맵 전체를 일관되게 수정할 수 있었던 이유다. ElasticFusion 논문은 ICL-NUIM 합성 데이터셋에서 당시의 강한 실내 재구성 결과를 보고했다. kt0·kt1·kt2 시퀀스의 ATE RMSE는 1.4cm 이하였고, 그중 kt0·kt1은 0.9cm였다(글로벌 루프 클로저가 발동하는 kt3은 예외적으로 큰 값). 이 수치는 KITTI나 TUM RGB-D가 아니라 논문이 사용한 ICL-NUIM 설정에 한정해 읽어야 한다. --- ## 9.5 BundleFusion: 오프라인 SfM 품질을 온라인으로 2017년 Dai, Nießner, Zollhöfer, Izadi, Theobalt가 ACM Transactions on Graphics에 발표한 [Dai et al. 2017. BundleFusion](https://doi.org/10.1145/3072959.3054739)은 다른 방향에서 문제에 접근했다. KinectFusion 계열이 실시간성을 타협하지 않으면서 품질을 높이려 했다면, Dai 팀은 GPU 연산을 최대한 투입해 온라인 시스템에서도 SfM 수준의 번들 조정을 실행하는 것을 목표로 삼았다. BundleFusion은 최적화를 세 층으로 나눴다. 가장 빠른 층에서는 현재 프레임과 이전 프레임 사이의 dense depth alignment로 초기 포즈를 잡는다. 그 위 층에서는 SIFT feature를 이용한 sparse frame-to-frame alignment로 보정하고, 세 번째 층에서 계층적 global bundle adjustment가 누적된 프레임들의 포즈를 재최적화한다. Bundle adjustment는 프레임이 누적될수록 과거 포즈도 재추정한다. "retroactive pose correction"이라고 불린 이 방식은 오프라인 SfM 파이프라인이 모든 데이터를 가진 뒤 정합하는 것과 유사한 효과를 온라인으로 달성하려 했다. 갱신된 포즈 시퀀스를 TSDF에 역투영해 재통합하므로, 추적 오류가 맵에 그대로 쌓이지 않는다. Dai 팀이 TUM RGB-D 벤치마크에서 보고한 수치는 ElasticFusion을 능가했다. 시각적 재구성 품질도 당시 기준으로 오프라인 COLMAP 파이프라인에 근접했다. > 📜 **예언 vs 실제.** BundleFusion은 real-time online global bundle adjustment를 "unprecedented speed"로 달성했다고 주장하며, 오프라인 SfM 품질을 온라인으로 끌어올리는 경로를 제시했다. GPU 연산력은 이후로도 빠르게 증가했지만, 2021년 이후 연구 관심의 한 축은 COLMAP으로 포즈를 얻고 NeRF로 장면을 표현하는 파이프라인으로 이동했다. 이는 TSDF 기반 online dense mapping의 직접 후계라기보다 novel-view synthesis와 장면 표현을 향한 우회였다. --- ## 9.6 하드웨어와 알고리즘의 공진화 KinectFusion에서 BundleFusion까지의 6년은 하드웨어와 알고리즘이 서로를 밀어붙인 과정이다. Kinect 1세대는 구조광 방식이었다. 깊이 정밀도는 미터 범위에서 수 밀리미터였지만 햇빛 아래에서는 IR 패턴이 잡히지 않았다. 2013년 출시된 Kinect 2는 ToF(Time-of-Flight) 방식으로 바꿨다. 정밀도가 올라갔고 동적 범위도 나아졌다. Intel의 RealSense 시리즈가 뒤를 이었다. 센서 선택지가 늘어날수록 알고리즘이 가정할 수 있는 깊이 품질이 달라졌고, 연구자들은 더 작은 노이즈를 활용하거나 더 큰 노이즈를 견디는 방식을 실험했다. GPU 쪽에서는 CUDA 생태계가 성숙했다. KinectFusion이 나온 2011년부터 BundleFusion이 나온 2017년 사이에는 GPU의 처리량과 메모리 성능도 발전했다. 배율을 비교할 때에는 GPU 모델과 연산 정밀도를 함께 지정해야 한다. Whelan이 ElasticFusion에서, Dai가 BundleFusion에서 점점 더 무거운 최적화를 실시간으로 실행할 수 있었던 것은 알고리즘만의 성과가 아니었다. Kinect의 $150 가격과 소비자 시장 보급은 연구실의 접근 문턱을 낮췄다. > 📜 **예언 vs 실제.** KinectFusion이 2011년에 보여준 512³ 고정 볼륨의 한계(공간 범위, 드리프트, 실외 부적합)는 이후 연구의 로드맵이 됐다. 볼륨 확장은 Kintinuous, ElasticFusion, BundleFusion이 차례로 공략했다. 반면 실외는 다른 결론에 도달했다. IR 구조광은 햇빛 아래에서 패턴이 잡히지 않는다. 초기 Kinect 기반 dense SLAM은 주로 실내에 머물렀고, LiDAR는 실외 dense mapping의 주요 센서가 되었다. 이 제한을 다른 방식의 RGB-D 센서 전체에 적용할 수는 없다. --- ## 9.7 dense-only의 퇴장 2011년부터 2017년 사이 dense RGB-D SLAM은 Visual SLAM의 주된 방향이 될 것처럼 보였다. 실제 전개는 그렇지 않았다. sparse backend는 계속 지배했다. [ORB-SLAM2](https://arxiv.org/abs/1610.06475)와 [VINS-Mono](https://arxiv.org/abs/1708.03852)로 대표되는 2015년 이후의 실용 SLAM 시스템들은 dense 맵을 기본으로 삼지 않았다. 이유는 복합적이었다. 512³ TSDF는 512MB 이상이 필요해 모바일 플랫폼이나 임베디드 시스템에서는 감당하기 어려웠다. 해시 블록에 TSDF를 저장하는 [Voxblox](https://arxiv.org/abs/1611.03631)와 octree에 점유 확률을 저장하는 [OctoMap](https://www.hrl.uni-bonn.de/papers/wurm10octomap.pdf)은 서로 다른 지도 표현으로 메모리 부담을 줄였지만 sparse 방식의 효율성과는 격차가 있었다. 실시간 dense 처리는 GPU를 전제했는데, 자율주행 차량의 임베디드 프로세서나 드론의 경량 플랫폼에서는 KinectFusion 수준의 파이프라인을 돌리기 어려웠다. Kinect의 IR depth가 실외에서 작동하지 않는다는 점도 발목을 잡았다. 자율주행과 드론처럼 상용화 요구가 큰 분야 대부분이 실외 환경이었다. 그 사이 dense map data structure의 계보는 KinectFusion의 512³ 고정 볼륨에서 여러 방향으로 갈라졌다. [Museth(2013)의 VDB](https://doi.org/10.1145/2487228.2487235)는 동적으로 커지는 sparse root와 얕고 넓은 B+ tree를 결합해, 활성 영역에만 노드를 할당하면서 빠른 임의 접근을 지원했다. OpenVDB로 공개된 이 구조는 대규모 sparse volumetric data를 다루는 대표적 기반이 됐다(Ch.17 LiDAR의 nvblox 계보와 비교할 수 있다). [Reijgwart et al.(2023)의 wavemap](https://arxiv.org/abs/2306.08125)은 wavelet 변환으로 occupancy를 압축해 해상도-메모리 트레이드오프를 재조정했다. Ramos와 Ott가 이끈 다른 계보는 아예 표현을 연속 함수로 넘겼다. [O'Callaghan과 Ramos(2012)의 GPOM(Gaussian Process Occupancy Map)](https://doi.org/10.1177/0278364911435991)은 깊이 측정을 Gaussian Process 회귀로 연결해 측정되지 않은 복셀까지 확률적으로 채웠고, [Ramos와 Ott(2016)의 Hilbert Map](https://doi.org/10.1177/0278364916684382)은 Hilbert space 특징을 logistic regression으로 학습시켜 스트리밍 가능한 확률적 occupancy를 제공했다. [Behley와 Stachniss(2018)의 SuMa](https://www.ipb.uni-bonn.de/wp-content/papercite-data/pdf/behley2018rss.pdf)는 ElasticFusion이 실내 RGB-D에서 쓴 surfel 표현을 outdoor LiDAR로 옮겨 KITTI에서 작동하는 surfel-based SLAM을 만들었다(→ Ch.17). KinectFusion이 방 한 칸에서 멈췄던 자리에서, 이 계보들이 outdoor·도시 규모·확률적 불확실성 쪽으로 각자의 방향을 열었다. 2020년을 전후해 NeRF가 등장하면서 고품질 dense 재구성을 원하는 수요는 NeRF와 3D Gaussian Splatting으로 이동했다. 지도 표현은 달라졌지만 RGB-D의 측정 깊이는 추적과 지도 geometry 학습 양쪽의 제약으로 계속 쓰였다. dense 시대는 짧았지만 흔적은 남았다. TSDF 표현은 자율주행용 occupancy map으로 이어졌고, ICP는 LiDAR SLAM의 표준 추적 수단이 됐다. 접근 방식은 퇴각했지만 부품들은 다른 시스템 안으로 흩어졌다. --- ## 🧭 아직 열린 것 **대규모 실외 dense 재구성.** IR 구조광의 햇빛 취약성은 active depth 센서 전반의 문제다. LiDAR는 더 먼 거리를 다루지만 색상과 세밀한 표면 정보가 빈약하다. RGB-D 방식으로 실외 대규모 환경을 dense하게 처리하는 방법은 2026년 기준으로 아직 없다. Stereo depth estimation이 학습 기반으로 빠르게 발전하고 있어 일부 연구들이 대안을 탐색 중이지만, 어두운 영역·반사면·원거리에서의 한계가 해결되지 않았다. **동적 장면의 dense 재구성.** KinectFusion부터 BundleFusion까지 모든 시스템이 정적 장면을 전제로 설계됐다. 사람이 걸어 다니는 공간을 dense하게 재구성하려면 동적 물체를 분리해야 하는데, 여기에는 semantic segmentation이나 geometry 잔차를 이용한 분리 등을 쓸 수 있다. [DynaSLAM](https://arxiv.org/abs/1806.05620), [MaskFusion](https://arxiv.org/abs/1804.09194) 등이 시도했지만 계산 비용과 robustness 모두에서 실용 배포 수준에 미치지 못한다. **복셀 지도의 메모리 효율.** Voxblox는 해시 구조에 TSDF를, OctoMap은 octree에 점유 확률을 저장해 서로 다른 지도 표현의 메모리 비용을 줄였다. 그러나 건물 층 단위, 도시 블록 단위의 dense 표현은 여전히 수십 기가바이트 수준이다. 어떤 해상도를 어느 영역에서 유지할지를 자동으로 결정하는 adaptive resolution dense map은 아직 범용 해법이 없다. [Instant-NGP](https://arxiv.org/abs/2201.05989)와 같은 implicit neural representation이 이 문제에 접근하고 있지만, 실시간 갱신과 쿼리 속도는 트레이드오프가 남아 있다. dense SLAM이 실내 방 한 칸을 메시로 채우는 동안, 그 방으로 돌아오는 문제는 별도의 계보가 맡고 있었다. KinectFusion에는 loop closure가 없었고, 장소를 기억하는 문제는 옥스퍼드의 별도 계보가 맡았다. --- # Ch.10 — Place Recognition의 평행선: FAB-MAP에서 NetVLAD까지, 그리고 AnyLoc까지 2003년 Davison이 웹캠 한 대로 실시간 3D 추적을 증명하던 무렵, Oxford 모바일 로보틱스 그룹의 Mark Cummins와 Paul Newman은 다른 질문을 붙잡고 있었다. "로봇이 이전에 지나간 장소를 어떻게 알아보는가?" VO(visual odometry)가 누적 drift에 시달리는 한, 이 질문에 답하지 못하면 어떤 SLAM 시스템도 루프를 닫을 수 없었다. Place recognition은 Visual SLAM의 나머지 구성 요소들과 평행하게, 그러나 독자적인 계보로 2000년대 내내 발전했다. FAB-MAP은 Josef Sivic의 BoW 아이디어를 로봇 공간으로 이식했고, DBoW2는 그것을 실용화했으며, NetVLAD는 학습 기반 표현으로 전환했다. 2023년 AnyLoc는 foundation model의 feature를 그대로 가져왔다. Place recognition은 tracking도 mapping도 아니고, 어느 쪽에서도 파생되지 않은 독립 문제였다. 그럼에도 feature-based, direct, dense mapping 계보 모두 루프 클로저 없이는 불완전했고, 그 루프 클로저의 "어디서 봤는가" 판단을 place recognition이 공급했다. --- ## 10.1 BoW 이전의 place recognition GPS가 없는 실내, 터널, 도심 협곡에서 로봇이 루프를 닫으려면 현재 관측과 과거 관측 사이의 유사도를 수천 장의 후보 이미지 중에서 빠르게 찾아야 한다. 픽셀 단위 비교는 선형 탐색이어서 O(N)이고, 이미지 수가 수만 장을 넘으면 실시간은 불가능하다. 2000년대 초 컴퓨터 비전에서 이 문제를 먼저 다룬 것은 Sivic과 Zisserman이었다. 2003년 ICCV에서 발표된 ["Video Google"](https://www.robots.ox.ac.uk/~vgg/publications/2003/Sivic03/sivic03.pdf)은 문서 검색의 TF-IDF를 이미지에 적용했다. SIFT 기술자를 k-means로 군집화해 "visual word"를 만들고, 이미지를 그 단어들의 빈도 벡터로 표현했다. inverted index는 질의에 등장한 visual word의 posting list만 조회하게 해 전체 영상을 훑는 비용을 줄였다. place recognition 연구자들은 이 아이디어를 곧바로 받아들였다. --- ## 10.2 FAB-MAP — 확률적 BoW와 Chow-Liu 트리 (2008) Mark Cummins와 Paul Newman은 Oxford 모바일 로보틱스 그룹에서 2008년 [Cummins & Newman. FAB-MAP: Probabilistic Localization and Mapping in the Space of Appearance](https://doi.org/10.1177/0278364908090961)를 발표했다. FAB-MAP(**Fast Appearance-Based Mapping**)은 "이 장면은 데이터베이스에 있는 장소인가, 아니면 전혀 새로운 곳인가?"를 묻는다. 단순 유사도 점수로는 이 판단을 내릴 수 없다. 비슷해 보이는 복도가 수십 개라면 가장 높은 유사도가 정답을 보장하지 않는다. Cummins와 Newman은 이를 베이즈 추론 문제로 구성했다. 관측 $z_t$(visual word의 발생 여부 집합)가 주어졌을 때, 현재 위치가 데이터베이스의 각 장소 $\ell_i$일 확률을 계산한다: $$P(\ell_i \mid z_t) \propto P(z_t \mid \ell_i) P(\ell_i)$$ 문제는 $P(z_t \mid \ell_i)$다. visual word들이 독립이라고 가정하면 naïve Bayes가 되지만, 실제로 visual word들은 상관된다. "문"이라는 word가 등장하면 "문손잡이"라는 word도 같이 등장하기 쉽다. 독립 가정은 확률 값을 왜곡한다. FAB-MAP은 **Chow-Liu tree**를 사용해 이 상관을 모델링했다. Chow-Liu tree는 word 간의 pairwise mutual information을 최대화하는 트리 구조 그래픽 모델이다. 두 word $e_i, e_j$ 사이의 mutual information은 $$I(e_i; e_j) = \sum_{e_i, e_j} P(e_i, e_j) \log \frac{P(e_i, e_j)}{P(e_i)P(e_j)}$$ 로 정의되고, Chow-Liu 알고리즘은 이를 엣지 가중치로 삼아 최대 스패닝 트리를 구성한다. 이 트리로 joint likelihood를 분해하면 $$P(z_t \mid \ell_i) = \prod_k P(z_t^k \mid z_t^{\text{pa}(k)}, \ell_i)$$ 가 된다. 여기서 $z_t^k \in \{0,1\}$은 $k$번째 word의 발생 여부이고, $\text{pa}(k)$는 트리에서 $k$의 부모 노드다. 나이브 베이즈(독립 가정) 대비 word 간 공동 발생 패턴을 반영하므로, 복도처럼 시각적으로 유사한 장소들에서 false positive를 낮출 수 있다. 학습 단계에서 대규모 이미지 집합으로 vocabulary와 tree를 함께 훈련한다. 또한 FAB-MAP은 현재 위치가 데이터베이스에 없는 새 장소일 가능성을 명시적으로 다룬다. "new place" 가설을 넣자 false positive가 줄었다. Loop closure에서 false positive는 catastrophic failure로 이어진다. > 🔗 **차용.** FAB-MAP의 visual word 방식은 Sivic & Zisserman의 "Video Google"(2003)에서 직접 이식되었다. 문서 검색의 inverted index 논리를 로봇의 장소 기억에 적용한 것이다. 2011년 Cummins와 Newman은 [FAB-MAP 2.0](https://www.robots.ox.ac.uk/~mjc/Papers/cummins_newman_ijrr_fabmap2_2010_preprint.pdf)을 발표했다. 처리 가능한 지도 규모를 1,000 km 수준으로 확장한 것이 목표였다. 실험적으로 도시 규모 데이터셋에서 작동함을 보였다. --- ## 10.3 DBoW2 — binary descriptor와 vocabulary tree (2012) FAB-MAP은 SIFT처럼 부동소수점 descriptor를 기반으로 했다. 2012년 무렵 SLAM 커뮤니티는 더 빠른 binary descriptor, 특히 BRIEF·ORB·BRISK 쪽으로 이동하고 있었다. SIFT vocabulary를 그대로 쓰는 것은 연산 비용이 문제였다. Dorian Gálvez-López와 Juan D. Tardós(Universidad de Zaragoza)는 2012년 [Gálvez-López & Tardós. Bags of Binary Words for Fast Place Recognition in Image Sequences](https://doi.org/10.1109/TRO.2012.2197158)를 발표했다. **DBoW2**는 binary descriptor를 사용하는 vocabulary tree로, Hamming distance 기반 비교로 SIFT보다 수십 배 빠른 word 배정이 가능했다. DBoW2의 구조는 계층적 k-means로 만든 vocabulary tree다. 이미지를 표현하는 BoW 벡터는 TF-IDF 가중치가 부여된 binary word 빈도 벡터다. $k$분기 $d$깊이 트리의 각 리프 노드 $w_i$에 TF-IDF 가중치 $$\eta_i = \frac{n_i}{n} \cdot \log \frac{N}{N_i}$$ 를 부여한다. 여기서 $n_i$는 해당 이미지에서의 word 빈도, $n$은 총 word 수, $N$은 데이터베이스 이미지 수, $N_i$는 $w_i$를 포함한 이미지 수다. 두 이미지 $a$, $b$의 유사도는 L1-norm $$s(\mathbf{v}_a, \mathbf{v}_b) = 1 - \frac{1}{2} \left\| \frac{\mathbf{v}_a}{|\mathbf{v}_a|} - \frac{\mathbf{v}_b}{|\mathbf{v}_b|} \right\|_1$$ 으로 계산한다. Inverted index는 query에 나타난 visual word의 posting list만 조회하게 해, 모든 데이터베이스 이미지를 순회하는 비용을 피한다. 실제 조회 비용은 word 수와 각 posting list의 길이에 좌우된다. > 🔗 **차용.** DBoW2의 vocabulary tree 개념은 Nistér & Stewénius의 2006년 ["Scalable Recognition with a Vocabulary Tree"](https://people.eecs.berkeley.edu/~yang/courses/cs294-6/papers/nister_stewenius_cvpr2006.pdf)(CVPR)에서 계보를 잇는다. DBoW2는 그 구조를 binary descriptor 세계로 이식하고, 가중치 체계를 SLAM에 맞게 조정했다. DBoW2의 영향은 알고리즘보다 배포에서 컸다. 오픈소스로 공개된 이 라이브러리는 ORB-SLAM(2015)의 loop closure 모듈로 채택되었고, ORB-SLAM2·ORB-SLAM3까지 같은 DBoW2를 썼다. 2015년부터 2020년대 중반까지 DBoW2는 ORB-SLAM 계열과 이를 활용한 시스템의 place recognition을 널리 담당했다. Gálvez-López와 Tardós의 파트너십 역시 주목할 만하다. Tardós는 이후 Mur-Artal, Campos와 함께 ORB-SLAM 삼부작을 이끈 인물이다. DBoW2는 그 프로젝트의 place recognition 계층을 미리 준비한 셈이었다. --- ## 10.4 NetVLAD — CNN 기반 VPR (2016) BoW 계열은 한 가지 근본 한계가 있었다. vocabulary는 특정 descriptor와 특정 환경에서 훈련된 것이었다. 조명이 바뀌거나 계절이 달라지거나 시점이 크게 달라지면 visual word의 분포가 달라지고, 미리 훈련된 vocabulary는 부정합을 일으켰다. Relja Arandjelović, Petr Gronat, Akihiko Torii, Tomáš Pajdla, Josef Sivic는 2016년 CVPR에서 [NetVLAD: CNN Architecture for Weakly Supervised Place Recognition](https://doi.org/10.1109/CVPR.2016.572)을 발표했다. 저자 중 Sivic은 2003년 "Video Google"의 그 Sivic이다. 2003년 ICCV에서 BoW를 이미지 검색에 도입한 사람이, 13년 뒤 그 방식의 한계를 넘는 논문에 공동저자로 이름을 올렸다. NetVLAD는 **VLAD(Vector of Locally Aggregated Descriptors)** aggregation을 미분 가능하게 만들었다. VLAD는 2010년 [Jégou et al.](https://inria.hal.science/inria-00548637/file/jegou_compactimagerepresentation.pdf)이 제안한 aggregation 방식으로, 각 local descriptor가 가장 가까운 cluster center(visual word)에 "잔차"로 얼마나 기여하는지를 누적해 이미지 전체를 표현한다. cluster center $k$에 대한 VLAD 부분 벡터는 $$\mathbf{V}(k) = \sum_{\mathbf{x}_i : \text{NN}(\mathbf{x}_i)=k} (\mathbf{x}_i - \boldsymbol{\mu}_k)$$ 이고, 전체 VLAD 벡터 $\mathbf{V} = [\mathbf{V}(1)^\top, \ldots, \mathbf{V}(K)^\top]^\top$는 이를 모든 cluster에 대해 연결(concatenate)한 뒤 L2-normalize한 것이다. $K$ clusters, $D$차원 descriptor라면 최종 벡터는 $KD$차원이다. VLAD 벡터는 BoW의 이진 할당보다 훨씬 풍부한 정보를 담는다. > 🔗 **차용.** NetVLAD의 aggregation 설계는 Jégou et al.의 "Aggregating Local Descriptors into a Compact Image Representation"(CVPR 2010)에서 VLAD를 직접 계승했다. NetVLAD는 VLAD의 hard assignment를 soft assignment로 바꾸고 전체 파이프라인을 end-to-end로 학습 가능하게 만들었다. NetVLAD layer는 기존 VLAD의 nearest-neighbor 할당을 softmax로 완화한다: $$\bar{a}_k(\mathbf{x}_i) = \frac{e^{\mathbf{w}_k^\top \mathbf{x}_i + b_k}}{\sum_{k'} e^{\mathbf{w}_{k'}^\top \mathbf{x}_i + b_{k'}}}$$ 여기서 $\mathbf{x}_i$는 CNN에서 추출한 local feature, $\mathbf{w}_k$와 $b_k$는 학습 가능한 파라미터다. 이 soft 할당으로 NetVLAD 벡터를 누적하면 $$\mathbf{V}(k) = \sum_i \bar{a}_k(\mathbf{x}_i)\,(\mathbf{x}_i - \boldsymbol{\mu}_k)$$ 이고, 전체 벡터 $\mathbf{V} = [\mathbf{V}(1)^\top, \ldots, \mathbf{V}(K)^\top]^\top$를 intra-normalization(각 부분 벡터 L2 정규화) 후 전체를 다시 L2-normalize하면 최종 VPR descriptor가 된다. hard assignment VLAD와 달리 gradient가 역전파되므로 CNN backbone과 함께 end-to-end 학습이 가능하다. 학습 방식도 달랐다. 저자들은 Google Street View Time Machine 데이터를 활용해 같은 장소의 다른 시점 이미지 쌍을 양성 예, 다른 장소를 음성 예로 삼는 weakly supervised triplet loss를 사용했다. GPS 위치를 약한 감독 신호로 삼아 수작업 장소 레이블 없이 학습할 수 있었다. Pittsburgh 250k, Tokyo 24/7 벤치마크에서 NetVLAD는 DBoW 계열과 이전 VLAD 기반 방법들을 큰 차이로 앞섰다. 조명·계절 조건 변화에 걸쳐 훨씬 강건했고, 시점 차에도 내성이 있었다. 그러나 실용 SLAM 파이프라인에 NetVLAD가 바로 통합되지는 않았다. 추론 속도와 메모리 요구가 DBoW2보다 무거웠고, 이미 ORB-SLAM 생태계가 DBoW2에 맞춰 구축되어 있었기 때문이다. --- ## 10.5 Patch-NetVLAD, MixVPR, AnyLoc (2020-2023) NetVLAD 이후 VPR(Visual Place Recognition) 연구는 일반화 성능을 개선하는 여러 방향으로 갈라졌다. 2021년 Hausler et al.은 [Patch-NetVLAD](https://arxiv.org/abs/2103.01486)를 내놓았다. Global descriptor 하나로 장소를 판단하는 NetVLAD 대신, 이미지를 패치로 분할해 각 패치의 NetVLAD 표현을 공간적으로 결합하는 방식이다. Tokyo 24/7에서 NetVLAD 대비 Recall@1을 약 10% 포인트 올렸다. 패치 단위 처리로 추론 비용도 함께 늘었다. 2023년 Ali-bey et al.의 [MixVPR](https://arxiv.org/abs/2303.02190)는 Transformer-style feature mixing으로 global feature를 생성했다. 경량화와 성능 사이의 균형이 목표였다. 이 시기 VPR 논문들은 공통으로 Mapillary Street Level Sequences(MSLS)와 Nordland 같은 계절 변화 데이터셋을 벤치마크로 삼았다. 극단적 조명·계절 조건이 공통의 장벽으로 떠올랐다. 2023년 Keetha et al.의 [AnyLoc: Towards Universal Visual Place Recognition](https://arxiv.org/abs/2308.00688)은 다른 방향을 택했다. DINOv2 기반 self-supervised feature를 fine-tuning 없이 그대로 place recognition에 쓰는 것이다. > 🔗 **차용.** AnyLoc의 feature 추출은 Oquab et al.의 [DINOv2](https://arxiv.org/abs/2304.07193)(Meta AI, 2023)에서 사전 학습된 ViT 표현을 가져온다. AnyLoc은 그 위에 VLAD aggregation을 얹었다. FAB-MAP에서 시작한 BoW-VLAD 계보가 foundation model 시대에 다시 합류한 형태다. DINOv2는 대규모 인터넷 이미지로 학습된 Vision Transformer(ViT)다. 여러 환경에 적용할 수 있는 범용 feature를 생성하지만, 사전학습 자료의 편향까지 없어지는 것은 아니다. Keetha et al.이 AnyLoc에서 주목한 건 DINOv2의 **facet** 개념이었다. ViT의 각 attention head는 query(Q), key(K), value(V) 행렬과 최종 token(patch feature)을 출력한다. Keetha et al.은 이 네 종류의 facet 중 value(V) facet이 place recognition에 가장 의미론적으로 안정된 표현을 제공함을 실험으로 확인했다. Q·K facet은 구조·기하 정보에, V facet은 의미론(semantics)에 더 집중되는 경향이 있어, 계절·조명에 걸친 일관된 장소 표현에 유리하다. Keetha et al.은 이 V facet 표현을 VLAD aggregation에 연결하면 세계 각지, 실내외, 지하, 항공 뷰 등 매우 다양한 환경에서 단일 모델이 동작함을 보였다. Pittsburgh, Tokyo, 실내 공장, 지하 주차장, 도서관 등 7개 이상의 환경에서 single-model이 이전 specialized 방법들과 경쟁하거나 앞섰다. 한편 modality 경계를 넘는 갈래도 나타났다. Lee et al.의 [(LC)²](https://arxiv.org/abs/2304.08660) (RA-L 2023)는 카메라 영상과 LiDAR 점군을 공통 2.5D depth image로 투영해, 2D 쿼리로 LiDAR 지도에서 장소를 조회하는 cross-modal retrieval을 시도했다. 이런 cross-modal 평가에는 Lee et al.의 [ViViD++](https://arxiv.org/abs/2204.06183) (RA-L 2022)처럼 visible·thermal·event·LiDAR·관성·depth를 실내외와 지하에서 동기화해 놓은 데이터셋을 활용할 수 있다. --- ## 10.6 place recognition과 metric localization의 통합 시도 (2024-2025) Place recognition 연구는 2000년대 초부터 SLAM의 나머지 구성 요소와 평행하게 달려왔다. ORB-SLAM이 DBoW2를 내장했지만 place recognition 모듈은 mapping·tracking으로부터 격리된 블랙박스였다. 이미지를 입력받아 루프 후보 ID를 출력했다. 2023-2024년에는 Berton et al.의 [EigenPlaces](https://arxiv.org/abs/2308.10832)(2023)와 Izquierdo & Civera의 [SALAD](https://arxiv.org/abs/2311.15937)(2023 arXiv / CVPR 2024)가 viewpoint 변화에 강한 global descriptor와 local feature aggregation을 발전시켰다. 두 방법의 직접 출력은 여전히 데이터베이스 이미지 검색 결과다. Metric pose가 필요하면 검색된 후보와의 기하 검증이나 별도의 localization 단계가 뒤따라야 한다. 2024년 전후로는 Gaussian map 표현과 place recognition을 결합하려는 시도들도 등장했다. 3DGS(3D Gaussian Splatting)가 지도 표현으로 올라온 흐름과 맞물린 방향이었다. > 📜 **예언 vs 실제.** Cummins와 Newman은 2011년 FAB-MAP 2.0 논문에서 1,000 km 규모 궤적에서의 appearance-only 루프 클로저를 시연하며 place recognition의 스케일 한계를 밀어올렸다. Oxford 캠퍼스와 도심 일부를 달리던 초기 FAB-MAP 실험 기준으로 두 자릿수 배율의 도약이었다. 이후 DBoW2와 대형 vocabulary를 쓴 도시 규모 실험들이 같은 스케일을 실용 SLAM에서 재현했다. 처리할 수 있는 규모는 커졌고, deep learning은 계절·조명 변화에 취약한 vocabulary 기반 표현을 보완할 도구를 더했다. 다만 규모와 외관 변화에 대한 일반화가 모든 환경에서 해결된 것은 아니다. > 📜 **예언 vs 실제.** Arandjelović et al.은 2016년 NetVLAD 논문 서론에서 place recognition을 풀기 위한 세 가지 도전(CNN 아키텍처, 충분한 학습 데이터, end-to-end 학습 절차)을 명시하고 각각에 대한 자신들의 기여를 제시했다. 아키텍처와 학습 절차 쪽은 NetVLAD로 직접 답했지만, 이후 7년간 외관 조건(계절·조명·시점) 일반화를 목표로 한 VPR 논문들이 연이어 나왔다. 2023년 AnyLoc은 fine-tuning 없는 foundation model feature로 다환경 단일 모델의 가능성을 보였다. 특화 모델에서 범용 모델 쪽으로 축이 옮겨간 것에 가깝다. --- ## 10.7 🧭 아직 열린 것 **계절·조명 극변.** Nordland(노르웨이 철도, 여름-겨울)와 Oxford RobotCar(1년치 계절 변화) 데이터셋에서 10년 넘게 같은 장벽이 보고된다. DINOv2 기반 방법들이 격차를 줄였지만, 눈이 쌓인 겨울과 나뭇잎이 무성한 여름을 비롯한 서로 다른 조건에서 한 모델이 일관된 정밀도와 재현율을 내는 문제는 아직 풀리지 않았다. 외관 변화가 심한 환경에서의 장소 인식은 2026년 기준으로도 열린 문제다. **Place recognition과 metric localization의 통합.** 현재 대부분의 SLAM 파이프라인에서 place recognition은 "어디서 봤는가"만 답하고, 이후 descriptor matching 등으로 기하 대응을 찾고 PnP 같은 방법으로 실제 pose를 추정한다. 두 과정을 하나의 표현으로 통합하려는 시도들이 2023-2025년에 등장했으나, 정밀도와 속도를 함께 만족하는지는 센서·장면·배포 조건별로 확인해야 한다. **인식 가능한 장소 표현의 프라이버시.** VPR 시스템이 저장하는 장소 표현은 복원 공격으로 원본 이미지나 3D 구조를 되살리는 데 쓰일 수 있다. 상업 로봇이 가정·병원·사무실 실내를 매핑할 때 이 문제는 현실이 된다. 성능 저하 없이 프라이버시를 보장하는 장소 표현 방식은 아직 없다. --- ORB-SLAM이 feature-based 파이프라인을 표준화하고, DSO가 photometric 이론을 완성하고, KinectFusion 계열이 dense mapping의 가능성과 한계를 드러내는 동안, place recognition은 그 어느 계보와도 다른 위치에 있었다. 컴퓨터 비전의 이미지 검색 문제에서 자라난 뒤, SLAM이 루프 클로저를 필요로 했을 때 이를 공급하는 역할을 맡았다. 이러한 독립성은 결과적으로 이점이 됐다. deep learning이 확산될 때 place recognition은 기존 SLAM 파이프라인보다 빠르게 새 도구를 흡수했다. 2023년 AnyLoc이 등장했을 때 Sivic의 이름은 참고문헌에 있었다. 그는 2003년 BoW를 이미지 검색에 도입했고, 2016년에는 NetVLAD로 그 한계를 넘은 논문의 공동저자였다. 그 계보의 끝에서 AnyLoc은 Sivic이 연 문을 foundation model 쪽으로 밀어 넘겼다. --- # Ch.11 — 깊이 추정의 부활: Eigen에서 Depth Anything까지 Feature-based, direct, RGB-D, place recognition 계통은 각자의 방식으로 성숙 단계에 올랐다. ORB-SLAM은 epipolar 기하로 세계를 재구성했고, DSO는 photometric consistency로, KinectFusion은 ICP로 표면을 쌓아 올렸다. 그 경계를 흔든 논문은 SLAM 연구자가 아니라 NYU의 컴퓨터 비전 대학원생에게서 나왔다. Monocular depth 추정은 컴퓨터 비전에서 가장 오래된 ill-posed 문제 중 하나였다. 한 장의 이미지에서 깊이를 복원한다는 것은 원론적으로 불가능하다. 카메라는 3D 세계를 2D로 투영하면서 깊이 정보를 버리기 때문이다. 그러나 인간은 원근감, 폐색, 텍스처 기울기, 표면의 음영으로 단안 깊이를 판단한다. 이것들을 통계적으로 학습할 수 있는가. 2014년 NYU의 David Eigen은 이 질문에 CNN을 적용했고, 그 실험은 10년 뒤 SLAM 파이프라인을 다시 쓰게 될 계보의 출발점이 되었다. --- ## 1. Eigen 2014 — 초기 CNN depth 2014년 이전에도 monocular depth 추정 연구는 있었다. Ashutosh Saxena(Make3D, Stanford)가 [2005년 SVM과 Markov Random Field를 결합해 단일 이미지에서 depth map을 예측하는 시스템](https://papers.nips.cc/paper/2921-learning-depth-from-single-monocular-images)을 발표했다. Make3D는 주로 실외 장면을 대상으로 평면 조각 단위의 거친 3D 구조를 예측했으며, CNN 이전 학습 기반 단안 깊이의 대표 사례였다. Eigen, Puhrsch, Fergus의 [Eigen et al. 2014](https://arxiv.org/abs/1406.2283)는 접근 자체를 바꿨다. coarse network가 전역 구조를 예측하고, fine network가 지역 세부를 보정하는 두 단계 CNN을 썼다. NYU Depth v2의 464개 실내 장면에서 얻은 약 120,000개 프레임을 학습에 사용했다. 성능은 당시 기준으로 Make3D보다 개선됐고, 깊이를 학습할 수 있음을 보였다. 그러나 결정적 약점이 하나 남았다. **scale ambiguity**다. 네트워크의 절대 스케일은 학습 데이터의 분포와 카메라 설정에 묶였다. NYU 실내에서 학습한 모델을 야외에 적용하면 스케일이 틀릴 수 있다. 이 문제는 2024년에도 단안 metric depth의 중심 과제로 남았다. > 🔗 **차용.** Eigen 2014는 깊이 추정 task 자체를 Make3D(Saxena 2005)에서 물려받았다. SVM과 MRF를 CNN으로 교체한 것이 핵심 교체였고, task 정의와 평가 지표(RMSE, threshold accuracy)는 이어받았다. --- ## 2. Garg → Godard — self-supervised depth supervised depth 학습의 병목은 데이터였다. Kinect는 실내에서 잘 작동하지만 야외 환경, 특히 햇빛 아래에서는 적외선 패턴이 날아간다. 대규모 실외 RGB-D 데이터셋 구축은 비용이 크다. 2016년, [Ravi Garg(UCL)는 다른 길을 열었다](https://arxiv.org/abs/1603.04992). stereo 이미지 쌍을 학습 신호로 쓰는 것이다. left 이미지를 보고 depth를 예측한 뒤, 그 depth와 카메라 baseline을 이용해 right 이미지를 재구성한다. right 이미지는 이미 존재하므로 photometric loss로 supervision이 가능하다. 라벨이 필요 없다. Clément Godard(UCL)는 2017년 이 아이디어를 [Godard et al. 2017](https://doi.org/10.1109/CVPR.2017.699)에서 **MonoDepth**로 체계화했다. left 이미지에서 예측한 depth와 right 이미지에서 예측한 depth가 서로 일치해야 한다는 양방향 제약이 left-right consistency다. 구조적 유사도(SSIM)를 photometric loss에 포함해 텍스처 없는 영역에서의 안정성을 높였다. 학습 시에만 stereo 쌍이 필요하고 추론에는 단일 이미지만 쓴다. 원 논문은 KITTI의 당시 비교 protocol에서 표에 실린 기존 self-supervised 방법보다 낮은 오차를 보고했다. > 🔗 **차용.** Garg와 Godard의 photometric loss는 stereo matching 문헌에서 온다. [Scharstein과 Szeliski가 정리한(2002)](https://vision.middlebury.edu/stereo/taxonomy-IJCV.pdf) disparity estimation의 intensity consistency 제약을 depth network의 학습 신호로 전용한 것이다. 2019년 Godard의 *MonoDepth2* ([Godard et al. 2019, ICCV](https://arxiv.org/abs/1806.01260))는 stereo 쌍 대신 monocular video를 쓰는 self-supervised 학습으로 나아갔다. depth network와 pose network를 동시에 학습한다. 연속 프레임 사이의 카메라 운동을 pose network가 예측하면, depth network의 출력으로 이전 프레임을 현재로 warping한다. warping 오차가 줄어드는 방향으로 두 네트워크가 함께 최적화된다. 두 가지 핵심 장치가 추가됐다. 첫째, **minimum reprojection loss**: 여러 소스 프레임 중 photometric error가 가장 낮은 것을 선택해 occluded 영역 오류를 줄인다. 둘째, **auto-masking**: 예측한 depth·pose로 warping한 오차가 source frame을 움직이지 않고 비교한 identity reprojection error보다 작지 않은 픽셀을 학습에서 제외한다. 이 조건은 카메라가 멈춘 프레임이나 카메라와 비슷한 속도로 움직여 인접 프레임에서 겉보기 위치가 거의 변하지 않는 물체를 거르는 효과가 있다. 깔끔한 구조였다. 그러나 여전히 문제가 있었다. 움직이는 물체와 반사 표면이 걸림돌이 됐고, 하늘처럼 텍스처가 없는 영역에서는 더 심했다. 이 영역에서 photometric consistency 가정이 무너진다. 그리고 스케일은 여전히 모호하다. video supervision은 스케일을 프레임 간 상대적으로만 풀어준다. --- ## 3. MiDaS — 데이터셋 혼합 Intel의 René Ranftl이 이끈 팀이 2020년 발표한 [Ranftl et al. 2020](https://doi.org/10.1109/TPAMI.2020.3019967) **MiDaS**(Mixing Datasets for Zero-shot Cross-dataset Transfer)는 다른 질문을 던졌다. 한 데이터셋이 아니라 여러 데이터셋을 동시에 학습하면 어떻게 될까? 문제는 데이터셋마다 depth의 단위와 스케일이 다르다는 것이다. NYU는 미터 단위 실내, KITTI는 LiDAR 포인트 실외, ReDWeb은 stereo 영화, MegaDepth는 SfM 재구성. 이것들을 그대로 섞으면 네트워크가 혼란스러워진다. Ranftl은 **affine-invariant loss**를 사용했다. 학습 중 loss를 계산할 때 각 이미지의 depth prediction과 정답을 affine transformation(스케일 + 시프트)으로 정규화한다. 구체적으로, 예측과 정답 각각에서 중앙값을 빼 shift를 제거하고, 중앙값으로부터의 절대편차를 평균한 값으로 나눠 scale을 제거한 뒤 비교한다. 이 scale-and-shift invariant 정규화 덕분에 데이터셋 간 단위 불일치가 사라진다. 이렇게 하면 네트워크는 "상대적으로 어느 것이 더 멀리"를 배운다. 절대 거리는 아니다. 초기 MiDaS는 서로 다른 여러 데이터셋을 함께 학습해, 학습에 쓰지 않은 데이터셋으로도 이전되는 zero-shot cross-dataset 성능을 보였다. 이후 공개된 MiDaS 계열은 학습 혼합을 최대 12개 데이터셋으로 넓혔다. 야외와 실내처럼 분포가 다른 영상에서도 relative depth를 추정했지만 절대 스케일을 복원하는 모델은 아니었다. 이후 Ranftl 팀은 2021년 [**DPT**(Dense Prediction Transformer)](https://arxiv.org/abs/2103.13413)를 별도 발표해 MiDaS backbone을 ViT 기반으로 교체했다. MiDaS v3부터 DPT가 기본 backbone이 됐고, v3.1(2022)은 그 개선판이었다. 성능이 크게 올랐다. > 🔗 **차용.** MiDaS v3의 DPT는 ViT 기반 encoder를, 이후 Depth Anything은 DINOv2로 사전학습한 encoder를 깊이 추정에 활용했다. backbone 교체만으로 성능이 크게 오르는 현상은 foundation model 시대의 일반적 패턴이지만, DPT(Ranftl 2021)는 depth estimation에서 그 효과를 보인 초기 사례였다. --- ## 4. Depth Anything — foundation 규모 2024년 1월, TikTok Research의 Lihe Yang 팀이 발표한 [Yang et al. 2024](https://arxiv.org/abs/2401.10891) **Depth Anything**은 규모로 문제를 풀었다. 1.5M개의 labeled 이미지(기존 데이터셋 통합)와 62M개의 unlabeled 이미지를 썼고, unlabeled 이미지에는 pseudo-label을 생성해 학습에 포함했다. pseudo-label 품질을 높이기 위해 semantic segmentation feature를 auxiliary supervision으로 썼다. Depth Anything은 MiDaS를 포함한 이전 방법들을 KITTI, NYU, ScanNet, DIODE 등 모든 주요 벤치마크에서 앞질렀다. 모델 크기는 ViT-L 기반 335M 파라미터였고, 추론 속도는 실시간과 거리가 있었다. 같은 해 나온 [**Depth Anything v2**](https://arxiv.org/abs/2406.09414)는 Virtual KITTI, Hypersim 등의 합성 데이터로 teacher를 학습하고, 이 teacher가 실제 영상에 붙인 pseudo-label로 student 모델을 학습했다. 합성 데이터는 반사·투명 표면처럼 실제 데이터에서 annotation이 어려운 영역을 커버한다. v2는 v1보다 edge 세부와 얇은 구조 표현에서 눈에 띄게 개선되었다. 그러나 Depth Anything도 여전히 relative depth로, scale은 없다. [**ZoeDepth**(Shariq Farooq Bhat et al. 2023)](https://arxiv.org/abs/2302.12288)는 relative-depth 사전학습에 데이터 도메인별 metric-bin head와 자동 routing을 결합했다. [**Depth Anything v2**(2024)](https://arxiv.org/abs/2406.09414)는 강한 relative-depth backbone을 metric label로 별도 fine-tuning한 모델을 제공했다. [**Metric3D v2**(2024)](https://arxiv.org/abs/2404.15506)는 서로 다른 초점거리에서 생기는 모호성을 canonical camera space로 변환하고, 추론 뒤 camera parameter로 metric depth를 되돌리는 방식을 택했다. 세 방법은 모두 relative depth만 내던 계보를 metric prediction으로 넓혔지만, 학습 도메인과 카메라 설정을 넘어서는 일반화는 여전히 별개의 문제다. --- ## 5. SLAM으로의 역수입 2017년 CNN-SLAM부터 monocular depth 모델을 SLAM 안에 결합하는 시도가 있었다. 이후에는 초기화에도 학습한 깊이를 활용했다. Monocular SLAM은 구조상 초기화가 까다롭다. 두 프레임에서 triangulation을 하려면 baseline이 충분해야 하고, scale은 첫 단계부터 모호하다. depth prior를 첫 프레임에 주입하면 초기화가 빨라지고 장면 구조의 초기값을 제공할 수 있다. metric depth 모델이거나 알려진 기준으로 scale을 보정한 prior라면 metric scale의 초기값도 줄 수 있다. [Teed와 Deng이 2021년 발표한 DROID-SLAM](https://arxiv.org/abs/2108.10869)은 recurrent optical flow와 BA를 묶은 구조인데, 이 계통에서 나온 후속 연구들이 monocular depth prior를 geometric initialization에 붙이는 방식을 실험했다. scale recovery 쪽은 더 직접적이었다. monocular visual odometry(VO)는 달리면서 scale drift가 쌓인다. metric으로 보정된 depth prediction을 주기적인 scale anchor로 쓰면 이 drift를 억제할 수 있다. 반면 MiDaS 같은 relative-depth 출력만으로는 외부 metric 기준 없이 절대 스케일을 정할 수 없다. > 📜 **예언 vs 실제.** Eigen은 2014년 논문에서 surface normal 등 3D geometry 정보와의 결합을 자연스러운 확장 방향으로 언급했다. joint multi-task learning은 이후 PAD-Net·VPD 등으로 부분 실현됐다. 그러나 2024년 시점 실질적 영향은 task를 합친 것보다 ViT backbone 공유로 왔다고 볼 여지가 크다. 예측한 방향과 실제 경로는 달랐다. > 📜 **예언 vs 실제.** MiDaS(2020)는 scale-and-shift invariant loss로 절대 스케일을 포기하고 상대 깊이에 집중했다. 이후 ZoeDepth와 Depth Anything v2의 metric 모델은 metric label을 이용한 fine-tuning으로, Metric3D v2는 camera model 차이를 canonical space에서 보정하는 방식으로 이 문제를 공략했다. metric prediction의 범위는 넓어졌지만, 보지 못한 카메라와 도메인에서의 scale 일반화는 아직 진행형이다. --- ## 🧭 아직 열린 것 **반사·투명 표면의 depth.** 유리, 물, 금속 반사면은 카메라가 포착하는 것이 실제 표면이 아니다. 물리 광학 수준의 문제다. 합성 데이터로 학습을 늘려도 real-world 반사 장면에서의 일반화는 여전히 불안정하다. [ClearGrasp(Sajjan et al. 2020)](https://arxiv.org/abs/1910.02550) 같은 specialized 접근이 있으나 general solution은 없다. Foundation 규모 모델에서도 이 영역의 오차는 구조적으로 크다. **Dynamic scene에서 ego-depth와 object-depth의 분리.** 자동차·사람·자전거가 움직이는 장면에서 photometric consistency는 근본적으로 위반된다. self-supervised 방법들은 moving object를 masking해 우회했다. 우회였지 해법은 아니었다. 움직이는 물체의 depth를 에고모션과 분리해 동시에 풀어야 하는 문제는 [Ranjan et al.(2019)](https://arxiv.org/abs/1805.09806)을 비롯한 여러 후속 연구가 시도했으나 실용 수준에서는 여전히 난제다. **Metric scale의 일반화.** ZoeDepth와 Depth Anything v2의 metric variant는 metric label로 scale을 배우고, Metric3D v2는 camera parameter 차이를 명시적으로 다룬다. 그러나 intrinsic이나 신뢰할 만한 메타데이터를 모르는 CCTV·아카이브 영상도 흔하다. 학습 도메인과 카메라에 독립적인 metric depth는 foundation model 규모에서도 쉽지 않다. 이것이 2025년 시점 monocular depth의 남은 핵심 질문이다. --- 2024년 Depth Anything이 벤치마크 기록을 바꾸던 같은 시기, Cambridge의 한 논문은 이미 9년째 SLAM 커뮤니티의 미완성 숙제로 남아 있었다. [PoseNet](https://arxiv.org/abs/1505.07427)은 특징 추출이나 최적화 없이 한 장의 이미지에서 절대 pose를 바로 출력했다. 지도는 처음부터 존재하지 않았다. --- # Ch.12 — End-to-end 좌절 단안 카메라 하나로 depth를 복원하는 일이 가능해졌다. Eigen의 망은 픽셀에서 metric depth를 꺼냈고, SfMLearner는 라벨 없이도 기하학적 supervision을 만들어냈다. 그렇다면 pose 추정과 루프 클로저를 포함한 SLAM 전체를 하나의 망으로 끝낼 수 있을까. 2015년부터 2018년 사이, 이 물음은 답을 찾지 못했다. 2015년, Cambridge Computer Laboratory 박사과정 학생 Alex Kendall은 Roberto Cipolla 지도교수 아래 Google Street View 이미지로 학습한 신경망에 사진 한 장을 넣고 6-DoF pose를 출력하는 연구를 완성했다. [Kendall et al. 2015. PoseNet](https://doi.org/10.1109/ICCV.2015.336)이라 명명한 이 논문은 Santiago de Chile에서 열린 ICCV에서 발표됐다. SLAM의 30년짜리 방정식(특징 추출, 매칭, 최적화, 지도 관리)을 단일 CNN으로 압축할 수 있다면? 이 물음은 2015년부터 2018년까지 수십 편의 논문을 낳았고, 후속 연구들은 비슷한 한계를 반복해서 보고했다. --- ## 12.1 PoseNet PoseNet이 물려받은 것은 [AlexNet(Krizhevsky et al. 2012)](https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks)이었다. Kendall은 ImageNet에서 분류 task로 학습된 깊은 CNN이 고수준 시각 표현을 형성한다는 사실을 확인하고, 그 feature hierarchy를 pose estimation으로 전용했다. > 🔗 **차용.** PoseNet의 backbone은 [GoogleNet(Inception, Szegedy et al. 2014)](https://arxiv.org/abs/1409.4842) 구조다. Classification head를 제거하고 7차원 회귀 head(x, y, z, quaternion 4개)를 붙인 것이 전부다. ImageNet 학습으로 얻은 feature hierarchy를 localization에 직접 이식했다. Kendall이 직접 수집한 Cambridge Landmarks 데이터셋(킹스 칼리지 예배당, 거리, 옛 병원 등 여러 야외 장면)에서 PoseNet은 장면에 따라 위치 오차 2 m 안팎, 방향 오차 5-8° 수준을 기록했다(원 논문 §5 기준). 2015년 기준으로는 인상적인 수치였다. GPU 한 장으로 5 ms 이내에 답이 나왔다. 특징 추출도, RANSAC도, 지도 조회도 없었다. 후속 연구가 이어졌다. [Bayesian PoseNet(Kendall & Cipolla 2016)](https://arxiv.org/abs/1509.05909)은 Monte Carlo Dropout으로 자세 불확실성을 추정하려 했다. LSTM PoseNet은 시퀀스 정보를 통합했다. Geometric loss를 추가한 변형이 등장했다. Kendall 자신도 2017년에 재귀 구조와 photometric loss를 결합한 버전을 냈다. 그러나 비교 기준이 올라갈수록 격차가 드러났다. 같은 장면에서 [Active Search(Sattler et al. 2012)](https://www.graphics.rwth-aachen.de/media/papers/sattler_eccv12_preprint_1.pdf)나 DenseVLAD는 위치 오차 0.2 m 수준을 달성했다. PoseNet 계열은 수 미터 오차에서 좀처럼 벗어나지 못했다. 이미지 한 장에서 절대 자세를 회귀하는 접근에는 원론적 한계가 있었다. --- ## 12.2 DeepVO PoseNet의 한계 중 하나가 단일 이미지 입력이라면, 시퀀스를 입력하면 어떨까. Sen Wang(에든버러 Heriot-Watt)과 공저자들은 2017년 [Wang et al. 2017. DeepVO](https://arxiv.org/abs/1709.08429)를 ICRA에 발표했다. FlowNet에서 영향을 받은 CNN으로 연속 프레임 쌍의 optical flow feature를 추출하고, LSTM으로 시간 맥락을 축적해 VO를 직접 출력하는 구조였다. > 🔗 **차용.** DeepVO의 훈련 라벨은 KITTI의 GPS/IMU ground truth다. feature 추출 설계는 [FlowNet(Dosovitskiy et al. 2015)](https://arxiv.org/abs/1504.06852)의 optical flow CNN 구조에서 직접 차용했다. "deep VO"는 고전 센서 측정과 이전 딥러닝 연구 양쪽에 동시에 기댔다. LSTM이 temporal modeling을 맡으면서 drift 억제를 기대했다. KITTI 시퀀스 일부에서 DVO-SLAM이나 VISO2-M 대비 낮은 drift를 보이는 결과가 논문에 실렸다. 하지만 조건이 있었다. 훈련 시퀀스와 비슷한 주행 패턴, 비슷한 조명 조건, 비슷한 도시 풍경, 비슷한 속도 프로파일. 조건이 어긋나면 LSTM이 축적한 "맥락"은 오히려 편향이 되었다. Tinghui Zhou(UC Berkeley)가 같은 해에 발표한 [Zhou et al. 2017. SfMLearner](https://arxiv.org/abs/1704.07813)는 다른 각도로 접근했다. 자기지도(self-supervised) 학습으로 depth와 ego-motion을 동시에 추정하되, photometric reprojection loss를 학습 신호로 썼다. 라벨 없이 학습 가능하다는 점이 강점이었다. > 🔗 **차용.** SfMLearner의 photometric loss는 고전 direct SLAM이 사용하는 intensity residual과 수학적으로 동일하다. [DSO(Engel et al. 2018)](https://arxiv.org/abs/1607.02565)의 photometric 원리를 미분가능 학습 프레임워크로 옮겼다. SfMLearner의 self-supervision 아이디어는 MonoDepth2를 거쳐 결국 DROID-SLAM의 전제 조건 중 하나가 된다. 다만 SfMLearner의 짧은 구간 VO 평가는 장기 궤적과 루프 클로저까지 포함한 ORB-SLAM의 시스템 성능과 구별해서 읽어야 한다. --- ## 12.3 실패 원인 세 가지 2019년에서 2020년 사이, end-to-end pose regression과 learned odometry를 둘러싼 논의에서는 세 가지 한계가 반복해서 드러났다. **첫 번째: inductive bias의 부재.** 고전 SLAM은 수십 년에 걸쳐 축적된 기하학적 제약을 알고리즘 구조 안에 새겨 넣었다. epipolar constraint, rigid body motion 가정, scale invariance, 공간의 연속성. CNN은 이것들을 데이터로부터 새로 배워야 했다. ImageNet의 고양이와 자동차 사진이 3D 공간의 metric geometry를 가르쳐주지는 않는다. 회귀 망이 pose를 맞추는 것처럼 보여도, 실제로 그것이 3D 공간을 이해해서인지 아니면 특정 조명·색상·질감 조합을 외워서인지 구별하기 어려웠다. **두 번째: 일반화 실패.** 훈련 집합 밖으로 나가면 성능이 급락했다. 한 장면에서 학습한 PoseNet은 보지 못한 장면의 절대 위치를 그대로 추정할 수 없었고, KITTI에서 학습한 learned odometry도 카메라·주행 환경·영상 통계가 달라지면 오차가 커졌다. 고전 ORB-SLAM은 특징 검출에 실패하거나 조명이 극단적으로 변하면 추적을 잃었지만, 실패를 명시적으로 감지하고 재초기화하는 절차를 둘 수 있었다. end-to-end 모델은 잘못된 추정과 실패 감지를 함께 해결해야 했다. **세 번째: 불확실성 정량화의 부재.** SLAM이 단순한 pose 추정기로 끝나지 않는 이유는 downstream 시스템(경로 계획, 장애물 회피)이 위치 추정의 공분산을 요구하기 때문이다. EKF와 factor graph는 공분산을 자연스럽게 전파한다. Bayesian PoseNet이 dropout으로 분산을 추정하려 했지만, 그 분산이 실제 위치 오차와 calibrated 관계를 맺는지 검증하기 어려웠다. 특히 훈련 분포 밖 입력에서 Bayesian PoseNet은 오히려 자신만만한 틀린 답을 냈다. 틀린 것보다 자신만만하게 틀리는 것이 로봇 시스템에는 더 위험하다. --- ## 12.4 반성의 기록 박사학위를 마친 뒤인 2019년, Kendall은 Wayve로 자리를 옮겨 자율주행용 imitation learning과 world model 연구로 방향을 틀었다. 이후 연구 방향은 달라졌지만, 학습 기반 localization 자체를 포기한 것은 아니었다. Federico Tombari 그룹(TU Munich, 이후 Google)도 앞선 2017년에 [CNN-SLAM(Tateno et al. 2017)](https://arxiv.org/abs/1704.03489)을 시도했다. CNN이 예측한 dense depth를 직접(direct) monocular SLAM의 깊이 측정과 융합하려는 접근이었다. 학습 부분이 dense depth에 국한되었다는 점에서 완전한 end-to-end는 아니었지만, "CNN이 단안 SLAM의 스케일·저텍스처 문제를 해결해 줄 수 있지 않을까"라는 기대의 한 갈래였다. 성능은 장면에 따라 들쭉날쭉했고, 정확도에서 일관된 우위를 보이지 못했다. > 📜 **예언 vs 실제.** Kendall은 PoseNet 논문(2015)에서 불확실성 추정, temporal 정보 통합, 더 넓은 규모의 장면으로의 확장을 다음 과제로 꼽았다. Bayesian PoseNet(2016), LSTM PoseNet(2016), 복수의 outdoor 확장 실험으로 세 방향 모두 실행되었다. 그러나 각 시도가 새 벽에 부딪혔고, 이 접근법은 결국 주류에서 멀어졌다. 예언이 합리적이어도 기반 접근 자체의 한계가 크면 소용없다. 일부 시도는 다른 방향으로 살아남았다. SfMLearner의 photometric self-supervision은 MonoDepth2(Godard 2019) 같은 후속 단안 깊이 연구로 이어졌다. DROID-SLAM(Teed & Deng 2021)은 미분 가능한 기하 구조를 활용하지만, pose와 optical flow의 감독 신호로 학습한다. DeepVO가 보여준 LSTM 기반 temporal modeling은 시각-관성 학습 연구에서 변형된 형태로 재등장했다. 아이디어의 용도가 바뀌었을 뿐이다. > 📜 **예언 vs 실제.** Zhou는 SfMLearner 논문(2017)에서 dynamic object 처리와 photometric noise에 대한 강건성을 남은 과제로 제시했다. [GeoNet(Yin & Shi 2018)](https://arxiv.org/abs/1803.02276)을 비롯한 후속 self-supervised 연구들이 부분적으로 이 방향을 밀었다. 그러나 self-supervised VO 단독으로 SLAM을 대체하는 경로는 주류에 합류하지 못했다. photometric self-supervision 자체는 계보를 이어갔지만, end-to-end VO라는 목표는 주류가 되지 못했다. --- ## 12.5 교훈의 정착 2020년을 전후해 연구의 무게중심은 "geometry는 알고리즘, learning은 feature와 prior"라는 방향으로 옮겨갔다. > 🔗 **차용.** 이 원칙의 실천은 13장에서 다루는 CodeSLAM(Bloesch 2018)과 DROID-SLAM(Teed & Deng 2021)에서 구체화된다. 두 시스템 모두 factor graph 또는 bundle adjustment라는 기하학적 뼈대를 유지하고, 학습 부분은 CodeSLAM에서 depth 표현을, DROID-SLAM에서 dense correspondence와 반복 update를 맡는다. PoseNet이 버린 뼈대가 사실 포기할 수 없는 것이었다는 확인이다. 고전 파이프라인이 학습 기반 대안에 일관되게 우월한 것이 아니었다. ORB-SLAM도 textureless 환경에서, 야간에서, 비에서 자주 실패했다. 문제는 end-to-end의 오류가 더 불투명하고 더 예측 불가능하다는 데 있었다. 실패를 데이터셋이나 아키텍처 하나만으로 설명할 수는 없었다. 이미지에서 바로 pose로 직결하는 경로에 30년짜리 기하학 지식이 통째로 빠져 있었다. --- ## 🧭 아직 열린 것 **어떤 inductive bias를 어떻게 주입할 것인가.** "geometry는 알고리즘으로"라는 원칙은 맞지만, 어떤 기하학을 어느 수준에서 코드화해야 하는지는 여전히 열린 질문이다. rigid body motion인가, epipolar constraint인가. foundation model 시대에 이 경계는 다시 흐려지고 있다. GaussianSLAM이나 3DGS 기반 시스템이 geometry를 학습 표현 안에 녹이는 방식을 실험하고 있다. **Learned uncertainty의 calibration.** Bayesian PoseNet의 실패 이후에도 이 문제는 해결되지 않았다. 딥러닝 기반 uncertainty estimate가 실제 오차와 얼마나 calibrated 관계를 가지는지(특히 out-of-distribution 입력에서)는 2026년 기준으로도 열려 있다. 자율주행이 이 질문에 실용적 압력을 가하고 있다. **"End-to-end"의 의미 재정의.** PoseNet이 정의한 end-to-end(이미지→pose, 학습만으로)는 실패했다. 그러나 foundation model이 등장한 2023년 이후 end-to-end의 의미가 바뀌고 있다. SLAM의 어느 모듈을 학습으로 채우고 어느 모듈을 알고리즘으로 유지할 것인가. 이 분할선 자체가 재협상 중이다. "geometry는 알고리즘으로, learning은 feature로"라는 원칙이 이 시기에 굳어졌다. 2018년 Kensington의 Imperial College London, Andrew Davison 연구실에서 그 원칙을 구현한 CodeSLAM이 나왔다. --- # Ch.13 — Hybrid 승리: CodeSLAM에서 DROID-SLAM까지 Michael Bloesch가 2018년 CVPR에 CodeSLAM을 발표했을 때, 그의 소속은 Imperial College London의 Dyson Robotics Lab이었고 지도교수는 Andrew Davison이었다. 같은 연구실에서 2011년 Richard Newcombe가 DTAM을 만들었고, Jan Czarnowski가 2020년 DeepFactors를 내놓았으며, Edgar Sucar와 Tristan Laidlow가 계보를 이었다. Davison이 2002년부터 쌓아온 확률론적 SLAM 계보와 2010년대 중반 딥러닝의 "표현을 배울 수 있다"는 발상이 2018년 Bloesch의 논문에서 만났다. --- ## 13.1 CodeSLAM — latent code와 지도 전통적인 monocular SLAM에서 depth는 추정의 대상이었다. 수백 개의 sparse landmark이든, [DTAM](https://www.doc.ic.ac.uk/~ajd/Publications/newcombe_etal_iccv2011.pdf)(Newcombe et al. 2011)처럼 모든 픽셀이든, depth는 결국 최적화 변수였다. 그 변수 공간의 차원은 이미지 해상도에 비례했다. keyframe 한 장의 dense depth map은 640×480 해상도에서 307,200개의 독립 변수를 의미한다. 최적화는 무겁고, 초기화는 민감하고, prior를 넣기가 어렵다. [Bloesch et al. 2018. CodeSLAM](https://doi.org/10.1109/CVPR.2018.00271)은 depth map 자체가 아니라, depth map을 생성하는 저차원 잠재 벡터(**latent code**)를 최적화했다. Variational autoencoder(VAE)를 훈련해 실제 depth 분포를 학습시키면, 그 bottleneck latent space는 "사실적인 depth map"들이 사는 다양체를 근사한다. 최적화는 그 다양체 위에서만 움직인다. 변수가 수십만 개에서 수백 개로 줄어든다. > 🔗 **차용.** CodeSLAM의 latent depth 표현은 [Kingma & Welling 2013. VAE](https://arxiv.org/abs/1312.6114)에서 확립된 encoder-decoder 잠재 공간 구조를 차용했다. 학습 단계에서는 VAE 틀을 따르되, SLAM 추론 시에는 stochastic sampling 없이 **z**를 직접 MAP 최적화 변수로 다룬다. 생성 모델의 잠재 공간 도구가 약 5년 뒤 SLAM 최적화의 저차원 표현으로 쓰였다. keyframe마다 VAE encoder가 이미지에서 latent code **z**를 추출한다. Decoder는 **z**에서 dense depth map을 재구성한다. Camera pose와 **z**는 jointly 최적화된다. photometric loss가 consistency를 강제하고, latent prior가 **z**를 사전 분포 근방에 머물도록 regularize한다. 목적함수는 다음과 같다: $$E(\mathbf{z}, T) = \sum_{i,j} \rho\bigl(I_j(\pi(T_{ij}, D_\mathbf{z}(u_i), u_i)) - I_i(u_i)\bigr) + \lambda \|\mathbf{z}\|^2$$ $D_\mathbf{z}$는 decoder, $\pi$는 projection, $\rho$는 robust cost, $T_{ij}$는 keyframe 간 상대 pose다. latent prior 항 $\lambda\|\mathbf{z}\|^2$은 표준 정규 prior $p(\mathbf{z}) = \mathcal{N}(0, I)$의 negative log-likelihood에 해당하며, Gaussian prior 가정 아래 MAP inference에서 자연스럽게 등장하는 regularizer다. > 🔗 **차용.** Factor graph(Dellaert & Kaess의 [GTSAM](https://gtsam.org/tutorials/intro.html))는 DeepFactors의 backend 골격을 제공했다. CodeSLAM이 joint optimization으로 처리한 pose-latent 결합 구조를 Czarnowski는 명시적 factor graph로 재정식화했다. Learning이 만든 잠재 변수가 전통적인 pose node 옆에 또 하나의 graph 변수로 편입된 것은 DeepFactors에 이르러서다. 두 세계의 인터페이스가 graph의 edge였다. sparse 입력에서 geometry를 채워 넣는 능력이 기존 방법을 앞섰다. 그러나 CodeSLAM 자체는 실시간이 아니었다. VAE 추론과 최적화 루프가 느렸고, 논문에도 이 한계가 명시됐다. > 📜 **예언 vs 실제.** CodeSLAM은 compact learned representation을 dense SLAM 안에 들이는 가능성을 보였지만 속도와 규모 양쪽에서 여지를 남겼다. 이어진 DeepFactors(2020)가 같은 Imperial 그룹에서 실시간 쪽으로 한 발 더 나아갔으나 상용 배포 수준은 되지 못했고, monocular·stereo·RGB-D를 아우르는 범용 성능은 결국 다른 팀(Teed·Deng, Princeton)이 학습된 frontend + dense BA라는 다른 설계로 달성했다. --- ## 13.2 DeepFactors — Imperial Dyson Lab, factor graph 통합 2020년, Jan Czarnowski도 Davison 지도 아래 Imperial Dyson Robotics Lab에서 [Czarnowski et al. 2020. DeepFactors](https://doi.org/10.1109/LRA.2020.2969036)를 발표했다. Czarnowski의 목표는 CodeSLAM의 아이디어를 실제 SLAM 파이프라인 안으로 끌어들이는 것이었다. DeepFactors는 CodeSLAM의 latent depth 표현과 joint optimization을 명시적인 factor graph로 확장하면서 tracking과 mapping을 명시적으로 분리하고, keyframe 선택 기준을 도입했다. NVIDIA GTX 1080 위에서 keyframe 대비 tracking은 약 250Hz로 돌았으나, network Jacobian 계산이 keyframe당 수백 밀리초를 차지해 전체 파이프라인의 병목이었다. 방향은 보여주었으되 상용 배포 수준의 실시간에는 미치지 못했다. DeepFactors는 learned representation을 factor graph의 한 노드로 넣을 수 있고, geometry optimization이 그 latent space 위에서 작동할 수 있음을 보였다. Czarnowski는 파이프라인 일부를 학습 가능한 모듈로 바꾸는 경로를 택했다. 같은 시기 TU Munich의 Daniel Cremers 그룹도 같은 원칙에 도달했다. 출발점이 Imperial과 달랐을 뿐이다. Davison 계보가 CodeSLAM의 VAE latent 위에 factor graph를 쌓았다면, Cremers 그룹은 자신들이 2016년 내놓은 direct sparse odometry([DSO](https://arxiv.org/abs/1607.02565))를 뼈대로 두고 거기에 neural prediction을 주입했다. [Yang, Wang, Stückler, Cremers 2018. DVSO](https://arxiv.org/abs/1807.02570)는 단안 DSO에 neural depth를 "가상 스테레오"로 주입해 두 번째 카메라에 해당하는 관측을 만들었고, [Yang, von Stumberg, Wang, Cremers 2020. D3VO](https://arxiv.org/abs/2003.01060)는 self-supervised로 학습된 depth·pose·uncertainty 세 종류의 neural prediction을 DSO의 factor graph에 추가 factor로 넣었다. [Wimbauer et al. 2021. MonoRec](https://arxiv.org/abs/2011.11814)과 [Wimbauer et al. 2023. Behind the Scenes](https://arxiv.org/abs/2301.07668)는 같은 계보를 dynamic scene dense reconstruction과 single-view density field 쪽으로 이어갔다. 인적 계보는 Imperial 그룹과 분리되어 있지만, "neural prediction을 고전 optimization 구조 안으로 흡수한다"는 설계 원칙은 수렴했다. 그 원칙은 2021년 Princeton에서 또 다른 방식으로 다시 나타났다. --- ## 13.3 RAFT — recurrent optical flow Zachary Teed와 Jia Deng(Princeton)은 2020년 ECCV에 [Recurrent All-Pairs Field Transforms(RAFT)](https://arxiv.org/abs/2003.12039)를 발표했다. RAFT는 SLAM 논문이 아니었다. optical flow 추정 논문이었다. RAFT의 설계는 이후 DROID-SLAM에 들어갔다. 구조는 세 부분으로 나뉜다. 1. Feature encoder: CNN이 두 이미지에서 feature map을 추출한다. 2. Correlation volume: 모든 픽셀 쌍 간의 유사도를 4-level pyramid의 4D volume으로 구성한다. 3. Update operator: Gated Recurrent Unit(GRU)이 correlation volume을 lookup하며 flow field를 반복해서 refinement한다. 이름에 들어간 all-pairs가 이 구조의 차별점을 요약한다. 모든 후보 위치를 동시에 고려하고, 고정 해상도에서 flow field를 점진적으로 refinement한다. 기존 coarse-to-fine 방법(PWC-Net 등)과 달리 flow field를 원영상의 가로·세로 1/8인 단일 해상도로 유지한 채 correlation pyramid를 lookup하고, 최종 출력을 upsampling한다. RAFT 논문은 당시 공개 최고 결과와 비교해 KITTI의 F1-all 오류를 16%, Sintel final pass의 endpoint error를 30% 상대 감소시켰다고 보고했다. RAFT가 처음 풀던 과제는 SLAM이 아니라 optical flow였다. 그러나 Teed는 같은 update operator 구조가 SLAM의 iterative bundle adjustment와 구조적으로 유사하다는 점을 DROID-SLAM 설계에 활용했다. flow field를 refinement하는 GRU가 pose와 depth를 refinement하는 최적화 스텝과 얼마나 다른가. --- ## 13.4 DROID-SLAM — update operator와 BA 2021년 NeurIPS, [Teed & Deng. DROID-SLAM](https://arxiv.org/abs/2108.10869). 제목의 DROID는 "Differentiable Recurrent Optimization-Inspired Design"의 약자다. 아키텍처를 따라가면 hybrid 설계의 의도가 드러난다. Frontend는 RAFT와 동일한 구조다. CNN encoder가 feature map을 추출하고, all-pairs correlation volume을 구성하고, GRU update operator가 dense flow를 반복 추정한다. 차이는 keyframe graph의 모든 edge에서 동시에 flow를 추정한다는 점이다. Backend는 Dense Bundle Adjustment(DBA)다. Pose와 inverse depth가 optimization 변수다. flow 추정이 제공하는 2D correspondence를 제약으로 사용해 pose-depth를 jointly 최적화한다. Schur complement trick으로 선형 시스템을 효율적으로 푼다. **DBA layer**가 두 모듈을 연결한다. GRU가 추정한 flow와 uncertainty가 DBA에 입력된다. DBA가 pose·depth를 업데이트하면, 그 결과가 다음 GRU iteration의 reference를 갱신한다. > 🔗 **차용.** Dense BA라는 아이디어 자체는 10년 전으로 거슬러 올라간다. Newcombe의 DTAM(2011)은 모든 픽셀을 사용한 photometric bundle adjustment의 선구자였다. DROID-SLAM은 그 아이디어를 learned flow라는 더 강건한 입력과 결합했다. Newcombe와 Teed의 affiliations는 다르지만 논리적 계보는 이어진다. > 🔗 **차용.** DROID-SLAM의 update operator는 같은 저자(Teed·Deng)의 RAFT에서 직접 이식했다. optical flow를 위해 설계된 all-pairs recurrent refinement가 bundle adjustment의 반복 최적화와 구조적으로 호환된다는 통찰이 핵심이었다. 두 논문의 저자가 같아 차용 관계도 직접적이다. EuRoC MAV 데이터셋에서 DROID-SLAM은 당시 최고 수준인 ORB-SLAM3보다 낮은 RMSE ATE를 기록했고, TartanAir(합성)와 실제 실내외 시퀀스 양쪽에서 특히 조명 변화와 texture 부족 상황에 feature-based 방법보다 강건했다. Teed가 이후 Handbook 회고에서 EuRoC V1_02 시퀀스에 대해 보고한 수치가 인상적이다. frontend만 돌렸을 때 16.5cm이던 ATE가 global optimization을 거치면 1.2cm로 떨어졌다. learned correspondence가 공급한 제약 위에서 고전 BA가 한 자릿수 cm까지 수렴하는 장면이다. PoseNet은 geometry constraint 없이 pose를 직접 회귀했고 일반화에 실패했다. Teed와 Deng은 역할을 나눴다. dense correspondence 추정은 학습에 맡기고, geometry 제약 강제는 BA가 담당했다. feature 추출·dense matching에서는 신경망이, consistency enforcement·uncertainty 전파에서는 geometry optimizer가 각자 강점을 살렸다. 인간이 설계한 feature를 학습된 feature로 교체하면서도 optimization structure는 그대로 유지됐다. 이 지점에서 2021년의 hybrid는 2015년의 end-to-end와 갈라졌다. --- ## 13.5 Imperial Dyson Lab 계보도 CodeSLAM에서 DROID-SLAM까지의 흐름은 Imperial Dyson Robotics Lab의 인적 계보를 따라가면 제대로 보인다. Andrew Davison은 2002년 MonoSLAM 이후 20년 동안 Imperial에서 SLAM 연구를 이끌었다. 제자와 협력자들이 차례로 분기점을 만들었다. - **Richard Newcombe** (Davison 지도, Imperial): DTAM(2011), [KinectFusion](https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ismar2011.pdf)(2011). 이후 Oculus→Meta Reality Labs - **Michael Bloesch** (Davison 지도, Imperial): CodeSLAM(2018), touch·inertial SLAM 연구 - **Jan Czarnowski** (Davison 지도, Imperial): DeepFactors(2020) - **Edgar Sucar** (Davison 그룹, Imperial): [iMAP](https://arxiv.org/abs/2103.12352)(2021), 이후 NeRF-SLAM 계보로 연결 - **Tristan Laidlow** (Davison 그룹, Imperial): dense 3D reconstruction, 이후 neural implicit SLAM 계보로 연결 이 계보는 "학파"라는 단어가 과장이 아닌 경우 중 하나다. factor graph + uncertainty 철학이 단안 sparse에서 dense latent로, dense latent에서 implicit representation으로 모양을 바꾸며 이어졌다. Davison은 [FutureMapping](https://arxiv.org/abs/1803.11288)(2018)과 [FutureMapping 2](https://arxiv.org/abs/1910.14139)(2019, Ortiz 공저)에서 Spatial AI 시스템이 갖춰야 할 계산 구조와 표현을 직접 지도 위에 스케치했다. 다양한 geometric·semantic 표현을 하나의 확률 그래프 위에 묶자는 주장이었다. CodeSLAM과 DeepFactors는 그 스케치의 첫 번째 실험들이었다. Teed와 Deng은 Princeton에서 독립적인 경로를 걸었다. 그러나 DROID-SLAM이 채택한 dense BA의 논리적 전임자는 DTAM이고, DTAM은 Newcombe·Davison의 작품이다. 계보는 인적 연결을 건너뛰고도 논리로 이어진다. --- ## 13.6 2023-2025 — DROID 이후의 확장 DROID-SLAM 이후 몇 년간 여러 후속 연구가 나왔다. [GO-SLAM](https://arxiv.org/abs/2309.02436)(Zhang et al. 2023)은 DROID-SLAM의 tracking을 확장해 online loop closing과 full bundle adjustment를 얹고, mapping은 Instant-NGP 계열 neural implicit 표현(multi-resolution hash encoding)으로 돌렸다. tracking은 DROID 계열 dense flow + BA, map은 implicit representation. hybrid의 두 번째 층이다. [NICER-SLAM](https://arxiv.org/abs/2302.03594)(Zhu et al. 2023)은 다른 길을 갔다. tracking과 mapping을 하나의 hierarchical neural implicit representation 위에서 동시에 풀었다. RGB-only dense SLAM이라는 목표는 공유하지만 경로가 다르다. DROID 계보의 외곽에서 같은 문제에 부딪히는 방식이다. [SplaTAM](https://arxiv.org/abs/2312.02126)(Keetha et al. 2024)은 map representation을 3D Gaussian Splatting으로 바꾸고, tracking도 silhouette-guided differentiable rendering 기반으로 다시 짰다. 3DGS 계보와의 결합이지 DROID 계보의 직접 확장은 아니다. [DPV-SLAM](https://arxiv.org/abs/2408.01654)(Lipson, Teed, Deng 2024)은 DROID-SLAM과 같은 Princeton 그룹에서 나왔다. [DPVO](https://github.com/princeton-vl/DPVO)(Deep Patch Visual Odometry)를 기반으로, 근접 기반 loop closure와 CUDA block-sparse BA를 추가해 DROID-SLAM 대비 약 2.5배 빠르고 메모리 footprint가 작은 시스템을 만들었다. 핵심은 patch 기반 sparse 표현 + 효율적 loop closure다. DROID 계보 바깥에서는 Naver Labs가 열어놓은 [DUSt3R](https://arxiv.org/abs/2312.14132)(Wang et al. 2023) 계보에서 2024-2025년 확장이 이어졌다. DUSt3R가 두 이미지에서 pointmap을 직접 출력해 SfM의 절차 자체를 재정의한 뒤, 같은 Revaud 그룹이 [Cabon et al. 2025. MUSt3R](https://arxiv.org/abs/2503.01661)에서 symmetric multi-view 확장과 working memory를 도입해 이미지 쌍 단위였던 구조를 다수 프레임으로 늘렸다. offline SfM과 online VO/SLAM을 같은 네트워크로 처리할 수 있게 한 시도다. DROID 계열 도구들도 이 생태계 안에서 재활용된다. [Li et al. 2024. MegaSAM](https://arxiv.org/abs/2412.04463)은 DROID-SLAM의 differentiable dense BA를 dynamic scene과 uncalibrated 영상 쪽으로 밀어붙여 camera intrinsic까지 inference 도중 공동 최적화했다. NVIDIA의 [Huang et al. 2025. ViPE](https://arxiv.org/abs/2508.10934)는 DROID-SLAM의 dense flow network와 cuvslam의 sparse point, monocular depth network까지 세 종류의 제약을 하나의 dense BA로 결합해 유튜브 규모의 wild video annotation 파이프라인으로 산업화했다. learned frontend + classical backend라는 2021년 DROID의 설계가 2025년에는 calibration-free와 dynamic scene이라는 더 어려운 조건 위에서 반복되고 있다. 확장은 단일 경로를 따르지 않았다. GO-SLAM처럼 DROID tracking 위에 neural map을 얹는 길, DPV-SLAM처럼 patch odometry로 가볍게 재설계하는 길, NICER-SLAM이나 SplaTAM처럼 implicit/splatting 표현 위에서 tracking을 새로 쓰는 길, MegaSAM·ViPE처럼 DROID의 dense BA를 uncalibrated·dynamic 조건으로 밀어붙이는 길이 동시에 진행됐다. 2021년 Teed와 Deng의 learned frontend + classical backend는 여러 후속 시스템의 출발점이 됐지만, NICER-SLAM과 SplaTAM은 같은 문제를 별도 표현에서 풀어간 경로였다. > 📜 **예언 vs 실제.** DROID-SLAM은 differentiable BA를 end-to-end 학습과 결합한 hybrid의 기준점을 세웠다. 같은 그룹이 3년 뒤 내놓은 DPV-SLAM은 그 기준점을 efficiency 쪽으로 이어받았다. 반면 GO-SLAM·NICER-SLAM·SplaTAM 계열은 map representation을 implicit 혹은 Gaussian splatting으로 갈아끼우는 쪽으로 갈라져 나갔다. "learned frontend + classical backend"라는 DROID의 설계가 여러 갈래로 변주되는 중이며, 어느 갈래가 범용 해법이 될지는 2026년 현재 아직 결론이 나지 않았다. --- ## 🧭 아직 열린 것 Learned prior의 분포 밖 일반화가 첫 번째 문제다. CodeSLAM과 DeepFactors의 VAE는 훈련 데이터의 depth 분포를 학습한다. 완전히 다른 환경(실외 open-world, 비균질 texture, 야간)에서는 learned prior가 오히려 최적화를 잘못된 방향으로 당길 수 있다. DROID-SLAM의 flow estimator도 훈련 도메인 밖에서 성능이 떨어진다. 2026년 현재, "어떤 환경에서도 작동하는 learned SLAM"은 아직 없다. TartanAir처럼 다양한 합성 데이터로 훈련하는 접근이 있으나 sim-to-real gap이 남는다. 실시간 제약도 여전하다. DROID-SLAM은 NVIDIA RTX 2080Ti 기준으로 평균 10-15 fps 수준이다. keyframe graph 크기에 따라 더 느려진다. dense BA가 병목이다. 모바일 로봇이나 AR/VR처럼 실시간(30Hz+), 저전력 배포가 필요한 응용에서는 2026년 현재도 실용적이지 않다. 경량화 시도들(keyframe 수 줄이기, approximate BA)이 있으나 성능 trade-off가 따른다. Loop closure의 learned 통합도 미해결이다. DROID-SLAM은 loop closure를 명시적으로 다루지 않는다. Teed는 이후 Handbook 회고에서 "DROID-SLAM doesn't include any relocalization module, so large loops with lots of drift cannot be closed"고 명시했다. 국소 frontend와 달리 backend는 전체 keyframe에 걸쳐 global BA를 수행한다. 다만 별도의 relocalization이 없어 drift가 큰 루프에서 필요한 대응을 복구하는 데 한계가 있다. learned loop closure(10장의 place recognition 연구들)를 DROID의 factor graph에 통합하는 시도가 일부 있으나 아직 단일 시스템으로 수렴하지 않았다. 10장 NetVLAD 계보와 13장 DROID 계보가 만나는 지점이 아직 열려 있다. --- DROID-SLAM의 inverse depth map은 2021년의 강한 dense representation이었다. 그러나 2020년 [NeRF](https://arxiv.org/abs/2003.08934)(Neural Radiance Field)는 장면을 포인트나 메시가 아니라 연속 함수로 표현하는 다른 가능성을 제시했다. 렌더링이 미분 가능하다면 photometric consistency를 새로운 방식으로 강제할 수 있다. 2021년 Imperial College의 Edgar Sucar는 이를 rendering이 아니라 SLAM의 문제로 다뤘다. MLP가 TSDF voxel grid를 완전히 대체할 수 있는가. 14개월 동안 개발한 iMAP이 그 질문에 대한 답이었다. --- # Ch.14 — NeRF 충격과 SLAM 접목: iMAP→NICE-SLAM DROID-SLAM은 learned representation을 SLAM의 tracking에 직접 넣었다. MLP나 recurrent network가 feature를 만들고, 그 feature 위에서 포즈를 최적화했다. Imperial College의 Edgar Sucar는 2021년 iMAP에서 learned representation을 *지도 자체*에 사용했고, 그 재료를 SLAM 바깥의 NeRF에서 가져왔다. 2020년 3월, Ben Mildenhall과 동료들이 arXiv에 올린 [Mildenhall et al. 2020. NeRF](https://arxiv.org/abs/2003.08934)는 pose를 아는 여러 입력 영상으로 연속 volumetric scene function을 최적화하고, 보지 않은 시점의 영상을 합성했다. NeRF는 처음에는 novel-view synthesis 문제의 해법으로 제시됐다. 2021년 ICCV에서 Sucar가 iMAP을 발표하면서, 이 표현을 렌더링뿐 아니라 지도와 pose를 함께 추정하는 데 쓸 수 있다는 점이 드러났다. iMAP은 KinectFusion(Ch.9)의 dense mapping 문제를 implicit neural field로 다시 풀어 본 초기 시스템이었다. NeRF가 허공에서 나온 것은 아니었다. 2019년 한 해 동안 coordinate-based MLP로 3D를 표현하는 세 갈래가 거의 동시에 나왔다. [Park et al.의 DeepSDF](https://arxiv.org/abs/1901.05103)는 좌표를 넣으면 signed distance를 출력하는 MLP로 물체 표면을 암묵적으로 기술했고, [Mescheder et al.의 Occupancy Networks](https://arxiv.org/abs/1812.03828)는 같은 좌표 입력에서 occupancy 확률을 출력하게 만들었으며, [Sitzmann et al.의 SRN](https://arxiv.org/abs/1906.01618)은 좌표마다 scene feature vector를 저장해 differentiable ray marching으로 이미지를 합성했다. 세 연구는 좌표를 넣으면 field 값을 출력하는 같은 수학적 틀을 공유했다. Mildenhall et al. 2020 NeRF는 이 틀에 volume rendering 적분과 positional encoding을 더해 view synthesis까지 완성했다. iMAP이 이어받은 것은 그 1년짜리 계보 전체였다. --- ## NeRF: MLP 기반 공간 표현 NeRF는 하나의 MLP로 3D 공간을 암묵적으로 표현한다. 공간 좌표 $(x, y, z)$와 시선 방향 $(\theta, \phi)$을 입력하면 그 위치의 색상 $(r, g, b)$과 밀도 $\sigma$를 출력한다. 이 출력을 광선마다 적분해 장면을 렌더링한다. 렌더링은 volume rendering 방정식으로 이루어진다. 카메라 원점 $\mathbf{o}$에서 방향 $\mathbf{d}$로 나간 광선을 $t$ 매개변수로 샘플링한다: $$\hat{C}(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma\!\left(\mathbf{r}(t)\right) \mathbf{c}\!\left(\mathbf{r}(t), \mathbf{d}\right)\, dt$$ 여기서 $T(t) = \exp\!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right)$는 광선이 거기까지 막히지 않고 도달할 누적 투과율이다. 실제로는 이 적분을 구간별 리만 합으로 근사한다. > 🔗 **차용.** Volume rendering 방정식은 [Kajiya & Von Herzen(1984)](https://courses.cs.duke.edu/cps296.8/spring03/papers/RayTracingVolumeDensities.pdf)의 고전 그래픽스 논문에서 왔다. 40년 가까이 오프라인 렌더링의 물리 기반 도구였던 것을 Mildenhall은 역방향 최적화의 손실 함수로 전환했다. MLP가 고주파 공간 신호를 학습하지 못하는 문제를 Mildenhall et al.(2020) NeRF 논문은 positional encoding으로 풀었다. 좌표 $(x, y, z)$를 사인·코사인 함수로 여러 주파수에 걸쳐 투영하면, 네트워크가 세밀한 텍스처와 날카로운 경계를 학습할 수 있다: $$\gamma(p) = \left(\sin(2^0 \pi p),\, \cos(2^0 \pi p),\, \ldots,\, \sin(2^{L-1} \pi p),\, \cos(2^{L-1} \pi p)\right)$$ > 🔗 **차용.** NeRF의 positional encoding은 Mildenhall et al.(2020) 원 논문에 포함된 것이다. 같은 해 [Tancik et al.(2020)](https://arxiv.org/abs/2006.10739) "Fourier Features Let Networks Learn High Frequency Functions"가 NTK(neural tangent kernel) 이론으로 이 기법의 작동 원리를 설명했다. NeRF의 학습은 렌더링 과정의 역방향 최적화로 이뤄진다. 알고 있는 카메라 포즈에서 찍은 이미지들과 렌더링 결과를 비교해 픽셀 단위 L2 손실을 최소화한다. 최적화가 끝나면 MLP 가중치 자체가 장면의 geometry와 appearance를 저장한다. 복셀도, 메시도, 포인트클라우드도 쓰지 않는다. 공간은 네트워크 파라미터 안에 있다. 그러나 원래 NeRF에는 뚜렷한 약점이 있었다. 학습에 수 시간이 걸렸고, 한 장면에 특화되었으며, 카메라 포즈는 COLMAP 같은 외부 SfM으로 미리 구해야 했다. 이것을 SLAM에 이식하려면 포즈 추정과 지도 학습을 실시간에 가깝게 동시에 수행해야 한다. --- ## iMAP: 초기 neural implicit SLAM Imperial College Dyson Robot Learning Lab의 Edgar Sucar가 2021년 ICCV에 발표한 [Sucar et al. 2021. iMAP](https://doi.org/10.1109/ICCV48922.2021.00612)은 그 시도였다. **iMAP**(Implicit MAP)은 RGB-D 카메라의 입력을 받아 단일 MLP를 지도로 쓰면서 포즈를 동시에 최적화했다. 두 개의 교번 최적화 루프가 동작한다. *mapping* 루프는 현재 키프레임과 과거 랜덤 샘플 키프레임에서 광선을 샘플링해 MLP를 업데이트한다. *tracking* 루프는 MLP를 고정하고 현재 프레임의 포즈를 렌더링 손실로 최적화한다. 두 루프는 공유된 단일 MLP 위에서 동작한다. 손실 함수는 두 가지다. 색상 손실 $\mathcal{L}_{\text{color}} = \|\hat{C} - C\|_2^2$과 깊이 손실 $\mathcal{L}_{\text{depth}} = \|\hat{D} - D\|_2^2$. RGB-D를 쓰므로 depth supervision이 있어 geometry 학습이 안정적이었다. iMAP은 개념 증명이었다. 소규모 실내 장면에서 동작했지만 두 가지 구조적 문제가 있었다. 첫째, 단일 MLP는 새로운 영역을 학습할 때 공유 가중치 전체가 바뀌므로 이전 영역의 표현이 훼손될 수 있었다. Sucar는 keyframe replay로 이를 완화했으나 지역별로 독립된 메모리를 제공하지는 않았다. 둘째, 장면이 커질수록 하나의 MLP가 더 많은 공간 변화를 같은 파라미터 집합에 담아야 했다. 각 점의 질의 비용이 파라미터 수와 무관한 것은 아니며, 핵심 제약은 표현과 업데이트에 공간적 지역성이 없다는 데 있었다. > 📜 **예언 vs 실제.** Sucar는 iMAP 논문 Conclusion에서 "future directions for iMAP include how to make more structured and compositional representations that reason explicitly about the self similarity in scenes"라고 적었다. 구조화·합성적 표현 방향은 실제로 후속 연구의 중심 줄기가 되었다. 5개월 뒤 ETH 취리히의 NICE-SLAM 사전공개는 multi-resolution voxel feature grid로 공간을 계층적으로 쪼갰고, Wang et al.의 Co-SLAM(2023)은 hash grid와 coordinate encoding을 합성해 RTX 3090 Ti에서 초당 10-17프레임을 보고했다. 다만 "self-similarity를 명시적으로 추론하는" 쪽은 NeRF-SLAM 본류에서 크게 발전하지 않았고, 단일 MLP를 정교화하는 계보 역시 중심에서 밀려났다. --- ## NICE-SLAM: 계층 격자와 scalability iMAP의 단일 MLP 문제에 대한 직접적인 답은 ETH 취리히의 Zihan Zhu·Songyou Peng이 2022년 CVPR에서 발표한 [Zhu et al. 2022. NICE-SLAM](https://arxiv.org/abs/2112.12130)에서 나왔다. **NICE-SLAM**(Neural Implicit Scalable Coding for SLAM)은 단일 MLP 대신 multi-resolution voxel feature grid와 작은 MLP decoder를 결합했다. 공간을 명시적 복셀 격자로 나누고 각 복셀에 학습 가능한 feature vector를 둔다. 렌더링 시 샘플 좌표 주변 복셀들의 feature를 trilinear interpolation으로 결합한 뒤 작은 MLP에 통과시켜 색상과 occupancy를 얻는다. MLP는 크지 않아도 된다. 공간 정보의 대부분은 격자에 담겨 있기 때문이다. NICE-SLAM은 세 단계 해상도 격자를 계층적으로 쌓았다. 거친 격자부터 세밀한 격자까지 geometry를 여러 수준으로 표현하고, 색상은 별도의 color feature grid와 decoder로 표현한다. 새로운 영역이 추가되면 해당 복셀의 feature만 업데이트하면 되므로 다른 영역의 catastrophic forgetting이 크게 줄어든다. tracking에서 NICE-SLAM은 iMAP과 유사하게 MLP와 격자 feature를 고정하고 포즈를 최적화했다. mapping에서는 격자 feature를 업데이트했다. Replica·ScanNet 데이터셋에서 iMAP보다 넓은 공간을 다뤘고 세부 표현 품질도 높았다. 그러나 한계가 있었다. 격자 자체의 메모리가 해상도의 세제곱으로 증가했다. 실내 방 한두 개는 다룰 수 있었지만 복층 건물이나 야외로의 확장은 여전히 미해결이었다. 속도도 실시간과 거리가 있었다. Thomas Müller의 [Müller et al. 2022. Instant-NGP](https://nvlabs.github.io/instant-ngp/)는 2022년 SIGGRAPH에서 이 병목을 다른 각도에서 공략했다. Multi-resolution hash encoding과 GPU 구현을 결합해 학습을 크게 가속했고, 공개 구현은 지원되는 NVIDIA GPU의 일부 데모 장면에서 수 초 훈련을 시연했다. Instant-NGP는 SLAM 논문이 아니었지만, 이후 여러 NeRF-SLAM 시스템이 이 hash encoding을 채용했다. > 🔗 **차용.** NICE-SLAM의 multi-resolution feature grid는 Instant-NGP의 hash encoding과 시기적으로 겹치며 독립적으로 설계되었지만, 실제 NeRF-SLAM 구현에서는 Instant-NGP의 hash grid가 NICE-SLAM 격자를 빠르게 대체했다. TSDF를 격자에 저장하던 KinectFusion(Ch.9)의 논리적 후계가 feature를 격자에 저장하는 방식으로 이어진 계보이기도 하다. --- ## Co-SLAM과 NeRF-SLAM: 두 가지 통합 방향 iMAP·NICE-SLAM 이후 2022년 말부터 여러 시스템이 두 갈래로 나뉘었다. 한 방향은 implicit representation을 더 효율적으로 만드는 것, 다른 방향은 전통 SLAM의 강건한 backend를 NeRF map과 결합하는 것이었다. UCL의 [Wang et al.(2023) **Co-SLAM**](https://arxiv.org/abs/2304.14377)은 전자에 속한다. joint coordinate·parametric encoding을 써서 multi-resolution hash grid와 one-blob 인코딩을 결합했다. 두 표현이 서로 보완하도록 설계해 빠른 수렴과 surface completeness를 함께 노렸다. hash grid가 관측된 dense 영역을 빠르게 채우고, coordinate encoding이 미관측 영역에 smooth prior를 제공하는 방식이었다. 논문은 Replica 데이터셋과 RTX 3090 Ti 환경에서 초당 15-17프레임의 처리량을 보고했다. NeRF 기반 SLAM이 준실시간 영역에 닿은 사례였다. 같은 해 같은 CVPR에서 Idiap/EPFL의 [Johari et al.의 **ESLAM**](https://arxiv.org/abs/2211.11704)은 비슷한 문제를 다른 각도에서 풀었다. 3D feature grid 대신 multi-scale axis-aligned feature plane을 써 메모리 증가를 $O(n^3)$에서 $O(n^2)$로 낮추고, volume density 대신 TSDF를 decoding 목표로 삼아 수렴을 가속했다. Antoni Rosinol(MIT)이 2022년에 공개한 [**NeRF-SLAM**](https://arxiv.org/abs/2210.13641)은 다른 접근이었다. 포즈와 dense depth, 그 불확실성은 DROID-SLAM이 제공하고, 별도의 Instant-NGP 기반 mapping 모듈이 이를 받아 radiance field를 학습했다. > 🔗 **차용.** NeRF-SLAM은 DROID-SLAM의 recurrent update와 Dense Bundle Adjustment가 만든 pose·depth 추정을 가져오고, 지도 표현에 Instant-NGP를 결합했다. Neural radiance field가 기존 pose 추정기를 대체한 것이 아니라, 그 출력 위에 실시간 dense map을 쌓는 구성이었다. NeRF-SLAM은 모듈성을 택했다. Radiance field가 tracking을 직접 맡게 하지 않고, dense monocular SLAM이 내놓은 pose·depth·불확실성과 neural mapping을 분리했다. 이 구성의 장점은 추정기와 지도 표현을 각각 개선할 수 있다는 데 있었다. --- ## iMAP의 구조적 한계와 그 의미 iMAP은 단일 MLP가 전체 장면을 기억하고, 그 MLP를 실시간에 가깝게 업데이트하면서 포즈까지 최적화하는 초기 neural implicit SLAM 시스템이었다. 단일 MLP의 근본 문제는 지역성(locality)의 부재다. 공간의 어떤 부분을 렌더링하든 MLP 전체를 통과한다. 결과가 두 가지다. 첫째, 새 영역을 학습하면 가중치 전체가 바뀌어 기존 영역의 표현이 훼손된다(catastrophic forgetting). 둘째, 장면이 커질수록 단일 MLP가 담아야 할 공간 다양성이 늘어나 더 큰 네트워크, 더 많은 이터레이션이 필요해진다. 표현 용량은 파라미터 수에 선형으로 묶여 있는데 장면 복잡도는 공간 부피에 따라 커진다. 지역성 없는 표현은 규모가 커질수록 불리하다. NICE-SLAM의 격자, Instant-NGP의 hash encoding, Co-SLAM의 이중 인코딩은 모두 이 지역성 문제의 답이었다. 공간을 국소적으로 나눠 각 부분이 자신의 영역만 기억하게 하면, 새 정보 추가가 기존 기억을 덜 침범하고, 특정 영역 렌더링 비용이 전체 장면 크기와 분리된다. --- ## 🧭 아직 열린 것 **실시간 NeRF-SLAM.** 2023년 기준 iMAP·NICE-SLAM은 실시간과 거리가 있었고, Co-SLAM이 RTX 3090 Ti에서 초당 10-17프레임의 준실시간 처리량을 보고했지만 모바일·로봇 임베디드 환경의 실시간에는 여전히 미치지 못했다. Gaussian Splatting(Ch.15)이 명시적 표현으로의 복귀를 통해 속도 문제를 다른 방식으로 해결했지만, implicit neural field 자체의 고전적 실시간 SLAM(30fps 이상, 소비자 GPU 없이)은 미완으로 남아 있다. Instant-NGP가 렌더링 속도를 극적으로 높였음에도 동시 추적·지도 구축 루프의 전체 처리량은 여전히 제약이 있다. **대규모 야외 환경.** [Block-NeRF](https://arxiv.org/abs/2202.05263)(2022, Tancik et al.)처럼 공간을 여러 국소 NeRF로 분할하는 시도는 있었지만, SLAM의 루프 클로저·전역 일관성 요구와 매끄럽게 맞물리지 못했다. 도시 규모 NeRF-SLAM은 개방형 문제다. **semantic·편집 가능한 implicit 지도.** NeRF map은 렌더링에 최적화되어 있어 semantic label 삽입이나 사후 편집이 어렵다. "이 물체를 지도에서 지워라"나 "이 영역을 다른 용도로 분류하라"는 조작이 TSDF나 포인트클라우드 대비 훨씬 불편하다. [LERF](https://arxiv.org/abs/2303.09553)는 언어 feature를 공간에 정렬해 의미 질의를 가능하게 한 사례이며, 그 자체가 물체 삭제·편집 방법은 아니다. [Nerfstudio](https://arxiv.org/abs/2302.04264) 같은 도구 위에서 의미 표현과 편집 연구가 진행되지만 SLAM 파이프라인과의 실시간 통합은 2026년 현재 연구 단계다. --- iMAP·NICE-SLAM이 implicit field를 극한까지 밀어붙이는 동안, 명시적 표현으로 되돌아가는 다른 방향도 등장했다. 지도를 MLP 가중치나 feature grid 안에 암묵적으로 가두는 대신, 공간에 명시적으로 배치된 수백만 개의 작은 타원체로 흩뿌리면 렌더링은 빠르고 편집은 직관적일 수 있었다. Splatting 자체에는 2001년 EWA splatting 같은 선행 기법이 있었다. Bernhard Kerbl의 SIGGRAPH 2023 논문은 학습 가능한 3D Gaussian과 빠른 미분 가능 렌더러를 결합해 이 방향의 성과를 보였다. --- # Ch.15 — Gaussian Splatting 시대: 3DGS에서 GS-SLAM까지 iMAP과 NICE-SLAM은 MLP로 공간을 기억하는 방법의 가능성을 보여주었다. 그러나 표현을 직접 편집하기는 어려웠다. iMAP에서는 새 관측에 대한 전역 MLP 갱신이 다른 영역에도 영향을 주었고, NICE-SLAM은 지역 feature grid로 이 문제를 완화했다. NICE-SLAM의 RTX 3090에서 1fps 아래 처리 속도는 "실시간 SLAM"이라는 말과 공존하기 어려운 수치였다. 장면은 네트워크 파라미터 안에 갇혀 있었고, 그 안을 들여다볼 방법이 없었다. 2023년 8월 SIGGRAPH에서 Bernhard Kerbl(INRIA), Georgios Kopanas, Thomas Leimkuhler, George Drettakis가 [논문](https://arxiv.org/abs/2308.04079)을 발표했다. Kerbl은 NeRF가 3년에 걸쳐 쌓은 implicit representation 패러다임을 버리지 않으면서, 다른 선택을 했다. iMAP·NICE-SLAM·Co-SLAM이 MLP와 voxel grid로 장면을 잠그는 동안, Kerbl은 수백만 개의 작은 타원체, 즉 Gaussian primitive를 공간에 흩뿌렸다. 이후 6개월 동안 이 표현을 활용한 SLAM 논문들이 잇따라 나왔다. Kerbl의 선택은 Matthias Zwicker의 EWA splatting(2001)이라는 20년 된 그래픽스 기법을 뿌리로 삼고, NeRF의 differentiable rendering 정신은 그대로 계승했다. 차이는 표현의 형식에 있었다. --- ## 3DGS의 구조 Kerbl은 장면을 explicit한 Gaussian 집합으로 표현했다. 각 Gaussian은 위치(mean) $\boldsymbol{\mu} \in \mathbb{R}^3$, 공분산 행렬 $\boldsymbol{\Sigma} \in \mathbb{R}^{3 \times 3}$, 학습되는 불투명도 $o \in (0,1]$, 그리고 구면조화함수(spherical harmonics) 계수로 표현된 색상을 가진다. 공분산은 학습 안정성을 위해 스케일 벡터 $\mathbf{s}$와 단위 쿼터니언 $\mathbf{q}$로 분해한다: $$\boldsymbol{\Sigma} = \mathbf{R}\mathbf{S}\mathbf{S}^\top\mathbf{R}^\top$$ 렌더링은 projected 2D Gaussian을 깊이 순서대로 알파 블렌딩(alpha-blending)한다. 각 Gaussian의 유효 불투명도 $\alpha_i$는 학습되는 불투명도 $o_i$와 픽셀 위치에서 평가된 2D Gaussian $G_i(\mathbf{x})$의 곱이다. 이 장에서 학습되는 불투명도를 $o_i$로 쓰면 픽셀 색상 $C$는 $$C = \sum_{i \in N} c_i \alpha_i \prod_{j
🔗 **차용.** 3DGS의 래스터화 기반 splatting은 Zwicker et al.의 [EWA splatting (2001)](https://www.cs.umd.edu/~zwicker/publications/EWAVolumeSplatting-VIS01.pdf)을 직접 계승한다. Zwicker는 점 구름을 렌더링하기 위해 각 점에 타원형 가중 평균 커널을 씌웠다. Kerbl은 그 커널을 learnable Gaussian으로 교체하고 GPU 타일 래스터라이저로 가속했다. --- ## 3DGS와 SLAM의 구조적 적합성 implicit representation은 SLAM에 적용할 때 여러 제약이 있었다. MLP 기반 NeRF는 새 관측이 들어올 때마다 전체 네트워크를 재학습해야 했고, catastrophic forgetting 탓에 incremental update가 어려웠다. 지도 확장은 네트워크 크기 재조정을 뜻했다. NICE-SLAM의 voxel grid는 이 문제를 완화했지만, 해상도와 메모리의 트레이드오프를 피할 수 없었다. Gaussian은 공간에 명시적으로 있는 객체여서, 새 키프레임이 들어오면 해당 영역에 Gaussian을 추가하기만 하면 된다. Densification 절차가 keyframe 추가와 자연스럽게 맞물렸고, 렌더링 품질은 NeRF 수준을 유지했다. 실시간 처리도 가능했다. 2023년 후반부터 GS-SLAM 논문들이 잇따라 나왔다. --- ## GS-SLAM: 초기 시도 Chi Yan(홍콩대)과 공동 연구자들은 2023년 11월 [Yan et al. 2023. GS-SLAM](https://arxiv.org/abs/2311.11700)을 arXiv에 게시했다. 비슷한 시기에 여러 3DGS-SLAM 원고가 공개된 가운데, 3DGS를 tracking과 mapping에 결합한 초기 시스템 중 하나였다. GS-SLAM의 구조는 전통 SLAM 프레임워크를 따른다. Tracking은 현재 프레임의 포즈를 추정하고, Mapping은 Gaussian 지도를 갱신한다. Yan의 기여는 두 가지였다. adaptive Gaussian expansion은 새 키프레임이 추가될 때 coverage가 낮은 영역에 Gaussian을 삽입한다. geometry-aware Gaussian selection은 렌더링 손실 역전파 시 기여가 큰 Gaussian만 골라 최적화해 속도를 확보한다. Tracking은 포즈를 렌더링 포토메트릭 손실로 최적화한다. GS-SLAM의 tracking 손실은 sampled pixel에 대한 L1 색 손실이다: $$\mathcal{L}_{track} = \sum_m \|\mathbf{C}_m - \hat{\mathbf{C}}_m\|_1$$ Mapping 단계에서 Yan은 color L1과 depth L1을 가중합해 쓴다. 한편 3DGS 원 논문의 training 손실인 $(1-\lambda)\mathcal{L}_1 + \lambda\mathcal{L}_{D\text{-}SSIM}$ ($\lambda=0.2$) 조합은 Gaussian 지도를 학습할 때 상속되는 기본형이다. 이 손실을 포즈에 대해 미분할 수 있는 이유가 3DGS의 differentiable rasterizer다. Replica 데이터셋에서 NICE-SLAM 대비 PSNR을 유지하면서 처리 속도를 높였다. 한계도 명확했다. RGB-D 카메라를 전제했고, 실외 대규모 환경에서는 검증하지 않았다. --- ## SplaTAM: silhouette 기반 densification Nikhil Keetha(카네기멜론대)의 [Keetha et al. 2024. SplaTAM (CVPR)](https://arxiv.org/abs/2312.02126)은 설계 철학에서 GS-SLAM과 달랐다. Keetha는 복잡한 선택 메커니즘 대신 실루엣 마스크 기반 단순한 densification을 택했다. Keetha는 **silhouette mask**를 이용해 현재 뷰에서 기존 Gaussian으로 설명되지 않는 영역을 찾고, 렌더링된 마스크의 빈 부분에 새 Gaussian을 추가했다. Gaussian이 없는 위치를 찾아 채우는 규칙이다. Tracking은 포즈, Mapping은 Gaussian 파라미터를 각각 최적화한다. 추적할 때는 지도를 고정하고, 지도 갱신 단계에서 Gaussian을 조정한다. GS-SLAM도 pose와 지도 변수를 나누므로, 분리 자체를 SplaTAM만의 차이로 볼 수는 없다. > 🔗 **차용.** SplaTAM의 키프레임 기반 지도 관리 구조는 PTAM(Klein & Murray, 2007)의 아이디어가 새 표현 위에서 재동작하는 사례다. PTAM이 keyframe을 선택적으로 삽입해 지도를 유지하던 방식이, SplaTAM에서는 Gaussian densification의 트리거로 변환되었다. Replica 데이터셋 기준으로 SplaTAM은 PSNR 34.11 dB를 기록했다. 같은 논문의 표에서 NICE-SLAM은 24.42 dB였다. 렌더링 품질 격차는 명확했다. Keetha는 2024 CVPR 논문의 Limitations & Future Work에서 motion blur·depth noise·공격적 회전에 대한 민감성, 그리고 known intrinsics와 dense depth 의존을 제거하는 방향을 다음 과제로 들었다. 확장성 개선도 언급했다. > 📜 **예언 vs 실제.** Keetha는 SplaTAM(2024) 한계 절에서 motion blur·depth noise·공격적 회전 민감성, 그리고 known intrinsics/dense depth 의존 제거를 명시적 과제로 제시했다. 같은 해 Matsuki의 MonoGS(2024 CVPR)는 별도의 단안 RGB 시스템을 제시했다. intrinsics-free와 대규모 스케일 이슈는 2024-2025년 현재도 진행 중이다. --- ## MonoGS: monocular RGB Hidenobu Matsuki(Imperial College Dyson Robotics Lab)의 [Matsuki et al. 2024. MonoGS (CVPR)](https://arxiv.org/abs/2312.06741)는 depth sensor 없이 단안(monocular) RGB 카메라만으로 3DGS SLAM을 구동했다. monocular 설정에서는 scale이 문제다. depth 없이 절대 메트릭 스케일을 복원하는 것은 SfM에서도 풀리지 않은 문제다. Matsuki는 Gaussian geometry를 직접 최적화하고 등방성 손실을 추가했지만, 단안 궤적 자체는 scale-free다. $$\mathcal{L}_{iso} = \sum_k \| \mathbf{s}_k - \bar{s}_k \mathbf{1} \|_1$$ 여기서 $\mathbf{s}_k \in \mathbb{R}^3$는 k번째 Gaussian의 스케일 벡터, $\bar{s}_k = \frac{1}{3}\sum_j s_{k,j}$는 세 축 스케일의 평균이다. 이 등방성(isotropy) 정규화는 Gaussian이 지나치게 얇은 판 형태로 퇴화하는 것을 막는다. monocular에서 depth 감독 없이 Gaussian이 카메라 평면에 달라붙는 현상을 억제한다. Tracking에서 Matsuki는 포즈를 Gaussian 렌더링 photometric loss로 직접 최적화했다. 새 Gaussian을 만들 때 해당 픽셀에 렌더링된 depth가 있으면 이를 쓰고, 없으면 렌더링된 depth의 중앙값을 사용한다. 이 절차와 등방성 손실은 내부 지도의 상대적 스케일을 일관되게 유지하지만, 절대 메트릭 스케일을 복원하지는 않는다. > 🔗 **구분.** MonoGS는 pretrained monocular depth predictor나 외부 tracking 모듈을 사용하지 않는다. 따라서 MonoDepth2를 직접 차용했다는 계보는 논문에서 확인되지 않는다. Matsuki는 Imperial Dyson Robotics Lab 출신이다. Davison이 지도한 연구실에서 나온 Sucar(iMAP), Bloesch(CodeSLAM)와 같은 계보다. Lab의 관심이 implicit MLP → Gaussian explicit 표현으로 이동한 흐름을 MonoGS가 대표한다. TUM-RGBD 데이터셋에서 MonoGS는 monocular 평균 ATE RMSE 4.44cm, RGB-D 1.58cm를 기록했다. 단안 ATE는 정답 궤적에 scale alignment를 한 뒤 평가한 값이고, RGB-D는 rigid alignment만 적용했다. 렌더링 품질은 RGB-D 설정 Replica 기준 평균 PSNR 37.50 dB로 동세대 GS-SLAM 계열과 견줄 만한 수준이었다. --- ## RTG-SLAM과 실시간 처리 GS-SLAM 계열이 풀어야 할 다음 문제는 속도였다. GS-SLAM과 SplaTAM은 실시간이라 부르기 어려웠다. Zhejiang University의 Peng Zhexi 연구팀이 2024년 발표한 [RTG-SLAM (SIGGRAPH 2024)](https://arxiv.org/abs/2404.19706)은 명시적으로 실시간을 목표로 삼았다. RTG-SLAM의 전략은 Gaussian의 수를 제어하는 것이다. 현재 카메라 뷰에서 기여도가 큰 Gaussian만 골라 최적화한다. Gaussian을 surfel(표면 원소) 기반으로 초기화해 geometry를 유지하면서 수를 줄였다. Replica 데이터셋에서 실시간에 근접한 처리 속도를 냈다. --- ## 3DGS 이후 달라진 표현 선택 2024년 이후 3DGS 논문이 빠르게 늘면서 SLAM mapping 연구의 표현 선택지가 넓어졌다. TSDF·occupancy grid는 embedded·planning·안전 응용에서 계속 쓰이고, NeRF와 3DGS도 목표에 따라 선택된다. 3DGS 논문들은 렌더링 속도와 명시적 primitive 갱신을 장점으로 내세웠지만, 이 흐름만으로 다른 표현이 주류에서 사라졌다고 단정할 수는 없다. 이것은 표현의 전환이면서 하드웨어 친화성의 전환이다. GPU 래스터라이저는 GPU 레이마처(ray marcher)보다 훨씬 잘 최적화되어 있다. 3DGS가 기존 그래픽스 파이프라인 위에서 자연스럽게 돌아간다는 점이 NeRF 대비 채택 속도를 높였다. > 🔗 **차용.** 3DGS의 differentiable rendering 정신은 NeRF에서 직접 계승한다. 장면 표현을 gradient로 최적화한다는 아이디어, photometric loss로 관측과 렌더링을 연결하는 방식은 Mildenhall et al.(2020)의 유산이다. Kerbl은 표현(implicit MLP → explicit Gaussian)을 교체하면서 패러다임은 계승했다. > 📜 **예언 vs 실제.** Kerbl et al.은 3DGS(2023) §7.4 Limitations에서 관측이 부족한 영역의 elongated artifact와 popping, 정규화(regularization) 부재, 메모리 소비(훈련 중 20GB 초과, 대규모 씬 렌더링 시 수백 MB)를 한계로 꼽았다. Future work로는 antialiasing, 더 원칙적인 culling, point-cloud 압축 기법 차용을 제안했다. 메모리 축(압축)은 [Compact 3DGS](https://arxiv.org/abs/2311.13681) 계열과 [Niedermayr et al.](https://arxiv.org/abs/2401.02436)이 2024년에 직접 응답했다. dynamic scene 확장([4DGS](https://arxiv.org/abs/2310.08528), [Deformable 3DGS](https://arxiv.org/abs/2309.13101))과 생성·편집([DreamGaussian](https://arxiv.org/abs/2309.16653), [GaussianEditor](https://arxiv.org/abs/2311.14521))은 원 논문이 직접 거론하지 않은 영역이지만 2024년 전후로 별도 계통으로 갈라져 나왔다. --- ## 🧭 아직 열린 것 **Memory scaling.** Gaussian의 수는 장면 크기에 따라 선형으로 증가한다. 실내 Replica 데이터셋에서 수십만 개로 충분하던 것이 outdoor 도시 구역에서는 수천만 개로 늘어난다. Gaussian pruning과 level-of-detail 계층화가 연구되고 있지만, 대규모 환경에서 메모리와 렌더링 품질의 트레이드오프를 합리적으로 관리하는 방법은 아직 합의가 없다. Compact 3DGS 계열(Lee et al. 2024, Niedermayr et al. 2024)이 압축 방향을 탐색 중이다. **Semantic 통합.** 2023년 [LERF](https://arxiv.org/abs/2303.09553)가 NeRF에 언어 feature를 결합했고, [LangSplat](https://arxiv.org/abs/2312.16084)은 Gaussian 표현에 언어 feature를 결합했다. 이후 semantic Gaussian을 SLAM에 결합한 연구도 이어졌지만, 서로 다른 장면과 장비에서 실시간 갱신·tracking 품질·semantic 정확도를 함께 비교할 공통 protocol은 정착되지 않았다. semantic과 geometry를 공동 최적화할 때의 interference가 핵심 평가 항목이다. **Dynamic scene.** 4DGS와 Deformable 3DGS는 시간 차원을 Gaussian에 추가하는 방향을 제안했다. SLAM 설정에서 dynamic object는 배경과 다른 움직임을 가지므로 별도로 처리해야 한다. GS-SLAM(Yan et al. 2023), SplaTAM(Keetha et al. 2024), MonoGS(Matsuki et al. 2024) 모두 정적 세계 가정을 유지한다. Dynamic SLAM에서 Gaussian이 어떻게 이동하는 객체를 표현하고 추적할 것인가는 2025년 기준으로 열려 있다. 그러나 3DGS가 남긴 또 다른 질문이 있었다. Gaussian은 SfM point cloud에서 오는가, 아니면 depth sensor에서 오는가. 포즈를 알아야 Gaussian을 놓을 수 있고, Gaussian이 있어야 포즈를 추정할 수 있다. 이 닭-달걀 문제에 대응하는 초기화는 시스템마다 달랐다. RGB-D 방법은 측정 깊이를 활용했고, MonoGS는 외부 depth predictor 없이 내부에서 초기 깊이 가설을 만들었다. Ch.16에서 다루는 DUSt3R와 그 후계들은 다른 출발점을 선택했다. geometry 자체를 처음부터 학습하는 길이다. --- # Ch.15b — 정적 세계 가정이 무너지는 자리: Dynamic과 Deformable SLAM 2015년 Javier Fuentes-Pacheco는 동료 Ruiz-Ascencio, Rendón-Mancha와 함께 [*Visual simultaneous localization and mapping: a survey*](https://link.springer.com/article/10.1007/s10462-012-9365-8)를 *Artificial Intelligence Review*에 발표했다. 그 서베이의 마지막 절이 "Dynamic and Deformable Environments"였다. 그 이전에도 움직이는 물체를 다룬 논문은 있었지만 대부분 RANSAC이 걸러내야 할 outlier로 취급했다. 이 서베이는 동적 환경을 독립 주제로 묶어 다룬 이른 문헌이었다. 11년이 지난 2026년, *The SLAM Handbook*은 이 주제에 37페이지를 할애했다. 저자는 Lukas Schmid, José María Martínez Montiel, Shoudong Huang, Daniel Cremers, José Neira, Javier Civera 여섯 명이다. 정적 세계는 SLAM의 출발점이었지만, 거리의 자율주행차와 집안의 서비스 로봇, 장기의 내시경은 모두 그 가정 밖에서 작동해야 했다. --- ## 15b.1 세 개의 축 Schmid et al.이 Handbook Ch.15 §15.1에서 그린 프레임은 이전의 "dynamic SLAM" 정의를 다시 쓴다. 환경이 동적인지 정적인지는 환경의 속성이 아니라 *관측의 속성*이다. 같은 물리적 운동이 한 로봇에게는 short-term dynamic, 다른 로봇에게는 long-term dynamic이 된다. 관측률 $\text{Obs}$와 변화율 $\text{Dyn}$의 비율이 결정한다. $\text{Dyn} \ll \text{Obs}$이면 프레임 사이에서 움직임이 보이고, $\text{Dyn} \gg \text{Obs}$이면 방문 사이에서 장면이 변해 있다. Observation axis는 short-term과 long-term을 가른다. Reconstruction axis는 pose만 추정할지, scene geometry까지 복원할지, 4D spatio-temporal 이해까지 갈지를 정한다. Time axis는 online과 offline을 가른다. 이전 서술은 "동적 객체를 어떻게 제거할 것인가"라는 단일 질문으로 필드를 압축했는데, 이 과제는 3축 분류 공간의 한 구석에 지나지 않는다. 서로 다른 영역을 연구한 사람들이 같은 단어를 서로 다른 문제에 써 온 이유다. --- ## 15b.2 Short-term: 마스킹에서 multi-object SLAM으로 첫 해법은 단순했다. 움직이는 것을 지웠다. Zaragoza에서 박사과정을 하던 Berta Bescos는 2018년 [DynaSLAM](https://arxiv.org/abs/1806.05620)을 RA-L에 발표했다. ORB-SLAM2의 frontend에 Mask R-CNN을 끼워 넣어 사람·자동차를 사전 마스킹하는 시스템이었다. 마스킹된 영역은 keypoint 추출에서 제외됐다. 간단했지만 동작했다. TUM-RGBD의 walking sequence에서 ATE가 한 자릿수 cm로 내려갔다. 같은 시기 UCL의 Martin Rünz는 다른 선택을 했다. 움직이는 물체를 지우지 말고 따로 추적하자. Lourdes Agapito 지도하에 [Co-Fusion(Rünz & Agapito, 2017)](https://arxiv.org/abs/1706.06629)과 이듬해 [MaskFusion(Rünz et al., 2018)](https://arxiv.org/abs/1804.09194)을 연달아 내놓았다. 각 객체에 독립된 surfel 모델을 할당해, 카메라 궤적과 객체 궤적을 동시에 추정했다. Edinburgh의 Raluca Scona와 Imperial의 Stefan Leutenegger가 2018년 ICRA에 낸 [StaticFusion](https://arxiv.org/abs/1806.05628)은 또 다른 경로였다. Semantic segmentation 없이 residual clustering만으로 dynamic region을 분리했다. Segmentation 오류에 의존하지 않는 방향이다. 다음 단계에서는 움직이는 객체를 state에 포함시켜 함께 추정했다. QUT의 Jun Zhang이 이끈 [VDO-SLAM(Zhang et al., 2020)](https://arxiv.org/abs/2005.11052)은 각 동적 객체를 factor graph의 변수로 올렸다. 카메라 포즈 $T_i^w \in SE(3)$와 객체 $k$의 포즈 $T_{k,i}^w \in SE(3)$가 같은 그래프에 공존했다. Constant-velocity factor가 객체의 선속도·각속도에 연속성 제약을 걸었다. 카메라 SE(3)와 객체 SE(3)의 product manifold 위에서 joint optimization이 돌아갔다. Zaragoza의 Bescos는 2021년 [DynaSLAM II(Bescos et al., 2021)](https://arxiv.org/abs/2010.07820)에서 ORB-SLAM2 기반으로 같은 아이디어를 구현했다. CMU의 Yuheng Qiu가 2022년 RA-L에 발표한 [AirDOS](https://arxiv.org/abs/2109.09903)는 인간처럼 관절이 있는 객체까지 articulated body로 확장했다. > 🔗 **차용.** VDO-SLAM의 factor graph 확장은 Ch.6 graph SLAM에서 Dellaert와 Kaess가 세운 iSAM 전통을 직접 계승한다. 변수를 하나 늘리고 factor를 하나 더 다는 것이, dynamic SLAM에서는 움직이는 자동차 하나를 지도에 올리는 일로 바뀌었다. 관성 쪽에서는 KAIST URL의 Song·Lim·Lee·Myung이 2022년 RA-L에 [DynaVINS](https://arxiv.org/abs/2208.11500)를 발표했다. DynaVINS는 semantic mask도 multi-object tracking도 쓰지 않았다. IMU preintegration이 준 pose prior와 어긋나는 관측은 bundle adjustment에서 factor weight를 낮추는 식으로, 동적 특징이 joint state로 새어 들어가는 경로를 끊었다. 같은 그룹이 2024년 RA-L에 낸 [DynaVINS++](https://arxiv.org/abs/2410.15373)는 이 아이디어를 adaptive truncated least squares로 다시 짜, dynamic feature가 IMU bias 추정으로 역전파되며 발산하는 실패 양상까지 억제했다. Handbook은 이 계보를 §15.2.3 "Dense Dynamic SLAM"으로 정리하면서 Schmid 본인의 [Dynablox(Schmid et al., 2023)](https://arxiv.org/abs/2304.10049)를 LiDAR MOS의 현재형으로 배치한다. 2025년의 [AnyCam](https://arxiv.org/abs/2503.23282)은 transformer 기반으로 일상 영상에서 직접 4D를 추정한다. Rünz가 2017년 문을 연 "simultaneous tracking + reconstruction" 계통의 2025년판이다. --- ## 15b.3 Long-term: 시간을 가로지르는 지도 Short-term이 프레임 사이의 운동이라면, long-term은 방문 사이의 변화다. 어제 본 의자가 오늘은 옆으로 밀려 있다. 이 문제는 다른 계보에서 자랐다. Sherbrooke의 Mathieu Labbé가 Michaud 지도 아래 2013년부터 개발한 [RTAB-Map](https://introlab.github.io/rtabmap/)은 인간 기억 모델에서 직접 빌려왔다. short-term, working, long-term memory의 계층을 두고, 시간과 관측 빈도에 따라 노드를 옮겼다. 한 세션 안에서는 작동 메모리에 남고, 자주 방문하지 않으면 장기 메모리로 내려가고, 의미가 없어지면 폐기되는 구조다. 2019년 JFR 논문에서 Labbé는 이 구조가 다중 세션 SLAM에서 어떻게 스케일하는지 정리했다. 한국과학기술원 김아영 팀의 임현준이 2021년 발표한 [ERASOR](https://arxiv.org/abs/2103.04316)는 지도를 깨끗이 만드는 문제를 scene differencing으로 풀었다. 같은 장소를 두 번 지나갔을 때 사라진 점을 찾아낸다. Handbook §15.3 전체를 관통하는 구분이 하나 있다. **absence of evidence vs evidence of absence**. 의자가 없는 것인지, 내가 못 본 것인지를 구별해야 한다. 이 구분이 빠지면 map cleaning은 정당한 객체를 지우고, change detection은 가려진 영역을 잘못 판정한다. Schmid가 2022년 RA-L에 발표한 [Panoptic Multi-TSDF](https://arxiv.org/abs/2109.10165)는 이 문제를 submap 구조로 풀었다. 각 객체를 독립 submap으로 관리하고, local consistency 하에서 active와 inactive를 구분했다. 같은 그룹이 2024년 낸 [Khronos](https://arxiv.org/abs/2402.13817)는 graduated non-convexity로 association을 견고화하고, loop closure 이후에도 deformable geometric change detection을 돌려, 각 객체의 변화 시점까지 추정한다. Metric-semantic 지도가 4D spatio-temporal 지도로 바뀐다. > 🔗 **차용.** Panoptic Multi-TSDF는 여러 TSDF submap을 함께 관리하는 volumetric submapping과, 객체별 모델을 두는 object-centric mapping 계열을 결합했다. 논문의 관련 연구와 참고문헌에서는 ORB-SLAM Atlas를 직접 계승했다는 근거가 확인되지 않는다. 같은 질문은 LiDAR 쪽에서 별도 계보로 전개됐다. KAIST URL의 Jang·Lee·Nahrendra·Myung이 2026년 공개한 [Chamelion](https://arxiv.org/abs/2602.08189)은 dual-head 네트워크 위에 scene-mixing augmentation을 얹어, 공사 현장이나 재배치가 잦은 실내처럼 구조가 자주 바뀌는 transient 환경에서 change detection을 ground truth 없이 돌린다. Khronos가 RGB-D·panoptic 쪽에서 4D를 세웠다면, Chamelion은 포인트 클라우드 위에서 long-term map maintenance 쪽으로 그 질문을 끌고 간다. 이 계보의 또 다른 축에는 반복성을 다루는 연구가 있다. 스웨덴 Örebro의 Tomáš Krajník과 Achim Lilienthal이 2014년부터 발전시킨 **frequency maps**는 주기적 사건(출근길 차량 흐름, 낮과 밤의 조명 변화)을 Fourier 기반으로 모델링한다. Stockholm Royal Institute of Technology의 Martin Magnusson 그룹이 2019년 정리한 Maps of Dynamics(MoD)는 *typical motion pattern*을 지도에 직접 인코딩했다. "이 복도에서는 사람이 왼쪽으로 걷는다"가 지도의 일부가 된다. 2023년 발표된 [Changing-SLAM(Schmid et al., 2023)](https://arxiv.org/abs/2301.09479)은 ORB-SLAM 확장 위에 Kalman 필터로 short-term을, semantic class 매칭으로 long-term을 동시에 다룬 시도다. --- ## 15b.4 Deformable: 형상이 변할 때 배경조차 움직이면 어떻게 되는가. Zaragoza의 Civera와 Montiel은 이 문제를 지속적으로 다뤘다. 시작은 다른 곳이었다. 2015년 CVPR best paper는 Microsoft Research의 Newcombe, Fox, Seitz가 발표한 [DynamicFusion](https://grail.cs.washington.edu/projects/dynamicfusion/)이었다. KinectFusion의 canonical TSDF에 embedded deformation graph를 얹어, 카메라 앞에서 변형하는 객체(얼굴, 몸통)를 실시간 비강체로 복원했다. 회전·이동이 노드마다 할당된 변형 그래프가 매 프레임 최적화됐다. 같은 계열에서 TU München의 Matthias Innmann이 2016년 [VolumeDeform](https://arxiv.org/abs/1603.08161)으로 색 정보를 더했고, 2017년 Miroslava Slavcheva가 낸 [KillingFusion](https://campar.in.tum.de/pub/slavcheva2017cvpr/slavcheva2017cvpr.pdf)은 Killing vector field 정칙화를 들여와 위상 변화(손이 몸통과 붙었다 떨어지는)까지 허용했다. MIT에서 Tedrake의 지도를 받은 Wei Gao가 2019년에 발표한 [SurfelWarp](https://arxiv.org/abs/1904.13073)는 TSDF 대신 surfel을 골라 exploration 친화성을 확보했다. > 🔗 **차용.** DynamicFusion의 embedded deformation graph는 컴퓨터 그래픽스에서 Sumner, Schmid, Pauly가 2007년 발표한 ED graph를 직접 가져왔다. 메시 변형을 위한 희소 제어 그래프였던 것이, 실시간 비강체 SLAM의 변수 표현이 되었다. Montiel 지도 아래 박사를 한 Juan Lamarca가 2021년 [DefSLAM](https://arxiv.org/abs/1908.08918)을 RA-L에 발표했다. isometric NRSfM으로 keyframe마다 template를 다시 계산하고, ORB frontend와 Lucas-Kanade optical flow를 섞어 trace를 유지했다. 평면 토폴로지를 가정하는 한계가 있었다. 같은 그룹의 Juan J. Gómez Rodríguez는 2023년 [NR-SLAM](https://arxiv.org/abs/2308.04036)으로 그 한계를 없앴다. dynamic deformable graph로 임의 토폴로지를 다루고, visco-elastic 모델로 시간 방향 정칙화를 넣었다. Handbook §15.4.2가 이 계보를 "deformable SLAM의 monocular 계통"으로 정리한다. Tsinghua의 Song이 2018년 낸 [MIS-SLAM](https://ieeexplore.ieee.org/document/8458232)은 stereo endoscopy로 수술 중 장기의 변형을 추적했다. Children's National의 Jayender 그룹이 개발한 EMDQ(Expectation Maximization + Dual Quaternion)는 SURF feature 위에서 부드러운 deformation field를 추정했다. 이들 시스템이 겨냥하는 것은 minimally invasive surgery의 실제 환경에서 intra-operative navigation을 돌리는 일이다. Handbook §15.4.1은 **Floating Map Ambiguity**를 강조한다. 비강체 객체의 rigid motion과 카메라의 rigid motion은 prior 없이는 구별되지 않는다. 손이 30cm 움직인 것인지 카메라가 30cm 움직인 것인지, 관측만으로는 어느 쪽도 말할 수 없다. 이 운동 분해의 모호성은 단안 SLAM의 오래된 scale ambiguity와는 성격이 다르다. Scale만이 아니라 trajectory와 deformation이 동시에 결합하여 ill-posed가 된다. DefSLAM과 NR-SLAM이 isometric prior, visco-elastic prior로 이 ambiguity를 부분적으로 깨지만, 원리적 해법은 2026년 기준에도 없다. > 📜 **예언 vs 실제.** Newcombe는 DynamicFusion(2015) §7 Future Work에서 "extension to larger scenes and topology changes"와 "integration with loop closure"를 다음 과제로 꼽았다. 토폴로지 변화는 2017년 KillingFusion이 다뤘고, 대규모 scene은 surfel 기반 SurfelWarp(2019)가 일부 확장했다. 2024년 Khronos도 deformation graph, loop closure, change detection을 한 장기 scene-mapping 시스템에 결합했다. 다만 Khronos 논문은 DynamicFusion을 인용하지 않으므로, 이를 DynamicFusion의 Future Work에 대한 직접 응답이라고 단정할 근거는 없다. --- ## 15b.5 연결을 읽는 한 가지 방식 이 연구들을 세 개의 고정된 학파로 나누기는 어렵다. 다만 논문과 공저 관계, 연구자의 이동을 따라가면 세 계보가 겹쳐 보인다. **Zaragoza 계보**(Montiel, Neira, Civera, Lamarca, Rodríguez)는 MonoSLAM(Ch.5)과 ORB-SLAM(Ch.7)에서 DynaSLAM·DefSLAM·NR-SLAM으로 이어지며 단안 기하를 밀어붙였다. **Dense dynamic reconstruction 계보**는 KinectFusion에서 DynamicFusion으로, SLAM++에서 Co-Fusion·MaskFusion으로 이어졌고 Davison·Newcombe·Agapito·Rünz·Cremers의 연구가 여러 기관에서 맞물렸다. **Schmid의 이동 경로**는 TUM의 Cremers 그룹에서 MIT의 Carlone 그룹을 거쳐 JPL로 이어지며 Dynablox·Panoptic Multi-TSDF·Khronos를 연결한다. 이는 Handbook이 선언한 공식 분류가 아니라, 연구 계보를 읽기 위한 이 장의 해석이다. --- ## 🧭 아직 열린 것 **Absence vs evidence of absence.** 지도에서 객체가 사라졌는지, 가려서 못 봤는지를 구별하는 문제는 long-term SLAM의 근원적 난제로 남아 있다. Schmid의 Panoptic Multi-TSDF는 active submap 구조로 부분 답을 내놓았다. Handbook은 가림이 70%를 넘는 경우를 여전히 풀기 어려운 extreme-environment 사례로 든다. 다만 이를 Panoptic Multi-TSDF의 보편적인 오차 임계값으로 보고한 것은 아니다. 2026년 기준, 이 문제에 원리적 해법을 주장한 논문은 없다. **Floating Map Ambiguity.** Deformable SLAM에서 카메라의 rigid motion과 객체의 rigid motion을 분리하는 문제는 isometric·visco-elastic prior로만 우회되고 있다. prior 없이 두 motion을 식별하는 조건이 무엇인지, 어떤 관측이 ambiguity를 깨는지는 미해결이다. Lamarca의 [2023년 IJRR 논문](https://arxiv.org/abs/2302.03710)이 관측 조건을 일부 정리했지만 일반 이론은 아직 없다. **Online deformable SLAM.** DefSLAM과 NR-SLAM은 실시간에 근접했지만, Khronos 수준의 change-aware 통합을 단안 RGB에서 online으로 돌리는 시스템은 없다. Optimization 계산량이 실시간 처리 범위를 넘어선다. GPU 가속과 learned prior가 가능성을 열고 있으나 검증된 파이프라인이 아직 없다. **의료 MIS의 실세계 격차.** MIS-SLAM과 NR-SLAM이 phantom과 ex vivo 데이터에서는 동작하지만, 실제 수술 환경의 혈액·연기·도구 가림·급격한 조명 변화 앞에서는 견고성이 떨어진다. 2024년 EndoGS 같은 Gaussian 기반 시도가 나오고 있지만 배포 수준에 도달한 시스템은 보고되지 않았다. 이 장이 따라온 질문은 변하는 세계를 어떤 표현으로 담을 것인가였다. GS-SLAM·SplaTAM·MonoGS는 장면 표현의 속도와 밀도를 바꿨지만 정적 세계 가정은 유지했다. Ch.16의 DUSt3R와 후속 계보는 representation pipeline을 다듬는 대신, 기하 prior 자체를 학습하는 길로 들어간다. --- # Ch.16 — Foundation 3D: DUSt3R에서 VGGT까지 Naver Labs Europe의 Philippe Weinzaepfel과 Jerome Revaud는 2022년 CroCo를 발표하면서, 두 이미지가 같은 장면을 찍었다는 사실을 단서 삼아 visual representation을 학습하는 cross-view self-supervised pretraining 방식을 제안했다. CroCo는 feature learning 연구로 발표됐다. 1년 뒤 같은 팀이 CroCo의 구조 위에서 calibration 없이 pointmap을 직접 출력하는 시스템을 만들었고, DUSt3R는 multi-view geometry 전체를 재정의했다. Naver Labs Europe에서 시작한 계보가 Oxford의 VGG 그룹으로 이어지며, 2026년 현재 "SfM이 무엇인가"라는 질문 자체를 다시 쓰고 있다. --- ## 16.1 DUSt3R — learned pointmap 2013년부터 10년간 3D 재건은 동일한 절차를 따랐다. 특징점을 찾고, 매칭하고, 카메라 내부 파라미터와 외부 파라미터를 추정하고, triangulation으로 점군을 만들고, bundle adjustment로 전체를 정제한다. [COLMAP(Schönberger & Frahm, 2016)](https://openaccess.thecvf.com/content_cvpr_2016/html/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.html)이 이 파이프라인의 가장 완성된 형태였다. 오차는 줄었지만 절차의 구조는 바뀌지 않았다. [Shuzhe Wang et al. 2023. DUSt3R: Geometric 3D Vision Made Easy](https://arxiv.org/abs/2312.14132)는 이 절차를 우회한다. 두 이미지를 입력으로 받아 각 픽셀에 대한 3D 좌표를 직접 출력한다. 내부 파라미터(focal length, principal point)를 요구하지 않는다. pointmap이라 불리는 이 출력은, 이미지 좌표계가 아닌 공통된 3D 공간에서의 좌표다. 카메라가 어떤 렌즈를 달고 있는지 몰라도 된다. DUSt3R의 transformer는 CroCo에서 물려받은 encoder-decoder 구조를 쓴다. 각 이미지는 독립적으로 encoding된 뒤, decoder에서 cross-attention을 통해 다른 이미지의 encoder 출력을 참조한다. self-attention이 단일 이미지 내 픽셀 관계를 처리한다면, cross-attention은 두 이미지 사이의 대응을 암묵적으로 학습한다. 어느 픽셀이 어느 픽셀과 같은 3D 점을 보는지, 이 대응 관계를 대규모 데이터에서 패턴으로 흡수한다. 논문은 Habitat·MegaDepth·ARKitScenes·Static Scenes 3D·BlendedMVS·ScanNet++·CO3D-v2·Waymo에서 850만 이미지 쌍을 구성했다. 정답 표시는 합성 데이터, SfM 소프트웨어의 재건 결과, 전용 센서 측정에서 왔다. 따라서 고전 SfM은 학습 시대의 ground truth 전부가 아니라 그 일부를 공급했다. > 🔗 **차용.** DUSt3R의 backbone은 ViT([Dosovitskiy et al. 2020](https://arxiv.org/abs/2010.11929))에서 가져온다. 그러나 결정적 발판은 Naver Labs Europe 내부의 선행 작업인 CroCo([Weinzaepfel et al. 2022](https://arxiv.org/abs/2210.10716))다. CroCo는 두 이미지에서 한 쪽의 masking된 영역을 다른 이미지의 정보로 복원하는 cross-view self-supervised pretraining을 제안했다. DUSt3R는 CroCo의 encoder-decoder 구조를 그대로 물려받아 태스크만 "pointmap 예측"으로 바꿨다. 두 이미지의 pointmap은 이미 첫 카메라의 공통 좌표계에 있다. 상대 카메라 포즈는 예측한 3D 점과 두 번째 이미지의 픽셀 대응으로 PnP-RANSAC을 풀어 구할 수 있다. 여러 이미지 쌍에서는 별도의 global alignment로 pointmap과 카메라들을 조정한다. pose estimation이 pointmap의 파생물이 된다. 세 장, 열 장의 이미지로 확장할 때 DUSt3R는 global alignment를 푼다. 모든 이미지 쌍의 pointmap을 하나의 공통 좌표계로 정합하는 최적화 문제다. 이때 비로소 bundle adjustment와 유사한 무언가가 등장하지만, 피처 매칭이나 카메라 모델 없이 진행된다. --- ## 16.2 매칭을 삼킨다: MASt3R DUSt3R의 결과는 novel view synthesis보다 reconstruction에 가깝다. 그런데 재건에서 중요한 서브태스크(두 이미지 사이의 정밀한 픽셀 대응 찾기, 즉 feature matching)를 DUSt3R는 암묵적으로만 처리한다. SuperPoint+SuperGlue, LightGlue가 수행하는 명시적 매칭을 대체하려면 추가 장치가 필요했다. [Vincent Leroy et al. 2024. Grounding Image Matching in 3D with MASt3R (ECCV)](https://arxiv.org/abs/2406.09756)는 DUSt3R에 matching head를 추가한다. pointmap과 함께 각 픽셀의 feature descriptor를 출력하도록 훈련하되, 3D 위치와 feature가 일관되도록 joint learning한다. 이렇게 나온 feature는 3D 공간에 anchored되어 있다. 매칭은 이 feature descriptor를 nearest neighbor 검색하는 것으로 단순화된다. > 🔗 **차용.** MASt3R의 3D-anchored matching은 SuperGlue([Sarlin et al. 2020](https://arxiv.org/abs/1911.11763))가 풀려던 문제(2D descriptor의 모호성을 context로 해소)를 다른 방향에서 공략한다. SuperGlue는 그래프 신경망으로 2D 매칭의 모호성을 줄였다. MASt3R는 3D 구조를 직접 학습함으로써 3D에 근거한 대응으로 매칭의 모호성을 줄인다. MASt3R 공개 이후 수개월 내에 SLAM 커뮤니티에서 SuperPoint+SuperGlue 조합을 MASt3R로 교체하는 실험이 여러 그룹에서 보고되었다. 2024년 말 [Riku Murai, Eric Dexheimer, Andrew Davison](https://arxiv.org/abs/2412.12392)(Imperial College London의 Davison 그룹)이 MASt3R-SLAM을 공개했을 때, 이 시스템은 MASt3R의 매칭을 frontend로, 그래프 기반 global optimization을 backend로 사용했다. 고전적 SLAM 아키텍처의 모양은 유지한 채 내부 부품이 거의 전부 교체된 형태다. MASt3R의 강점은 ground-truth calibration 없이도 dense 매칭이 가능하다는 점이다. 2026년 현재 연구자들은 COLMAP 기반 SfM 파이프라인의 초기화나 매칭 단계에 DUSt3R·MASt3R를 삽입하는 구성을 시험하고 있다. > 📜 **예언 vs 실제.** DUSt3R 논문 자체는 별도의 "Future Work" 절을 두지 않았지만, pair-wise + global alignment라는 구조 자체가 sequence 처리와 실시간 구동을 다음 과제로 암시한다. Spann3R는 2024년 8월, MASt3R-SLAM은 2024년 말 나왔다. 두 후속 작업이 각각 sequential extension과 SLAM 통합 문제에 6-12개월 내에 응답했다. 후속 작업이 나온 간격은 짧았다. --- ## 16.3 Spann3R — sequential 처리 그런데 batch 처리에는 근본적인 제약이 있다. SLAM은 이미지가 미리 다 갖춰지지 않는다. DUSt3R와 MASt3R는 이미지 집합을 입력받아 일괄 처리한다. 가방 속 이미지들을 한 번에 펼쳐 놓고 정합하는 방식이다. SLAM은 다르다. 이미지가 시간 순서로 들어오고, 시스템은 각 프레임마다 지도를 갱신해야 한다. [Hengyi Wang & Lourdes Agapito 2024. 3D Reconstruction with Spatial Memory (Spann3R)](https://arxiv.org/abs/2408.16061)는 DUSt3R의 구조를 sequential 처리에 맞게 고쳤다. 이미 처리한 프레임의 정보를 spatial memory에 저장하고, 새 프레임이 들어오면 이 memory bank에 cross-attention을 수행한다. 새 이미지의 각 픽셀이 과거 프레임의 어떤 정보와 연관되는지는 attention이 결정한다. > 🔗 **차용.** Spann3R의 spatial memory 메커니즘은 개념적으로 cross-attention memory와 유사하다. 구조적으로 DUSt3R의 사전학습된 ViT encoder-decoder를 그대로 물려받되, 디코더 출력(geometric feature)과 이미지 feature를 결합한 memory key를 두어 appearance와 distance를 동시에 반영한 메모리 조회를 구현한다. DUSt3R가 학습한 기하 표현이 그대로 sequential 메모리의 색인으로 재활용되는 경로다. Spann3R는 calibrated 카메라 없이도 작동하는 DUSt3R의 특성을 그대로 가져간다. 순차 이미지가 들어올 때마다 현재까지의 지도를 점진적으로 갱신한다. 완전한 실시간은 아니지만, DUSt3R의 일괄 처리 방식보다 SLAM 적용에 한 발 더 가깝다. --- ## 16.4 VGGT — multi-view joint inference Spann3R는 sequential 처리를 가능하게 했다. 그러나 pair-wise pointmap + global alignment라는 DUSt3R의 기본 골격은 그대로였다. 이 골격은 Oxford VGG 그룹의 후속 연구에서 바뀌었다. Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotny는 2025년 초 임의의 다중 이미지를 동시에 입력받아 카메라 포즈와 깊이, 점군을 한 번의 forward pass로 출력하는 방식을 제시했다. [Jianyuan Wang et al. 2025. VGGT: Visual Geometry Grounded Transformer](https://arxiv.org/abs/2503.11651)는 DUSt3R의 pair-wise 처리를 진정한 multi-view joint inference로 바꿨다. DUSt3R에서 N장의 이미지를 전수 조합하면 N(N-1)/2쌍의 pointmap을 구한 뒤 global alignment를 풀어야 한다. VGGT는 N장을 한꺼번에 transformer에 통과시킨다. attention이 모든 이미지 쌍 사이의 관계를 동시에 처리한다. > 🔗 **차용.** DUSt3R 훈련 정답의 일부를 SfM 재건이 공급했다는 점에는 역전이 있다. 고전 SfM이 명시적 알고리즘으로 나누었던 pair-wise geometry estimation·graph construction·global optimization의 기능을, VGGT는 한 모델의 추론 안에서 함께 수행한다. 고전 파이프라인의 단계가 foundation model 안에 다른 형태로 흡수된 셈이다. DUSt3R와의 정량 비교에서 VGGT는 카메라 포즈 추정 정확도와 점군 품질 면에서 일관된 우위를 보였다. 처리 속도도 global alignment 최적화가 없으므로 더 빠르다. --- ## 16.5 pose estimation과 reconstruction의 경계 소멸 전통 컴퓨터 비전은 두 문제를 구분했다. 지도 기반 localization은 이미 알려진 지도에서 현재 위치를 찾는 것이고, 3D 재건은 알려지지 않은 환경의 기하를 복원하는 것이다. SLAM은 이 둘을 동시에 풀기 때문에 어려웠다. DUSt3R부터 VGGT까지의 시스템은 기하와 카메라 추정에 공통 학습 표현을 쓴다. DUSt3R는 pointmap에서 별도 연산으로 pose를 복원하고 여러 뷰를 global alignment로 묶는다. VGGT는 한 번의 forward pass에서 카메라와 기하를 함께 출력한다. 같은 방향에 서 있지만, 추론에 필요한 후처리까지 같지는 않다. Multi-view geometry는 폐기되지 않는다. DUSt3R·MASt3R·VGGT는 SfM 재건과 센서 데이터에서 얻은 정답을 포함한 기하 감독으로 pose·depth·pointmap 같은 출력을 학습했고, 그 결과 다중 시점의 기하적 규칙성을 활용한다. 다만 epipolar constraint, triangulation, bundle adjustment라는 특정 알고리즘이 transformer weight 안에 그대로 학습됐다고 논문이 입증한 것은 아니다. 폐기된 것은 명시적 파이프라인의 일부다. 연구자는 Schönberger의 COLMAP 코드를 디버깅하던 방식으로 DUSt3R를 디버깅할 수 없다. 어디서 실패했는지, 왜 실패했는지가 attention weight 안에 묻혀 있어 해석 가능성 문제가 새로운 형태로 등장한다. > 📜 **예언 vs 실제.** MASt3R 논문은 결론부를 짧게 맺으며 ground-truth calibration이 없는 매칭이 여러 downstream 태스크에 열려 있다고 시사했다. 명시적 파이프라인 재편 예언은 아니었다. 2026년 현재 여러 photogrammetry 소프트웨어가 DUSt3R/MASt3R를 initialization 단계로 채택하는 것을 평가 중이며, hybrid 삽입의 형태로 자리 잡고 있다. Naver Labs Europe은 CroCo(2022) → DUSt3R(2023) → MASt3R(2024)의 단계를 2년 내에 밟았다. 한 팀이 pretraining 방법론부터 매칭 시스템까지의 스택을 연속해서 발표했다. 이 계보의 출발점은 Google Brain, DeepMind, Meta AI가 아니라 Weinzaepfel·Revaud·Leroy가 속한 Naver Labs Europe이었다. SLAM 단계로 옮기는 일은 Imperial College London의 Davison 그룹(MASt3R-SLAM)이 이어받았다. --- ## 16.6 다른 갈래 — semantic foundation이 지도로 들어온다 DUSt3R·MASt3R·VGGT는 pointmap·카메라 포즈·기하 구조를 다루는 geometric foundation 계보다. 2022년 전후로 "foundation 3D"라는 단어는 SLAM 문헌에서 두 갈래로 쓰이기 시작했다. 한쪽은 Naver Labs Europe에서 출발한 geometric 계보고, 다른 쪽은 CLIP·DINO·SAM을 지도 안으로 끌어들이는 semantic 계보다. 전자는 calibration을 없앴고, 후자는 dictionary를 없앴다. semantic 갈래의 시작은 MIT의 Luca Carlone 그룹에서 나왔다. [Nathan Hughes et al. 2022. Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization](https://arxiv.org/abs/2201.13360)는 Kimera(Rosinol 2020)의 metric-semantic mesh 위에 objects → places → rooms → buildings의 hierarchical scene graph를 online으로 얹었다. closed-set 분류기를 쓰는 한 handbook이 "100-1000 labels predefined dictionary"라고 못 박은 제약 안이었지만, Hydra는 hierarchical map이 실시간으로 굴러간다는 것을 처음 보여줬다. dictionary의 벽은 foundation model이 허물었다. [Songyou Peng et al. 2023. OpenScene: 3D Scene Understanding with Open Vocabularies (CVPR)](https://arxiv.org/abs/2211.15654)가 ETH/Pollefeys 그룹에서, 곧이어 [Qiao Gu et al. 2024. ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning (ICRA)](https://arxiv.org/abs/2309.16650)가 Montréal-MIT 협업으로 발표됐다. OpenScene은 CLIP feature를 3D 점군에 distillation해서 "이 점은 의자와 얼마나 가까운가"를 자연어 질의로 풀 수 있게 했다. ConceptGraphs는 한 걸음 더 나아갔다. class label 대신 VLM이 생성한 language description을 node attribute로 달고, object 사이 관계를 LLM이 서술한다. 이들 연구는 고정된 class 사전 밖의 개념을 3D 표현에 연결하는 길을 열었다. ConceptGraphs는 Hydra의 계층 구조를 직접 확장한 방법과는 구별된다. [Dominic Maggio et al. 2024. Clio: Real-time Task-Driven Open-Set 3D Scene Graphs](https://arxiv.org/abs/2404.13696)는 이 계보를 task 쪽으로 돌렸다. 로봇이 받은 자연어 task를 information bottleneck으로 해석해서, 그 task에 필요한 추상화 수준만 scene graph에 남긴다. "커피 머신 근처 청소"라는 지시에서 커피 머신과 그 주변 객체는 보존되고, 무관한 디테일은 묶인다. hierarchical graph의 어느 층을 노출할지가 task에 따라 달라지는 것이다. > 🔗 **차용.** Clio는 Hydra를 바탕으로 task-driven abstraction을 추가한 Carlone 그룹의 직접 후속 계보다. ConceptGraphs는 open-vocabulary object graph를 만든 별도 계열이며, Hydra의 objects-places-rooms 계층을 그대로 물려받았다고 단정할 근거는 확인되지 않는다. 지도에 semantic을 싣는 문제는 Ch.18 §18.4가 2017-2019년 object-as-landmark 계보의 축소를 짚은 뒤 이어지는 별개 궤적이다. semantic SLAM은 hierarchical scene graph라는 형태로 귀환했지만, geometric foundation(DUSt3R 계보)과 semantic foundation(Hydra → Clio와 별도의 ConceptGraphs 계열)은 2026년 현재 아직 본격적으로 만나지 않았다. VGGT의 pointmap에 CLIP feature를 붙인 end-to-end 시스템, 또는 Clio의 scene graph에 DUSt3R의 calibration-free 기하를 결합한 구성은 아직 보고되지 않았다. 두 계보의 합류는 아직 열린 문제다. --- ## 16.7 SLAM에서 무엇이 남는가 MASt3R-SLAM은 고전 SLAM의 아키텍처를 빌려 쓴다. keyframe 선택과 loop closure, map management는 새 표현 위에서도 그대로 필요했다. DUSt3R 계열이 feature matching과 reconstruction의 내부를 교체했지만, SLAM 시스템 수준의 판단들은 고전 방법이 해결한 방식 그대로 재사용한다. 이 관찰은 5부 전체에 걸쳐 반복되는 패턴과 일치한다. NeRF-SLAM이 NeRF를 map 표현으로 채택하면서도 keyframe 기반 tracking을 유지했다. 반면 초기 3DGS-SLAM 시스템인 MonoGS는 명시적 loop closure를 구현하지 않았고 논문의 future work로 남겼다(Ch.15). 표현(representation)을 바꾼다고 시스템 수준의 기능이 자동으로 따라오는 것은 아니다. Foundation 3D의 경우에도 이 패턴이 반복된다. 2025년 MIT의 Dominic Maggio·Hyungtae Lim·Luca Carlone이 [VGGT-SLAM](https://arxiv.org/abs/2505.12549)을 공개했다. VGGT가 local submap을 재건하면, 시스템은 순차 submap 사이의 15자유도 projective transform을 \(\mathrm{SL}(4)\) 위에서 최적화하고 loop-closure 제약까지 함께 반영한다. transformer가 local 기하를 만들지만 global graph 최적화는 남는다. Revaud도 *The SLAM Handbook* Ch.13에서 대규모 재건에는 결국 "a form of factor graph is still necessary"라고 적었다(p.386). 전개는 빨랐지만 최종 형태는 여전히 열려 있다. 실시간 대규모 sequence를 foundation 3D가 어디까지 감당할지, 그리고 16.6에서 언급한 semantic 갈래와 어디서 합류할지는 2026-2027년의 관찰 대상이다. --- ## 🧭 아직 열린 것 **대규모 sequence 처리의 벽.** DUSt3R와 VGGT의 transformer는 이미지 수에 quadratic하게 메모리를 요구한다. 100장까지는 현실적이지만 1,000장, 10,000장은 다른 문제다. Spann3R의 incremental 방식이 partial answer지만, 대규모 outdoor 환경의 매끄러운 처리는 미해결이다. 여러 그룹이 sparse attention, hierarchical global alignment를 탐색하고 있으나 합의된 방법이 없다. **Loop closure를 이 프레임에서 어떻게 정의하는가.** 고전 SLAM에서 loop closure는 이전에 방문한 장소를 인식하고 누적 오차를 교정하는 메커니즘이다. DUSt3R 계열에서 "이전에 방문한 장소"를 어떻게 표현하고, pointmap 기반 지도에서 교정을 어떻게 propagate하는가. MASt3R-SLAM이 기존 방식으로 처리하지만, 이것이 최선인지 원리적 해법인지 알 수 없다. **Metric scale의 일반화.** DUSt3R의 pointmap은 relative scale이다. 두 이미지 사이의 깊이 비율은 복원하지만 절대 스케일은 모른다. Metric3D나 Depth Anything v2가 metric depth를 목표로 했듯, foundation 3D에서도 metric scale을 일반화하는 문제가 남는다. 카메라 독립적 metric은 foundation 규모에서도 쉽지 않다. GPS나 IMU 없이 absolute scale을 결정하는 물리적 제약은 데이터 규모와 무관하게 존재한다. **이 흐름이 SLAM의 미래인가, 별개 갈래인가.** 15장의 3DGS처럼 foundation 3D도 SLAM 커뮤니티가 흡수하는 중이다. MASt3R-SLAM과 VGGT-SLAM이 2024-2025년에 연달아 등장하며 흡수의 경로는 윤곽이 잡혔다. 그러나 실시간 대규모 sequence 구동, 그리고 §16.6이 짚은 semantic 갈래(Hydra → Clio와 별도의 ConceptGraphs 계열)와의 합류 지점은 여전히 불분명하다. geometric foundation과 semantic foundation이 한 시스템에서 만나는 형태는 Ch.19 열린 문제의 핵심 축이다. --- NeRF든 foundation model이든 표현을 바꾸면 reconstruction과 localization의 내부가 바뀐다. 그러나 SLAM 시스템 수준의 구조(keyframe, loop closure, map management)는 새 표현 위에서도 살아남는다. 같은 시기에 LiDAR 연구자들은 같은 문제를 다른 센서와 다른 문화로 풀고 있었다. --- # Ch.17 — LiDAR 평행 우주: LOAM에서 FAST-LIO까지 사진측량에서 Foundation 3D에 이르는 계보는 센서가 카메라라는 공통 전제 위에 서 있다. MonoSLAM·PTAM·ORB-SLAM·DSO·DUSt3R는 모두 픽셀로 세계를 읽는 전통 안에 있다. 같은 시간, 같은 로봇공학 커뮤니티 안에서 전혀 다른 계보가 자라고 있었다. LiDAR 계보는 카메라 진영의 keypoint·photometric consistency·feature descriptor와 무관하게, ICP의 뼈대 위에서 자체적인 문법을 만들어냈다. 두 계보는 인용 관계가 드물었고, 주로 쓰는 벤치마크와 발표 무대도 달랐다. Ji Zhang은 2014년 RSS에서 LOAM을 발표했다. 그해 Visual 진영에서는 LSD-SLAM 계열이 전개되고 있었고, ElasticFusion은 이듬해인 2015년에 발표됐다. LOAM은 카메라 기반 방법론이나 코드와 직접 연결되지 않았고, 연구 커뮤니티의 겹침도 적었다. 두 계보는 같은 로보틱스라는 이름 아래에서 접점이 드문 채 10년을 달렸다. LOAM은 [ICP (Besl·McKay, 1992)](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf)의 오래된 뼈대 위에 섰다. 다만 graph SLAM에는 Lu–Milios의 laser pose network처럼 LiDAR 쪽의 오래된 계보도 있었다. 평행 우주는 제한된 교류 속에서 성숙했다. --- ## 17.1 LOAM: edge와 plane, 그리고 KITTI의 점령 2014년, Google의 Waymo 전신 프로그램이 도로 위를 달리고 있었고, DARPA Urban Challenge의 여파가 채 가시지 않은 때였다. Velodyne HDL-64E는 한 대에 75,000달러였다. 당시 LiDAR 연구는 이 장비를 감당할 수 있는 대형·고예산 연구 그룹에 주로 한정됐다. CMU Robotics Institute의 Autonomous Mobile Robot Lab(Sanjiv Singh 교수 연구실)이 그중 하나였다. LOAM 이전에도 LiDAR로 지도를 쌓는 시도는 있었다. [Lu & Milios 1997. "Globally Consistent Range Scan Alignment for Environment Mapping" (Autonomous Robots)](https://doi.org/10.1023/A:1008854305733)은 2D 레인지 스캔을 노드로 두고 스캔 간 상대 제약을 edge로 묶어 전체 궤적을 동시 최적화하는 방식을 제안했고, 이 "network of poses" 발상은 훗날 pose-graph SLAM의 원형이 된다(Ch.6 참조). 매칭 자체는 Besl·McKay의 ICP 외에도 [Biber·Straßer 2003. "The Normal Distributions Transform" (IROS)](https://doi.org/10.1109/IROS.2003.1249285)가 제안한 NDT(셀 단위 가우시안 분포에 정렬하는 분포 기반 매칭)와 이후 Magnusson의 3D 확장이 ICP 대안으로 공존했다. 이 모두가 2D 또는 오프라인 3D였다. LOAM의 몫은 실시간 3D였다. Ji Zhang은 Singh의 지도 아래 [Zhang & Singh 2014. "LOAM: Lidar Odometry and Mapping in Real-time" (RSS)](https://www.roboticsproceedings.org/rss10/p07.pdf)를 냈다. LiDAR 포인트를 두 종류의 feature로 분류했다. **edge point**는 smoothness $c$가 높은 지점(곡률 높음), **planar point**는 $c$가 낮은 지점(곡률 낮음). ICP처럼 포인트 전체를 등록하지 않고 이 두 feature 집합만 매칭한다. edge point는 이웃 scan의 edge line에, planar point는 이웃 scan의 local plane에 point-to-line·point-to-plane 거리로 제약을 건다. 계산 비용이 낮아진다. 실시간 가능성이 열린다. 알고리즘 구조는 두 단계로 나뉜다. Lidar Odometry는 스캔 간 6-DoF 변환을 10Hz에서 추정한다. Lidar Mapping은 더 낮은 주파수(1Hz)에서 전체 맵과 정합해 오차를 보정한다. 고주파 odometry와 저주파 mapping을 분리해 drift를 억제하면서도 실시간성을 유지한다. 이 two-tier 구조는 이후 LiDAR SLAM의 기본 문법이 된다. KITTI 공개 비교표에서 LOAM은 발표 뒤 선두권에 들었다. 시퀀스 00의 relative translation error로 널리 인용되는 값은 0.78%이고, 전체 시퀀스 평균은 0.84%다. 이는 당시의 경쟁력을 보여주지만, 센서와 평가 조건이 다른 visual odometry에 대한 보편적 우위를 뜻하지는 않는다. > 🔗 **차용.** LOAM의 feature-based 포인트 등록은 Besl·McKay(1992)의 ICP에서 출발한다. 차이는 edge와 planar feature만 선택적으로 매칭한다는 점이다. 고전 등록을 선별적으로 재사용함으로써 속도와 정밀도 모두를 얻었다. --- ## 17.2 LeGO-LOAM: 땅을 먼저 잘라낸다 LOAM의 문제는 지면(ground plane)을 명시적으로 다루지 않는다는 점이었다. 실외 자율주행 환경에서 포인트 클라우드의 상당 비율은 도로면이 차지한다. 이를 edge/planar feature로 한데 처리하면 매칭 노이즈가 생긴다. Stevens Institute of Technology의 Robust Field Autonomy Lab에서 Tixiao Shan과 지도교수 Brendan Englot은 [Shan & Englot 2018. LeGO-LOAM](https://doi.org/10.1109/IROS.2018.8594299)에서 ground segmentation을 첫 단계로 분리했다. 포인트 클라우드를 range image로 투영한 뒤, 지면 포인트를 먼저 분리하고 비지면 포인트를 다시 클러스터링한다. Ground는 roll·pitch 추정에, 클러스터는 yaw·translation 추정에 각각 사용된다. 두 단계 최적화다. 결과는 LOAM 대비 연산 절감이었다. 원래 LOAM이 Velodyne VLP-16에서 실시간 동작이 버거웠다면, LeGO-LOAM은 동일 센서에서 임베디드 플랫폼(NVIDIA Jetson)에서도 돌아간다. 경량화의 대가는 있다. 포인트 희소 환경이나 지면 구조가 불규칙한 환경(레이저가 가리는 구간, 울퉁불퉁한 야지, 건물 내부)에서는 segmentation이 실패하고 odometry가 흔들린다. LeGO-LOAM은 "센서 입력을 구조화된 모듈로 전처리한 뒤 odometry를 돌린다"는 설계 원칙을 남겼다. FAST-LIO와 LIO-SAM이 뒤에 이 원칙을 받아들인다. LeGO-LOAM과 같은 시기, Bonn 대학의 Jens Behley와 Cyrill Stachniss는 edge/plane feature가 아니라 **surfel**(surface element)을 outdoor LiDAR에 가져왔다. [Behley & Stachniss 2018. "Efficient Surfel-Based SLAM using 3D Laser Range Data in Urban Environments" (RSS)](http://www.roboticsproceedings.org/rss14/p16.pdf)의 **SuMa**는 각 포인트 이웃을 원반 모양 surfel로 요약해 scan-to-model 등록을 수행했고, 후속 [Chen et al. 2019. "SuMa++" (IROS)](https://doi.org/10.1109/IROS40897.2019.8967704)는 semantic segmentation을 결합해 움직이는 물체를 surfel 수준에서 걸러냈다. Kinect 실내 RGB-D 계보(Ch.9)에서 ElasticFusion이 쓰던 surfel representation이 outdoor Velodyne으로 건너온 순간이다. feature 선택(LOAM), segmentation 선행(LeGO-LOAM), surfel 누적(SuMa)의 세 갈래가 2018년 전후로 동시에 경쟁하고 있었다. --- ## 17.3 FAST-LIO — tightly coupled LiDAR-IMU LiDAR의 스캔 주파수는 10-20Hz다. 그 사이사이에서 빠른 움직임이 있으면 포인트 클라우드에 motion distortion이 생긴다. 스캔이 끝나는 순간의 센서 위치와 시작 순간의 위치가 다르기 때문이다. 고속 이동체에서 LOAM 계열이 흔들리는 주된 이유가 여기에 있다. IMU는 100-400Hz로 동작한다. LiDAR의 틈을 채우기에 충분하다. 그런데 LiDAR와 IMU를 어떻게 결합하느냐에 따라 성능이 갈린다. loosely coupled는 각각 독립적으로 추정한 뒤 fusion하고, tightly coupled는 한 추정기 안에서 원시 관측이나 잔차 수준의 제약을 함께 사용한다. tight coupling은 센서 사이의 제약을 더 많이 보존할 수 있지만, 모델·보정·구현 조건과 무관하게 항상 우월한 것은 아니다. Hong Kong University(HKU) MaRS Lab의 Wei Xu와 지도교수 Fu Zhang은 2021년 RA-L에 [**FAST-LIO**](https://arxiv.org/abs/2010.08196)를 발표했다. 드론 제어 연구실에서 나온 논문으로, 로터 진동이 심하고 기동이 빠른 UAV에서도 LiDAR odometry가 버텨야 했다. 이들이 사용한 도구는 **iterated Extended Kalman Filter(iEKF)**였다. iEKF는 측정 업데이트 단계에서 선형화 점을 현재 추정치로 반복 갱신한다. 한 번의 linearization으로 끝내는 기본 EKF보다 측정 모델의 선형화 오차를 줄일 수 있지만, 개선 폭은 초기 추정과 운동·관측 조건에 달려 있다. 이듬해 TRO에 발표한 **FAST-LIO2**([Xu et al. 2022](https://doi.org/10.1109/TRO.2022.3141876))는 ikd-Tree를 추가했다. 기존 kd-Tree는 포인트가 추가될 때마다 재구성 비용이 크다. ikd-Tree는 부분 재구성만 수행하는 incremental 방식이다. 맵 포인트가 수백만 개에 달해도 실시간 nearest-neighbor 탐색이 가능하다. 실험에서는 UAV·핸드헬드·자율주행차에서 일관된 성능이 나왔다. 드론 환경에서도 drift가 낮게 유지됐다. FAST-LIO 계보의 다음 수는 motion distortion을 아예 없애는 쪽이었다. 같은 MaRS Lab에서 나온 [He et al. 2023. "Point-LIO: Robust High-Bandwidth Light Detection and Ranging Inertial Odometry" (Advanced Intelligent Systems)](https://doi.org/10.1002/aisy.202200459)는 LiDAR 포인트가 들어올 때마다 state를 갱신한다. point-by-point 관측 업데이트다. 각 포인트를 자기 시각에서 바로 fusion해 왜곡이 발생할 틈을 지운다. 고기동 플랫폼에서 FAST-LIO2보다 drift가 줄어든 것이 보고됐다. > 🔗 **구분.** FAST-LIO는 raw IMU를 이용한 상태 전파와 LiDAR 측정의 iterated Kalman update를 결합한다. [Forster et al. 2016. "On-Manifold Preintegration" (TRO)](https://doi.org/10.1109/TRO.2016.2597321)의 preintegrated factor를 iEKF로 재구현한 구조가 아니다. 두 방법은 IMU를 결합하지만 추정기 형식이 다르다. --- ## 17.4 LIO-SAM: LiDAR와 IMU를 factor graph로 묶다 같은 시기, visual-inertial 시스템에서는 factor graph가 널리 쓰였고, LiDAR 쪽에도 Lu–Milios의 pose network처럼 더 오래된 graph-SLAM 계보가 있었다. [GTSAM (Dellaert·Kaess, 2012)](https://gtsam.org/)은 여러 센서의 제약을 factor로 함께 구성하는 도구를 제공했다. Tixiao Shan이 LeGO-LOAM 이후 낸 [Shan et al. 2020. LIO-SAM](https://doi.org/10.1109/IROS45743.2020.9341176)은 GTSAM의 factor graph를 LiDAR-IMU 시스템의 backend로 명시적으로 채택했다. [공개 구현](https://github.com/TixiaoShan/LIO-SAM)은 두 그래프를 유지한다. 장기 mapping 그래프는 LiDAR odometry, GPS, loop closure 제약을 keyframe 사이에 쌓는다. 별도의 IMU preintegration 그래프는 관성 측정과 LiDAR odometry로 상태·bias를 추정하며, 계산량을 제한하려고 주기적으로 초기화된다. > 🔗 **차용.** LIO-SAM은 GTSAM의 factor graph에 IMU preintegration·LiDAR odometry·GPS·loop closure factor를 함께 구성했다. 다만 graph SLAM 자체는 Lu–Milios의 laser pose network를 포함한 LiDAR 계보에도 이미 있었으므로, 이를 Visual SLAM에서 LiDAR로 일방적으로 건너온 구조라고 볼 수는 없다. LIO-SAM은 FAST-LIO2와 달리 loop closure를 포함해 장기 궤적의 누적 drift를 보정할 수 있다. 반면 계산 비용이 높고 GPS나 추가 sensor input이 없으면 factor graph의 강점이 줄어든다. 두 시스템은 설계 목표가 다르다. FAST-LIO2는 실시간 LiDAR-IMU 구성의 속도와 정밀도를, LIO-SAM은 다중 센서 long-term mapping의 일관성을 우선한다. 역설적인 회귀도 있었다. LOAM 이후 10년 가까이 LiDAR odometry는 feature 선택·surfel·neural descriptor로 점점 복잡해지는 쪽을 달렸는데, 2023년 Bonn 대학의 [Vizzo et al. 2023. "KISS-ICP: In Defense of Point-to-Point ICP" (RA-L)](https://doi.org/10.1109/LRA.2023.3236571)은 반대 방향을 냈다. feature 추출도, 학습된 descriptor도 없이, 적응형 threshold로 튜닝이 거의 필요 없는 point-to-point ICP 하나로 KITTI에서 경쟁력 있는 odometry를 보였다. 이름 그대로 Keep It Small and Simple이다. 이 결과는 motion compensation, sampling, robust correspondence 처리 같은 구현 선택을 잘 결합하면 고전 등록법도 경쟁력이 있음을 보였다. 이를 LOAM이 등장한 역사적 원인에 대한 저자들의 평가로 읽을 필요는 없다. --- ## 17.5 센서 가격 하락과 보급: 2007–2024 LiDAR SLAM의 역사에서 기술 논문 못지않게 중요한 것이 센서 가격이다. 2007년 DARPA Urban Challenge에서 주요 팀들이 장착한 Velodyne HDL-64E는 대당 75,000달러였다. 자율주행 연구팀이나 국방 프로젝트가 아니면 접근하기 어려운 장비였다. 2012년에도 HDL-32E가 30,000달러 수준. LOAM이 발표된 2014년에는 VLP-16이 7,999달러로 내려왔지만 여전히 연구 예산의 상당 부분이었다. 그 뒤 LiDAR의 가격대가 크게 넓어졌다. 중국 업체 Livox(DJI 계열)는 2019년 [Mid-40의 미국 소매가를 599달러로 발표했다](https://www.livoxtech.com/news/1). 같은 해 Ouster의 128채널 OS1-128은 18,000달러였고, Ouster는 2020년 [양산 프로그램용 solid-state ES2의 2024년 목표 가격을 600달러로 제시했다](https://investors.ouster.com/news-releases/news-release-details/ouster-announces-first-high-performance-true-solid-state-digital). 수백 달러대 제품과 양산 목표가가 등장한 것은 분명한 변화였지만, 이를 모든 LiDAR의 일률적인 100배 인하로 읽을 수는 없다. 채널 수·시야각·거리 성능이 다르고 소매가와 대량 양산 목표가도 서로 다른 기준이기 때문이다. 센서가 보급되어도 알고리즘 문제가 함께 해결된 것은 아니었다. Solid-state LiDAR는 spinning 타입보다 시야각(FoV)이 제한적이고, 제품에 따라 비반복 스캔 패턴을 쓴다. 360° 회전 스캔을 전제로 설계된 원래 LOAM은 그대로 적용하기 어렵다. 반면 FAST-LIO와 FAST-LIO2는 기계식·solid-state LiDAR를 함께 다루는 방향으로 설계됐고, 작은 FoV와 불규칙한 sampling에서도 작동하는 사례를 보고했다. 저가 센서의 확산은 알고리즘을 없앤 것이 아니라 관측 범위와 scan pattern에 맞춘 새 과제를 만들었다. --- ## 17.6 Visual-LiDAR 계보 분리의 원인 Visual SLAM과 LiDAR SLAM은 동시대에 발전했지만 두 커뮤니티의 교류는 오랫동안 제한적이었다. 이유는 여러 층에 걸쳐 있었다. 첫째는 센서 자체다. 카메라는 texture와 color를 보고, LiDAR는 range와 geometry를 본다. 카메라 기반 방법이 keypoint·descriptor·photometric consistency를 중심으로 발전할 때, LiDAR는 edge·plane·range image로 분화했다. 문제 공식 자체가 달랐다. 학회도 달랐다. CVPR·ICCV는 카메라 기반 방법의 주 발표 무대였고, ICRA·IROS·RSS는 LiDAR SLAM이 주로 나왔다. 연구자 집단의 겹침도 적었다. Velodyne이 구글과 자율주행 업계에 공급되던 2010년대 초중반에 LiDAR SLAM 연구자 집단은 자율주행 로봇공학 쪽에 밀집했다. Place recognition 방법도 달랐다. 카메라는 DBoW2·NetVLAD처럼 visual appearance를 사용한다. LiDAR는 [Scan Context(Kim·Kim, 2018)](https://gisbi-kim.github.io/publications/gkim-2018-iros.pdf)나 [PointNetVLAD](https://arxiv.org/abs/1804.03492) 같이 3D point cloud의 구조적 특징을 활용한다. 동일 장소라도 인식하는 신호 자체가 다르다. 수렴의 첫 신호는 2020년대 초에 나타났다. LiDAR-Camera 융합을 다루는 논문이 CVPR에 올라오기 시작했고, Tixiao Shan이 낸 [LVI-SAM (2021)](https://arxiv.org/abs/2104.10831)은 LIO-SAM에 visual-inertial 서브시스템을 붙인 시도였다. 저자들은 이를 tightly coupled factor-graph 시스템으로 규정했다. 다만 구현은 서로 정보를 주고받고 필요하면 독립 동작도 하는 visual-inertial subsystem(VIS)과 lidar-inertial subsystem(LIS)으로 나뉜다. --- ## 17.7 Visual-LiDAR 수렴 시도: 2024-2025 2024년을 기점으로 융합 시도가 늘었다. Foundation model이 센서와 무관하게 feature를 뽑는 방향으로 발전하면서, 카메라와 LiDAR를 하나의 프레임에서 처리하는 연구가 이어졌다. 갈래는 둘이다. 하나는 multi-modal pretrained feature다. LiDAR와 카메라를 같은 embedding space로 align한다. [CLIP(Radford et al., 2021)](https://arxiv.org/abs/2103.00020)이 image-text alignment를 해낸 것처럼, LiDAR-image contrastive learning을 사용하는 접근이다. 2023-2024년 여러 그룹에서 실험 단계다. 다른 하나는 unified sensor abstraction이다. 센서 출력을 geometric primitive나 neural field로 통합한 뒤 단일 backend에서 처리한다. 이쪽은 아직 연구 논문 단계이고 실시간 동작을 보인 시스템은 드물다. 어느 방향도 아직 LiDAR SLAM과 Visual SLAM을 실질적으로 통합한 단일 계보를 만들지 못했다. FAST-LIO2와 ORB-SLAM3는 여전히 독립적으로 쓰인다. --- ## 17.8 Radar는 본 책의 scope 밖이다 LiDAR 평행 우주 바로 옆에는 또 하나의 평행 우주가 있다. Radar SLAM은 spinning radar(Navtech CIR 계열)와 SoC 기반 4D mmWave radar라는 두 하드웨어 분기 위에, Doppler radial velocity를 직접 측정해 correspondence-free odometry가 가능하다는 점, 그리고 speckle·multipath·receiver saturation 같은 전파 고유의 noise 모델 위에서 독립 subfield로 성숙했다. [Cen & Newman 2018](https://doi.org/10.1109/ICRA.2018.8460687)의 Oxford 계열 radar localisation에서 출발해 Adolfsson·Magnusson의 **CFEAR**, 그 후속 **TBV-SLAM**, Burnett·Barfoot의 continuous-time ICP까지 계보가 이어졌고, Oxford Radar RobotCar·Boreas·MulRan 같은 전용 데이터셋이 이 영역의 벤치마크 기반을 이룬다. 악천후와 연기 관통성이라는 실용 동기는 분명하지만, 본 책이 추적해 온 photogrammetry → SfM → Visual SLAM → learning → 3D foundation의 계보와는 접점이 얇다. radar는 "앞으로 합류할 이웃"으로 남겨 두고, 이 책은 별도 역사를 쓰지 않는다. 상세는 *The SLAM Handbook*(2026) Ch.9 참조. --- ## 📜 예언 vs 실제 > Zhang·Singh는 2014년 LOAM 논문 Conclusion에서 다음 두 가지를 명시적 future work로 꼽았다. 첫째, loop closure를 도입해 drift를 보정하는 것. 둘째, IMU 출력을 Kalman filter로 자신들의 방법과 결합하는 것. 두 방향 모두 이후 10년 안에 실현됐다. IMU 결합은 FAST-LIO(2021)·FAST-LIO2(2022)가 iEKF로 tightly coupled 방식으로 정리했고, loop closure는 LIO-SAM(2020)이 factor graph backend로 통합했다. 저자들이 스케치한 경로는 꽤 정확히 구현됐다. 그러나 이 두 축 너머에는, 논문 Conclusion에는 등장하지 않았지만 실무 현장에서 꾸준히 부각된 과제가 있었다. dynamic object 처리다. LiDAR 포인트에서 움직이는 보행자·차량을 실시간 분리하는 작업은 2026년 현재도 주로 deep learning segmentation에 의존하고, SLAM 내부의 geometry 기반 방법도 있지만, 다양한 동적 환경을 포괄하는 일반 해법은 정착되지 않았다. --- ## 🧭 아직 열린 것 **Visual+LiDAR 완전 융합.** LVI-SAM 이후에도 두 센서를 하나의 상태 추정기에서 다루는 tightly coupled 설계 가운데 널리 받아들여진 공통형은 없다. 안개·강우에서 카메라가 약해질 때 LiDAR가 빈자리를 채워야 하는 시나리오는 자율주행에서 명확한 요구다. 알고리즘과 센서 캘리브레이션의 난이도가 장벽이고, 2024-2025년의 transformer 기반 융합도 아직 연구 시제품 단계다. **Solid-state LiDAR에 최적화된 알고리즘.** 원래 LOAM은 360° spinning LiDAR의 scan line과 특징 구조를 전제로 했다. 일부 Livox 제품의 비반복 스캔과 solid-state 센서의 제한된 FoV는 관측 가능성과 motion distortion의 양상을 바꾼다. FAST-LIO2의 direct point-to-map 방식과 Livox LOAM처럼 이를 다루는 시스템은 이미 있지만, 서로 다른 FoV와 scan pattern을 한 설정으로 포괄하는 방법은 아직 확립되지 않았다. **동적 물체 처리.** 이 문제는 LOAM의 두 가지 명시적 future work와는 별도로 남은 과제다. 정적 환경 가정은 SLAM의 오래된 전제이고, LiDAR도 예외가 없다. 움직이는 물체를 포인트 클라우드에서 실시간 분리하는 작업은 흔히 segmentation network에 맡긴다. SLAM 내부의 geometry 기반 방법은 연산 비용과 안정성 문제를 함께 안고 있다. 실제 제품의 처리 계통은 대체로 비공개이며, 공개 연구에서도 하나의 일반 해법이 널리 받아들여지지는 않았다. --- LiDAR 계보는 Visual 주축과 접점이 적은 채로 성숙했다. 두 계보는 각자의 언어를 갖추었고, 그 언어들 사이의 번역은 아직 진행 중이다. --- # Ch.18 — 실패 사례와 사라진 계보 LOAM과 FAST-LIO2가 성숙해 가던 같은 시간, 로보틱스 커뮤니티 안에는 카메라 계보도 LiDAR 계보도 아닌 다른 방향으로 걷던 사람들이 있었다. 그 접근들이 주류가 되지 못했다고 해서 역사에 없던 것은 아니다. 각 계보의 출발점은 분명했다. Milford와 Wyeth의 RatSLAM(2004)은 O'Keefe와 Dostrovsky가 1971년에 보고한 place cell 연구를 cognitive map 이론을 거쳐 받아들였다. Event SLAM은 ETH Zürich INI의 Lichtsteiner·Posch·Delbruck가 2006년 ISSCC에서 공개하고 2008년 JSSC 논문으로 확장한 DVS와 silicon retina 계보를 이었다. Salas-Moreno 등의 SLAM++(2013)은 1990년대 object-level scene understanding을 SLAM state 안으로 옮겼다. 이들 유산이 어떤 공학적 벽을 만났는지를 되짚어 보면 기술의 성패를 가른 제약 조건이 뚜렷해진다. SLAM의 역사는 성공한 계보만으로 이루어지지 않는다. 매 10년마다 충분한 논문과 초기 결과를 갖추고도 주류로 진입하지 못한 접근들이 있었다. 공학적 확장이 막히거나, 더 실용적인 대안이 먼저 자리를 잡은 경우였다. 기술적 실패와는 다른 문제였다. --- ## 18.1 RatSLAM — place cell 기반 위상 지도 2004년 ICRA에서 [Milford et al. 2004](https://doi.org/10.1109/ROBOT.2004.1302555)가 발표한 RatSLAM은 장소 인식 문제에 전혀 다른 방식으로 접근했다. 쥐의 해마 안에 있는 **place cell**과 **head direction cell**의 발화 패턴을 모방해, 로봇이 환경을 탐색하면서 자연스럽게 장소 표현을 형성하게 했다. 계산 모델의 이름은 **Continuous Attractor Network(CAN)**이었다. 뉴런들의 활성화 상태가 2D 격자 위에서 연속적인 활성화 'bump'를 형성하고, 로봇의 속도·회전 입력(path integration)에 따라 그 bump가 격자를 따라 이동하는 구조다. 시각 입력이 들어오면 저장된 장소 표현과 비교해 bump 위치를 보정(correction)한다. 이 loop(이동으로 인한 bump 전파, 시각 매칭으로 인한 보정)이 RatSLAM의 핵심 동작 원리다. > 🔗 **차용.** [O'Keefe와 Dostrovsky(1971)](https://pubmed.ncbi.nlm.nih.gov/5124915/)의 place cell 발견은 신경과학에서 시작해 인지 지도(cognitive map) 이론으로 이어졌다. RatSLAM은 그 생물학적 메커니즘을 공학 시스템으로 옮긴 초기의 완성도 높은 시도였다. 다만 이 공학적 계보는 크게 확장되지 않았다. Milford와 Gordon Wyeth는 Queensland University of Technology(QUT) 로보틱스 연구실을 거점으로, 2004년부터 2008년 사이에 브리즈번 교외 도로에서 실외 주행 실험을 반복했다. 실험 차량은 지붕에 카메라를 달고 교외 주택가를 달렸다. RatSLAM은 그 이미지 스트림을 받아 이미 지나온 길을 알아보고 loop를 닫았다. [Milford & Wyeth 2008](https://doi.org/10.1109/TRO.2008.2004520) IEEE T-RO 논문에는 66km 경로에서 수만 장의 이미지를 처리한 결과가 실렸다. 같은 시기 기하학적 SLAM 시스템들이 몇 백 미터 단위에서 고전하던 때였으니, 숫자만 보면 RatSLAM이 앞서 있었다. 그러나 공학적 확장은 거기서 멈췄다. CAN은 장소 수가 늘수록 계산 복잡도가 올랐다. 더 깊은 문제는 정밀도였다. RatSLAM이 만드는 위상 지도(topological map)는 "여기 왔던 적 있다"는 판단은 했지만, 미터 단위의 metric 위치 추정은 안정적으로 내놓지 못했다. 자율주행과 조작(manipulation)이 요구하는 것은 정확한 좌표였다. 인지 지도는 그 요구에 맞지 않았다. > 📜 **예언 vs 실제.** Milford·Wyeth는 2008년 T-RO 논문 Conclusion에서 RatSLAM이 "vision-only SLAM의 대안적 접근"이며, 기존 state-of-the-art SLAM에게는 도전이 될 만한 환경(장거리 경로, 큰 누적 오차, 시각적 모호성)에서 반복적이고 신뢰도 높은 loop closure를 수행한다고 주장했다. 대체가 아니라 대안이라는 주장이었다. 실제로 이 주장은 부분적으로 맞았다. RatSLAM은 특정 benchmark에서 경쟁력을 보였다. 그러나 이후 분야 전체의 흐름에서는 2012년 이후 graph-based SLAM과 visual odometry가 정확도·속도 모두에서 앞서 나갔고, 위상 지도는 지금도 일부 place recognition 연구에 등장하지만, metric-topological 통합이라는 RatSLAM의 원래 야망은 다른 방식으로 이어지지 않았다. RatSLAM이 남긴 것은 "장소 표현이 기하학 없이도 가능하다"는 아이디어였다. 그 아이디어는 place recognition 문헌에 스며들었다. 2012년 [SeqSLAM](https://doi.org/10.1109/ICRA.2012.6224623)이 같은 Milford 그룹에서 나왔고, 이미지 시퀀스 비교 기반 장소 인식은 visual place recognition 벤치마크의 한 축이 됐다. 계보 자체는 살아남았고, 다만 형태가 달라졌다. --- ## 18.2 biologically-inspired SLAM의 공학적 한계 RatSLAM은 biologically-inspired SLAM의 가장 완성된 사례였지만 혼자가 아니었다. 2000년대 중반부터 2010년대 초반까지 인지 지도, entorhinal grid cell, hippocampal replay를 모방한 SLAM 변형들이 꾸준히 나왔다. 모두 비슷한 문제를 안고 있었다. 생물학적 모델은 뇌가 *어떻게* 공간을 표현하는지 기술한다. 그것이 *왜* 그 방식인지, 그 방식이 공학적 목적에도 맞는지는 다른 질문이다. 쥐의 해마가 작동하는 생물학적 조건과 로봇이 평가받는 센서·계산·정확도 조건은 서로 다르다. 공학적 SLAM은 미터 이하의 위치 추정 정확도, 실시간 처리, 새로운 환경에 대한 빠른 적응, 검증 가능한 오류 경계를 요구한다. 인지 모델은 이 조건들을 보장하기 어려웠다. 신경과학과 로봇공학은 서로에게서 영감을 얻을 수 있지만, 그 간격은 짧지 않았다. 2020년대 들어 이 논의는 다시 열릴 여지가 생겼다. Foundation model이 스스로 형성한 representation은 인지 지도와 비교할 만한 질문을 던졌지만, 그 닮음을 같은 기제로 설명할 수 있는지 아니면 비유에 그치는지는 아직 모른다. --- ## 18.3 Event SLAM — 하드웨어와 알고리즘 성숙 격차 Patrick Lichtsteiner, Christoph Posch, Tobi Delbruck가 ETH Zürich Institute of Neuroinformatics(INI)에서 개발한 Dynamic Vision Sensor(DVS)는 [ISSCC 2006 논문](https://tilde.ini.uzh.ch/users/tobi/public_html/papers/lichtsteiner_ISSCC2006_D27_09.pdf)에서 처음 공개됐고, [JSSC 2008 논문](https://doi.org/10.1109/JSSC.2007.914337)으로 확장됐다. 각 픽셀이 독립적으로 대수(log) 광도 변화를 임계값과 비교해 양(ON) 또는 음(OFF) 극성의 이벤트를 비동기(asynchronous)로 출력하는 구조다. 전역 셔터 없이 픽셀별로 발화 시점을 마이크로초 단위로 기록한다. 프레임이 없는 카메라였다. > 🔗 **차용.** DVS event sensor는 생물학적 망막의 변화 감지 메커니즘에서 착안한 하드웨어였다. Event SLAM은 이 센서를 손에 쥐고 시작했다. 센서 공개 뒤 event 기반 3D SLAM과 odometry가 이어졌지만, 그 간격을 하나의 정확한 기간으로 단정하기는 어렵다. 이벤트 카메라의 장점은 명확했다. μs 단위의 시간 해상도, 고속 운동에서 블러 없음, 고동적 범위(HDR)로 터널과 햇빛 직사 환경 모두 대응, 전력 소비는 기존 카메라의 수십 분의 일. 2014년 ICRA에서 [Weikersdorfer et al. 2014](https://doi.org/10.1109/ICRA.2014.6906882)는 event 기반 3D SLAM을 발표했다. 같은 해 다른 그룹에서도 event-based optical flow와 depth 추정이 나왔다. 2016-2018년 사이에 Henri Rebecq(Davide Scaramuzza 그룹, University of Zurich RPG 연구실)가 [EVO](https://doi.org/10.1109/LRA.2016.2645143)(RA-L 2017)와 [ESIM](https://proceedings.mlr.press/v87/rebecq18a.html)(CoRL 2018) 등을 발표하면서 event SLAM 파이프라인이 구체화됐다. 현실 환경에서는 제약이 컸다. 문제는 두 곳에 있었다. 첫째는 해상도였다. 초기 DVS 센서는 128×128 픽셀이었다. 기존 VGA 카메라와 비교할 수 없는 수준이었고, feature matching과 map building이 해상도에 직접 의존하는 SLAM에서 이 제약은 컸다. 둘째는 알고리즘 패러다임 자체였다. 기존 프레임 기반 알고리즘을 이벤트 스트림에 그대로 쓸 수 없었다. 새로운 방식이 필요했고, 그 개발에 시간이 걸렸다. 2014년부터 2018년까지 event SLAM 논문들은 controlled 환경과 각 논문이 정한 데이터에서 결과를 보고했다. 이들 실험은 기존 visual-inertial odometry와 동일한 센서·환경·평가 조건을 공유하지 않으므로, 분야 전체의 보편적 우열을 뒷받침하지는 않는다. 그 사이 event 접근의 적용 범위는 odometry 밖으로도 넓어졌다. [EventVLAD](https://ieeexplore.ieee.org/document/9635907/)(Lee & Kim, IROS 2021)는 event stream에서 복원한 edge 이미지를 NetVLAD descriptor로 묶어, 급격한 조명 변화와 모션 블러 조건에서도 장소 재인식이 가능함을 보였다. frame 기반 VPR이 취약한 환경에 event를 적용한 시도였다. --- ## 18.4 Semantic SLAM — object-as-landmark 경로의 축소 2017년부터 2019년까지 CVPR, ECCV, IROS에서 "semantic"을 다룬 연구가 빠르게 늘었다. 딥러닝이 instance segmentation과 object detection에서 연속으로 돌파구를 열던 시기였다. 이 semantic 이해를 SLAM에 통합하려는 연구가 이어졌지만, 실행은 담론을 따라가지 못했다. [Salas-Moreno et al. 2013](https://doi.org/10.1109/CVPR.2013.178)의 **SLAM++**가 그 계보의 첫 대형 선언이었다. Imperial College의 Salas-Moreno와 지도교수 Andrew Davison 그룹은 기존의 포인트나 패치 대신 *사물(object)*을 지도의 기본 단위로 삼았다. 의자, 책상, 모니터 같은 사전 정의된 3D 객체 모델을 데이터베이스에 저장하고, SLAM 실행 중 RGB-D 입력에서 ICP(Iterative Closest Point) 기반 정합으로 그 객체들을 인식해 지도에 올렸다. 포인트 수천 개 대신 객체 수십 개로 지도를 표현하면, 맵 크기가 줄고 장소 인식과 loop closure가 더 의미론적으로 이루어질 수 있었다. > 🔗 **차용.** SLAM++의 object-level representation은 그래픽스의 scene graph 표현과 컴퓨터 비전의 model-based recognition을 결합한 것이었다. 2020년대 LERF와 LangSplat도 지도 표현에 의미 정보를 결합했지만, 두 논문이 SLAM++를 직접 계승했다는 인용 계보는 확인되지 않는다. SLAM++ 이후 [SemanticFusion](https://arxiv.org/abs/1609.05130)(McCormac et al., 2017, ICRA)과 [MaskFusion](https://arxiv.org/abs/1804.09194)(Rünz et al., 2018, ISMAR)이 semantic 정보를 지도와 동적 객체 분리에 사용했다. 같은 시기 [SuperPoint](https://arxiv.org/abs/1712.07629)(DeTone et al., 2018) 기반 시스템도 등장했지만, SuperPoint는 semantic feature가 아니라 keypoint detector와 local descriptor를 함께 학습하는 방법이다. 두 흐름은 모두 학습을 썼지만 문제와 표현 단위가 달랐다. Semantic SLAM 쪽의 주장은 semantic 이해를 기하 파이프라인에 통합하면 환경 변화에 더 강건한 지도를 만들 수 있다는 것이었다. 실제 전개는 달랐다. 2019년까지 널리 비교되던 geometric 파이프라인에는 ORB-SLAM2와 VINS-Mono가 있었고, LIO-SAM은 2020년에 발표됐다. 당시 semantic 시스템들은 서로 다른 실내·객체·데이터셋 설정에서 평가돼 보편적 우열을 직접 비교하기 어려웠다. 다만 고정된 segmentation model과 class vocabulary에 의존한다는 한계는 각 논문에서 반복됐다. > 📜 **예언 vs 실제.** Salas-Moreno는 SLAM++ 논문 Conclusion에서 자신들의 방식이 "보다 일반적(generic) SLAM 방법으로 가는 첫 걸음"이라며, 낮은 차원의 형상 변이를 갖는 객체, 나아가 장기적으로는 스스로 객체 클래스를 분할·정의하는 시스템으로 확장되기를 기대했다. 논문 도입부는 이에 더해 객체 단위 표현이 "맵 저장량의 큰 압축"과 "효율성·견고성 이득"을 준다고 주장했다. 실제 전개는 일부만 적중했다. Object-level map은 AR과 특정 manipulation 응용에서 자리를 찾았고, 압축·효율 측면의 이점은 실내 반복 객체 환경에서 재확인됐다. 그러나 주류 geometric SLAM은 2026년 기준에도 sparse point와 keyframe 기반 graph를 유지하고 있고, 객체를 스스로 segmentation·정의하는 단계는 도달하지 못했다. Object-as-landmark의 확산은 제한적이었지만, semantic은 SLAM 내부의 동적 영역 분리와 하류 태스크인 semantic mapping·task planning에서 다른 역할을 찾았다. semantic-first SLAM이 주류가 되지 못한 원인은 두 곳에 있었다. 하나는 의존성이었다. Semantic SLAM은 segmentation이 정확해야 했는데, segmentation이 틀리면 지도 전체가 오염됐다. 기하학적 파이프라인은 feature matching이 부분적으로 실패해도 robust estimation으로 버텼다. 다른 하나는 일반화였다. 특정 객체 클래스로 훈련한 semantic prior는 그 클래스 밖에서 쓸모가 없었다. SLAM이 들어가야 할 환경은 그 prior가 상정한 세계보다 훨씬 넓었다. 축소된 것은 object-as-landmark 경로였다. 같은 시기 다른 경로가 살아남았다. [SuMa++](https://doi.org/10.1109/IROS40897.2019.8967704)(Chen et al., IROS 2019)가 LiDAR point cloud에 semantic class를 덧씌워 동적 물체를 걸러냈고, [Kimera](https://doi.org/10.1109/ICRA40945.2020.9196885)(Rosinol et al., ICRA 2020)가 metric-semantic mesh와 3D scene graph를 묶었다. [Hydra](https://doi.org/10.15607/RSS.2022.XVIII.050)(Hughes et al., RSS 2022)는 그 scene graph를 실시간·계층적으로 확장했고, [ConceptGraphs](https://doi.org/10.1109/ICRA57147.2024.10610243)(Gu et al., ICRA 2024)와 [Clio](https://doi.org/10.1109/LRA.2024.3451395)(Maggio et al., RA-L 2024)에 이르러 open-vocabulary foundation feature가 그 위에 얹혔다. Semantic은 지도의 상위 layer로 올라가 살아남았다. 이 계보는 2026년까지 진행 중이고, [Ch.15b](chapter_15b_dynamic.md)(Dynamic·static 분리의 semantic 귀환), [Ch.16](chapter_16_foundation_3d.md)(foundation 3D·metric-semantic 본체), [Ch.19 §19.7](chapter_19_open_problems.md#197-semantic-표현의-귀환과-open-world)(Semantic의 귀환)에서 이어 다룬다. --- ## 18.5 Manhattan-World 가정 — 적용 범위와 보조 제약으로의 존속 비슷한 시기에는 Manhattan-world assumption을 이용한 SLAM도 시도됐지만, 이후 연구 범위가 축소됐다. 가정 자체는 단순했다. [Coughlan & Yuille 1999](https://doi.org/10.1109/ICCV.1999.790349)의 Manhattan world 개념을 이어받아, 실내 환경은 대부분 세계 좌표계의 세 직교 축(x, y, z)에 정렬된 구조라고 봤다. 벽, 바닥, 천장이 그 방향을 만든다. 이미지 속 평행선 묶음은 소실점(vanishing point)으로 수렴하며, 각 소실점은 카메라의 회전 행렬 R과 방향 벡터 d의 관계 `v = K R d`로 기술된다(K: 카메라 내부 행렬). 세 직교 소실점을 찾으면 R의 세 열을 직접 복원할 수 있다. IMU나 feature matching 없이 기하학적 제약만으로 drift를 억제할 수 있다는 말이다. 이 아이디어를 visual odometry와 결합하려는 시도들이 이 시기에 등장했다. 긴 복도와 직사각형 방에서는 구조적 prior가 drift를 줄일 수 있었다. 문제는 그 밖이었다. 야외로 나가거나, 둥근 구조물이 있거나, 불규칙한 산업 환경에 들어서면 Manhattan-world 가정 자체가 성립하지 않았다. 그래서 이 가정은 적용 가능한 환경에서 선택적으로 쓰는 보조 제약으로 남았다. 조사한 논문만으로는 general-purpose visual-inertial odometry의 성숙이 이 계보를 쇠퇴시켰다거나 독립 연구 계보가 사라졌다는 인과를 뒷받침할 수 없다. --- ## 18.6 소멸 계보의 재발견 패턴 계보가 죽는다는 것이 무엇을 뜻하는지는 사례마다 다르다. RatSLAM의 topological map 아이디어는 SeqSLAM으로 이어졌고, 그 후예가 visual place recognition 분야에서 살아 있다. SLAM++와 [LERF](https://arxiv.org/abs/2303.09553)(Kerr et al., 2023)·[LangSplat](https://arxiv.org/abs/2312.16084)(Qin et al., 2023)은 모두 지도 표현에 의미 정보를 결합하지만, 직접 인용 계보라기보다 서로 다른 표현에서 나타난 공통 목표로 보는 편이 정확하다. Event camera SLAM은 경로가 달랐다. 초기에는 하드웨어의 제약이 컸다. 2022년 이후 640×480 이상의 event camera가 시장에 나왔고, 고속 드론과 HDR 환경에서의 필요가 분명해졌다. [Guillermo Gallego](https://arxiv.org/abs/1904.08405)(TU Berlin)를 중심으로 한 event vision 커뮤니티는 2020-2024년 사이에 event-based depth estimation과 ego-motion 추정에서 경쟁력 있는 결과를 냈다. 영감이 좋아도 공학이 따라오는 데 시간이 걸리고, 센서가 새로워도 알고리즘은 따로 만들어야 한다. 그 간격을 메우는 데 얼마나 걸리느냐는 알고리즘의 성숙도와 하드웨어의 실용화 속도에 달렸다. 그 사이에 더 나은 대안이 먼저 자리를 잡느냐도 변수였다. --- ## 🧭 아직 열린 것 **Biologically-inspired SLAM.** Foundation model이 대규모 비지도 학습으로 공간 표현을 형성하는 방식은 인지 지도와 구조적으로 닮은 특성이 있다. 그러나 transformer 내부 표현을 place cell과 같은 기제로 볼 수 있는지는 아직 검증되지 않았다. RatSLAM류의 계보가 foundation model 패러다임 안에서 다른 이름으로 돌아올 가능성은 가설로 남아 있다. **Event camera SLAM의 주류화.** 2022년 이후 상업용 고해상도 event camera가 보급되면서 연구 기반이 넓어졌다. 그러나 event 데이터를 효과적으로 처리하는 알고리즘 패러다임은 아직 안정적인 공통 프레임워크를 갖추지 못했다. Frame 기반 pipeline과의 통합과 새로운 event representation, 그리고 real-world benchmark의 다양화와 평가 기준 정립이 동시에 진행 중이다. 주류화 여부는 2026년 기준에도 판단이 이르다. **"Semantic map" 개념의 향방.** 2017년 semantic SLAM의 빠른 증가세가 둔화한 뒤, semantic 표현은 동적 영역 분리와 downstream task 등으로 역할을 넓혔다. 2023년부터 LERF와 언어 기반 Gaussian splatting이 언어 feature를 밀도 있는 scene representation과 결합하면서 다른 형태가 나왔다. 내부화로 이어질지, 다시 downstream으로 남을지는 모른다. geometry가 먼저 옳아야 semantic이 쓸모 있다는 패턴이 이번에도 반복될지, 아니면 representation 자체의 변화가 그 순서를 바꿀지가 관건이다. 2026년 기준으로도 다음 주류 형태는 아직 정해지지 않았다. --- # Ch.19 — 오늘의 지도와 내일의 공란 2026년에는 AR 레이어가 벽에 고정되고, 실내 배송 로봇이 지도 없이 주방과 회의실을 구분하며, DUSt3R 계열에 사진 몇 장을 입력하면 수 초 안에 3D 구조가 나온다. SLAM이 풀렸다는 인식은 이 풍경에서 나온다. 풀린 것은 2003년의 문제다. 정적 장면, 안정된 조명, 제한된 공간, 단안 카메라의 기하학이라는 가정들 위에서 EKF가 작동했고, graph SLAM이 루프를 닫았으며, ORB-SLAM이 keyframe을 관리했다. 각 답은 진짜 답이고, 각 가정은 진지하게 선택된 단순화였다. 그러나 각 계보가 풀었다고 선언한 자리 바로 옆에는 아직 열린 문제가 남아 있다. 앞선 장들에서 표시한 문제를 한곳에 모으면, 서로 다른 계보가 같은 제약과 마주치는 지점이 드러난다. 여기서 새 문제를 만드는 것이 아니라 이미 남아 있던 문제의 구조를 읽는다. --- ## 19.1 조명과 환경 변화: 카메라가 감당하지 못하는 현실 Visual SLAM이 실외로 나온 순간부터 따라다닌 문제가 있다. 현장에는 카메라의 측광 모델이 감당하지 못하는 조건이 늘 존재했다. Learned descriptor는 훈련 도메인에선 ORB를 능가하지만 underwater·thermal·low-light에서 일관성이 없고, 2026년에도 우열 합의가 없다(Ch.2 §2.7 참조). Ch.5의 저조도·동적 추적 실패는 여전하다. 2007년 PTAM이 "Small AR Workspaces"로 스스로 범위를 제한한 이유도 대부분의 feature-based SLAM에 지금도 암묵 가정으로 남아 있다(Ch.5 §🧭 참조). Direct method에서 이 문제는 더 구조적이다. 밝기 보존이라는 근본 전제가 자동 노출, 역광, 터널-야외 전환에서 즉각 붕괴하고, 조명 모델을 동적으로 추정하는 완전한 해법은 없다(Ch.8 §🧭 참조). Place recognition에서도 같은 장벽이 10년째 같은 자리다. [DINOv2](https://arxiv.org/abs/2304.07193) 기반 방법이 격차를 줄였어도, [Nordland](https://nikosuenderhauf.github.io/projects/placerecognition/)·[Oxford RobotCar](https://robotcar-dataset.robots.ox.ac.uk/)의 계절·조명 극변에서 한 모델이 조건 전반에 걸쳐 일관된 정밀도와 재현율을 내는 문제는 풀리지 않았다(Ch.10 §10.7 참조). ORB-SLAM의 장기 지도 재사용도 같은 경계에 막힌다. Atlas가 멀티맵을 가능하게 했지만 아침에 만든 지도로 저녁의 같은 장소를 인식할 때 조명이 크게 달라지면 실패할 수 있다(Ch.7 §🧭 참조). Ch.2·5·7·8·10이 같은 장벽을 각자의 언어로 보고했을 뿐이다. --- ## 19.2 정적 세계 가정: 가장 오래된 단순화의 한계 정적 세계 가정은 SLAM의 가장 오래된 단순화이며, 여러 계보가 이 가정의 한계를 보고했다. SfM 계보에서 동적 물체는 COLMAP 같은 주류 정적 장면 파이프라인의 취약점이고, 2026년 기준 COLMAP 수준의 범용성을 가진 Dynamic SfM 구현체는 확인되지 않는다(Ch.3 §3.7 참조). KinectFusion부터 BundleFusion까지 모두 정적 장면 전제 위에 있고, DynaSLAM·MaskFusion의 실시간 segmentation 결합 시도는 비용·robustness 모두에서 실제 배포 수준에 못 미친다(Ch.9 §🧭 참조). Monocular depth에서는 self-supervised가 moving object를 masking으로 회피한다(Ch.11 §🧭 참조). 3DGS SLAM은 2025년에도 정적 세계 가정 위에 있고, [4DGS](https://arxiv.org/abs/2310.08528)·[Deformable 3DGS](https://arxiv.org/abs/2309.13101)가 시간 차원을 탐색 중이지만 SLAM 설정의 통합된 방식은 없다(Ch.15 §🧭 참조). LiDAR SLAM도 면제되지 않는다. LOAM의 명시적 future work와 별개로 동적 처리 문제는 남아 있고, 자율주행 기업의 production stack은 대부분 사유 기술이라 공개 연구 시스템과 직접 비교하기 어렵다(Ch.17 §🧭 참조). 다섯 장에서 같은 질문이 반복된다. long-term dynamic/deformable 문제도 같은 층위다. **Absence vs evidence of absence**(객체가 사라졌는가, 가려졌는가)는 [Schmid의 Panoptic Multi-TSDF](https://doi.org/10.1109/LRA.2022.3148854)(2022)가 active submap으로 부분 답을 냈다. Handbook은 70%를 넘는 가림을 여전히 어려운 extreme-environment 사례로 들지만, 이를 Panoptic Multi-TSDF의 보편적 오차 임계값으로 보고하지는 않는다. **Floating Map Ambiguity**(카메라 rigid motion과 객체 rigid motion 분리)는 isometric·visco-elastic prior로 우회될 뿐 prior 없는 식별 조건은 미해결이다. Monocular RGB에서 Khronos 수준의 change-aware 온라인 통합 시스템은 없고, 의료 MIS는 phantom·ex vivo를 넘어 실제 수술 환경에서 견고성이 떨어진다. 이 네 항목은 여전히 열려 있다. --- ## 19.3 Scale과 표현 메모리: 크기가 달라지면 문제가 달라진다 SLAM 시스템이 방 한 칸에서 건물로, 건물에서 도시로 확장될 때마다 같은 질문이 새로운 형태로 돌아왔다. Monocular scale은 1980년대 SfM 이론이 이미 밝힌 기하학적 모호성이다. IMU·depth 같은 추가 정보가 없으면, 단안 영상의 투영 기하만으로는 궤적과 장면의 metric scale을 정할 수 없다(Ch.5 §🧭 참조). Ch.11에서는 같은 질문이 다른 언어로 재등장한다. [Metric3D v2](https://arxiv.org/abs/2404.15506)·[Depth Anything v2](https://arxiv.org/abs/2406.09414)가 intrinsic 조건부 metric depth를 내놓았지만, intrinsic을 모르는 상황(스마트폰, CCTV, 아카이브, 위성)이 흔하고 카메라 독립적 metric depth는 foundation scale에서도 쉽지 않다(Ch.11 §🧭 참조). TSDF 계보에서 메모리 문제는 표현의 한계로 드러났다. [Voxblox](https://arxiv.org/abs/1611.03631)·[OctoMap](https://octomap.github.io/)이 비용을 줄였어도 dense 표현의 메모리는 지도 범위와 voxel 해상도에 따라 빠르게 늘고, 어느 영역에 어느 해상도를 둘지 자동 결정하는 범용 adaptive-resolution 방식은 확립되지 않았다(Ch.9 §🧭 참조). NeRF-SLAM도 같은 천장에 막혔다. 도시 규모는 개방형이다(Ch.14 §🧭 참조). Gaussian Splatting에서도 장면 범위와 세부도가 커지면 필요한 Gaussian 수가 늘며, [Compact 3DGS](https://arxiv.org/abs/2311.13681)(Lee et al. 2024) 계열이 압축을 탐색하지만 합의된 방법은 없다(Ch.15 §🧭 참조). Foundation 3D에서 이 문제는 transformer의 물리적 한계로 재정의된다. 모든 image token 사이의 attention은 token 수에 대해 quadratic한 메모리를 요구하므로 긴 시퀀스로 확장하기 어렵고, Spann3R의 incremental 방식은 부분 답이다(Ch.16 §🧭 참조). 표현이 바뀌어도 크기의 장벽은 같은 자리에 있다. 크기 문제의 다른 얼굴은 **데이터 이동 비용**이다. 용량이 아니라 프로세서-메모리 사이 비트 이동의 물리적 비용이 전력을 먹는다. Davison은 Handbook Ch.18 §18.8에서 12번째 SLAM 지표로 "on-device data movement, measured in bits × millimetres"를 제안하며 metric을 하드웨어 공학의 언어로 재정의한다(p.547). Handbook Ch.16의 식도 표현별 복잡도를 구분한다. label별 flat voxel map은 $O(L\cdot V/\delta^3)$이고(Eq. 16.34), metric-semantic hierarchy는 $O(V/\delta^3+N_\text{objects}+N_\text{rooms}+\cdots)$이며(Eq. 16.35), topological 3D scene graph는 $O(N_\text{sub-sym}+N_\text{objects}+N_\text{rooms}+\cdots)$다(Eq. 16.36). Davison의 12번째 지표 재정의가 얼마나 받아들여질지는 결론이 없다. --- ## 19.4 학습 기반 시스템의 불확실성 calibration Julier와 Uhlmann이 Ch.4에서 EKF의 inconsistency를 증명한 이래, SLAM 시스템이 "자신이 어디 있는지 모른다는 것을 얼마나 정확하게 아는가"는 이 분야의 물음으로 남아 있다. 비가우시안 불확실성은 EKF의 핵심 가정에 닿는다. 현실 센서 오류는 다중 모드·heavy-tail이 흔하고, Stein particle·normalizing flow·learned uncertainty가 시도되나 실시간 검증은 제한적이다(Ch.4 §4.8 참조). Graph SLAM에서 robust cost function 선택도 직관에 기댄다. Huber·Cauchy·Geman-McClure 중 환경·센서에 맞는 kernel을 사전 결정하는 원칙적 방법이 없다(Ch.6 §🧭 참조). [Ch.6b](chapter_06b_certifiable.md)의 tightness 경계도 같은 층위다. SE-Sync의 exact recovery는 노이즈 $\beta$ 이하라는 충분조건만 주고, 실제 인스턴스에서 $\beta$를 사전 계산하는 방법은 없다. Visual SLAM·VIO로 certifiable을 확장하는 문제, 새 측정이 들어올 때 SDP를 다시 풀어 certificate를 갱신하는 online certification도 열린 채 남는다. 학습 기반 방법에서 문제는 더 날카롭다. Bayesian PoseNet 실패 이후에도 learned uncertainty가 OOD 입력에서 calibrated인지는 열려 있다(Ch.12 §🧭 참조). DROID-SLAM 계보에서 확인했듯 learned prior는 훈련 도메인 밖에서 성능이 저하되어도 그 정도가 겉으로 드러나지 않을 수 있다. learned 실패는 그럴듯한 모양으로 나타날 수 있다. geometric 방법에도 잘못된 지역해나 과도한 확신처럼 겉으로 드러나지 않는 실패가 있다. [TartanAir](https://arxiv.org/abs/2003.14338) 같은 합성 데이터로도 sim-to-real gap이 남는다(Ch.13 §🧭 참조). Foundation 3D에서는 이 문제가 loop closure 재정의로 이어진다. DUSt3R 계열에서 pointmap 기반 교정 propagate는 MASt3R-SLAM이 기존 방식으로 처리하지만 원리적 해법인지는 불확실하다(Ch.16 §🧭 참조). 자율주행·의료 로봇에서 calibrated uncertainty가 필수인데 그 수준의 시스템은 드물다. Davison은 Handbook Ch.18에서 문제를 재정식화한다. *"100장으로 3D 모델을 만든 네트워크에 이미지 1장이 추가되면 전체를 다시 돌려야 하는가"*(p.528). 장기 표현과 fusion을 인정하는 순간 probabilistic state estimation과 modular scene representation이 필요해진다. 대안으로 제시된 [GBP Learning](https://arxiv.org/abs/2312.14294)(Nabarro et al.)은 신경망 weight를 factor graph의 random variable로 넣어 *"training time"*과 *"test time"*의 구분을 지우는 방향이다(§18.7, pp.543-545). 이것이 원리적 답인지 문제 이관인지는 판단이 이르다. --- ## 19.5 센서 융합과 새 모달리티: 통합의 미완 Visual SLAM과 LiDAR SLAM은 같은 시기에 같은 문제를 다른 언어로 풀었다. 통합 시도에도 두 계보는 하나로 수렴하지 않았다. LVI-SAM은 visual-inertial subsystem과 lidar-inertial subsystem을 factor graph로 결합했고, 저자들은 이를 tightly coupled 시스템으로 규정했다. 두 subsystem이 서로 정보를 주고받으면서도 필요하면 독립 동작하는 구조라는 점은 함께 구분해야 한다. 안개·강우 같은 자율주행 시나리오에서 다중 센서 융합의 알고리즘·캘리브레이션 난이도는 여전히 장벽이다(Ch.17 §🧭 참조). Solid-state LiDAR 보급이 가져온 알고리즘 공백도 같은 층위다. 원래 LOAM의 360° scan-line 가정과 달리 제한된 FoV와 비반복 pattern을 쓰는 센서는 다른 관측 조건을 만든다. FAST-LIO2와 Livox LOAM이 일부를 다루지만 센서 계열 전반에 대한 일반화는 미흡하다(Ch.17 §🧭 참조). 융합 문제는 wide-baseline 매칭에서도 나타난다. 큰 시점 변화에서 Harris·ORB 같은 hand-crafted local feature의 성능이 저하되지만, 모든 데이터에 적용되는 45도 임계값은 확인되지 않는다. DUSt3R는 전통적인 sparse matching 단계를 두지 않지만 이것이 descriptor 문제의 종말인지 우회인지는 판단이 이르다(Ch.2 §2.7 참조). Place recognition과 metric localization의 통합도 파이프라인 수준의 단절이다. 두 과정을 하나의 표현으로 통합하는 2023-2025년 시도들이 있었지만 정밀도·속도를 동시에 달성한 방법은 없다(Ch.10 §10.7 참조). Event camera는 모달리티가 새로울 때 알고리즘이 얼마나 뒤따르는지를 보여준다. 2022년 이후 상업 고해상도 event camera가 보급되었지만 frame 기반 pipeline과의 통합, event representation, real-world benchmark가 동시 진행 중이다(Ch.18 §🧭 참조). Kinect가 2010년 출시되고 1년 뒤 KinectFusion이 나왔던 순서와 같다. 이 책이 범위 밖으로 둔 모달리티가 있다. **4D imaging radar**와 **legged/proprioceptive SLAM**이다. Radar는 안개·강우처럼 카메라와 LiDAR의 성능이 함께 저하될 수 있는 조건을 보완한다. Oxford Radar RobotCar(2019), NuScenes의 radar 채널, Arbe·Mobileye를 비롯한 4D imaging radar 개발은 이 모달리티를 자율주행 연구와 제품 개발의 한 축으로 넓혔다. Legged SLAM은 ANYmal·Spot·Unitree의 2020년대 실외 배포와 함께 kinematic·contact prior 융합의 별도 계보를 열었다. 둘 다 visual·LiDAR·foundation 3D와 다른 원류·벤치마크를 가지며, 각자의 역사서가 필요한 크기다. --- ## 19.6 계산 구조와 하드웨어의 재결합 2020년대 후반 Davison Handbook Ch.18은 알고리즘의 그래프 구조와 실리콘의 그래프 구조를 정합시키는 문제를 전면에 놓았다. Dennard scaling 붕괴로 단일 코어 clock speed가 2000년대 중반 약 4GHz에서 정체됐다는 논의가 *"this has stopped being true"*라는 문장에 이어지고(Handbook Ch.18, pp.528-529), 착용형 Spatial AI의 제약은 안경 한 짝(65g, <1W)으로 남아 있다. 이 간극이 **heterogeneous·specialized·parallel** 아키텍처로 분야를 밀어넣는다. Handbook은 서로 다른 시기의 구체 실리콘 사례를 나란히 놓는다. [Apple Vision Pro R1](https://www.apple.com/uk/newsroom/2023/06/introducing-apple-vision-pro/)(2023)은 카메라·센서의 새 이미지를 12ms 안에 디스플레이로 전달하도록 설계됐고, [Meta Aria Gen 2](https://ai.meta.com/blog/aria-gen-2-research-glasses-under-the-hood-reality-labs/)(2025)는 on-device machine perception용 custom coprocessor를 쓴다. [Graphcore IPU](https://www.graphcore.ai/products/ipu)는 수천 코어가 로컬 메모리와 메시지 패싱으로 연결되고, Manchester [SCAMP5](https://personalpages.manchester.ac.uk/staff/p.dudek/papers/carey-iscas2013.pdf)는 256×256 per-pixel in-plane processing을 1.2W에 처리하며, [SpiNNaker](https://apt.cs.manchester.ac.uk/projects/SpiNNaker/)는 ARM 코어 최대 100만 개의 neuromorphic 구조로 동작한다. 각자 다른 graph topology를 요구하고, 어느 실리콘에 어떻게 매핑할지에 대한 체계적 이론은 아직 없다. 이 축 위에서 Davison의 후기 연구 주제인 **Gaussian Belief Propagation**이 자리를 잡았다. [Ortiz et al.](https://arxiv.org/abs/2203.11618)(2022)은 IPU에서 GBP로 Bundle Adjustment를 CPU 대비 30× 가속했고, [Murai et al. Robot Web](https://arxiv.org/abs/2306.04620)(2024)은 여러 로봇이 Wi-Fi로 factor graph 조각을 공유해 asynchronous message passing으로 수렴하는 다중 로봇 SLAM을 보였다. *"We must get away from the idea that a 'god's eye view' of the whole structure of the graph will ever be available"*(Handbook Ch.18, p.541)가 이 계보의 철학이다. Factor graph를 master representation으로 두고 full posterior를 포기한 채, 메시지가 그래프 위를 "bubble"하며 국지적으로 수렴한다. 이 접근이 MASt3R-SLAM 같은 transformer 기반 시스템과 결합할지, 끝까지 다른 줄기로 남을지는 아직 답이 없다. Davison이 제안한 12개 지표 중 11번 "power usage"와 12번 "on-device data movement"가 하드웨어 공학의 새 지표다. 정확도만큼 **전력과 장치 안에서 이동하는 데이터의 비트량·거리**로 평가하라는 제안이고, TUM·KITTI·EuRoC 같은 주류 벤치마크로 흡수될지는 합의가 없다. 알고리즘 중심인 이 책의 편향 바깥 영역이며, 그 편향 자체가 2020년대 후반 새로 문제화되고 있다. --- ## 19.7 Semantic 표현의 귀환과 Open-World Semantic이 landmark 자리에서 축소됐다는 [Ch.18 §18.4](chapter_18_dead_ends.md#184-semantic-slam--object-as-landmark-경로의-축소)의 판정은 좁은 의미에서 사실이다. ORB-SLAM3도 MASt3R-SLAM도 object-level primitive를 쓰지 않는다. 그러나 같은 시기에 semantic은 **지도의 상위 layer**로 올라가 실질적 성과를 냈다. 흐름은 뚜렷하다. [Kimera](https://doi.org/10.1109/ICRA40945.2020.9196885)(2020)가 metric-semantic mesh와 3D scene graph를 묶고, [Hydra](https://doi.org/10.15607/RSS.2022.XVIII.050)(2022)가 이를 실시간·계층적으로 확장했다(*"first online system to produce fully hierarchical scene graphs that included objects, places, and rooms"*, Handbook Ch.16, §16.4.2). 그 위에 foundation feature가 얹혔다. [ConceptFusion](https://arxiv.org/abs/2302.07241)·[VLMaps](https://arxiv.org/abs/2210.05714)(2023)가 CLIP을 dense map에, [ConceptGraphs](https://doi.org/10.1109/ICRA57147.2024.10610243)(2024)가 open-vocabulary object node에, [Clio](https://doi.org/10.1109/LRA.2024.3451395)(2024)가 task-driven hierarchy에, [LERF](https://arxiv.org/abs/2303.09553)·[LangSplat](https://arxiv.org/abs/2312.16084)이 radiance field와 Gaussian splatting에 CLIP을 실었다. Semantic SLAM은 표현 층위를 올렸다. 그러나 이 흐름이 해결한 것보다 연 것이 더 많다. Hughes/Carlone이 꼽은 open problem은 *"performing uncertainty quantification in hierarchical representations mixing discrete and continuous variables is still a largely unexplored problem"*(p.488). object category·room ID 같은 discrete 변수와 pose·surface 같은 continuous 변수가 섞인 그래프의 불확실성 전파는 아직 충분히 탐구되지 않았다. Outdoor·unstructured로 scene graph를 확장하는 문제도, task-driven hierarchy의 동적 재구성(Clio의 Information Bottleneck, Handbook Ch.17 Eq. 17.8)의 일반화도 열려 있다. 더 큰 질문은 "지도가 여전히 필요한가"다. Handbook Ch.17 §17.4.2 "Revisiting the Question of the Need for Maps"에서 Paull과 편집자들이 직접 다룬다. long-context VLM에 과거 프레임을 다 넣으면 explicit scene graph 없이 planning이 가능한가? [OpenEQA](https://open-eqa.github.io/)와 [Mobility VLA](https://arxiv.org/abs/2407.07775)(2024)의 결과는 map-free가 단기·단순 과제엔 작동하지만 공간·시간 범위가 길어지면 실패한다는 것이다. *"the need for an explicit map representation ... largely depend[s] on the spatial and temporal horizons of the considered tasks and remains an active area of research"*(p.515). 풀렸다는 선언도, 불필요하다는 선언도 나오지 않았다. SLAM과 생성형 로봇 정책의 관계도 같은 질문으로 이어진다. [RT-2](https://robotics-transformer2.github.io/)(2023)·[OpenVLA](https://arxiv.org/abs/2406.09246)(2024)·[π₀](https://www.physicalintelligence.company/blog/pi0)(2024) 같은 VLA 모델이 SLAM을 대체하는가, 위에 서는가. Handbook Ch.17의 결론부는 *"true generalization and scalability to compositional tasks ... could be achieved through some form of explicit structure that is learned through a process such as SLAM. ... these two paradigms ... are entirely complementary"*라고 정리한다(Paull/Carlone, §17.4.4, p.520). 합의에 가장 가까운 입장이지만 "complementary"가 어떤 아키텍처 결합인지는 열려 있다. --- ## 19.8 열린 질문의 구조 열린 문제들이 같은 방식으로 남아 있는 것은 아니다. Ch.5의 monocular scale ambiguity는 SfM 이론에서 이미 증명된 기하학적 사실이고, 2026년에도 같은 정식화로 남아 있다. 반면 정적 세계 가정의 한계는 형태를 바꾸면서 20년 동안 되돌아왔다. Ch.3의 SfM 언어로, Ch.9의 dense SLAM 언어로, Ch.15의 Gaussian 언어로, Ch.17의 LiDAR 언어로 각각 다르게 나타났다. Foundation 3D 계보에서 loop closure를 어떻게 재정의할 것인지, learned uncertainty를 어떻게 calibrate할 것인지는 비교적 최근에 부각된 문제로 역사가 짧다. Ch.0은 SLAM이 풀렸다고 여겨지는 시대를 묘사했다. 그 묘사는 정확하다. 같은 2026년 *The SLAM Handbook*의 Epilogue 조언 모음에 *"If someone tells you 'SLAM is solved,' don't listen to them"*이라는 문장이 실린 것도 같은 풍경을 내부에서 본 것이다. SLAM의 역사는 언제 무엇을 놓아줘야 하는지 배우는 과정이었다. 어떤 가정을 놓아주는 순간, 이전에 닫혔던 문제가 새로운 형태로 돌아온다. EKF의 국소 선형화와 단일 Gaussian 근사에서 벗어나자 particle filter가 뒤를 이었고, sparse feature를 놓자 dense method가, geometric prior를 놓자 learned prior가 그 자리를 채웠다. 각 전환은 새로운 가정 체계로 넘어가는 일이었다. 2026년에 풀렸다고 여기는 것도 대부분 이 순환 어딘가에 있다. 지금 확신하는 가정이 흔들릴 때 공란이 다시 생긴다. --- ## 19.9 계보 약도 연도는 논문이 처음 공개된 시점을 기준으로 적었다. 따라서 ORB-SLAM3은 arXiv 공개 연도인 2020년, MASt3R-SLAM은 첫 공개 연도인 2024년으로 표시한다. ```mermaid graph TD PM[사진측량 1858] BA[Bundle Adjustment
Brown 1958] SfM[Photo Tourism 2006] COLMAP[COLMAP 2016] SC[Smith-Cheeseman 1986] Mono[MonoSLAM 2003] PTAM[PTAM 2007] ORB[ORB-SLAM 2015] ORB3[ORB-SLAM3 2020] LSD[LSD-SLAM 2014] DSO[DSO 2016] VIDSO[VI-DSO 2018] LM[Lu-Milios 1997] FG[Factor Graph
Dellaert 2000s] iSAM[iSAM 2008] iSAM2[iSAM2 2012] g2o[g2o 2011] Forster[Preintegration
Forster 2015] VINS[VINS-Mono 2018] Kinect[KinectFusion 2011] Elastic[ElasticFusion 2015] SESync[SE-Sync 2019] TEASER[TEASER 2020] LOAM[LOAM 2014] FAST[FAST-LIO 2021] NeRF[NeRF 2020] iMAP[iMAP 2021] NICE[NICE-SLAM 2021] GS3D[3DGS 2023] Spla[SplaTAM 2023] MonoGS[MonoGS 2023] DROID[DROID-SLAM 2021] DPV[DPV-SLAM 2024] DUSt3R[DUSt3R 2023] MASt[MASt3R 2024] VGGT[VGGT 2025] MASlam[MASt3R-SLAM 2024] Kimera[Kimera 2020] Hydra[Hydra 2022] Clio[Clio 2024] PM --> BA --> SfM --> COLMAP SC --> Mono --> PTAM --> ORB --> ORB3 PTAM -.-> LSD --> DSO --> VIDSO LM --> FG --> iSAM --> iSAM2 FG --> g2o FG -.-> SESync Forster --> VINS Forster --> ORB3 Forster --> VIDSO Kinect --> Elastic LOAM -.-> FAST NeRF --> iMAP --> NICE GS3D --> Spla GS3D --> MonoGS DROID --> DPV COLMAP -.-> DUSt3R --> MASt DUSt3R -.-> VGGT MASt --> MASlam Kimera --> Hydra --> Clio click PM "#chapter-1" "Ch.1 선사시대 — 사진측량" click BA "#chapter-1" "Ch.1 선사시대 — Bundle Adjustment" click SfM "#chapter-3" "Ch.3 Structure from Motion" click COLMAP "#chapter-3" "Ch.3 SfM — COLMAP" click SC "#chapter-4" "Ch.4 EKF-SLAM — Smith-Cheeseman" click Mono "#chapter-5" "Ch.5 MonoSLAM·PTAM" click PTAM "#chapter-5" "Ch.5 MonoSLAM·PTAM" click ORB "#chapter-7" "Ch.7 ORB-SLAM 계열" click ORB3 "#chapter-7" "Ch.7 ORB-SLAM3" click LSD "#chapter-8" "Ch.8 Direct Methods — LSD-SLAM" click DSO "#chapter-8" "Ch.8 Direct Methods — DSO" click VIDSO "#chapter-8" "Ch.8 Direct Methods — VI-DSO" click LM "#chapter-6" "Ch.6 Graph SLAM — Lu-Milios" click FG "#chapter-6" "Ch.6 Graph SLAM — Factor Graph" click iSAM "#chapter-6" "Ch.6 Graph SLAM — iSAM" click iSAM2 "#chapter-6" "Ch.6 Graph SLAM — iSAM2" click g2o "#chapter-6" "Ch.6 Graph SLAM — g2o" click Forster "#chapter-7" "Ch.7b IMU Preintegration (Ch.7 뒤)" click VINS "#chapter-7" "Ch.7 — VINS-Mono" click Kinect "#chapter-9" "Ch.9 RGB-D — KinectFusion" click Elastic "#chapter-9" "Ch.9 RGB-D — ElasticFusion" click SESync "#chapter-6" "Ch.6b Certifiable (Ch.6 뒤)" click TEASER "#chapter-6" "Ch.6b Certifiable — TEASER" click LOAM "#chapter-17" "Ch.17 LiDAR — LOAM" click FAST "#chapter-17" "Ch.17 LiDAR — FAST-LIO" click NeRF "#chapter-14" "Ch.14 NeRF-SLAM" click iMAP "#chapter-14" "Ch.14 NeRF-SLAM — iMAP" click NICE "#chapter-14" "Ch.14 NeRF-SLAM — NICE-SLAM" click GS3D "#chapter-15" "Ch.15 Gaussian Splatting" click Spla "#chapter-15" "Ch.15 — SplaTAM" click MonoGS "#chapter-15" "Ch.15 — MonoGS" click DROID "#chapter-13" "Ch.13 Hybrid — DROID-SLAM" click DPV "#chapter-13" "Ch.13 — DPV-SLAM" click DUSt3R "#chapter-16" "Ch.16 Foundation 3D — DUSt3R" click MASt "#chapter-16" "Ch.16 — MASt3R" click VGGT "#chapter-16" "Ch.16 — VGGT" click MASlam "#chapter-16" "Ch.16 — MASt3R-SLAM" click Kimera "#chapter-16" "Ch.16 §16.6 Semantic Foundation — Kimera" click Hydra "#chapter-16" "Ch.16 §16.6 Semantic Foundation" click Clio "#chapter-16" "Ch.16 §16.6 — Clio" ```
# Ch.0 — SLAM Solved? In 2026, you pick up a phone and an AR layer sticks to the wall. Indoor delivery robots tell the kitchen from the conference room without being handed a map. Give a few photos to a [DUSt3R](https://arxiv.org/abs/2312.14132)-family model and a 3D structure emerges in seconds. By now, these are products rather than demos, part of the background. SLAM often looks like a more-or-less solved problem. --- Go back to 2003 and the scene is different. Andrew Davison, in a lab at Imperial College London, demonstrated real-time 3D tracking with one laptop and one webcam. The system, called [MonoSLAM](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf), ran at 30 Hz on a desktop, tracked about ten features per frame, and maintained a sparse map of a few dozen landmarks. It covered one desk in one room; when the camera left the desk, the map diverged. That scale and limitation were representative of real-time monocular SLAM at the time. Those systems tracked far fewer features per frame than a phone AR session does today. How did the tracking systems of 2003 develop into the AR systems of 2026? --- SLAM's history is not a single development curve. It traces four traditions that ran independently before colliding and absorbing one another. Photogrammetrists solved bundle adjustment by hand a century ago. Roboticists began treating maps in the language of probability with [Smith-Cheeseman](https://arxiv.org/abs/1304.3111)'s 1986 stochastic spatial-relations framework, and the name "SLAM" was attached to this problem setting nine years later, in [Durrant-Whyte & Leonard's 1995 survey](https://ieeexplore.ieee.org/document/476131). Computer vision researchers focused on real-time feature tracking. The deep learning community of the 2020s is trying to absorb all of it into a single network. The book asks not "how" but "why this way." Was the replacement of EKF-based SLAM by graph-based SLAM a natural technical evolution, or a contingency decided by a few people? Was the split between feature-based and direct methods foreseen from the start? Why has deep learning been so slow to replace the geometry pipeline? Counterfactuals matter only when the alternatives actually existed. Here, they did. --- Tracing that path needs tools. A list of years gives a chronicle; an explanation of techniques gives a textbook. This history uses two lenses: lineage and prediction. Where did an idea come from? How did the future researchers expected differ from what actually unfolded? Four devices recur throughout the book and guide each chapter. **Lineage openings** sit in the first paragraph or two of a chapter. They show, through names and years, which intellectual inheritance the chapter's protagonist took on. No idea in SLAM was born in a vacuum. Follow the lineage and the terrain of borrowing becomes visible. **🔗 Borrowed boxes** are margin annotations that state in one or two sentences where a specific technique came from: "ORB-SLAM's structure here came from Strasdat 2011." Researchers cite their sources, but often leave the lineage implicit. The box makes it explicit. **📜 Prediction vs. outcome boxes** contrast what the original paper's Conclusion, Future Work, or Summary section anticipated with what actually happened. In §12, "Summary and Recommendations," the [Triggs 1999](https://dblp.org/rec/conf/dagstuhl/TriggsMHF99.html) bundle adjustment (BA) synthesis made the exploitation of large-scale sparse structure a central recommendation. In the 2010s, [COLMAP](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf) approached that problem from another direction, turning SfM over tens of thousands of images into an open-source production tool. The direction was right; the route was not. This device examines the gap between the future a researcher expected and the one that arrived. **🧭 Still open** sits at the end of the chapter. It lists questions on that chapter's subject that remain unresolved as of 2026, drawing out the open problems hidden by the perception that SLAM is solved. Ch.19 gathers these items from every chapter and reassembles them by theme. --- The book runs in six parts. **Part 1: Prehistory** traces the tools that photogrammetry and classical computer vision built up before SLAM was born in robotics. Why is bundle adjustment still the skeleton of every optimization backend? **Part 2: Classical SLAM** follows probabilistic mapping and the limits of the EKF through MonoSLAM, PTAM, graph-based SLAM, and optimality certification. How did joint mapping and estimation expand from filtering to optimization? **Part 3: Maturity** covers ORB-SLAM, inertial preintegration, continuous-time estimation, direct methods, RGB-D, and place recognition. It traces how systems expanded beyond early real-time demonstrations to larger environments and more varied sensors. **Part 4: Learning Fusion** covers monocular depth estimation, end-to-end SLAM, and hybrid methods that combine geometry with learning. Which computations did learning take over, and where did geometric constraints remain? **Part 5: Representation** covers Neural Radiance Fields, 3D Gaussian Splatting, dynamic-scene representations, and 3D foundation models. Changes in how maps and scenes are represented also changed the relationship between tracking and reconstruction. **Part 6: Dead Ends and Open Problems** pulls out the failed routes in SLAM's history and the structural unresolved problems still sitting behind today's perception that it is "solved." --- A useful map needs boundaries. The relationship between foundation models and SLAM is treated as an open question arising from current research, without predicting a settled outcome. The book's material is what happened in the past and why. Nor does it set out to declare earlier research wrong. It asks what a choice meant under the constraints of its moment. Homogeneous coordinates, epipolar geometry, and EKF formulas are assumed knowledge. The task here is to trace lineage rather than explain concepts; choosing a camera or LiDAR belongs to another book. For a systematic account of the equations, theorems, and proofs, see the [SLAM Handbook](https://github.com/SLAM-Handbook-contributors/slam-handbook-public-release). Edited by Carlone, Kim, Barfoot, Cremers, and Dellaert and published by Cambridge University Press in 2026, its 18 chapters cover the current theory and systems of SLAM. The present history records the path to that point. The five editors close the Handbook with the line *"If someone tells you 'SLAM is solved,' don't listen to them."* The tendency to treat SLAM as solved, noted at the opening of this chapter, is a phenomenon within the field rather than its consensus. --- When Davison stood in front of his webcam in 2003, he did not know exactly what he was starting. That demo video is still on the internet: the shaky frame, the blinking landmark dots, and a sparse map containing only dozens of points. The history between that room and today's systems is the subject of this book. The record starts well before MonoSLAM. Before the acronym "SLAM" settled in the 1990s, and even before Smith-Cheeseman expressed a probabilistic map in equations, photogrammetrists were already recovering 3D structure from cameras. The next chapter traces that prehistory. --- # Ch.1 — Photogrammetry and Bundle Adjustment: The 100 Years Before Triggs The skeleton of today's SLAM optimization backend was born in German surveying. In the early twentieth century, Carl Pulfrich's method for hand-computing two-view triangulation on glass plates combined with Albrecht Meydenbauer's photogrammetric system to form a single surveying tradition. That tradition passed through Duane C. Brown's numerical formulation in 1958, and in 1999 Bill Triggs, Philip McLauchlan, Richard Hartley, and Andrew Fitzgibbon translated it into the language of computer vision. Their synthesis did not invent bundle adjustment; it made a century-old surveying inheritance usable by the computer vision community. Triggs et al. (1999) inherited the parallax principle from Pulfrich's geometry and the reprojection formulation from Brown's (1958) military surveying. Levenberg-Marquardt supplied the solver skeleton. --- ## 1. Early twentieth-century glass plates and stereophotogrammetry In 1901, [Carl Pulfrich](https://en.wikipedia.org/wiki/Carl_Pulfrich) presented the **stereocomparator**, built by the Zeiss optical works, at the Hamburg conference of natural scientists (this was the formal unveiling, following a prototype stereoscopic rangefinder shown in Munich in 1899). The device photographed the same point from two camera viewpoints and computed distance by reading the coordinate difference on the glass plates. The principle was simple: the parallax between two views is inversely proportional to depth. The mathematics was Greek-era trigonometry; what was new was the precision of the optical instrument. A generation earlier, [Albrecht Meydenbauer](https://www.uni-marburg.de/de/fotomarburg/histfoto/gliederung/messbilder/messbildverfahren) had systematized **architectural photogrammetry** for the preservation of buildings. In 1858, after nearly falling while surveying the exterior of Wetzlar Cathedral, he conceived of using photographs in place of direct measurement. In 1885 he founded the Royal Prussian Photogrammetric Institute (Königlich Preussische Messbild-Anstalt). These two streams merged in twentieth-century aerial surveying. In aerotriangulation, an aircraft photographed terrain from above and pairs of images yielded three-dimensional maps. It was the age of hand calculators. > 🔗 **Borrowed.** Modern SLAM's stereo depth estimation runs on the same principle as Pulfrich's stereocomparator. Depth comes from the baseline between two cameras and the parallax. A pixel array has replaced the glass plate of 125 years ago. --- ## 2. 1958, Brown, and numerical bundle adjustment What Pulfrich and Meydenbauer had solved with optical instruments, Brown moved into equations. [Duane C. Brown](https://digital.hagley.org/08206139_solution) was a surveying engineer in the United States Air Force ballistic missile development program. He worked on jointly estimating satellite orbits and ground coordinates, simultaneously optimizing many camera viewpoints and many ground control points. In his 1958 report "A Solution to the General Problem of Multiple Station Analytical Stereotriangulation" (RCA-MTP Data Reduction Technical Report No. 43, AFMTC-TR-58-8), Brown left one of the early documents that formulated **bundle adjustment (BA)** numerically (Helmut Schmid is named as a co-inventor from the same period). The core is the **reprojection error**. Minimize the difference between the 2D image coordinate $x_{ij} \in \mathbb{R}^2$ observed in camera $i$ and the predicted coordinate $\pi(K_i, R_i, t_i, X_j)$ obtained by projecting the 3D point $X_j \in \mathbb{R}^3$ through the intrinsic matrix $K_i$ and extrinsic matrix $[R_i | t_i]$: $$E = \sum_{i,j} \| x_{ij} - \pi(K_i, R_i, t_i, X_j) \|^2$$ The name "bundle" comes from the bundle of rays extending out from each camera center to the observed 3D points. Camera poses and point locations are adjusted together so that those rays meet at the 3D points. It took forty years for a technique that began in military and intelligence applications to be absorbed into academia. > 🔗 **Borrowed.** Bundle techniques from satellite geolocation entered the computer vision community in the 1990s. While military classification kept those methods out of reach, academia independently rediscovered the same problem. Triggs 1999 is where the two streams met. --- ## 3. Levenberg and Marquardt — pioneers of nonlinear optimization Brown supplied the objective function; the tool that solved it came from somewhere else. Reprojection error minimization is a nonlinear least-squares problem. There is no analytical solution, so iterative numerical optimization is needed. In 1944, [Kenneth Levenberg](https://cs.uwaterloo.ca/~y328yu/classics/levenberg.pdf) published a method that interpolated between Gauss-Newton and steepest descent with a damping parameter $\lambda$. Larger $\lambda$ moves toward steepest descent for safe convergence; smaller $\lambda$ uses the fast convergence of Gauss-Newton. In its basic form, the strategy adds $\lambda \mathbf{I}$ to the matrix $\mathbf{J}^{\top}\mathbf{J}$ of the linearized normal equations, improving numerical stability. It was twenty years ahead of computer vision. In 1963, [Donald Marquardt](https://epubs.siam.org/doi/10.1137/0111030) independently rediscovered the same idea and formulated it more explicitly. The name settled as the **Levenberg-Marquardt (LM) algorithm**. Another thirty-five years passed before the LM algorithm became the standard BA solver in computer vision. The delay came from the walls between fields, not from missing technology. --- ## 4. 1999, Triggs et al. — a hundred years of inheritance integrated Triggs and colleagues brought this numerical tool and the computational structure of BA together in a systematic synthesis for computer vision readers. At the 1999 Vision Algorithms Workshop, Bill Triggs, Philip McLauchlan, Richard Hartley, and Andrew Fitzgibbon presented ["Bundle Adjustment — A Modern Synthesis"](https://link.springer.com/chapter/10.1007/3-540-44480-7_21). The paper translated and synthesized BA theory scattered across twentieth-century surveying and aerial photogrammetry into the language of computer vision. Triggs et al. contributed two things. First, they made the structural properties of sparse BA explicit. Using the sparse block structure of the Hessian matrix (the Schur complement trick), joint camera-point optimization can be performed far more efficiently. Second, they treated gauge freedom (the arbitrariness of the reference frame) explicitly. Seven years after this paper, Noah Snavely's [Photo Tourism (2006)](https://phototour.cs.washington.edu/Photo_Tourism.pdf) automatically reconstructed famous landmarks such as Notre-Dame and the Trevi Fountain from hundreds of photographs scattered across the Internet. Ten years later, Johannes Schönberger's [COLMAP (2016)](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf) released an open-source implementation of robust incremental structure from motion (SfM) for tens to hundreds of thousands of images. It turned a research line that had already reached the million-image range into a tool others could reproduce. Without Triggs's language, that path would have been much slower. --- ## 5. Reprojection Error — the formation of the concept Triggs et al. described the error function being minimized. The function itself had a separate history. There were two transitions before this error function took its present shape. Early twentieth-century aerial triangulators measured error as "distance difference in the ground coordinate frame." Because they compared points directly in 3D space, a misaligned camera lens or poor calibration could be absorbed into the ground-coordinate residual and disappear from view. Brown moved the comparison to the image plane in his 1958 report. He matched the projected location of a 3D point to the actual image observation in pixel units. Calibration error, lens distortion, and extrinsic-parameter error then surface together in one residual. The formulation is also cleaner statistically. Camera image noise can be modeled as an isotropic Gaussian in pixel units, and under that model reprojection-error minimization becomes maximum-likelihood estimation. Triggs et al. (1999) standardized that formulation in the language of computer vision textbooks. As of 2026, reprojection-error minimization remains a core optimization problem for visual observations in factor graph-based SLAM backends. > 🔗 **Borrowed.** The observation model for a visual landmark in SLAM, $z = \pi(K, T, p) + \epsilon$, directly inherits Brown's (1958) reprojection formula. A SLAM backend that minimizes this with Gauss-Newton has the same mathematical structure as a 1958 aerial triangulation solver. --- ## 6. The skeleton of the SLAM backend — through 2026 The acronym "SLAM" came into broad use through 1990s literature that included [Durrant-Whyte and Leonard's 1995 survey](https://ieeexplore.ieee.org/document/476131), but the backend mathematics inherits Brown's 1958 reprojection formulation almost unchanged. Modern systems still show that inheritance. ORB-SLAM3 jointly optimizes SE(3) poses and 3D landmark locations through g2o. LIO-SAM runs the LM algorithm on top of GTSAM's factor graph. DROID-SLAM gets its update direction from GRU-based optical flow, but its final bundle adjustment layer still uses the Schur complement trick. Lie groups and factor graphs replaced the matrix notation of 1999, and neural networks took over descriptor computation, but the common structure of minimizing observation residuals persists. Visual BA uses reprojection errors from multiple viewpoints to estimate camera poses and the map, while LiDAR systems use sensor-specific constraints such as distance or point-to-plane residuals. Pulfrich's glass plate has become a pixel array, and hand calculation has moved to the GPU. The underlying estimation problem remains the same. This continuity is both a strength and a weakness. The field inherits a hundred years of convergence proofs and practical validation. When BA's assumptions (static world, point features, Gaussian noise) fail in real environments, however, there is no ready alternative. --- > 📜 **Prediction vs. outcome.** Triggs et al. (1999) named scaling BA to large problems (thousands of cameras, millions of points) as the main challenge. Over the next twenty years, systems reached that scale. In 2006, Snavely's Photo Tourism reconstructed landmarks from hundreds of Internet photographs; in 2016, COLMAP standardized the robust incremental SfM implementation of that line. Later scaling required more than simply enlarging the optimization problem. The result depended on an engineering layer: incremental BA and visibility graph pruning, with vocabulary-tree loop closure on top. --- ## 🧭 Still open **Global optimum guarantees for nonlinear BA.** The LM algorithm converges to a local minimum. With poor initialization, it converges to the wrong structure. Initialization methods such as the 5-point algorithm, PnP, and epipolar-geometry estimation appeared in turn, but the geometric solvers should be distinguished from the robust estimation procedures around them. RANSAC is used as an outer procedure to reject outliers, and iterative refinement can be applied to the resulting estimates. Convex-relaxation approaches that guarantee a global optimum in large-scale environments remain under study, but they are not yet practical at the speed and scale of real-time SLAM. **The gap between photogrammetric accuracy and Visual SLAM.** Accuracy in aerial photogrammetry is assessed against the imaging conditions and the requirements of the delivered product. It uses calibrated cameras and high-quality GCPs (ground control points), and the optimization runs offline. Real-time Visual SLAM uses the same formulation under the constraints of GPS-denied environments, low-resolution cameras, and online estimation. Comparing the two fields requires distinguishing image-coordinate errors from ground-coordinate errors and stating the range, resolution, control-point conditions, and components being evaluated. --- BA's assumptions (static world, point features, Gaussian noise) begin to break when the camera meets a moving object. Surveyors measured bridges, not robot soccer fields. The inheritance passed to computer vision in the form of a single question: which pixels in a moving image should serve as the "corresponding points" that bundle adjustment would later receive? --- # Ch.2 — The Classical CV Toolbox: Harris to SIFT, and on to ORB Bundle adjustment requires "corresponding points," the same physical location found independently in two or more images. The surveyor planted targets in the field by hand; computer vision had to hand that role to an algorithm. Feature detection and description began with that handoff. In the late 1970s, Hans Moravec tried to locate salient points in the environment with a camera on the Stanford Cart project. The work was written up in his 1980 Stanford doctoral thesis, ["Obstacle Avoidance and Navigation in the Real World by a Seeing Robot Rover"](https://frc.ri.cmu.edu/~hpm/project.archive/robot.papers/1975.cart/1980.html.thesis/index.html). Moravec supplied a quantitative criterion for selecting trackable points from intensity changes in neighboring patches. In 1988, Chris Harris and Mike Stephens formalized that intuition in terms of the eigenvalues of the autocorrelation matrix. Lucas and Kanade had laid down the framework for pixel tracking seven years earlier. Lowe absorbed both ideas and built a descriptor invariant to scale and rotation. Rublee produced a faster, patent-free alternative. The SLAM front end runs on this lineage. --- ## 2.1 The idea of a corner: from Moravec to Harris A point where an image patch changes substantially under a small camera motion is called a "corner." Moravec's (1977) criterion was simple. If the Sum of Squared Differences (SSD) against neighboring pixels is large in every direction (up, down, left, and right), the point counts as a corner. Harris and Stephens replaced this with continuous differentiation at the 1988 Alvey Vision Conference in ["A Combined Corner and Edge Detector"](https://www.bmva.org/bmvc/1988/avc-88-023.html). For image $I$, shifting a window $W$ around point $(x,y)$ by $(\Delta x, \Delta y)$ and approximating the intensity change gives: $$M = \sum_{(x,y) \in W} \begin{pmatrix} I_x^2 & I_x I_y \\ I_x I_y & I_y^2 \end{pmatrix}$$ The two eigenvalues $\lambda_1, \lambda_2$ of $M$ classify the point: two large eigenvalues indicate a corner, one large eigenvalue indicates an edge, and two small eigenvalues indicate a flat region. Harris avoided the eigenvalue decomposition altogether by using the score $R = \det(M) - k \cdot \text{tr}(M)^2$. $k$ is typically 0.04–0.06. > 🔗 **Borrowed.** Harris's (1988) autocorrelation-matrix idea refined Moravec's (1977) SSD-based corner search through continuous differentiation. The prototype of the concept was in the Stanford Cart report. In 1994, Jianbo Shi and Carlo Tomasi showed in ["Good Features to Track"](https://cecas.clemson.edu/~stb/klt/shi-tomasi-good-features-cvpr1994.pdf) (CVPR 1994) that using $\min(\lambda_1, \lambda_2)$ directly, in place of the Harris score, is more stable for optical-flow tracking. This criterion is the Shi-Tomasi corner detector. Thirty years later, OpenCV still exposes it under the same `goodFeaturesToTrack` function name. --- ## 2.2 The archetype of tracking: Lucas-Kanade and KLT Harris's matrix $M$ finds the point. Finding the same point again in the next frame is a separate problem. Bruce Lucas and Takeo Kanade, in the 1981 paper ["An Iterative Image Registration Technique"](https://www.ijcai.org/Proceedings/81-2/Papers/017.pdf), formulated inter-frame pixel motion as a minimization problem under the brightness constancy assumption. The brightness-constancy assumption states that the intensity of pixel $(x,y)$ is the same before and after the motion. $$I(x, y, t) = I(x + u, y + v, t + 1)$$ A Taylor expansion followed by linearization gives: $$I_x u + I_y v + I_t = 0$$ One equation, two unknowns. Lucas-Kanade adds the assumption that pixels inside a $3\times3$ or $5\times5$ window move with the same $(u,v)$, producing an overdetermined system solved by least squares. $$\begin{pmatrix} \sum I_x^2 & \sum I_x I_y \\ \sum I_x I_y & \sum I_y^2 \end{pmatrix} \begin{pmatrix} u \\ v \end{pmatrix} = -\begin{pmatrix} \sum I_x I_t \\ \sum I_y I_t \end{pmatrix}$$ The matrix on the left is the same structure matrix $M$ as Harris's. Corner detection and optical flow rest on the same mathematics. Tomasi and Kanade, in the 1991 tech report ["Detection and Tracking of Point Features"](https://cecas.clemson.edu/~stb/klt/tomasi-kanade-techreport-1991.pdf), gave a concrete implementation that selects tracking-window quality by the eigenvalue criterion and refines displacement through Newton-Raphson iteration. Bouguet (Intel, 2000) later added an image-pyramid-based coarse-to-fine strategy so the tracker would converge under large motion, and this combination became the KLT (Kanade-Lucas-Tomasi) tracker. Real-time VIO systems such as [VINS-Mono](https://arxiv.org/abs/1708.03852) (2018) still run a front end descended from this work. A least-squares tracker from 1981 runs inside the VIO of a smartphone-class drone more than forty years later. > 🔗 **Borrowed.** Lucas-Kanade (1981) → KLT tracker → Qin et al.'s VINS-Mono (2018): optical flow proposed 37 years earlier survives unchanged as the feature-tracking backbone of real-time VIO. --- ## 2.3 SIFT — invariance and the patent KLT fits the case of a single camera moving a little at a time. Connecting the same point across images taken by different cameras, on different days, is a different problem altogether. A change in viewpoint alters the patch's shape, size, and orientation for the same point, and a plain pixel comparison no longer works. That is why a **descriptor** is needed. David Lowe (UBC) presented the idea at ICCV 1999. The talk was titled "Object Recognition from Local Scale-Invariant Features," and the demo compared 128-dimensional vectors to match the same object across different photographs. Five years later, Lowe published the full account, ["Distinctive Image Features from Scale-Invariant Keypoints"](https://www.cs.ubc.ca/~lowe/papers/ijcv04.pdf), in IJCV. It is the paper now cited as SIFT. SIFT (Scale-Invariant Feature Transform) runs in two stages. **Detection stage.** Compute DoG (Difference of Gaussians) at several scales and select local extrema as keypoints. DoG is an approximation of the Laplacian of Gaussian. Let $L(x,y,\sigma) = G(x,y,\sigma) * I(x,y)$ denote the Gaussian-smoothed image: $$D(x, y, \sigma) = L(x, y, k\sigma) - L(x, y, \sigma)$$ Here $k$ is the ratio between adjacent scales (typically $2^{1/s}$, where $s$ is the number of scales per octave). Searching for extrema across multiple octaves allows the same point to be detected under a change of scale. **Descriptor stage.** A $16\times16$ window around the keypoint is divided into $4\times4$ blocks, and the 8-bin gradient-orientation histogram in each block is concatenated into a 128-dimensional vector. Rotating the patch into the keypoint's dominant gradient direction also provides rotation invariance. The result was a 128-dimensional descriptor robust to scale, rotation, and partial affine deformation. Before KITTI and standardized SLAM benchmarks, that robustness made SIFT useful for matching images across changes in scale and viewpoint. Lowe filed a patent on SIFT in March 2000, and it was granted in March 2004 (US6711293B1, with priority from March 1999). The patent imposed licensing fees for commercial use, and until it expired in March 2020 it was one of the motivations for efforts to replace SIFT. > 📜 **Prediction vs. outcome.** In "9 Conclusions" of the 2004 SIFT paper, Lowe listed the descriptor's possible extensions as "view matching for 3D reconstruction, motion tracking and segmentation, robot localization, image panorama assembly, epipolar calibration." Most of those predictions proved accurate: SfM, SLAM, panoramas, and early vision-based robot localization leaned on SIFT in the late 2000s. SIFT's position in long-term correspondence weakened after CNNs arrived. Following AlexNet in 2012, object-recognition work shifted to CNNs, and learned descriptors such as SuperPoint and R2D2 gradually took over the local-descriptor role in SLAM. The application domains were predicted correctly; the descriptor technology changed. --- ## 2.4 SURF — a speed–accuracy compromise SIFT's 128-dimensional descriptor was accurate but slow, taking hundreds of milliseconds per image on the desktop CPUs of the time. It was not usable for real-time SLAM. Herbert Bay (ETH Zürich) presented ["SURF: Speeded-Up Robust Features"](https://people.ee.ethz.ch/~surf/eccv06.pdf) at ECCV 2006. The method relied on two ideas. SURF detects keypoints with the *determinant of the Hessian matrix* instead of DoG. It approximates the second Gaussian derivatives with box filters on an integral image to speed up computation. The descriptor is 64-dimensional, half of SIFT's. The neighborhood of the keypoint is split into $4\times4$ subregions, and in each subregion four values from Haar wavelet responses $d_x, d_y$, $(\sum d_x,\, \sum d_y,\, \sum|d_x|,\, \sum|d_y|)$, are concatenated into a $4\times4\times4=64$-dimensional vector. A 128-dimensional extension (SURF-128) exists, but the default is 64-dimensional. SURF was 3–7 times faster than SIFT. But accuracy comparisons depended on the detectors, descriptor designs, and evaluation conditions as well as dimensionality, and Bay could not avoid a patent either (ETH Zürich patent). SIFT lost ground on speed; SURF lost ground on both accuracy and patent restrictions. ORB addressed both problems at once. > 🔗 **Borrowed.** Lowe's (1999/2004) DoG scale-space → Bay's (2006) Hessian integral image: two answers for achieving scale invariance. DoG is theoretically elegant; the Hessian approximation is engineered to be fast. --- ## 2.5 ORB — binary descriptor and release from the patent In 2011, Ethan Rublee (Willow Garage), Vincent Rabaud, Kurt Konolige, and Gary Bradski presented ["ORB: An Efficient Alternative to SIFT or SURF"](https://www.gwylab.com/download/ORB_2012.pdf) at ICCV. Willow Garage was also the birthplace of ROS, and the title reflected the goal: a feature that robotics researchers could use in practice. ORB combines and improves two existing techniques. **Detection.** [FAST](https://www.edwardrosten.com/work/rosten_2006_machine.pdf) (Features from Accelerated Segment Test, Rosten & Drummond 2006) tests a 16-pixel circle around a candidate and declares it a corner if a contiguous arc is sufficiently brighter or darker. It is more than 10 times faster than SIFT's DoG. ORB adds a Harris score on top of FAST and keeps only the strong responses. **Descriptor.** [BRIEF](https://www.cs.ubc.ca/~lowe/525/papers/calonder_eccv10.pdf) (Binary Robust Independent Elementary Features, Calonder et al. 2010) compares the intensities of randomly chosen point pairs in the patch around a keypoint to produce a 256-bit string by default. Matching uses Hamming distance instead of Euclidean distance, so the distance is computed by XOR followed by a population count of the differing bits. BRIEF's weak point was the lack of rotation invariance. Rublee built **rBRIEF (rotated BRIEF)** by rotating the patch to align with the FAST corner's intensity centroid. This supplied the missing orientation invariance. $$\theta = \text{atan2}(m_{01},\, m_{10}), \quad m_{pq} = \sum_{x,y} x^p y^q I(x,y)$$ ORB ran 100 times faster than SIFT, carried no patent restriction, and entered OpenCV immediately. [ORB-SLAM](https://arxiv.org/abs/1502.00956) (Mur-Artal et al. 2015), as the name indicates, was built on ORB, and the line continued through the trilogy. ORB-SLAM3 still used the same front end in 2021. > 🔗 **Borrowed.** Calonder et al.'s (2010) BRIEF → Rublee et al.'s (2011) ORB: adding intensity-centroid-based orientation estimation to a binary descriptor secured rotation invariance. --- ## 2.6 Learned descriptors If ORB is the practical peak, the next question follows: are learned features better than hand-designed ones? Yi et al.'s 2016 [LIFT](https://arxiv.org/abs/1603.09114) (Learned Invariant Feature Transform, ECCV 2016) tried to replace the three stages of detection, orientation estimation, and description with CNNs. It connected three separately trained networks in a pipeline. In 2018, DeTone et al.'s [SuperPoint](https://arxiv.org/abs/1712.07629) (CVPRW 2018) trained keypoint detection and a 256-dimensional descriptor jointly under a self-supervised scheme called homographic adaptation. It was pretrained on synthetic data and adapted to real images, and later became one of the learned local features widely tested in SLAM research. Even so, traditional descriptors have not disappeared as of 2026. ORB is faster than SuperPoint on embedded devices, and its behavior is more predictable than that of learned descriptors, which can generalize unpredictably on out-of-domain images. DINOv2-based features have entered place recognition through systems such as [AnyLoc](https://arxiv.org/abs/2308.00688) (Keetha et al. 2023), but ORB-SLAM3 has used ORB since its 2021 release. Moravec's 1977 intuition still runs on robots in the 2020s. --- ## 2.7 🧭 Still open **Generalization limits of learned descriptors.** SuperPoint, R2D2, DISK, and others outperform classical methods inside the training domain, but behave inconsistently in new environments (underwater, thermal, low-light). There is no consensus on which family is more reliable. The question remains open in 2026. **Failure modes of wide-baseline matching.** Harris- or ORB-based matching can degrade under large viewpoint changes; the extent depends on the scene, rotation axis, and matching conditions. Affine-covariant detectors (ASIFT, MSER) patched part of the gap, but there is no complete solution. [DUSt3R](https://arxiv.org/abs/2312.14132) (Wang et al. 2023) opened a path by bypassing matching itself, though it is still too early to judge whether this is the end of the descriptor problem or a detour around it. --- Harris's intuition and Lowe's invariance supplied the foundation, and Rublee's speed optimization made it practical. Together they formed a usable toolbox. Each technique was designed to operate on one or two images. Connecting dozens or hundreds of images simultaneously and with geometric consistency required another layer. --- *References* - Harris, C. & Stephens, M. (1988). A Combined Corner and Edge Detector. *Proc. Alvey Vision Conference*. - Lucas, B. D. & Kanade, T. (1981). An Iterative Image Registration Technique with an Application to Stereo Vision. *IJCAI*. - Shi, J. & Tomasi, C. (1994). Good Features to Track. *CVPR*. - [Lowe, D. G. (2004). Distinctive Image Features from Scale-Invariant Keypoints.](https://doi.org/10.1023/B:VISI.0000029664.99615.94) *IJCV 60(2)*. - Bay, H., Tuytelaars, T. & Van Gool, L. (2006). SURF: Speeded-Up Robust Features. *ECCV*. - Calonder, M. et al. (2010). BRIEF: Binary Robust Independent Elementary Features. *ECCV*. - [Rublee, E. et al. (2011). ORB: An Efficient Alternative to SIFT or SURF.](https://doi.org/10.1109/ICCV.2011.6126544) *ICCV*. - DeTone, D., Malisiewicz, T. & Rabinovich, A. (2018). SuperPoint: Self-Supervised Interest Point Detection and Description. *CVPRW*. [arXiv:1712.07629](https://arxiv.org/abs/1712.07629) --- # Ch.3 — Structure from Motion: From Longuet-Higgins to COLMAP While Harris and Lowe were refining how to detect salient points in one image, a different lineage asked what could be recovered when those points appeared in two images. Feature *detection* and spatial *reconstruction* developed side by side over the same period; only in the mid-2000s did they merge into one pipeline. In 1981, H.C. Longuet-Higgins, a theoretical psychologist at Cambridge, published a three-page paper in *Nature*. The title was "[A Computer Algorithm for Reconstructing a Scene from Two Projections](https://cseweb.ucsd.edu/classes/fa01/cse291/hclh/SceneReconstruction.pdf)." He showed that eight coordinate pairs for the same points in two photographs were enough to solve simultaneously for camera motion and the scene's three-dimensional shape. He was neither a roboticist nor a computer vision researcher. Structure from Motion (SfM) began in those three pages, and the mathematics reached a widely used engineering system in 2016, when Johannes Schönberger released COLMAP. --- ## 3.1 Essential Matrix and the 8-point Algorithm Longuet-Higgins began with a simple constraint. When two cameras capture the same point, an algebraic relation holds between the two image coordinates. Once the coordinate system is normalized, the relation can be written as a single matrix. He defined it as the **essential matrix** $\mathbf{E}$. Let the two camera centers be $\mathbf{O}_1$ and $\mathbf{O}_2$, and let the corresponding points in normalized coordinates be $\mathbf{x}_1$ and $\mathbf{x}_2$. The constraint is: $$\mathbf{x}_2^\top \mathbf{E} \mathbf{x}_1 = 0$$ $\mathbf{E}$ factors through the rotation $\mathbf{R}$ and translation $\mathbf{t}$ between the cameras as $\mathbf{E} = [\mathbf{t}]_\times \mathbf{R}$, where $[\mathbf{t}]_\times$ is the skew-symmetric matrix of $\mathbf{t}$. Once scale ambiguity is removed, the essential matrix has five degrees of freedom. Before the nonlinear 5-point algorithm ([Nistér 2004](http://www.cad.zju.edu.cn/home/gfzhang/training/SFM/2004-PAMI-David%20Nister-An%20Efficient%20Solution%20to%20the%20Five-Point%20Relative%20Pose%20Problem.pdf)) solved it with five correspondences, the standard approach fixed one of the nine matrix entries as unit scale, treated the remaining eight as unknowns, and solved a linear system from eight correspondences. The rank-2 and unit-scale constraints were then enforced. This is the **8-point algorithm**. Longuet-Higgins himself gave a procedure that produced a unique solution from exactly eight points. The implementation was simple and computationally inexpensive. Numerical stability was the problem. When image coordinates run into the hundreds or thousands of pixels, the entries of the coefficient matrix span sharply different scales, and the SVD becomes unstable. > 🔗 **Borrowed.** Hartley's 1997 normalized 8-point algorithm ([In Defense of the Eight-Point Algorithm](https://www.cse.unr.edu/~bebis/CS485/Handouts/hartley.pdf)) applied a linear transform to image coordinates so that their mean was zero and their average distance was $\sqrt{2}$, and then estimated the fundamental matrix before transforming it back to the original coordinates. This serves a different purpose from normalizing coordinates with camera intrinsics. The geometry of Longuet-Higgins was left untouched; only the numerical conditioning was fixed. The normalization was widely adopted in later multiple-view-geometry texts and implementations. The fundamental matrix $\mathbf{F}$ generalizes the essential matrix. Even without knowing the camera intrinsics $\mathbf{K}$, the relation $\mathbf{x}_2^\top \mathbf{F} \mathbf{x}_1 = 0$ holds. With intrinsics $\mathbf{K}_1$, $\mathbf{K}_2$ for the two cameras, the relationship is $\mathbf{F} = \mathbf{K}_2^{-\top} \mathbf{E} \mathbf{K}_1^{-1}$. For images from the same camera ($\mathbf{K}_1 = \mathbf{K}_2 = \mathbf{K}$) it simplifies to $\mathbf{F} = \mathbf{K}^{-\top} \mathbf{E} \mathbf{K}^{-1}$. In an SfM pipeline, when $\mathbf{K}$ is unknown $\mathbf{F}$ is estimated first; when $\mathbf{K}$ is known, $\mathbf{E}$ is solved directly. --- ## 3.2 Tomasi-Kanade Factorization For ten years after 1981, SfM was studied mostly as the geometry between two photographs. Processing many photographs at once was a separate problem. Carlo Tomasi and Takeo Kanade at CMU provided one answer in 1992 with the **[factorization method](https://people.eecs.berkeley.edu/~yang/courses/cs294-6/papers/TomasiC_Shape%20and%20motion%20from%20image%20streams%20under%20orthography.pdf)**. Given $F$ frames observing $P$ points, the image coordinates stack into a $2F \times P$ matrix $\mathbf{W}$. After subtracting each frame's point centroid to remove translation, $\mathbf{W}$ has rank at most three under an orthographic (scaled orthographic) camera model. The original paper (Tomasi & Kanade 1992) started from this assumption. Then: $$\mathbf{W} = \mathbf{M} \mathbf{S}$$ where $\mathbf{M}$ is a $2F \times 3$ motion matrix and $\mathbf{S}$ is a $3 \times P$ structure matrix. Keeping only the top three singular values of $\mathbf{W}$ through SVD gives a rank-3 factorization. Because $\mathbf{M}\mathbf{A}$ and $\mathbf{A}^{-1}\mathbf{S}$ produce the same $\mathbf{W}$, a metric-upgrade step then determines the invertible matrix $\mathbf{A}$ from orthogonality and equal-norm constraints on the camera rows. The procedure estimated every frame's motion and every point's 3D position through one low-rank SVD followed by the metric upgrade. Its cost is dominated by the SVD of a $2F \times P$ matrix, and its structure is simpler to implement than iterative nonlinear bundle adjustment. > 🔗 **Borrowed.** Later literature places Nistér, Naroditsky, and Bergen's 2004 CVPR paper "Visual Odometry" at the point where real-time ego-motion estimation became an applied branch of this lineage. Instead of using Tomasi-Kanade's batch factorization directly, the work solved relative pose between frames inside a short window, trading batch accuracy for lower latency. The limitation was the orthographic/affine assumption. An affine camera ignores perspective distortion. The model holds only when depth variation in the scene is small relative to the distance from the camera, as with small, distant objects. Error grew for close scenes, wide-angle lenses, and scenes with large foreground–background depth differences. From the late 1990s, researchers pursued perspective-camera extensions from several directions, leading back to bundle adjustment. --- ## 3.3 Hartley & Zisserman and the canonization Tomasi-Kanade's factorization framed the multiple-view problem. The remaining tasks were to extend it to perspective cameras and unify the scattered mathematics in one language. In 2000, Richard Hartley and Andrew Zisserman's 680-page textbook *[Multiple View Geometry in Computer Vision](https://www.robots.ox.ac.uk/~vgg/hzbook/)* appeared. It consolidated the SfM mathematics scattered from 1981 through the 1990s into the language of projective geometry. Hartley & Zisserman did more than compile earlier results. They brought the essential matrix, fundamental matrix, homography, camera calibration, and bundle adjustment into one projective-geometry framework. Their treatment showed within one text that concepts developed separately shared the same foundation. Bundle adjustment received particular attention. Hartley & Zisserman placed the reprojection-error minimization problem reviewed in Ch.1 through the synthesis by Triggs et al. (1999) inside the projective-geometry framework and included an explicit *robust cost function* $\rho$. Huber or Cauchy losses downweighted outliers that would otherwise break optimization on real data. The solver was Levenberg-Marquardt, and the sparsity of the Jacobian reduced computation. Most SLAM and visual odometry (VO) papers in the early 2000s cited this textbook as their standard reference. With the definitions unified in one source, large-scale applications such as Photo Tourism could focus on implementation rather than redefining the basics. --- ## 3.4 Photo Tourism and Bundler — Internet-scale SfM In 2006, Noah Snavely, Steven Seitz, and Richard Szeliski published the SIGGRAPH paper "[Photo Tourism](https://doi.org/10.1145/1179352.1141964)." They gathered photographs of tourist sites uploaded to the internet (the Florence Duomo, the Trevi Fountain in Rome) and reconstructed them in 3D. The data were uncontrolled. Cameras, weather, and composition varied, and some images were unrelated indoor shots. This was not a systematically captured dataset but thousands of images uploaded in no particular order by thousands of people. Snavely's pipeline first used SIFT detection and matching to find correspondences between image pairs. RANSAC with the fundamental matrix removed geometrically inconsistent matches. Incremental SfM started from highly connected image pairs and added cameras one at a time. After each addition, bundle adjustment reoptimized the full set of poses and points. The datasets reported in the paper included the Notre Dame Cathedral (597 registered out of 2,635 candidates), the Trevi Fountain in Rome (360 out of 466), Yosemite Half Dome (325 out of 1,882), the Great Wall (82 out of 120), and Trafalgar Square (278 out of 1,893), with an average reprojection error of about 1.5 pixels on 1,611×1,128 images. The work became a turning point in assembling hundreds of uncontrolled internet photographs into consistent reconstructions. Bundler implemented this pipeline. Snavely released it as open source, and it became the default starting point for SfM researchers. --- ## 3.5 COLMAP — engineering maturity > 📜 **Prediction vs. outcome.** The "Discussion and future work" section of Snavely et al. 2006 states, "Ultimately, we wish to scale up our reconstruction algorithm to handle millions of photographs," and lists better image-registration ordering, lens-distortion modeling, repeated-structure handling, and disconnected-structure reconstruction as remaining problems. COLMAP (Schönberger 2016) and OpenSfM pursued scale, reaching tens to hundreds of thousands of images. The SLAM lineage addressed real-time and online processing separately through fixed-lag smoothers and loop closure, not through incremental refinement. This was progress toward scale, although the sizes cited here do not establish that the goal of millions of photographs was met. In 2016, Johannes Schönberger and Jan-Michael Frahm published the CVPR paper "[Structure-from-Motion Revisited](https://openaccess.thecvf.com/content_cvpr_2016/papers/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.pdf)." Despite the modest title, the paper systematically redesigned the pipeline around ten years of improvements since Bundler. COLMAP differed from Bundler most in three respects. The first change was the order in which cameras were added. Bundler started from pairs with high connectivity but had no systematic criterion for which pair to extend first. COLMAP automated the choice of initial image pair and camera-registration order using triangulation angle, feature-track length, and visibility score. Reconstruction became substantially more stable. The second change was the bundle-adjustment cadence. Running a full bundle adjustment after every camera addition is expensive. COLMAP alternated local bundle adjustment (optimizing only the recently added camera together with cameras that shared many points with it) with periodic global bundle adjustment. The third change was geometric verification. For each pair of matched feature points, COLMAP ran RANSAC with two models in parallel: the fundamental matrix and homography. The fundamental matrix covered general non-planar scenes; the homography covered planar scenes or pure rotation. COLMAP compared the two models' inlier counts to classify the scene and filtered out matches that fit neither. It was more robust than Bundler to poor matches and planar degeneracy. > 🔗 **Borrowed.** COLMAP's incremental bundle adjustment strategy modularized Snavely's Bundler pipeline and added quality control at each stage. The core mathematics of the algorithm (essential matrix estimation, triangulation, Levenberg-Marquardt) came from the Hartley & Zisserman textbook. COLMAP's contribution was the systematization of engineering judgment rather than new mathematics. COLMAP became widely used for reasons beyond performance. The codebase was well organized, the documentation was adequate, and CUDA acceleration supported large image collections. After NeRF appeared in 2020, many NeRF training codebases took COLMAP's output (camera poses + sparse point cloud) as input, and the reference 3D Gaussian Splatting implementation used the same preprocessing. COLMAP became a common front end for 3D reconstruction research. --- ## 3.6 The split between SfM and SLAM SfM and SLAM use the same mathematics yet solve different problems. The distinction came into sharp relief in the early 2000s. SfM is *offline*. All images are gathered before processing, so there is no time constraint, and global bundle adjustment can be run multiple times with full access to the whole dataset. If a camera pose is wrong, the system can go back and recompute it. SLAM is *online*. Sensor data streams in real time, and the robot's current position must be available immediately. The system cannot retain and revisit past data indefinitely, computation grows with the map, and accumulated drift must be corrected when the robot returns to a previously visited place. The fields differ in the constraints of online processing. SLAM must detect revisits during motion and correct accumulated drift, potentially across many poses connected by the loop. SfM and SLAM share tools such as bundle adjustment and image retrieval, but differ in when and which states must be updated. Uncertainty propagation differed as well. SLAM tracks the uncertainty of the current pose in real time and updates it with each new observation. A probabilistic representation in the form of EKF or factor graph is needed. In SfM, covariance can be computed after optimization finishes, and real-time tracking is not required. Davison's [MonoSLAM (2003)](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf) called itself "real-time SfM." But its structure, which kept camera pose and landmarks together in an EKF state vector, differed from SfM's global batch optimization. Over the 2000s, the two fields split into independent lineages, each with its own problem setting. --- ## 3.7 🧭 Still open **SfM with dynamic objects.** Mainstream general-purpose SfM pipelines, including COLMAP, assume a static world. Bundle adjustment is solved on the premise that scene points are stationary, so contaminated matches distort optimization around cars or pedestrians. RANSAC filters some of them, while Dynamic SfM research models segmentation or per-object motion. Among these public implementations, none has settled into a general-purpose tool with COLMAP's scope. **The blurring boundary between SfM and SLAM.** In 2023, [DUSt3R](https://arxiv.org/abs/2312.14132) (Wang et al.) took two images into a single pretrained network and produced a dense point map and camera poses at once. It needed no feature matching, RANSAC, or bundle-adjustment initialization. Extended as [MASt3R](https://arxiv.org/abs/2406.09756) (2024), it handled tens of images. Modules of the traditional SfM pipeline are now being replaced one at a time. COLMAP became the front end for NeRF and 3DGS; the DUSt3R line is trying to replace that front end. Whether it will displace COLMAP or prevail only in specific domains remains unknown. --- While SfM refined precise offline reconstruction, another question grew urgent in robotics. A moving robot had to estimate its pose immediately from images that had not yet been gathered. Under that pressure, the question Randall Smith and Peter Cheeseman [posed in 1986](https://people.csail.mit.edu/brooks/idocs/Smith_Cheeseman.pdf), how to propagate uncertain spatial relations, grew into a separate field called SLAM. --- # Ch.4 — Smith-Cheeseman and the Rise and Fall of EKF-SLAM The photogrammetry, SfM, and bundle adjustment of Part I shared one assumption: either the camera remained still or there was time after capture to process every image offline as a batch. Hartley-Zisserman's geometry, RANSAC's robust estimation, and Levenberg-Marquardt's iterative optimization could measure the world, but they did not ask where a moving robot was *right now*. Part II begins with that question. The robot must build a map, locate itself, and maintain the estimate as uncertainty accumulates. Probabilistic mapping began with a small memo from SRI International. In 1986 Randall Smith and Peter Cheeseman set out to formalize uncertainty in a robot's spatial measurements. Their work at SRI International inherited Kalman's (1960) filter mathematics but extended it from a single state estimate to an entire *network of spatial relationships*. Several years later, Hugh Durrant-Whyte in Sydney and John Leonard at MIT coupled that mathematics to a new problem statement: a robot estimates its own position while building a map. The acronym "SLAM" emerged from that union. --- ## 4.1 The mathematics of uncertain spatial relationships — Smith, Self, Cheeseman (1988) In 1986 Randall Smith, Matthew Self, and Peter Cheeseman at SRI International set out to express how error propagates when a robot accumulates measurements across several places. Their work was presented at UAI 1986 and published in 1988 as ["Estimating Uncertain Spatial Relationships in Robotics"](https://arxiv.org/abs/1304.3111). The question was clear: when a robot measures B from A and then C from B, how is the uncertainty from A to C computed? The [Kalman filter](https://www.cs.unc.edu/~welch/kalman/kalmanPaper.html) already existed. It had been used since 1960 for radar tracking, ballistic calculation, and satellite-orbit correction. Smith, Self, and Cheeseman reformulated Kalman's covariance-propagation equations for the composition of spatial transforms. They placed the robot pose $\mathbf{x}_r$ and landmark positions $\mathbf{m}_i$ in one state vector and maintained the joint covariance $\mathbf{P}$ over the entire state. $$\mathbf{x} = [\mathbf{x}_r^\top,\ \mathbf{m}_1^\top,\ \ldots,\ \mathbf{m}_N^\top]^\top$$ $$\mathbf{P} = \begin{bmatrix} \mathbf{P}_{rr} & \mathbf{P}_{rm} \\ \mathbf{P}_{mr} & \mathbf{P}_{mm} \end{bmatrix}$$ The off-diagonal block $\mathbf{P}_{rm}$ records the crucial relation: robot-pose uncertainty and landmark-position uncertainty are *correlated*. Only by tracking that correlation can the estimate remain consistent. The paper demonstrated this explicitly, establishing a starting point for the SLAM field. > 🔗 **Borrowed.** Smith-Self-Cheeseman's (1988) spatial-relationship mathematics inherits directly from Kalman's (1960) covariance propagation. A technique for tracking a single moving object became a framework for tracking a robot and every element of its map at once. --- ## 4.2 How the name "SLAM" settled in There is no "SLAM" in the 1988 Smith-Self-Cheeseman paper. In the early 1990s, Hugh Durrant-Whyte, who had moved from Oxford to Sydney, and John Leonard at MIT used different names for the same problem in their respective labs. Once the groups began citing each other, they needed a shared term, and "SLAM" gradually became standard. Researchers' memories differ on which document used it first. No canonical first-use paper exists. Leonard and Durrant-Whyte's 1991 paper, ["Simultaneous Map Building and Localization for an Autonomous Mobile Robot"](https://doi.org/10.1109/IROS.1991.174711), is often cited as an early mainstream robotics paper to state the problem in its title. The title captured the intuition before the acronym existed: mapping and localization are inseparable and must be performed simultaneously. The name was "Simultaneous Localization and Mapping," abbreviated SLAM. For the next ten years, the field converged around it. > 🔗 **Borrowed.** [Bar-Shalom's multi-target tracking](https://archive.org/details/trackingdataasso0000bars) (multi-target tracking, collected as a 1988 monograph) supplied a framework for estimating the states of many objects at once. Leonard and Durrant-Whyte can be read as mapping "target position" to "landmark position" and "tracker position" to "robot pose" inside that framework. This was a case of radar technology translated into indoor robot mapping. --- ## 4.3 The EKF-SLAM formulation The Extended Kalman Filter (EKF) was already standard in nonlinear system estimation before 1988, making its application to SLAM unsurprising. It runs in two stages: predict and update. During the predict stage, the motion model $f(\cdot)$ predicts the state as the robot moves, and the Jacobian $\mathbf{F}$ propagates the covariance. $$\hat{\mathbf{x}}^- = f(\hat{\mathbf{x}}, \mathbf{u})$$ $$\mathbf{P}^- = \mathbf{F}\mathbf{P}\mathbf{F}^\top + \mathbf{Q}$$ During the update stage, a sensor measurement $\mathbf{z}$ arrives, and the Jacobian $\mathbf{H}$ of the observation model $h(\cdot)$ yields a Kalman gain $\mathbf{K}$ that updates the state and covariance. $$\mathbf{K} = \mathbf{P}^-\mathbf{H}^\top(\mathbf{H}\mathbf{P}^-\mathbf{H}^\top + \mathbf{R})^{-1}$$ $$\hat{\mathbf{x}} = \hat{\mathbf{x}}^- + \mathbf{K}(\mathbf{z} - h(\hat{\mathbf{x}}^-))$$ $$\mathbf{P} = (\mathbf{I} - \mathbf{K}\mathbf{H})\mathbf{P}^-$$ These two stages define EKF-SLAM. The structure is simple, and that simplicity imposed a scalability ceiling from the start. The problem is state dimension. A state containing a 6DOF pose and $N$ 3D landmarks has dimension $6 + 3N$; its covariance matrix contains $(6+3N)^2$ entries, an $O(N^2)$ structure. Even with a fixed observation dimension, updating the full covariance costs $O(N^2)$. The Kalman gain requires multiplying the state–observation covariance by the inverse of $\mathbf{S} = \mathbf{H}\mathbf{P}^-\mathbf{H}^\top + \mathbf{R}$; the size of $\mathbf{S}$ depends on the observation dimension. 100 landmarks gives $306 \times 306 \approx$ 94k entries; 1,000 landmarks gives $3006 \times 3006 \approx$ 9M. A regular PC in the early 2000s could maintain only tens to low hundreds of landmarks in real time. [Andrew Davison's MonoSLAM (2003)](https://www.doc.ic.ac.uk/~ajd/Publications/davison_iccv2003.pdf) was limited to a few dozen landmarks in its live demos because EKF-SLAM's $O(N^2)$ wall set the ceiling. --- ## 4.4 The scalability wall When Davison ran real-time 3D tracking from a single webcam at ICCV 2003, he mapped a desk-sized space with a few dozen features. With no commercial SLAM systems available, a real-time monocular demo was rare. Its ceiling came from the size of the covariance matrix. At 100 landmarks the covariance matrix is $306 \times 306$ (6DOF pose + 100 3D landmarks, state dimension $6 + 3 \times 100 = 306$). At 1,000 it is $3006 \times 3006$. Every time step that matrix has to be updated along with a matrix inversion. On top of that, because the EKF keeps the full joint distribution in one block, adding a new landmark immediately generates cross-correlations with every existing landmark. As the map grows, update cost grows quadratically. The main workaround through the mid-2000s was submapping: partition the map into small overlapping regions, run an EKF inside each one, and connect them through a separate structure. [Chong and Kleeman (1999)](http://www.cs.cmu.edu/afs/cs/Web/People/motionplanning/papers/sbp_papers/integrated1/chong_feature_map.pdf) proposed an early form. Information loss at submap boundaries, difficult loop closure, and implementation complexity made these approaches hard to deploy. > 🔗 **Borrowed.** Chong-Kleeman's (1999) submaps and local-window optimization in modern SLAM share the goal of limiting computation. ORB-SLAM's local map and VINS-Mono's sliding window reflect that concern. Their map partitioning and state marginalization differ, however, so the parallel does not establish direct inheritance of the same method. --- ## 4.5 The consistency problem: Julier-Uhlmann's counterexample A deeper flaw in EKF-SLAM surfaced at ICRA 2001. Simon Julier and Jeffrey Uhlmann analyzed EKF-based SLAM through numerical experiments and showed that the filter trusts itself too much. Their IEEE ICRA paper was titled ["A Counter Example to the Theory of Simultaneous Localization and Map Building"](https://doi.org/10.1109/ROBOT.2001.933257). The title was provocative, and the content matched it. Secondary literature summarizes the result this way: EKF-SLAM becomes asymptotically *overconfident*. Actual estimation error grows while the covariance (uncertainty) computed by the filter converges below its true value. This is inconsistency. The cause sits in linearization error. The EKF approximates nonlinear motion and observation models by a first-order Taylor expansion. When this approximation error accumulates step by step, the covariance begins to underestimate the real error. Once the robot becomes overconfident that "I am here," the filter trusts subsequent measurements less, and errors pile up without correction. In 2007 [Shoudong Huang and Gamini Dissanayake](https://doi.org/10.1109/TRO.2007.903811) analyzed the cause of this inconsistency more precisely. They found that basic constraints among Jacobians evaluated at the current state estimate break down, driving EKF-SLAM's inconsistency. As a result, the variance of the robot's heading angle (yaw) can wrongly converge to zero when it should remain nonzero. Later observability-based analyses start from this result: the system's observable degrees of freedom change with the linearization point, and the filter injects spurious information into unobservable directions. > 📜 **Prediction vs. outcome.** After Julier and Uhlmann's 2001 counterexample, researchers spent nearly a decade designing consistent estimators. Filter variants included the Unscented Kalman Filter (UKF), Invariant EKF, and robust covariance methods. From the vantage point of 2026, however, optimization-based estimation became another major approach to the problem. [iSAM](https://www.cs.cmu.edu/~kaess/pub/Kaess08tro.pdf) (Kaess et al., 2008), [g2o](http://ais.informatik.uni-freiburg.de/publications/papers/kuemmerle11icra.pdf) (Kümmerle et al., 2011), and GTSAM became widely used, while filter-based VIO continued to develop. Iterative optimization can relinearize retained states, but this alone does not guarantee consistency. Unobservable directions and the linearization used during marginalization must also be managed. --- ## 4.6 FastSLAM — divide and conquer [FastSLAM](https://cdn.aaai.org/AAAI/2002/AAAI02-089.pdf) attacked EKF-SLAM's $O(N^2)$ wall from another direction. Michael Montemerlo, Sebastian Thrun (Stanford), Daphne Koller, and Ben Wegbreit presented it at AAAI 2002. Rao-Blackwellization supplies the key observation. Given the robot path $x_{0:t}$, the position estimates of each landmark become *mutually independent*. The path can therefore be represented by a particle filter (each particle standing for one possible path), with a separate landmark EKF running independently for each particle. With $K$ particles and $N$ landmarks the per-step complexity is $O(K \log N)$, logarithmic in $N$ for fixed $K$ rather than quadratic as in EKF-SLAM (when using KD-tree-based landmark search). As landmark count rises, per-particle EKFs stay mutually independent, so there is no need to keep the full $N \times N$ covariance. $K$ is fixed at tens to hundreds, and the practical gain was large. FastSLAM maintained real-time operation with a few hundred landmarks in indoor environments and saw rapid adoption. Problems accumulated nonetheless. Particle depletion came first: as the map grows, most particles represent poor paths, and the effective sample count drops sharply. Reweighting paths during loop closure is difficult, and adding more particles did not solve drift accumulation in large-scale environments. [FastSLAM 2.0](https://www.ijcai.org/Proceedings/03/Papers/165.pdf) (Montemerlo et al. 2003) improved the proposal distribution, but the filter paradigm still imposed a scalability ceiling. Graph optimization eventually bypassed it. --- ## 4.7 The EKF's exit Graph-based approaches became practical from 2005 onward, and EKF-SLAM receded from the main line. [Feng Lu and Evangelos Milios's 1997 graph idea](https://doi.org/10.1023/A:1008854305733) combined with [Olson-Leonard-Teller's (2006)](https://april.eecs.umich.edu/pdfs/olson2006icra.pdf) efficient solver and then with the real-time factorization techniques of g2o, GTSAM, and iSAM2. The EKF's advantage of incremental updates was no longer distinctive. The difference appeared at loop closure, when map error must be corrected as the robot returns to its starting point. The EKF has to update the entire covariance matrix at that moment, at a cost of $O(N^2)$. Graph optimization adds one edge to the pose graph and refactors a sparse matrix, at far lower cost. Around 2010, choosing the EKF as the backend for a new SLAM system became uncommon. It survived chiefly under special constraints, such as very limited compute resources or a requirement for real-time filtering. > 📜 **Prediction vs. outcome.** Durrant-Whyte and Bailey's [2006 IEEE Robotics & Automation Magazine tutorial](https://people.eecs.berkeley.edu/~pabbeel/cs287-fa09/readings/Durrant-Whyte_Bailey_SLAM-tutorial-I.pdf) discussed SLAM's scalability and projected submap decomposition and the information filter as solutions for large-scale environments. The information filter (the EKF's inverse-covariance form) was expected to use a sparse information matrix to keep computation from slowing as landmark count grew. Development took another route. The information-filter family (SEIF and related methods) accrued marginalization error while forcing sparsity. Submaps entered some systems but did not become the mainstream solution. Factor graphs with iterative optimization dominated the 2010s. --- ## 4.8 🧭 Still open Filter vs. optimization coexistence. The EKF's retreat from backend primacy does not mean it disappeared. As of 2026, some autonomous-driving implementations still prefer filter-based backends. Optimization-based SLAM needs iterative convergence, which can make real-time guarantees difficult. Sparse EKFs and UKFs reappear in low-cost embedded systems. The mix depends on the use case and its constraints. Non-Gaussian uncertainty. The EKF assumes that uncertainty follows a Gaussian distribution. Real-world sensor errors are often multimodal or heavy-tailed. A single Gaussian severely oversimplifies actual uncertainty, especially when perceptual aliasing creates multiple location hypotheses (different places looking the same). Particle filters can represent non-Gaussian distributions in theory but become impractical in high-dimensional states. Stein particles, normalizing flows, and learning-based uncertainty estimation are being tried, but few forms have been validated inside real-time SLAM as of 2026. --- While EKF-SLAM reached its real-time ceiling at around 100 landmarks, Andrew Davison at Imperial College used that same limited landmark budget to prove that a single camera could track in real time without other sensors. The numerical limit remained; the design around it changed. --- # Ch.5 — MonoSLAM → PTAM: The Real-Time Daydream and the Split Revolution EKF-SLAM provided a way to estimate a map together with its uncertainty, but its covariance matrix hit a structural wall that grew as $O(N^2)$ in the number of landmarks $N$. The theory was not wrong; the limit followed from the design. Davison and Klein responded in different ways. In 2003, Davison plugged a single webcam into a laptop in an Imperial College lab. He carried over the probabilistic spatial-relations mathematics that Smith and Cheeseman had laid down in 1988 and the EKF-SLAM structure that Leonard and Durrant-Whyte had built on it, but used a single camera as the sensor. With no IMU or stereo rig, he combined the 1994 Shi-Tomasi corner detector with a Kalman predict-update loop and ran the system in real time. By the standards of the day, it was an unusual combination. Four years later, in 2007, Klein and Murray at Oxford offered a different answer. They split tracking and mapping into two threads, an architecture that became the backbone of Visual SLAM for the next ten years. --- ## 1. The 2003 demo At ICCV 2003, Davison's [Real-Time Simultaneous Localisation and Mapping with a Single Camera](https://doi.org/10.1109/ICCV.2003.1238654) stirred the room. The individual components were familiar; seeing them work together in real time was not. The mainstream of SLAM at the time used laser sensors. LiDAR delivered 2D ranges directly, and stereo cameras recovered depth at the pixel level. A monocular camera had no depth information to begin with. Estimating 3D structure from a single camera required at least two frames, and uncertainty in the initial depth estimate propagated through the entire EKF state vector. The theory existed; real-time implementation remained difficult. Davison chose a monocular camera for practical reasons. An IMU meant extra hardware, and stereo carried a calibration burden. He wanted "to prove it with one camera," reasoning that other sensors could be added later. That logic held. The EKF's capacity to absorb those additions did not. --- ## 2. The beauty and the wall of the EKF [MonoSLAM](https://doi.org/10.1109/TPAMI.2007.1049), published in IEEE PAMI in 2007 with Davison, Ian Reid, Nicholas Molton, and Olivier Stasse as co-authors, was the full account of the ICCV 2003 demo. MonoSLAM's state vector transplanted the formulation of [Smith-Self-Cheeseman (1988)](https://arxiv.org/abs/1304.3111) and [Leonard-Durrant-Whyte (1991)](https://ieeexplore.ieee.org/document/174711/) (Ch.4) directly onto a monocular camera. It packed a camera state $\mathbf{x}_v \in \mathbb{R}^{13}$ (position 3, quaternion orientation 4, velocity 3, angular velocity 3) and landmarks $\mathbf{y}_i \in \mathbb{R}^3$ into one vector $\mathbf{x} = (\mathbf{x}_v^\top, \mathbf{y}_1^\top, \ldots, \mathbf{y}_N^\top)^\top \in \mathbb{R}^{13+3N}$. The system maintained the full $(13+3N)\times(13+3N)$ covariance $\mathbf{P}$ frame by frame in a predict-update loop. The predict step propagated the covariance through the Jacobian $\mathbf{F}$ of the camera motion model $f$ ($\mathbf{P}^- = \mathbf{F}\mathbf{P}\mathbf{F}^\top + \mathbf{Q}$); the update step computed the Kalman gain from the Jacobian $\mathbf{H}_i$ of the projection function and refreshed the state and covariance. The EKF predict-update equations are identical to those in Ch.4 §4.3. The state vector now carried camera velocity and angular velocity along with pose because a moving camera needs a dynamics model. The dominant cost in the covariance update $(\mathbf{I} - \mathbf{K}_i\mathbf{H}_i)\mathbf{P}^-$ came from a $(13+3N)^2$ matrix multiplication, or $O(N^2)$ in the number of landmarks $N$. §III of the paper states that "about 100" features could be sustained at 30 Hz in real time. > 🔗 **Borrowed.** MonoSLAM's EKF state vector transplants the Smith-Cheeseman-Durrant-Whyte (1988-1991) probabilistic spatial-relations representation directly onto a monocular camera. The Kalman filter itself had existed since 1960, but Leonard and Durrant-Whyte established the practice of placing robot pose and landmarks in one "augmented state vector" in 1991. That number exposed the system's ceiling, which the paper acknowledged by proposing a submapping strategy as future work. Building a hierarchy inside an EKF was difficult because the covariance matrix carried every correlation between every pair of landmarks, with none omitted. MonoSLAM's choice of [Shi-Tomasi (1994)](https://doi.org/10.1109/CVPR.1994.323794) corners followed the same constraint. "Good Features to Track" selected points likely to remain trackable. Restricting the state vector to such corners made the EKF update more stable. The PAMI paper states that map management kept about 12 features stably visible per frame with a wide-angle lens. The EKF ran as long as that bounded set tracked well. > 🔗 **Borrowed.** The Shi-Tomasi 1994 corner detector was already in use in MonoSLAM, not first in PTAM. The design philosophy of "select good features, then track them" is the direct Shi-Tomasi → MonoSLAM → PTAM lineage. --- ## 3. 2007, the same year That bounded feature count exposed the EKF's ceiling. Klein confronted the same limit at Oxford. In 2007, Klein and Murray presented [Parallel Tracking and Mapping for Small AR Workspaces](https://doi.org/10.1109/ISMAR.2007.4538852) at ISMAR. The same year, PAMI carried the finished version of Davison's MonoSLAM. Their appearance in the same year reflected a direct lineage. Klein was then a doctoral student in Murray's group, which continued the Oxford Active Vision Laboratory where Davison had recently completed his doctorate under Murray. Klein saw in MonoSLAM not the EKF itself but proof that a monocular camera could run in real time. The next problem was scale, and Klein discarded the EKF. --- ## 4. The split PTAM separated tracking (camera pose tracking) from mapping (3D map construction) and ran them in two parallel threads. In the EKF, the two tasks were coupled inside one loop. Each frame ran a predict-update cycle: predict the state when the camera moves, then update it once landmarks are found in the image. PTAM separated the tasks. The tracking thread estimates only the camera pose in each frame. It matches the 2D projections of 3D points visible from the current keyframe set to the actual observations and computes the pose in real time. The mapping thread runs bundle adjustment whenever a new keyframe is added. Because tracking runs independently, slow mapping does not block it. The bundle adjustment on the mapping thread minimized the sum of reprojection errors over a keyframe set $\mathcal{K}$ and a 3D point set $\mathcal{P}$: $$\min_{\{\mathbf{T}_k\}, \{\mathbf{p}_j\}} \sum_{k \in \mathcal{K}} \sum_{j \in \mathcal{P}_k} \rho\!\left(\left\|\mathbf{z}_{kj} - \pi(\mathbf{T}_k,\, \mathbf{p}_j)\right\|^2_{\mathbf{\Sigma}_{kj}}\right)$$ where $\mathbf{T}_k \in SE(3)$ is the pose of keyframe $k$, $\mathbf{p}_j \in \mathbb{R}^3$ is a 3D point, $\pi$ is the camera projection function, $\mathbf{z}_{kj}$ is the observed pixel coordinate of point $j$ in keyframe $k$, $\mathbf{\Sigma}_{kj}$ is the measurement covariance, and $\rho$ is a robust kernel such as the Huber function. The mapping thread solved this optimization iteratively with Levenberg–Marquardt. Running asynchronously, it did not affect the real-time behavior of the tracking thread. > 🔗 **Borrowed.** The bundle adjustment on PTAM's mapping thread directly applies [Triggs et al. 1999 "Bundle Adjustment — A Modern Synthesis"](https://doi.org/10.1007/3-540-44480-7_21). The hundred-year photogrammetry tradition covered in Part I moved into the center of a real-time SLAM backend. Large joint updates were costly in EKF-SLAM because of covariance-matrix size; splitting the threads let keyframe BA run asynchronously from tracking. The split had large consequences. Because the mapping thread ran bundle adjustment asynchronously, the number of landmarks in the map was no longer bound by the EKF's $O(N^2)$ constraint. PTAM used hundreds of keyframes, each holding hundreds of patch features, compared with MonoSLAM's tens of landmarks. PTAM also changed initialization. As the user moved the camera slowly, the system estimated the essential matrix with a 5-point algorithm from the [Nistér 2004](https://doi.org/10.1109/TPAMI.2004.17) line (the PTAM paper cites the follow-up Stewénius·Engels·Nistér 2006) and recovered the initial 3D structure from the first keyframe pair. This too was borrowed. The essential matrix $\mathbf{E}$ is a $3\times 3$ matrix capturing the pure geometric relation between two camera frames, satisfying ${\mathbf{p}'}^\top \mathbf{E}\, \mathbf{p} = 0$ for corresponding point pairs $(\mathbf{p}, \mathbf{p}')$. $\mathbf{E}$ decomposes internally as $\mathbf{E} = \mathbf{t}_\times \mathbf{R}$ ($\mathbf{t}_\times$ is the skew-symmetric matrix of the translation, $\mathbf{R}$ is the rotation), so it has 5 degrees of freedom. Five point correspondences make the problem minimal, but they do not give a unique solution: the polynomial system has up to ten candidate solutions over the complex numbers. Nistér's contribution was to solve this system efficiently enough for a real-time RANSAC loop. PTAM used the solver during initialization to estimate the relative pose between the first two keyframes and triangulate the initial 3D point cloud. > 🔗 **Borrowed.** PTAM's 5-point essential-matrix initialization follows the minimal-solver lineage opened by David Nistér's 2004 "An Efficient Solution to the Five-Point Relative Pose Problem" (the PTAM paper directly cites the follow-up Stewénius·Engels·Nistér 2006 ISPRS). Nistér's solver used the minimum number of correspondences needed to build a monocular camera's initial map. PTAM placed it inside a RANSAC loop to estimate the relative pose of the first two keyframes at near-real-time speed. > 🔗 **Borrowed.** PTAM's keyframe structure traces back to the Leonard-Durrant-Whyte submap idea. The notion that "if the full map is hard to optimize at once, break it into regions" was expressed in PTAM as a set of keyframes. The covisibility graph of the subsequent ORB-SLAM is a more refined version of this keyframe management. --- ## 5. The diffusion of the new architecture PTAM was designed for AR (augmented reality) workspaces; the paper's title states "Small AR Workspaces" explicitly. Because the tracking thread ran reliably in real time, it could be integrated directly into AR applications. Commercial adoption was fast. In the early 2010s, Metaio (a German AR startup, acquired by Apple in 2015) and Qualcomm's Vuforia SDK adopted tracking/mapping split structures similar to PTAM's. These commercial SDKs helped spread stable planar AR on consumer smartphones. The academic effect was more direct. [ORB-SLAM](https://arxiv.org/abs/1502.00956), published in 2015 by Raul Mur-Artal, J.M.M. Montiel, and Juan D. Tardós, inherited PTAM's structure. It swapped patch features for ORB descriptors, refined keyframe management with a covisibility graph, and added loop closure on top. Without PTAM, ORB-SLAM's blueprint would have been different. Qin, Li, and Shen's [VINS-Mono](https://arxiv.org/abs/1708.03852) (2018) also uses two threads for sliding-window optimization and loop closure, extending the tracking/mapping split into Visual-Inertial Odometry (VIO). --- ## 6. Davison vs Klein & Murray — a view comparison Two papers came out in 2007. MonoSLAM PAMI was the finished version of the 2003 demo. PTAM came out the same year, with a new structure that broke through MonoSLAM's limits. One reason MonoSLAM retained the EKF was its explicit joint representation of state and uncertainty. The covariance matrix represented state uncertainty explicitly, tracking both the uncertainty of each landmark and the covariance between landmarks. From this viewpoint, bundle adjustment traded explicit uncertainty representation for scalability. Klein & Murray paid that price willingly. What mattered in AR applications was real-time tracking of the camera pose. There was no need to track map uncertainty at the centimeter level. Refining the map periodically through bundle adjustment was enough. The field accepted this trade. From the 2010s onward, graph-based optimization and bundle adjustment became mainstream, while EKF-SLAM receded except in applications with severely limited compute resources. MonoSLAM's concern with probabilistic consistency did not disappear. Rather than join the PTAM lineage directly, Davison's lab moved in stages toward factor-graph-based estimation and then Gaussian Belief Propagation (GBP) and the Robot Web. Twenty-three years later, in Ch.18 of the *SLAM Handbook*, Davison describes this trajectory as EKF → BA → factor graph → GBP. He does not name and assess MonoSLAM directly, but recasts the history around a general principle: each change of representation prompts a redesign of the system. --- ## 📜 Prediction vs. outcome > **Davison 2007 PAMI MonoSLAM**: In the Conclusion, Davison named larger indoor and outdoor environments, faster motion, and complex scenes with occlusion and lighting changes as the next tasks. He specifically proposed a submap strategy and CMOS cameras running above 100 Hz, and suggested extending the sparse map to a dense representation of "higher-order entities" (surfaces, etc.). > > These predictions met different fates. Submaps, PTAM's keyframe structure, and ORB-SLAM's covisibility graph share a concern with limiting computation, but this comparison alone does not establish direct inheritance. No system, however, achieved hierarchical scaling while retaining the EKF; that hierarchy arrived with the shift to BA-based architectures. High-frame-rate cameras took concrete form through 2010s event-camera research. Robustness to dynamic scenes remains open as of 2026. DynaSLAM and FlowSLAM are among the attempts, but no solution has yet entered the baseline pipeline. Davison did not flag IMU integration directly in Future Work (though the body cites related work), and the VIO boom of the 2010s pursued that direction. The concern with probabilistic consistency survived in factor graphs and GBP. Twenty-three years later, in Handbook Ch.18, Davison discusses system redesign through changes of representation. Placing MonoSLAM within that lineage is the interpretation of this history. > **Klein & Murray 2007 PTAM**: In §8 (Failure modes / Mapping inadequacies), Klein and Murray listed the system's limitations: corner-based tracking's vulnerability to motion blur, the geometric poverty of a point-cloud-centric map, and "not designed to close large loops in the SLAM sense." They stated plainly that global consistency across large loops was outside PTAM's design scope. > > In 2015 ORB-SLAM directly addressed those limitations. It added [DBoW2](http://doriangalvez.com/papers/GalvezTRO12.pdf)-based appearance loop closure and covisibility-graph-based keyframe management, and replaced patch features with ORB descriptors. ORB-SLAM took on the map-scaling task that PTAM had excluded. Klein & Murray did not explicitly identify appearance-based loop closure as the answer, but the limits they marked became the starting point for the subsequent lineage. --- ## 🧭 Still open **Monocular scale recovery.** From MonoSLAM to PTAM, every monocular system carries scale ambiguity. A single image cannot determine absolute distance; this is a geometric fact. Adding an IMU makes scale observable through gravity direction and accelerometer readings. In pure monocular systems without an IMU, however, scale recovery remains unsolved even in 2026. Learning-based monocular depth estimation ([MiDaS](https://arxiv.org/abs/1907.01341), [Depth Anything](https://arxiv.org/abs/2401.10891)) estimates relative depth from a single image, but converting it to metric scale still requires an external reference (a ground-plane assumption, a known object size, and so on). **Environmental generality of a single VO system.** MonoSLAM handled only indoor desktop scenes. PTAM explicitly limited its scope to "Small AR Workspaces." ORB-SLAM2 later tried to span indoor, outdoor, and RGB-D settings, but tracking still fails under extreme lighting changes or in low-texture spaces. As of 2026, no single pipeline robustly handles indoor corridors, outdoor downtowns, nighttime environments, and textureless white walls at once. Multimodal fusion (camera + LiDAR + IMU) covers some of this range, but the generality of a camera-only system remains unsettled. **Feature tracking in low light and dynamic scenes.** MonoSLAM assumed sufficient lighting and a static scene; PTAM did the same in 2007. As of 2026, most feature-based SLAM systems still carry these assumptions implicitly. ORB features can fail entirely in low light, while scenes crowded with moving people cause dynamic points to be misclassified as static. Learning-based optical flow and semantic segmentation attempt to address the problem, but no system has become a real-time, general-purpose solution. --- PTAM's tracking/mapping split left loop closure unresolved. As keyframes accumulated, so did error, becoming visible when the camera completed a loop. The PTAM paper itself declared loop closure, correcting accumulated error when the camera returned to a known place, out of scope. Elsewhere, graph-SLAM researchers had spent a decade preparing an answer. --- # Ch.6 — The Graph SLAM Revolution In a basement corridor at Carnegie Mellon in 1997, Feng Lu and Evangelos Milios were trying to align multiple laser scans into a globally consistent map. The EKF was the default, but they modeled relative measurements between poses as a graph and ran least-squares optimization on it. This produced global consistency without a Kalman filter. Lu and Milios were not alone. More than a decade earlier, [Chatila and Laumond (1985)](https://www.semanticscholar.org/paper/Position-referencing-and-consistent-world-modeling-Chatila-Laumond/c34a678e40a7d80cb3683f07fc837179fd9bf3ee) at LAAS had discussed reference frames and consistent world models for mobile robots in the language of smoothing. In 1999, [Gutmann and Konolige](https://www.semanticscholar.org/paper/Incremental-mapping-of-large-cyclic-environments-Gutmann-Konolige/3c1bda51b8ca59f1836ed1b96c485d905804989a) applied pose graph matching to incremental mapping of large cyclic environments, and in the early 2000s Thrun's group formalized the approach as the *full SLAM* problem and put it on a commercial trajectory. [Folkesson and Christensen (2004)](http://www.hichristensen.net/hic-papers/folkesson-icra2004.pdf), Konolige, and Dellaert followed with formulations of their own. Lu-Milios 1997 remains the most cited because it presented a complete pipeline ("laser scan matching plus batch least-squares"), not because it opened the direction alone. If Smith and Cheeseman supplied the mathematical basis for probabilistic mapping and Davison demonstrated real-time monocular SLAM, these contributors made parallel moves that reframed SLAM as graph inference. EKF-SLAM in the 2000s met an $O(N^2)$ covariance-update bottleneck as landmark count grew, while Klein and Murray's PTAM (2007) separately demonstrated real-time optimization through a BA-based keyframe structure and split tracking and mapping. Laboratories at CMU, LAAS, Stanford, and KTH had already been developing graph-smoothing alternatives to filtering. --- ## 6.1 From Laser Scans to Pose Graphs: Lu-Milios 1997 Before [Lu & Milios 1997, "Globally Consistent Range Scan Alignment"](https://doi.org/10.1023/A:1008854305733) appeared, alignment of successive laser scans was often handled by stitching together local matches from the ICP (Iterative Closest Point) family. ICP aligned two scans well locally, but as drift accumulated the map twisted after tens of meters. When the robot came back to close a loop, the starting point and the map no longer matched. Lu and Milios represented the robot's pose sequence $x_1, x_2, \ldots, x_n$ as nodes and each relative measurement between poses as an edge. Map building then becomes energy minimization on the graph. Each edge carries the relative transform $\hat{z}_{ij}$ between two poses and its uncertainty $\Omega_{ij}$. The full cost function is $$F = \sum_{(i,j) \in \mathcal{E}} e_{ij}^T \Omega_{ij} e_{ij}, \quad e_{ij} = z_{ij} - h(x_i, x_j)$$ where $h(x_i, x_j)$ computes the expected relative transform from the two poses, $z_{ij}$ is the actual measured relative transform, and $\Omega_{ij} = \Sigma_{ij}^{-1}$ is the information matrix, the inverse of the measurement uncertainty. Loop closures fit this formulation naturally. When the robot revisits a place and obtains a new relative measurement, one edge adds the constraint to the graph, and full optimization adjusts every pose accordingly. An EKF updated the covariance at $O(N^2)$ cost to close a loop; a pose graph expresses the constraint with an additional edge, but still requires reoptimization to update its estimates. > 🔗 **Borrowed.** The Lu-Milios formulation of pose graph optimization rests on the nonlinear least-squares algorithms of [Levenberg (1944)](https://www.ams.org/qam/1944-02-02/S0033-569X-1944-10666-0/) and [Marquardt (1963)](https://www.stat.cmu.edu/technometrics/70-79/VOL-14-03/v1403757.pdf). A numerical optimization technique developed decades earlier for nonlinear parameter estimation arrived at the backend of indoor laser mapping. The Lu-Milios solution was a batch linear system that solved for all poses at once, growing with the number of scans. It was closer to a proof of concept than a field-ready system, but it showed that optimization rather than filtering could provide global consistency. During the same period, Gutmann-Konolige emphasized incrementality, Folkesson-Christensen the robustness of data association, and Thrun's group application at real-world scale. Each developed a different part of the same conclusion. --- ## 6.2 Discovering Sparsity: The Information Matrix and the Pose-Graph Extension In the five years after the Lu-Milios idea was published, several groups pushed extensions in the same direction. The common discovery was the **sparsity** of the information matrix ($\Omega = \Sigma^{-1}$). The covariance matrix $\Sigma$ of EKF-SLAM is dense. Every time the robot observes a new landmark, its correlation with every existing landmark is updated. With the robot pose marginalized and $n$ 2D landmarks, $\Sigma$ is a $2n \times 2n$ matrix and the update cost is $O(n^2)$. That is why real-time performance collapsed around 100 landmarks. The information matrix of a pose graph is different. A nonzero term appears in the $(i,j)$ block of $\Omega$ only when poses $x_i$ and $x_j$ are directly connected by a measurement. Under continuous motion, edges connect only nearby poses; distant poses have no direct connection. $\Omega$ therefore has a banded sparse structure that reflects the graph topology. For a trajectory without loop closures, the structure is nearly tridiagonal. Sebastian Thrun's group, through the [Sparse Extended Information Filter (SEIF)](http://www.cs.cmu.edu/~thrun/papers/thrun.tr-seif02.pdf), and Edwin Olson began exploiting this sparsity explicitly. A sparse linear solver could reduce computation far below $O(n^2)$. Actual complexity depended on graph structure, but $O(n \log n)$ became possible in realistic scenarios where a robot moves within a bounded region. > 🔗 **Borrowed.** Thrun's sparse information filter (SEIF) and [Eustice's exactly sparse delayed-state filter](https://web.mit.edu/2.166/www/handouts/eustice_et_al_ieeetro_2006.pdf) showed that information-matrix sparsity was usable even in filter form. This insight set the stage for Dellaert's factor graph formulation and the Bayes tree data structure. At ICRA 2006, [Olson, Leonard, and Teller](https://april.eecs.umich.edu/pdfs/olson2006icra.pdf) presented a stochastic-gradient method for optimizing pose graphs. It offered no convergence guarantee but ran fast enough on graphs of hundreds of nodes, and Olson's implementation spread through the community. --- ## 6.3 Factor Graphs and Square Root SAM In 2006, Dellaert and his doctoral student Kaess published [Square Root SAM](https://doi.org/10.1177/0278364906072768), giving SLAM backends a new representation. Dellaert had worked on probabilistic graphical models at Georgia Tech. He treated SLAM as Bayesian inference and represented that inference on a factor graph. In a **factor graph**, a bipartite graph with variable and factor nodes, the variable nodes are robot poses and landmark positions, while factor nodes represent observations or priors. A factor $f_k(x_{i_1}, x_{i_2}, \ldots)$ expresses a probabilistic constraint among the variables it connects. The full joint probability is $$p(X) \propto \prod_k f_k(X_k)$$ and MAP estimation finds the $X^*$ that maximizes this probability. Under Gaussian factors, this becomes a nonlinear least-squares problem. The least-squares structure supplied Dellaert's key observation. Applying QR decomposition to the Jacobian matrix $J$ leaves an upper-triangular matrix $R$. Because $R^T R = J^T J = \Omega$, $R$ is the "square root information matrix." Its sparsity depends not only on the Jacobian but also on variable-elimination order and factor-graph topology. A good ordering (e.g., AMD, COLAMD) minimizes fill-in and yields a sparse $R$. This formulation is numerically more stable than the EKF covariance update. The full map of landmarks and poses can be optimized together in a consistent way, and a loop closure is expressed as the addition of a new factor. --- ## 6.4 iSAM and iSAM2: Online Incremental Inference Square Root SAM was a batch method. Recomputing the full decomposition of $J^T J$ whenever a new observation arrived added substantial work. Dense factorization costs $O(n^3)$; sparse costs depend on connectivity and elimination order. In 2008, [Kaess, Ranganathan, and Dellaert published **iSAM** (incremental Smoothing and Mapping)](https://www.cs.cmu.edu/~kaess/pub/Kaess08tro.pdf), which updated the factorization with Givens rotations. When a new variable and factor were added, iSAM appended only the new rows and updated $R$ rather than recomputing the QR decomposition. iSAM1's intrinsic limit was its relinearization schedule. The $R$ obtained by linearizing nonlinear factors is only a first-order approximation near the current estimate. As the robot moved away from the linearization point, approximation error accumulated. iSAM1 responded with **periodic full relinearization**: every few dozen steps, it relinearized the entire factor graph and recomputed the QR decomposition. Fill-in within $R$ after loop closures signaled when to run this batch step. An apparently incremental algorithm therefore reverted to batch processing every cycle. In 2012 [iSAM2](https://doi.org/10.1177/0278364911430419) solved this problem with a data structure called the Bayes tree. The Bayes tree is a tree built from the chordal Bayes net obtained by applying variable elimination to the factor graph. Its nodes are the cliques of the Bayes net, and its edges are the separators (shared variables between cliques). When a new factor is added, the cliques affected in the Bayes tree are identified, and only that subtree is turned back into a factor graph, relinearized, and reoptimized. The core is **fluid relinearization**. Only factors whose linearization error exceeds a threshold are selectively relinearized, and the effect propagates through Bayes tree separators only as far as needed. iSAM1's "everything, every cycle" schedule was replaced by "only the necessary factors, only the affected cliques". Even when a loop closure occurred, the set of connected cliques was often locally bounded, and full recomputation could be avoided. > 🔗 **Borrowed.** The Bayes tree extends the junction tree (join tree) lineage in probabilistic graphical models, represented by the elimination-order and chordal-graph inference techniques in textbooks such as Koller-Friedman's [*Probabilistic Graphical Models*](https://mitpress.mit.edu/9780262013192/probabilistic-graphical-models/). It brought a technique from AI inference into real-time robot SLAM. iSAM2 was packaged as the [GTSAM (Georgia Tech Smoothing and Mapping)](https://gtsam.org) library, with a C++ core and Python bindings. GTSAM development continued while Dellaert held his Georgia Tech position and also worked with Google. Its public documentation and examples cover applications including autonomous driving, drones, and robot-arm calibration. --- ## 6.5 g2o: A General Graph Optimizer in the ROS Ecosystem While the Georgia Tech group refined the theory, Rainer Kümmerle, Giorgio Grisetti, Hauke Strasdat, Kurt Konolige, and Wolfram Burgard at TUM (Munich) and Freiburg built a practical open-source implementation. At ICRA 2011, they presented [g2o](https://doi.org/10.1109/ICRA.2011.5979949) (general graph optimization), designed to "handle any kind of graph optimization in a plug-in manner." The author list bridged Burgard and Grisetti's Freiburg robotics tradition, Strasdat's monocular SLAM experience, and Konolige's industrial engineering perspective. The g2o design separates three concepts. Vertices (variable nodes) and edges (factors or constraints) form the graph, and a solver handles the sparse linear system. The user defines vertex types and the error function and Jacobian of edges, and g2o then runs the full optimization with Gauss-Newton or Levenberg-Marquardt. The sparse solver can be chosen among Cholmod, CSparse, and Eigen, or swapped out for an external library. As ROS (Robot Operating System) spread through mobile-robot research in the early 2010s, g2o became one of the leading implementations for graph-based SLAM. ORB-SLAM and LSD-SLAM adopted it, but ROS SLAM did not converge on one backend: gmapping belongs to the particle-filter line, and Cartographer uses Ceres. g2o was influential without being universal. --- ## 6.6 Why the Field Converged Here Chatila-Laumond (1985), Lu-Milios (1997), Gutmann-Konolige (1999), Folkesson-Christensen (2004), Thrun's group, Dellaert (2006), and Kaess (2012) started from different problems and developed tools for maintaining and solving graph constraints. The shift changed the model of the problem, not only the algorithm. EKF-SLAM maintains the best estimate of the current state and its uncertainty while marginalizing away the past. Past poses disappear, and accumulated error remains inside the current estimate. Closing a loop then requires a costly update to the current covariance. Graph SLAM retains the past. Poses, landmarks, and observations remain in the graph, and a loop closure becomes a new edge. Reoptimization adjusts the full trajectory consistently (for methods that use a time-continuous trajectory rather than discrete keyframes, see Ch.7c Continuous-Time SLAM). Keeping past poses revisable is the essential difference from filters. Costs differ as well. EKF's update cost is $O(N^2)$ (in the number of landmarks $N$), and information storage is $O(N^2)$. Graph methods can reduce that complexity substantially with sparse Cholesky (or QR) decomposition. Update cost depends on graph connectivity, fill-in during factorization, elimination order, and the region that must be recomputed. Motion within a bounded physical region alone does not guarantee $O(N \log N)$ updates. > 📜 **Prediction vs. outcome.** The limits of Dellaert's batch Square Root SAM (2006) led the same group toward incremental methods. iSAM handled the problem in 2008 with Givens-rotation updates, and iSAM2 improved loop-closure efficiency in 2012 with the Bayes tree. GTSAM, Ceres, and g2o all handle nonlinear least squares, but differ in their solvers and incremental data structures. The three papers resolved one problem in stages, largely along the predicted path. The flexibility of marginalization played a part as well. When an old pose in the graph is marginalized, its information is preserved as a linking factor among the remaining variables. Filters also remove past states while transferring their information to the current estimate. Both approaches must account for approximation and relinearization constraints introduced by this compression. Engineering trade-offs like sliding-window optimization and keyframe selection come in here. --- ## 6.7 Nonlinearity and Robustness: The Layer of Practical Engineering Implementation is less tidy than the theory. Much of 2010s SLAM engineering went into closing that gap. The first problem is dependence on the initial value. Gauss-Newton or LM optimization converges to a local minimum when the initial pose estimate is far from the truth. A wrong loop-closure correspondence corrupts that initialization, making verification and outlier rejection central pre-backend tasks. Ch.6b (Certifiable SLAM) treats convex-relaxation methods (SDP) that avoid this local-minimum problem and certify global optimality. Standard least squares is fragile to outliers, as practice quickly revealed. Robust Huber or Cauchy losses reduce the influence of wrong matches, and both g2o and GTSAM make them selectable. The choice depends on the environment and sensor, and as of 2026 still rests on the engineer's experience. The third problem is marginalization approximation. iSAM2's Bayes tree provides exact incremental inference, but the tree grows with the variable count. Real systems marginalize old poses to keep it manageable, and the resulting fill-in can make the information matrix dense. Implementation quality depends on how this fill-in is truncated and approximated with a prior factor. > 📜 **Prediction vs. outcome.** g2o's generality lies in allowing users to define states as vertices and observation constraints as edges. New line, plane, inertial, or object constraints can be implemented through that interface; examples must be checked against the estimator and edge implementation actually used by each system. Ch.7b describes preintegration, the standard way to place IMU factors in the graph. As of 2026, however, g2o itself prioritizes interface stability and compatibility with existing users over broad expansion of its built-in factors. Users commonly add new factor types through inheritance, forks, or wrappers. --- ## 🧭 Still open Which robust kernel to choose. Huber, Cauchy, Geman-McClure, DCS, and others are available, but no principled method decides in advance which kernel is optimal for a given environment and sensor. The choice still rests on engineering judgment. Researchers are learning cost functions themselves, but integration into an online incremental system remains unresolved. Representing non-Gaussian uncertainty inside a factor graph remains open. The basic continuous optimization in GTSAM and g2o starts from Gaussian residual models, although robust kernels, max mixtures, and hybrid factors provide extensions. Accurately representing loop-closure mismatch probabilities and multi-hypothesis poses in real time still lacks a general solution. The Bayes tree is most efficient when a new factor affects only local cliques. As loop-closure constraints become dense in a large map, the affected subtree and fill-in can grow, increasing computation and memory. Hierarchical management and submap partitioning are ways to control that growth. --- By the early 2010s, the backend debate had quieted. As tools such as g2o and GTSAM became widely used, attention moved to the layers above them. The question changed from "how to close a loop" to "which features can recognize one, and from how far away?" The front end became the new focus of competition. One question remained unresolved: do g2o and GTSAM actually return the global minimum? Ch.6b treats that question of certifiability before the front-end lineage resumes in Ch.7. --- # Ch.6b — Certifiable SLAM: Past the Local Minimum The lineage from Lu-Milios to g2o and GTSAM left one issue unresolved: pose graph optimization is non-convex, so Gauss-Newton and LM may return local minima. Practitioners relied on a rule of thumb, "with an odometry initial guess, it usually solves fine," yet some deployed backends converged to the wrong solution without raising an alarm. In 2015 Luca Carlone at MIT began replacing that rule of thumb with mathematics. Carlone's work on Lagrangian duality led to Rosen's SE-Sync in 2019, Briales-Gonzalez-Jimenez's Cartan-Sync, Yang-Carlone's TEASER, and Papalia's CORA. Together, this lineage moved the backend from non-convex optimization that usually worked to convex surrogates with certificates of global optimality. The tools came from outside SLAM: Shor relaxation from operations research, Burer-Monteiro factorization from mathematical optimization, Riemannian optimization from differential geometry, and Kirchhoff's Matrix-Tree Theorem from graph theory. --- ## 6b.1 The Old Anxiety of Local Minima Ch.6 §6.7 identified dependence on initial values as the first problem of a graph SLAM backend. The cost function is non-convex in rotation variables $\boldsymbol{R}_i \in \mathrm{SO}(3)$, so Gauss-Newton can enter the wrong basin when the initial estimate is far from the truth. The parking-garage example in Handbook §6.1 shows the symptom clearly: of four random initializations, one converges to the same global minimum as SE-Sync, while the other three settle at twisted local minima with the garage floor visibly folded. Through the late 2000s, the community followed two responses: trust odometry to provide a good initial value, or enforce loop-closure verification and outlier removal at the front end. Both worked, but neither determined whether the converged value was the true minimum. Huang and Dissanayake pointed to a simple issue around 2010: even with a good initial guess, ambiguous data can stop the optimizer at the wrong answer. PGO was also formalized as NP-hard around the same time, although g2o usually worked in the field. Backend theorists of the mid-2010s focused on this gap between worst-case theory and average practice. Convergence of Gauss-Newton does not imply global optimality. Under ideal second-order conditions a local minimum has zero gradient and a positive-semidefinite Hessian, but a numerical solver may also stop earlier because of tolerances or conditioning. Failure is least visible when the backend reports convergence. > 🔗 **Borrowed.** Ch.6's robust kernels (Huber, Cauchy) and this chapter's GNC share a root in the robust statistics and duality theorem of [Black & Rangarajan (1996)](https://cs.brown.edu/people/mjblack/Papers/ijcv1996.pdf). One branch changed cost weights to reduce outlier influence; the other used the same principle to navigate non-convexity. --- ## 6b.2 Shor Relaxation — A Tool From Operations Research PGO's non-convexity comes from the rotation constraint $\boldsymbol{R}_i \in \mathrm{SO}(d)$. Orthogonality, $\boldsymbol{R}^\top \boldsymbol{R} = \boldsymbol{I}$, is quadratic. In three dimensions, $\det(\boldsymbol{R})=+1$ is cubic as written, but right-handedness can instead be expressed with quadratic cross-product relations among the columns. Together with the quadratic objective, these constraints give a **QCQP** (Quadratically Constrained Quadratic Program). Operations research had used a convex relaxation for QCQP since 1987: [Naum Shor's relaxation](https://link.springer.com/article/10.1007/BF01582220). Using the identity $\boldsymbol{x}^\top \boldsymbol{M}\boldsymbol{x} = \mathrm{tr}(\boldsymbol{M}\boldsymbol{x}\boldsymbol{x}^\top)$, Shor introduces the lifted variable $\boldsymbol{X} \triangleq \boldsymbol{x}\boldsymbol{x}^\top$. The original QCQP then has a linear objective subject to $\boldsymbol{X} \succeq 0$ and rank 1. Dropping the rank-1 constraint leaves a convex **semidefinite program (SDP)**. The search space grows from $n$ to $n(n+1)/2$ in exchange for convexity. $$d^* = \min_{\boldsymbol{X}\in\mathbb{S}^n} \mathrm{tr}(\boldsymbol{C}\boldsymbol{X}) \;\; \text{s.t.} \;\; \mathrm{tr}(\boldsymbol{A}_i\boldsymbol{X})=b_i,\; \boldsymbol{X}\succeq 0.$$ The duality inequality $d^* \le p^*$ makes the relaxation useful. The SDP minimum is a lower bound on the original QCQP minimum. For any candidate $\hat{\boldsymbol{x}}$, the quantity $f(\hat{\boldsymbol{x}}) - d^*$ upper-bounds its suboptimality. This is the source of the term "certifiable": even without solving the original problem globally, one can bound the solution's error from optimality. If the SDP solution $\boldsymbol{X}^*$ has rank 1, then $\boldsymbol{X}^* = \boldsymbol{x}^*\boldsymbol{x}^{*\top}$ and $\boldsymbol{x}^*$ is a global minimizer of the original QCQP. The papers that follow ask how often this occurs in SLAM. Carlone's two papers at IROS and ICRA in 2015 ([Carlone et al. 2015 "Lagrangian duality in 3D SLAM"](https://arxiv.org/abs/1506.00746) and [Carlone & Dellaert 2015 "Planar pose graph optimization"](https://doi.org/10.1109/ICRA.2015.7139264)) are the starting point. They showed empirically that the duality gap is mostly zero in 2D PGO and suggested extension to 3D. Carlone had just finished his 2014 TRO survey of g2o/GTSAM initialization techniques and had seen how often, when odometry conflicted with loop closures, the optimizer halted at the wrong point. The 2015 paper reported "the empirical fact that the duality gap is typically zero" without giving a closed condition for when it holds. In the same period, [Briales & Gonzalez-Jimenez (2017)](https://arxiv.org/abs/1702.03235)'s Cartan-Sync extended the program to SO(3) synchronization. On the mathematical side, Boumal-Absil-Sepulchre were refining Riemannian optimization, and Burer-Monteiro's low-rank SDP factorization had existed since 2003. Rosen and colleagues assembled these materials in one paper in 2019. --- ## 6b.3 SE-Sync — What Rosen 2019 Assembled [Rosen, Carlone, Bandeira, Leonard's SE-Sync (IJRR 2019)](https://arxiv.org/abs/1612.07386) became the standard reference for certifiable SLAM. Rosen completed his doctorate with John Leonard at MIT; Leonard, together with Ch.4's Durrant-Whyte, had helped settle the name "SLAM" in the early 1990s. Afonso Bandeira, an expert in SDP and synchronization, contributed the theoretical proof for the globality of rank-deficient second-order critical points. The paper combined backgrounds in robotics, SLAM, mathematical optimization, and applied mathematics. It assembled Shor relaxation, translation elimination, Burer-Monteiro low-rank parameterization, and Boumal's Riemannian staircase around the single problem of PGO. The method proceeds in three steps. First, because translation becomes linear least squares once rotation is fixed, $\boldsymbol{t}$ is eliminated in closed form (Problem 6.2). Carlone had noted this in his 2014 TRO survey, and Rosen made it the first step of the convex relaxation. Second, Shor relaxation lifts the remaining rotation-only problem $\min_{\boldsymbol{R}\in\mathrm{SO}(d)^n} \mathrm{tr}(\tilde{\boldsymbol{Q}}\boldsymbol{R}^\top\boldsymbol{R})$ to an SDP (Problem 6.3). Third, because the $dn \times dn$-dimensional SDP would overwhelm interior-point methods at a few thousand poses, Burer-Monteiro reparameterization $\boldsymbol{Z} = \boldsymbol{Y}^\top \boldsymbol{Y}$ turns it into a low-dimensional unconstrained problem on the Stiefel manifold (Problem 6.4). Two theorems justify the method. Theorem 6.1 gives **exact recovery**: if measurement noise is below a constant $\beta$, the SDP relaxation has a unique solution $\boldsymbol{Z}^*$ of rank $d$ that recovers the global minimum of the original MLE. Unlike the rank-1 condition for a generic vector QCQP, SE-Sync uses $\boldsymbol{Z}=\boldsymbol{R}^\top\boldsymbol{R}$ with $\boldsymbol{R}\in\mathbb{R}^{d\times dn}$, so an exact solution has rank $d$. The caveat is that $\beta$ depends on the ground-truth matrix and is therefore unknown before the instance is solved. Theorem 6.2 applies a result of Boumal et al.: a rank-deficient second-order critical point on the Stiefel manifold is the global minimum. Together, the theorems enable the Riemannian Staircase. It starts at low rank, finds a second-order critical point, checks for rank deficiency, and increases the rank by one if the check fails. Once the rank reaches $dn + 1$, every $\boldsymbol{Y}$ is row-rank-deficient, so the process halts in finite steps. On practical datasets, one step usually suffices. On standard benchmarks such as sphere, torus, and garage, SE-Sync converged at g2o/GTSAM speed while returning an a posteriori certificate. g2o and GTSAM were fast but silent on when to trust the answer; Rosen's algorithm returns one additional number, a suboptimality bound. When that bound is zero, the solution is provably globally optimal. Twenty years after Lu-Milios, a backend could certify global optimality. A remaining gap alone, however, does not prove that the candidate is suboptimal. > 📜 **Prediction vs. outcome.** In §8.2 of the IJRR 2019 paper, Rosen wrote that "the algebraic simplification we have shown could be extended to anisotropic noise, outliers, and a variety of sensor modalities." The prediction partially held. Holmes-Barfoot's 2023 landmark-SLAM extension, Papalia's 2024 CORA range-measurement extension, and Yang-Carlone's TEASER line followed. The most ambitious extension, applying SE-Sync to the perspective projection of visual SLAM, had not arrived by 2026. Projection is a rational function and does not fold easily into polynomial optimization. > 🔗 **Borrowed.** The Burer-Monteiro factorization at the heart of SE-Sync is the low-rank SDP method of [Burer & Monteiro (2003)](https://link.springer.com/article/10.1007/s10107-002-0352-8), later sharpened by [Boumal-Voroninski-Bandeira (2016)](https://arxiv.org/abs/1605.08101)'s Riemannian proof of the globality of second-order critical points, which Rosen brought into SLAM. From pure math to the robot backend, sixteen years. --- ## 6b.4 The Unexpected Equivalence of Graph Laplacian and Fisher Information A different question remains after finding a global minimizer: how close is the estimate to the truth? The Cramér-Rao Lower Bound and Fisher Information Matrix provide the answer. In a simplified PGO model with fixed rotations, Rosen-Khosoussi-Barfoot showed that FIM is exactly a Kronecker product of the graph's weighted reduced Laplacian. $$\mathcal{I} = \boldsymbol{J}^\top \boldsymbol{\Sigma}^{-1} \boldsymbol{J} = \boldsymbol{L}_w \otimes \boldsymbol{I}_3.$$ The graph structure alone yields an approximation of estimation accuracy without actual measurements. By Kirchhoff's Matrix-Tree Theorem, the determinant of the reduced Laplacian equals the number of weighted spanning trees, which corresponds to D-optimality (determinant of the information matrix). Algebraic connectivity (Fiedler value) corresponds to E-optimality (worst-case variance). By the 2020s, Kirchhoff's 1847 theorem for electrical circuits had become a theoretical basis for measurement selection and active SLAM. In active SLAM, "maximize FIM" translates into spectral manipulation of the Laplacian. [The post-2014 work of Kasra Khosoussi and Timothy Barfoot](https://arxiv.org/abs/1709.08601) established this connection. Khosoussi did his doctorate in Sydney under Dissanayake and Huang, then went through MIT and Toronto. In the form generalized to 3D PGO, the Kronecker combination of the Laplacian and SE(3) adjoint representation appears, letting topological and geometric information be handled separately. That the "measurement selection criterion" can be approximated by a Laplacian six times smaller than the full FIM provides the mathematical basis for "loop closure selection," which Ch.6 left in place without developing. The EKF-SLAM consistency problem in Ch.4 §4.8 also connects to this result. Read through the CRLB, the overconfidence Julier-Uhlmann identified in 2001 means that linearization overestimates Fisher information. Handbook §6.2 therefore places the FIM chapter next to convex relaxation. Finding the global minimum and estimating its accuracy form a pair. > 🔗 **Borrowed.** [Kirchhoff's Matrix-Tree Theorem (1847)](https://en.wikipedia.org/wiki/Kirchhoff%27s_theorem), born as an analysis tool for electrical circuits, was transplanted through combinatorics into measurement-design literature and, in the 2010s through Khosoussi, became the language of SLAM active perception. A 180-year migration path. --- ## 6b.5 Extensions and Limits — TEASER, CORA, and Lasserre's Wall After SE-Sync, work expanded in two directions: certifiable estimators robust to outliers, and extended measurement models for range, landmarks, and anisotropic noise. Outliers came first. As Ch.6 §6.7 noted, imperfect loop-closure verification introduces mismatches into real-world pose graphs, and optimization collapses above a certain outlier ratio even with Huber or Cauchy kernels. By around 2017, certifiable methods needed to address them. [Yang, Shi, Carlone's TEASER (TRO 2020)](https://arxiv.org/abs/2001.07715) is representative, finding the global optimum in 3D point-cloud registration with up to 99% outliers. It solves truncated least squares inside a GNC wrapper, with SDP relaxation on the rotation subproblem, and returns a certificate. The method splits scale, translation, and rotation into separate certifiable subproblems while preserving a global-optimality guarantee for each. The follow-up [Yang & Carlone (2022)](https://arxiv.org/abs/2109.03349) generalized this through Lasserre moment relaxation as "certifiably robust estimation." [Papalia et al.'s CORA (2024)](https://arxiv.org/abs/2302.11614) extended the line to range-aided SLAM. Used directly, the range measurement $(\|\boldsymbol{t}_j - \boldsymbol{t}_i\| - \tilde r_{ij})^2$ contains the square root implicit in the norm and is not directly quadratic. Papalia introduced an auxiliary unit vector $\boldsymbol{b}_{ij} \in S^{d-1}$ and used bearing lifting to recast it as one. CORA showed that the relaxation is tight in the single-robot case but generally not exact in multi-robot settings, narrowing the conditions under which Shor relaxation works. On the landmark side, [Holmes & Barfoot (2023)](https://arxiv.org/abs/2308.05631) used the Schur complement to eliminate landmarks in advance, leaving a PGO that SE-Sync can solve directly. Holmes, Khosoussi, and Rosen later co-authored Ch.6 of the Handbook in 2025. Limits also appeared. Generalizing anisotropic noise and truncated-quadratic outliers to a POP (Polynomial Optimization Problem) calls for Lasserre's moment relaxation, but the derived SDP is **degenerate**: constraint qualification fails and the Riemannian Staircase no longer converges. Yang's 2022 sparse monomial basis offers a workaround, but its specialized solver remains slower than a general local solver. No algorithm is yet both fast and certifiable. Visual SLAM and VIO face a deeper limit, the structural incompatibility of perspective projection and IMU preintegration, treated in the 🧭 section. > 📜 **Prediction vs. outcome.** At ICRA 2015 Carlone wrote that "theoretical explanation for why most instances have a tight Lagrangian dual is needed." Ten years later, only part of the answer had arrived. The exact-recovery theorem of Rosen-Carlone-Bandeira-Leonard gave a sufficient condition, "noise below $\beta$," but no way to compute $\beta$ in advance for an actual SLAM instance. As of 2026, no **a priori** condition predicts when tightness breaks; per-instance certificates serve in its place. --- ## 🧭 Still open **The boundary where tightness breaks.** SE-Sync's exact-recovery theorem offered the sufficient condition "noise below $\beta$," but there is no way to compute $\beta$ on an actual instance. An a priori test for tightness would guide algorithm design. Systematic study of how relaxation fails under heavy outliers or extremely sparse graphs remains limited. **Integration with visual SLAM and VIO.** Perspective projection $\pi(\boldsymbol{X}) = [X/Z, Y/Z]$ is rational, not polynomial. Multiplying through the denominator adds a new variable and auxiliary constraint per feature, and ORB-SLAM3's thousands of map points push the SDP beyond real-time scale. Forster's 2015 IMU preintegration tangles the exponential map with bias drift, resisting incorporation into POP. As of 2026, the visual/VIO mainstream of Ch.7, Ch.8, and Ch.13 sits outside certifiable guarantees. This is the lineage's largest unresolved gap. **Online certification and scale.** SE-Sync is batch. Incremental certifiable SLAM, which resolves the SDP at each new measurement, is not mature. As iSAM2 did for conventional SAM, certifiable SLAM needs an incremental formulation. Warm starts, incremental rank increases, and composition of partial certificates remain open, while moment-relaxation solvers are still too slow at city scale. **Outlier-majority.** Estimators such as TEASER have shown strong registration results with majority outliers. Those results do not guarantee identifiability or certification for arbitrary SLAM problems and contamination patterns. When several solutions remain plausible, multiple-hypothesis approaches such as list-decodable regression are also relevant. Cheng-Shi-Carlone pursued this direction around 2024, but no tool has become a standard comparable to TEASER. --- The brief discussion of local-minimum convergence in Ch.6 §6.7 developed into a ten-year theoretical program. Ch.6 of *The SLAM Handbook*, co-authored by Carlone-Khosoussi-Rosen-Holmes-Barfoot-Dissanayake, devotes 34 pages to the subject, evidence of the lineage's current weight. During the same decade, the learning-based SLAM of Ch.12, Ch.13, and Ch.16 followed another path: one lineage tried to prove that a solution was global, while the other asked a neural network to predict it. Whether they will meet remains unanswered in 2026. Ch.19 groups this chapter's 🧭 items under "gaps in backend theory." Ch.7 returns to the front end, treating the backend as given and asking what runs on top of it. --- # Ch.7 — Feature-based Lineage: The ORB-SLAM Trilogy Ch.6's graph SLAM work established pose graph optimization as the standard language of SLAM. Kümmerle's g²o (2011) and Kaess's iSAM2 (2012) made iterative optimization feasible on large maps and reduced loop-closure cost to a practical level. Part 3 therefore turns to the front end. Which features should a system extract, and how should it track them? When Klein and Murray split tracking and mapping into two threads with PTAM in 2007, the result was a lab demo that broke down beyond small indoor scenes. At the University of Zaragoza in 2015, Raúl Mur-Artal combined that structure with Rublee's ORB descriptor (2011), Gálvez-López's DBoW2 visual vocabulary (2012), and Strasdat's Essential graph idea (2011). PTAM was a fast prototype; ORB-SLAM became a ten-year standard. --- ## 7.1 ORB-SLAM (2015): A Tripod of Design Choices [Mur-Artal, Montiel & Tardós 2015. ORB-SLAM](https://doi.org/10.1109/TRO.2015.2463671), published in *IEEE Transactions on Robotics*, describes SLAM built on ORB features. Each major component reflects a separate design choice. The system runs three threads: Tracking, Local Mapping, and Loop Closing. PTAM had used two (Tracking and Mapping); Mur-Artal added Loop Closing as a third. This thread recognizes places with DBoW2, optimizes the pose graph over the Essential graph, and finally runs global bundle adjustment (BA). The separation keeps Tracking real-time without making it wait for map edits. > 🔗 **Borrowed.** PTAM's Tracking–Mapping split (Klein & Murray, 2007) carried directly into ORB-SLAM's Tracking–LocalMapping structure. Mur-Artal acknowledged the debt in §3 of the paper. ORB-SLAM added a third thread and isolated loop closure as an independent module. Mur-Artal chose the ORB (Oriented FAST and Rotated BRIEF) descriptor for specific reasons. SIFT and SURF carried patent restrictions, while BRIEF was fast but weak under rotation. ORB added rotation invariance to FAST keypoints; Rublee et al. presented it at ICCV 2011. The original experiments reported roughly two orders of magnitude, or about 100 times, higher speed than SIFT, and its binary representation allows matching by Hamming distance in real time on a CPU. ORB obtains scale invariance from an image pyramid. The original image is shrunk by a scale factor $s$ (1.2 in ORB-SLAM) over 8 levels, and FAST keypoints are detected independently at each level. An intensity centroid defines each keypoint's orientation: the first-order moment of pixel intensity gives the patch center, and the orientation angle $\theta$ rotates the BRIEF bit-comparison pairs. The result is a 256-bit rotation-invariant descriptor. XOR followed by popcount computes the Hamming distance between two descriptors. > 🔗 **Borrowed.** The descriptor from [Rublee et al. 2011. ORB](https://doi.org/10.1109/ICCV.2011.6126544) gave the system its name. The Zaragoza team did not design ORB; Mur-Artal assembled an existing set of tools into a pipeline. That front-end choice remained in the system's name for the next ten years. The keyframe-selection policy departs from PTAM's. PTAM added keyframes aggressively; ORB-SLAM removes redundancy using a covisibility graph. In the **covisibility graph**, edge weights count the landmarks shared between keyframes. Two keyframes connect when they share 15 or more landmarks. Local Mapping uses this graph to select a local window and runs BA only within it. On KITTI sequence 00 (a full 4.5 km loop), ORB-SLAM recorded 1.2% translation drift. PTAM, the comparison target at the time, could not close the large loop. Absolute scale remained ambiguous in both monocular systems. The Essential graph and DBoW2 allowed ORB-SLAM to recognize the loop and absorb the drift. The **Essential graph** is a subgraph of the covisibility graph. It keeps only edges with 100 or more shared landmarks, the spanning tree, and the loop-closure edges. When a loop is detected the whole graph is optimized as a pose graph. Even with thousands of keyframes the edges of the Essential graph stay sparse. Optimization finishes within seconds. > 🔗 **Borrowed.** The Essential graph idea came from the hierarchical optimization structure of [Strasdat et al. 2011. Double Window Optimisation](https://doi.org/10.1109/ICCV.2011.6126517). Strasdat separated a local window from a global window to cut optimization cost. Mur-Artal generalized this into a sparse pose graph called the Essential graph. Place recognition for loop closure is handled by DBoW2. [Gálvez-López & Tardós 2012. DBoW2](https://doi.org/10.1109/TRO.2012.2197158) is a vocabulary tree for binary descriptors. ORB descriptors are hierarchically clustered with k-medians (k-means++ seeding) to build a tree-structured vocabulary. Once the branching factor $k_w$ and depth $L_w$ are fixed, the number of leaf nodes (words) becomes $k_w^{L_w}$. The DBoW2 paper reports an example with $k_w=10$, $L_w=6$ trained into a vocabulary of one million words, and the public ORB-SLAM implementation uses a vocabulary of similar size. Each word carries a TF-IDF (Term Frequency–Inverse Document Frequency) weight: the more frequently a given word appears across the entire keyframe database, the lower its IDF weight, so discriminative words carry more influence. A keyframe is represented by this weighted BoW vector and stored in an inverted index. When a new frame arrives, descending the vocabulary tree to determine the word takes O(log(k^L))=O(L), and the inverted index pulls up candidate keyframes directly. The whole map is never traversed. The Tracking thread estimates the current pose in every frame. After feature matching with the previous frame, motion-only bundle adjustment refines $\mathbf{T}_{cw} \in SE(3)$. The basic reprojection objective over 3D–2D correspondences $\{(\mathbf{X}_i, \mathbf{u}_i)\}$ is: $$\mathbf{T}^* = \arg\min_{\mathbf{T}} \sum_i \left\| \mathbf{u}_i - \pi(\mathbf{T}\mathbf{X}_i) \right\|^2$$ Here $\pi$ is the camera projection function, $\mathbf{X}_i$ is a map point in world coordinates, and $\mathbf{u}_i$ is its image observation. Pose optimization uses a robust loss and observation weights, holding map points fixed while changing only the current camera pose. EPnP and RANSAC supply initial pose hypotheses during relocalization. The separate Local Mapping thread performs local BA over neighboring keyframes and map points. --- ## 7.2 ORB-SLAM2 (2017) — Stereo/RGB-D ORB-SLAM (2015) was monocular only. A single camera cannot recover scale: image pixels alone cannot distinguish a 10 m corridor from a 100 m one. Mur-Artal and Tardós returned to this problem in 2016. [Mur-Artal & Tardós 2017. ORB-SLAM2](https://doi.org/10.1109/TRO.2017.2705103) addresses the problem by adding stereo and RGB-D. Stereo has a known baseline and triangulates depth directly; RGB-D provides a measured depth value. Both recover metric scale. The structure is the same three threads as mono. Only the front end changes with the sensor type. Stereo extracts ORB from a rectified image pair and computes depth by left-right matching. Features matched across the pair provide **stereo observations**, while features seen in only one image provide **monocular observations**. Points with depth estimates are further classified as close or far using a threshold proportional to the baseline length. **Stereo initialization** runs immediately from the first frame, unlike monocular initialization. The monocular mode builds a map from the Essential Matrix or Homography between two frames and retains scale ambiguity. Stereo computes depth at the first keyframe from the horizontal disparity $d$ between left and right images, the baseline $b$, and focal length $f$: $$Z = \frac{b \cdot f}{d}$$ Feature points with depth $Z$ below the threshold $Z_{\max}=40b$ are registered as 3D map points at once. RGB-D initialization works on the same principle. The depth value $Z$ at pixel $(u, v)$ is read from the depth image, and back-projection yields the 3D coordinate. In both cases, because scale is fixed, Local BA can run right after the first frame. On the Machine Hall 01 sequence of the EuRoC MAV (Micro Aerial Vehicle) dataset, ORB-SLAM2 (stereo) recorded an absolute translation error of 0.035 m in Table II. The same table uses Stereo LSD-SLAM as the comparison target, showing lower error for ORB-SLAM2 under those evaluation conditions. ORB-SLAM2 also ranked near the top among methods then published on KITTI odometry. On the day in May 2017 when the paper appeared in IEEE TRO, Mur-Artal and Tardós pushed the source to GitHub alongside it. Two people in the Zaragoza team released mono, stereo, and RGB-D modes on a single codebase. GitHub stars passed several thousand afterward, and ROS wrappers came out of the community. --- ## 7.3 ORB-SLAM3 (2021): Atlas and Visual-Inertial [Campos et al. 2021. ORB-SLAM3](https://doi.org/10.1109/TRO.2021.3075644), published in IEEE Transactions on Robotics in 2021, has a different author list. The first author is not Mur-Artal but Carlos Campos. Mur-Artal is listed as a coauthor alongside Tardós. Campos had done his PhD at the University of Zaragoza under Tardós. The lineage moved down a generation. ORB-SLAM3 added two core extensions: **Atlas** (multi-map) and **Visual-Inertial** mode. Atlas holds several separate maps simultaneously. When tracking fails, the existing map is suspended and a new one starts; if the system later revisits the same place, it merges the maps. ORB-SLAM and ORB-SLAM2 already attempted relocalization in the existing map. When recovery failed and a new map was started, however, they lacked Atlas's ability to retain and merge separate maps. ORB-SLAM3 presents Atlas as a response to that limitation. ORB-SLAM3 reinitializes after failure while retaining the previous map. Visual-Inertial (VI) mode integrates IMU data. Campos adopted the formulation that Forster et al. proposed at RSS 2015 as "IMU Preintegration on Manifold" and extended in IEEE TRO 2016 as [On-Manifold Preintegration for Real-Time Visual-Inertial Odometry](https://doi.org/10.1109/TRO.2016.2597321). The IMU bridges rapid motions that can make visual tracking fail. VI-SLAM also resolves a monocular camera's scale ambiguity: accelerometer measurements provide absolute scale together with the direction of gravity. The IMU measurements between keyframes $i$ and $j$ are integrated once. With accelerometer and gyroscope readings modeled as $\tilde{\mathbf{a}}_t = \mathbf{a}_t + \mathbf{b}_a + \mathbf{n}_a$ and $\tilde{\boldsymbol{\omega}}_t = \boldsymbol{\omega}_t + \mathbf{b}_g + \mathbf{n}_g$, where $\mathbf{b}$ is bias and $\mathbf{n}$ is noise, the relative rotation, velocity, and position increments are $$\Delta\mathbf{R}_{ij} = \prod_{k=i}^{j-1} \mathrm{Exp}\bigl((\tilde{\boldsymbol{\omega}}_k - \mathbf{b}_g)\Delta t\bigr)$$ $$\Delta\mathbf{v}_{ij} = \sum_{k=i}^{j-1} \Delta\mathbf{R}_{ik}\,(\tilde{\mathbf{a}}_k - \mathbf{b}_a)\Delta t$$ $$\Delta\mathbf{p}_{ij} = \sum_{k=i}^{j-1}\!\left[\Delta\mathbf{v}_{ik}\Delta t + \tfrac{1}{2}\Delta\mathbf{R}_{ik}\,(\tilde{\mathbf{a}}_k - \mathbf{b}_a)\Delta t^2\right].$$ Here $\mathrm{Exp}(\cdot)$ is the exponential map of $\mathfrak{so}(3)$. When bias shifts during BA, a first-order Jacobian correction avoids reintegration. ORB-SLAM3 adds these preintegrated terms as inertial edges in the factor graph and optimizes them jointly with the visual reprojection residual. Ch.7b gives the full derivation, from Lupton's Euler-angle attempt to Forster's manifold formulation. > 🔗 **Borrowed.** Campos used the Forster et al. On-Manifold Preintegration formulation (TRO 2016, originating at RSS 2015) as the core of ORB-SLAM3's inertial integration. Forster's formulas integrate continuous IMU measurements on the SO(3) manifold with bias correction. ORB-SLAM3 incorporated the formulation into factor graph optimization. On the mean RMSE ATE (Absolute Trajectory Error) across all 11 EuRoC MAV sequences, ORB-SLAM3 (mono-inertial) is reported at 0.043 m in Table II. In the same table VINS-Mono comes in at 0.110 m, and Kimera (stereo-inertial) at 0.119 m. Combined, VI mode and Atlas let a UAV or handheld device return to a previous map after lighting changes or lost tracking. These additions changed the character of the system, not only its version number. --- ## 7.4 Why It Is Still the Baseline in the 2020s In 2023, conference papers still included ORB-SLAM3 in comparison tables. New methods reported how much they improved on it. Although the algorithm had been stable since 2021, its benchmark role persisted. ORB features remain reasonably stable under illumination change, the binary descriptor is fast to compute, and a system can extract many of them in real time to reduce tracking failures. Learned features are more accurate on some datasets but can fail in new environments. ORB's behavior is more predictable. Reproducibility also matters. The code is public, ROS integration is solid, and thousands of real-world use cases are documented. Labs routinely run ORB-SLAM3 first when evaluating a new system. Because one codebase supports mono, stereo, RGB-D, and IMU, it provides a common baseline across several settings. The learned alternatives do not beat it consistently. DROID-SLAM (Teed & Deng, 2021) beats ORB-SLAM3 on several sequences. But as the paper itself reports, large sequences such as EuRoC and TartanAir need a 24 GB-class GPU, and on TartanAir it runs at 8 fps, not real time. ORB-SLAM3, by contrast, runs CPU-only, and community reports confirm basic operation on ARM and embedded platforms. --- ## 📜 Prediction vs. outcome > 📜 **Prediction vs. outcome.** Mur-Artal laid out two Future Work directions in Section IX-C of the 2015 ORB-SLAM paper. "Points at Infinity" proposed using distant points that lack parallax, and therefore cannot serve as ordinary map points, to estimate rotation. "Dense Map Reconstruction" suggested that compact keyframe selection could provide a skeleton for dense reconstruction. Ten years later, VI-SLAM and follow-on work had partly absorbed the first direction. The NeRF-SLAM and Gaussian Splatting line of the 2020s revisited the second with a "sparse skeleton + dense overlay" structure in different representations. The modality extensions the authors also identified (RGB-D, stereo, IMU) appeared in ORB-SLAM2 (2017) and ORB-SLAM3 (2021) under separate problem statements. > 📜 **Prediction vs. outcome.** In the Conclusions of the 2021 ORB-SLAM3 paper, Campos et al. identified low-texture environments as the system's main failure mode and proposed photometric techniques suited to the four data-association problems, citing endoscopic imagery as one example. Between 2023 and 2025, the community concentrated more heavily on learned front ends such as SuperPoint and LightGlue, while photometric integration continued separately in DSO and LDSO. The official ORB-SLAM3 repository still uses the traditional ORB descriptor in its main branch as of 2026. The authors' photometric direction and the community's learned-feature work diverged. --- ## 🧭 Still open Long-term map reuse. Atlas made multi-map maintenance possible, but map merging still fails under large lighting changes. A morning map and an evening revisit should merge as the same place, yet DBoW2 misses when appearance changes substantially. Groups working on long-term outdoor autonomy across seasonal change continue to study the problem. As of 2024, there is no complete answer. The place of the pure-vision baseline. Learned-feature systems have begun to beat ORB-SLAM3 on standard benchmarks. SuperPoint + SuperGlue, LightGlue, and DINOv2-based features show lower error on particular sequences, but generalization remains separate. Outside the training distribution, learned features sometimes perform worse than traditional ORB. Existing experiments are not broad enough to support a claim of consistent superiority. Drift at large outdoor scale. ORB-SLAM3 still lags LiDAR SLAM on urban driving and paths beyond several kilometers. Urban-scale localization in GPS-denied environments with a pure camera remains unsolved as of 2026. When changes in visual conditions, dynamic objects, and textureless stretches combine, drift accumulates. The gap to LiDAR survey precision is narrowing but has not closed. --- During the years when the ORB-SLAM trilogy set the feature-based standard, Newcombe and Engel took the opposite approach: use image brightness directly instead of extracting feature points. The two lineages developed side by side through the 2010s, and comparisons exposed their respective limits. ORB-SLAM3 led the EuRoC benchmark in 2021, while DSO beat ORB-SLAM2 in the TUM corridors. They followed the same timetable from different starting points. --- # Ch.7b — From a Shaking Sensor to a Constraint: The Invention of IMU Preintegration Sydney, 2009. At ACFR (the Australian Centre for Field Robotics), doctoral student Todd Lupton was working through a problem with his advisor, Salah Sukkarieh. When a drone moves aggressively, the IMU produces measurements at 200 Hz, too many to place individually in a factor graph. Keyframes arrive only a few times per second; how can the dozens or hundreds of IMU measurements between them become a single unit? Lupton's answer at IROS became the seed of preintegration. Six years later, at RSS 2015, Christian Forster, Davide Scaramuzza, Luca Carlone, and Frank Dellaert carried that seed onto the SO(3) manifold, making the IMU a first-class citizen of the factor graph. Those equations sit behind the line "we used Forster 2016" in Ch.7's ORB-SLAM3, Ch.8's VI-DSO, and Ch.17's LIO-SAM. FAST-LIO instead propagates raw IMU measurements and performs iterated filter updates. --- ## 7b.1 MEMS and the "democratization of sensors" Preintegration became necessary because IMUs became cheap. Strapdown inertial navigation has its roots in 1950s aerospace. Ring laser gyros for submarines and missiles cost tens of thousands of dollars, beyond the reach of most robotics laboratories. MEMS (Micro-Electro-Mechanical Systems) changed that. Analog Devices' ADXL and InvenSense's MPU series reduced the price of a six-axis IMU to a few dollars. The iPhone received an IMU in 2007, and by the early 2010s research drones and handheld devices carried MEMS units as a matter of course. This price collapse, driven by billions of smartphones, coincided with Visual SLAM's growing concern with monocular scale ambiguity (Ch.5 §🧭). The measurement model is simple. The accelerometer gives the specific force $\tilde{\mathbf{a}} = \mathbf{R}_w^b(\mathbf{a}^w - \mathbf{g}^w) + \mathbf{b}^a + \boldsymbol{\eta}^a$ with gravity included, and the gyroscope gives the angular velocity $\tilde{\boldsymbol{\omega}} = \boldsymbol{\omega}_b^b + \mathbf{b}^g + \boldsymbol{\eta}^g$. Here $\mathbf{b}$ is bias and $\boldsymbol{\eta}$ is white noise. Three consequences matter: gravity is always present, bias drifts slowly over time (random walk), and MEMS noise is high-frequency. An estimator must account for gravity and often aligns its world frame with that direction. It also needs a bias model that accounts for changes in temperature and power state. --- ## 7b.2 First attempt — Lupton & Sukkarieh (2009 / 2012) The problem was the factor graph's time axis. Kaess's iSAM2, covered in Ch.6, takes keyframe-rate poses as nodes, while the IMU produces dozens of measurements between keyframes. Making every measurement a node makes the graph unwieldy; discarding them loses information. In [Visual-Inertial-Aided Navigation for High-Dynamic Motion (IROS 2009, TRO 2012)](https://doi.org/10.1109/TRO.2011.2170332), Lupton and Sukkarieh integrated the IMU measurements between keyframes $i$ and $j$ *once* to build a relative increment. Treating that increment as one factor kept the raw measurements out of the graph. The name "pre-integration" came from this construction. Two obstacles limited the implementation. It represented rotation with Euler angles, which suffer from gimbal lock and do not form a manifold. Bias posed the larger problem. A changed bias estimate changes the increment, and full reintegration would require reprocessing hundreds of measurements per keyframe. Lupton already proposed first-order bias correction to reduce that work. Six years later, Forster formulated the correction on SO(3) while accounting for the manifold structure of rotation. --- ## 7b.3 The decisive turn — Forster-Carlone (2015 / 2017) At RSS 2015, Christian Forster, a doctoral student at ETH Zürich, joined Scaramuzza (UZH), Carlone (Georgia Tech, later MIT), and Dellaert (Georgia Tech, creator of GTSAM) to publish [IMU Preintegration on Manifold for Efficient Visual-Inertial Maximum-a-Posteriori Estimation](https://www.roboticsproceedings.org/rss11/p06.pdf). The extended version appeared in IEEE TRO 2017 as [On-Manifold Preintegration for Real-Time Visual-Inertial Odometry](https://doi.org/10.1109/TRO.2016.2597321). The paper united UZH's agile-drone experiments, Georgia Tech's GTSAM factor-graph language, and Carlone's optimization theory. They made three changes. First, they defined $\Delta\mathbf{R}_{ij}$ rigorously as a relative rotation on the SO(3) manifold and redefined $\Delta\mathbf{v}_{ij}, \Delta\mathbf{p}_{ij}$ to make them *independent of gravity and the initial state*. These are not physical increments but mathematically state-independent quantities, allowing the IMU factor to be evaluated without repeating preintegration. Its residual still depends on endpoint poses and velocities, gravity, and bias correction. Second, they propagated the covariance $\boldsymbol{\Sigma}_{ij}$ analytically with a right-Jacobian construction that moves noise to the end of the exponential map. The third change proved decisive: **linear correction via the bias first-order Jacobian**. When the bias shifts during BA, a first-order correction using precomputed partial derivatives replaces full reintegration. It is Lupton's Euclidean linearization applied on SO(3). The Jacobian is computed during the first integration between keyframes and reused through hundreds of graph-optimization iterations. A reintegration taking several milliseconds became a Jacobian-vector product taking several microseconds, allowing the IMU factor to enter real-time BA. Adoption accelerated when Forster's implementation entered GTSAM as a reference. Later systems did not rewrite the equations; they included `ImuFactor`. > 🔗 **Borrowed.** Forster's manifold preintegration uses the SO(3) right-Jacobian formalism later organized in [Barfoot 2017. *State Estimation for Robotics*](https://doi.org/10.1017/9781316671528). Exponential maps and Jacobians for small rotational variations were already the common language of robotic state estimation; Forster rewrote IMU preintegration in that language. Moving Lupton's Euler-angle formulation onto SO(3) removed the earlier limitation. > 🔗 **Borrowed.** [Lupton & Sukkarieh 2012](https://doi.org/10.1109/TRO.2011.2170332) first proposed the bias first-order Jacobian. Forster et al. TRO 2016 §VIII-B acknowledges the debt explicitly: "we follow [Lupton-Sukkarieh] but operate directly on SO(3)." Moving the Euclidean approximation onto the manifold made the computation real-time. --- ## 7b.4 The three schools of practical VIO Once Forster's formulation settled in, Visual-Inertial Odometry (VIO) systems branched into three lines between 2017 and 2022. The first line is the filter family, whose roots precede Forster. At ICRA 2007, UC Riverside's Anastasios Mourikis and Stergios Roumeliotis introduced the [MSCKF (Multi-State Constraint Kalman Filter)](https://doi.org/10.1109/ROBOT.2007.364024). It retains past camera poses in the filter state through stochastic cloning. Observed 3D points are not added to the state; projecting the linearized observation equations onto the left null space of the feature Jacobian removes their dependence on feature-position errors. This was an early influential real-time visual-inertial system built on an EKF without preintegration. In 2021, NASA JPL's Mars helicopter Ingenuity used an MSCKF-family estimator on Mars. Guoquan Huang's group at the University of Delaware open-sourced [OpenVINS](https://doi.org/10.1109/ICRA40945.2020.9196524) in 2020. The second line is the optimization family. Its representative system is [VINS-Mono](https://doi.org/10.1109/TRO.2018.2853729), published in TRO 2018 by Shaojie Shen's HKUST group with doctoral student Tong Qin. They placed Forster's formulation as an IMU factor inside tightly coupled sliding-window BA and provided a procedure for estimating scale and gravity direction separately during initialization. Released as code, it became the conference VIO baseline from 2019 to 2022. When Ch.7's ORB-SLAM3 reported an average ATE of 0.043 m over the eleven EuRoC sequences, VINS-Mono was the comparison method at 0.110 m in the same table. The third line is the direct family: VI-DSO (2018), Basalt (2019), and [DM-VIO (2022)](https://doi.org/10.1109/LRA.2021.3140129), covered in Ch.8. TUM's Cremers group placed Forster's inertial factor on top of DSO's photometric BA. DM-VIO added *delayed marginalization*. Premature marginalization before IMU initialization converges can lock in a wrong prior and cause long-term drift, so the method maintains two marginalization priors in parallel and merges them only after gravity and scale become observable. --- ## 7b.5 Observability — what cannot be seen Visual-Inertial systems do not see everything. Analyses by Huang's group from the early 2010s converged on one conclusion: the null space of an unanchored visual-inertial system is **four-dimensional**. It contains three dimensions of global position and one dimension of yaw around gravity. IMU and camera alone can never recover absolute coordinates or rotation about the gravity axis. GPS restores position; a magnetic field or external anchor restores yaw. Pure VIO cannot observe this four-dimensional subspace. *Roll and pitch*, however, are observable because the accelerometer provides a vertical reference through gravity. Adding an IMU also resolves the monocular scale ambiguity identified in Ch.5. Degenerate motion is harder. Global yaw remains unobservable regardless of motion. Insufficient parallax or variation in inertial motion can make depth, scale, and bias harder to distinguish. Hovering or driving straight at constant speed warrants checking for this lack of excitation. Takeoff, braking, and turning can add information, but no single maneuver guarantees a reliable scale estimate. > 📜 **Prediction vs. outcome.** Forster et al. in TRO 2017 §IX named three directions: integrating time synchronization with online extrinsic calibration, validating the bias random-walk assumption under long-term operation, and extending to asynchronous sensors such as event cameras and rolling shutter. As of 2026, the first has been standardized as VINS-Mono, Kalibr, and OpenVINS put the time offset onto the state vector; the second holds for navigation-grade IMUs but remains affected by temperature and power-supply variation on consumer MEMS; the third has found one branch of the answer in Le Gentil's GP continuous-time preintegration. The predictions were largely on target, but instead of the single extension the authors sketched, the line split into three. --- ## 7b.6 The branch into continuous-time At RSS 2021, Cédric Le Gentil and his advisor Teresa Vidal-Calleja at UTS (University of Technology Sydney) released [Continuous Integration over SO(3) for IMU Preintegration](https://roboticsproceedings.org/rss17/p075.pdf). Again in Sydney, only a few kilometers from Lupton's ACFR, they approached the same problem from another angle. Forster's preintegration uses discrete time. It assumes piecewise-constant IMU measurements between samples and applies Euler integration. Fusing asynchronous sensors such as LiDAR or event cameras requires evaluating states between IMU sampling times. Asynchrony itself does not invalidate the piecewise-constant approximation, but makes accurate temporal interpolation important. Le Gentil instead modeled the IMU as a **Gaussian Process**, treating angular velocity as a continuous function. The state can then be evaluated at any time $\tau$, allowing asynchronous measurements to enter naturally. This direction intersects the B-spline, STEAM, and GPMP lineages. --- ## 7b.7 Terrain of borrowings > 🔗 **Borrowed.** The skeleton that evaluates and optimizes an IMU factor on a factor graph is the [Dellaert GTSAM](https://gtsam.org/) tradition from Ch.6 unchanged. Forster's `ImuFactor` plugs into GTSAM's `NoiseModelFactor` interface and is optimized inside a single `Values` object alongside visual reprojection factors. The inheritance was of software structure, not of mathematics. > 🔗 **Borrowed.** The practice of treating bias as a random walk comes from the Kalman-filter state-propagation convention Ch.4 recorded. Well before Lupton, the navigation community used a model that "puts the bias in the state and gives it small process noise," and in the preintegration era this was reinterpreted as the bias random-walk factor. --- ## 🧭 Still open **Real-time detection of visual-inertial observability.** The four-dimensional null space and degenerate-motion table are theoretically settled, but runtime detection remains unfinished. The FEJ (First-Estimate Jacobian) line of Hesch, Li, and Huang preserves the null space at the linearization point, yet as of 2026 no broadly agreed method detects entry into and exit from a degenerate regime and feeds that information back to the control loop. When drone control and VIO estimation share compute resources, the controller may detect estimator failure too late. **Unification of preintegration and continuous-time.** Forster's discrete increments and Le Gentil's continuous GP representation solve the same problem in different mathematical languages. When combining LiDAR, event cameras, and frame cameras, the representation that should underpin the estimator remains an engineering choice. B-spline continuous-time BA offers a partial answer, but most deployed systems still use Forster's discrete factor. **Learning-based IMU bias models.** The bias random-walk assumption holds on navigation-grade IMUs, but consumer MEMS depart from it because of temperature hysteresis and power-supply transients. The TLIO and RoNIN line used LSTMs and Transformers to learn bias models for IMU-only odometry; more recent work models the bias distribution itself with conditional diffusion. How this approach fits inside a Forster factor, and how much of preintegration's mathematics remains when a learned model supplies the bias dynamics, is the next question. --- Lupton's idea remained limited by Euler angles for six years. Forster moved it and its bias-Jacobian correction to SO(3); Le Gentil, again in Sydney, extended it into continuous time. Three generations of work sit behind one line in ORB-SLAM3, one in VI-DSO, and one in LIO-SAM. Ch.7c follows the continuous-time branch, while Ch.8 continues the visual lineage with direct methods (DSO and VI-DSO) that use Forster's preintegration factor. --- # Ch.7c — When Time Must Flow Smoothly: Continuous-Time Trajectory Ch.7b's preintegration compressed IMU measurements into a relative factor between discrete keyframes. That compression presumes discrete endpoints: to fold the hundred inertial samples between two keyframes into one factor, each endpoint must have a definite timestamp. The assumption is harmless when a camera shutter opens and closes globally once per frame. Asynchronous sensors break it. In 2012, in Toronto, [Paul Furgale, Timothy Barfoot, and Gabe Sibley](https://asrl.utias.utoronto.ca/~tdb/bib/furgale_iros12.pdf) formalized the question in an IROS paper. In an image captured by a rolling shutter, each row is projected from a different pose in time. A vehicle travels several meters while a spinning LiDAR completes one rotation, and the IMU produces samples at 1 kHz while the camera runs at 30 Hz. Furgale, Barfoot, and Sibley represented pose as a function of time $t$ rather than a frame and chose the B-spline. That choice became the standard starting point for continuous-time trajectory estimation. Ten years later, the Handbook placed this branch alongside manifolds as one of SLAM's "two fundamental tools." Discrete keyframes cover most visual-inertial systems; continuous-time trajectory estimation addresses sensors that do not fit that clock. --- ## 7c.1 Limits of discrete-time Ch.7b's preintegration handles one mismatch: the IMU runs faster than the camera. It does not address four others. First, rolling shutter. A consumer CMOS camera reads one frame from top to bottom over tens of milliseconds. In a fast-moving camera, the first and last rows are captured from different poses. This distortion falls outside the photometric-consistency model assumed by Ch.8's DSO and LSD-SLAM. The Cremers group therefore added a B-spline trajectory to [Basalt](https://arxiv.org/abs/1904.06504) in 2019. Second, spinning LiDAR motion distortion. As Ch.17 notes, the Velodyne HDL-64E completes one rotation at 10 Hz. If a vehicle moves at 10 m/s during those 100 ms, the vehicle travels a total of 1 m during one scan, and each point is captured from a different pose along that interval. LOAM corrected the distortion indirectly inside the odometry loop; a continuous trajectory instead provides the pose at the instant each point was captured. Third, event cameras. The DVS described in Ch.18 produces asynchronous events at μs granularity per pixel. Events have no frame, so [Mueggler et al. 2015](https://arxiv.org/abs/1502.00796) formulated event SLAM on an SE(3) B-spline trajectory. Fourth, a high-rate IMU may be fused with several sensors running at different frequencies. When a system ingests a 200 Hz IMU, a 20 Hz camera, and a 10 Hz LiDAR, placing a discrete state node at every measurement time is impractical. The factor graph swells when the number of states grows with the number of measurements. The four problems share one structure: measurement time $t_i$ is not controlled. Observations arrive asynchronously, and the estimator must know the pose at each arrival. A continuous-time representation separates measurement time, estimation time, and query time. --- ## 7c.2 Parametric spline: the Furgale line Furgale, Barfoot, and Sibley chose the B-spline in 2012. The trajectory is written as a sum of basis functions, $\mathbf{p}(t) = \sum_k \Psi_k(t)\,\mathbf{c}_k$, with coefficients $\mathbf{c}_k$ as the optimization variables. Local support is the defining property: at any time $t$, only a handful of bases (usually four) are nonzero. Querying the pose at an arbitrary time $t_i$ therefore has constant cost and preserves factor-graph sparsity. > 🔗 **Borrowed.** The mathematical skeleton of the B-spline comes from [de Boor's 1978 *A Practical Guide to Splines*](https://link.springer.com/book/10.1007/978-1-4612-6333-3). Furgale lifted that structure onto SE(3) and placed the coefficients as variable nodes in the factor graph, transplanting a tool from numerical analysis into SLAM optimization. The trade-offs were clear. Closely spaced coefficients overfit, while wide spacing misses fast motion, and the choice depended on experience. A linear B-spline applied directly to SE(3) also produces interpolated poses outside the manifold. In 2013, Oxford's [Steven Lovegrove et al.](https://www.roboticsproceedings.org/rss09/p11.html) proposed the cumulative B-spline. Rearranging the basis as a cumulative product rather than a sum, $T(t) = \prod_k \exp\bigl(\tilde\Psi_k(t) \log(T_k T_{k-1}^{-1})\bigr) \cdot T_0$, keeps each factor on the Lie group. This became a standard representation in later rolling-shutter, event-camera, and VIO papers. Basalt, [Mueggler's event SLAM](https://arxiv.org/abs/1502.00796), and [Kerl et al. 2015 dense rolling shutter VO](https://doi.org/10.1109/ICCV.2015.172) all used the cumulative B-spline. The parametric spline remains common in real-time VIO and event systems because computation is light and the code is simple. Without a separate motion prior, smoothness in sparsely observed intervals depends on the basis and control points. Priors can be added to coefficients or derivatives; the GP-based branch instead starts by specifying a prior distribution over trajectories. --- ## 7c.3 SDE-based GP: the Barfoot line and STEAM In 2014, the Barfoot group in Toronto opened a second branch with ["Batch Continuous-Time Trajectory Estimation as Exactly Sparse Gaussian Process Regression"](https://www.roboticsproceedings.org/rss10/p01.pdf). Barfoot, Tong, and Särkkä treated the trajectory as a Gaussian process rather than a basis sum. A kernel $\mathcal{K}(t, t')$ defines the prior, and the posterior remains a conditional Gaussian when observations arrive. A dense GP has one problem: with $N$ observations, inverting the kernel matrix $K$ costs $O(N^3)$. Barfoot, Tong, and Särkkä identified a family of kernels that avoids this cost. When the trajectory is the solution of a linear time-invariant stochastic differential equation $\dot{\mathbf{x}}(t) = A\mathbf{x}(t) + L\mathbf{w}(t)$, the inverse $K^{-1}$ of its kernel has a block-tridiagonal structure. In factor-graph terms, binary factors connect only consecutive state nodes, not distant ones. > 🔗 **Borrowed.** Expressing an SDE-derived GP motion prior as a factor graph follows the SDE-GP connection in [Särkkä's 2013 *Bayesian Filtering and Smoothing*](https://users.aalto.fi/~ssarkka/pub/cup_book_online_20131111.pdf), which the Barfoot group brought into SLAM. The Rasmussen-Williams GP textbook writes the kernel in closed form, but real-time SLAM needs a sparse inverse. Särkkä's SDE representation supplied the bridge. The result was **STEAM** (Simultaneous Trajectory Estimation and Mapping). At RSS 2015, [Sean Anderson and Barfoot 2015, "Full STEAM Ahead"](https://www.roboticsproceedings.org/rss11/p45.pdf) formalized STEAM with a constant-velocity prior. It augments the state with pose $\mathbf{p}(t)$ and velocity $\mathbf{v}(t)$, with pose following from the white-noise integral of velocity. Anderson tightened the sparsity proof that same year, and it became the foundation of the Barfoot group's later continuous-time papers. STEAM's second advantage was GP interpolation. With only a small number of control poses, the estimator can query any intermediate pose as the posterior mean. Even when a spinning LiDAR captures 10,000 points at 10,000 instants within one scan, the model uses only one control point per scan. The number of estimated states can remain much smaller than the number of observations. Evaluating and accumulating the observation residuals still requires computation. In 2019, Tang and Barfoot's [open-source STEAM release](https://github.com/utiasASRL/steam) gave academia and industry a directly usable library. The same year, the Dellaert group's GTSAM received a GP continuous-time factor in contrib. The two paths had converged. --- ## 7c.4 Continuous-time on the Lie group Whether parametric or nonparametric, SLAM needs a trajectory on SE(3). Lifting a Euclidean spline or GP onto the group is not straightforward. The common approach works in the tangent space: interpolate linearly there, then return the result to the manifold with the exponential map. On the B-spline side, [Sommer, Demmel et al. 2020, "Efficient Derivative Computation for Cumulative B-Splines on Lie Groups"](https://arxiv.org/abs/1911.08860) derived the SE(3) cumulative-spline Jacobian in closed form. The CVPR paper supplied a standard B-spline trajectory formulation with real-time derivatives for rolling-shutter VIO, event cameras, and visual-inertial systems. Basalt and later work from the Cremers group used this result. On the GP side, Anderson and Barfoot proposed a "local variable" construction. Near each control pose $T_k$, they define a local perturbation $\xi_k(t) = \log(T(t)\,T_k^{-1})$ and run the GP on it. A GP is difficult to define directly on the global manifold, but a Euclidean GP can be defined in the tangent space around each control point. Crossing between control points introduces an adjoint, the same Lie-group operation that appears in Ch.7b's on-manifold preintegration. From 2015 onward, both tools used a common Lie-group grammar. > 🔗 **Borrowed.** [Anderson-Barfoot 2015 ICRA](https://doi.org/10.1109/ICRA.2015.7138984) systematically developed the use of a GP in a Lie-group local variable. Several later continuous-time LiDAR and VIO papers used the same construction: run a GP between two consecutive control points and apply the adjoint when crossing between them. One difference between the spline and GP implementations compared here is how they specify the motion prior. Spline coefficients can be estimated directly or given priors, as can their derivatives. A GP carries an SDE-derived prior, such as constant velocity or white jerk. Where observations are sparse, the prior supports the GP, while the spline relies on neighboring observations. Attempts to combine them (Johnson et al. 2020) have appeared, but the choice remains application-dependent. --- ## 7c.5 The line descends to applications: LiDAR and VIO The range of applications expanded beyond early calibration work. Around 2022, continuous-time became a common solution in three areas. First, LiDAR motion distortion. Paris's [Pierre Dellenbach et al. 2022, "CT-ICP"](https://arxiv.org/abs/2109.12979) parameterized each scan with two poses (a "start pose" and an "end pose") and interpolated linearly between them. Despite its simple model, CT-ICP beat prior LOAM and FAST-LIO accuracy on the KITTI, NCLT, and Newer College benchmarks. The same year, Toronto's [Keenan Burnett et al. 2022, "Are We Ready for Radar to Replace Lidar?"](https://arxiv.org/abs/2206.05432) and [STEAM-ICP](https://github.com/utiasASRL/steam_icp) applied GP-based continuous time to the Aeva FMCW LiDAR. The sensor reports Doppler velocity with each point, which maps directly to STEAM's velocity state. Without a continuous-time representation, that information could not enter the estimator directly. Second, rolling-shutter VIO. Basalt, the [Cremers group rolling-shutter VO](https://doi.org/10.1109/CVPR.2016.71), and follow-ups to [OKVIS](https://doi.org/10.1177/0278364914554813) query each image row's capture time on a B-spline trajectory. Rather than assuming a global shutter, they model the rolling shutter directly. Third, event cameras. After the 2010s difficulties described in Ch.18, several lines of event SLAM in the 2020s used continuous-time trajectories. Each event's μs timestamp is queried against a B-spline or GP to obtain the pose at that instant, and event-image consistency supplies the residual. Continuous-time trajectories fit event cameras because neither assumes frames. > 🔗 **Borrowed.** CT-ICP is a combination that lays intra-scan continuous-time linear interpolation on top of a point-to-plane objective from the ICP family descended from [Besl and McKay 1992 ICP](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf). Classic registration and Furgale's continuous-time spirit met inside one system, thirty years apart. --- ## 📜 Prediction vs. outcome > In the Future Work of their 2012 IROS paper, Furgale, Barfoot, and Sibley wrote two expectations. One was that "continuous-time representation will become the natural language for unifying rolling shutter and high-rate IMU sampling"; the other was "follow-up work proving compatibility with a sparse factor graph." Both were realized within a decade. Barfoot, Tong, and Särkkä 2014 completed the sparse GP proof, while rolling-shutter VIO and event SLAM of the 2020s use the cumulative B-spline as their standard representation. The authors did not anticipate one development. In 2012, the implicit division of labor placed discrete-keyframe ORB-SLAM in the mainstream and continuous time with specialty sensors. Development also moved in the opposite direction: when Burnett released STEAM-ICP using Doppler velocity from an FMCW LiDAR, continuous time became a way to exploit sensor-specific signals. --- ## 🔗 Borrowed (summary) Three further lineages underpin the chapter. Särkkä's SDE-GP textbook anchored the equations of Barfoot, Tong, and Särkkä 2014. De Boor's 1978 spline text supplied Furgale 2012 with its basis functions, and Anderson and Barfoot's 2015 local-variable technique brought GPs onto Lie groups. Continuous-time trajectory estimation brings numerical analysis, probability theory, and Lie-group differential geometry together within SLAM. --- ## 🧭 Still open **Learning-based continuous-time prior.** The motion prior that an SDE provides embeds physical assumptions such as constant-velocity or white-jerk. Real driving, walking, and UAV trajectories often violate these assumptions. In 2023-2024, attempts appeared to learn data-driven priors with neural SDE or neural ODE and plug them into the continuous-time factor graph. A system that layers a learned prior while keeping real-time sparse structure is still at the validation stage. **Integration of VIO and continuous-time.** Ch.7b's preintegration remains the dominant formulation in keyframe-based VIO. As of 2026, no consensus exists on whether continuous-time trajectories should replace preintegration or coexist with it in a hybrid. Le Gentil's [GP-augmented preintegration line](https://arxiv.org/abs/2007.04144) connects the two, but deployed systems such as ORB-SLAM3 and VINS-Fusion still rely on discrete-time preintegration. **Online sliding window for edge deployment.** STEAM- and B-spline-based systems slow as control points accumulate. Marginalizing past control points while preserving the consistency of the continuous-time posterior remains technically difficult. Continuous-time SLAM cannot become a long-lived standard on embedded platforms such as cars and drones without solving this problem. --- Ch.7b optimized discrete-time preintegration; this chapter followed the continuous-time alternative. The two tools need not compete. Since 2024, systems have placed an IMU preintegration factor and a continuous-time LiDAR factor side by side in one SLAM estimator. Ch.8 returns to the visual line, where DSO and VI-DSO use the Forster factor and the direct photometric approach develops separately from both preintegration supplements. --- # Ch.8 — The Direct Lineage: From DTAM to DSO Richard Newcombe was Andrew Davison's doctoral student. After seeing MonoSLAM's 30-landmark ceiling firsthand at Imperial College, he took the opposite approach in 2011: use every pixel. Davison had demonstrated real-time operation with the EKF premise that tracking a few points was sufficient. With a single GPU, Newcombe showed that a system could also operate in real time on the full image. DTAM descended directly from MonoSLAM while reversing its methodological choice. The ORB-SLAM lineage in Ch.7 extracted features first and tracked only those features. Harris corners and ORB descriptors reduced an image to a few hundred points; the remaining pixels were discarded. Direct methods use pixel-intensity differences as measurements rather than geometric errors between matched features. They include dense methods using all pixels and sparse methods selecting only a subset. In Munich that same year, Daniel Cremers was pursuing a related approach. He brought the variational machinery of computer vision, including Gauss-Newton image alignment and the formalism of optical flow, into a complete SLAM system. Cremers's student Jakob Engel introduced LSD-SLAM in 2014 and DSO in 2016. At different sampling densities, both methods compared pixel intensities directly instead of extracting features. --- ## 1. Every Pixel: DTAM [Newcombe, Lovegrove & Davison 2011. DTAM](https://doi.org/10.1109/ICCV.2011.6126513), presented at ICCV 2011, stands for "Dense Tracking and Mapping in Real-Time." It performs tracking and mapping simultaneously, using every pixel while operating in real time. The system has two parts. The tracking stage performs photometric alignment by comparing the entire current frame with a cost volume. It extracts no features and matches no descriptors, minimizing only pixel-intensity differences. The mapping stage estimates a depth map through multi-baseline stereo and maintains a smooth, dense 3D model with total variation regularization. $$E(\mathbf{u}) = \sum_{i} \rho\left( I_i\bigl(\pi(KT_i\mathbf{p}(\mathbf{u}))\bigr) - I_r\bigl(\pi(\mathbf{p}(\mathbf{u}))\bigr) \right) + \lambda \,\text{TV}(\mathbf{u})$$ Here $\mathbf{u}$ is the inverse depth map, $\mathbf{p}(\mathbf{u})$ the 3D point obtained by back-projecting $\mathbf{u}$, $K$ the camera intrinsic matrix, $T_i$ the rigid body transform of frame $i$ relative to the reference frame, $\pi$ the perspective projection, $\rho$ the Huber loss, and $\text{TV}(\mathbf{u}) = \|\nabla \mathbf{u}\|_1$ the total variation regularizer. Running this optimization in real time requires a GPU. DTAM used a single Nvidia GTX 480, the commodity configuration described in the paper's §3. > 🔗 **Borrowed.** DTAM's dense volumetric approach drew partly on depth-camera research, particularly the TSDF formulation of [Curless & Levoy 1996](https://doi.org/10.1145/237170.237269), but applied it to a monocular camera. [KinectFusion](https://doi.org/10.1109/ISMAR.2011.6092378) (2011, ISMAR), also led by Newcombe, developed the corresponding depth-sensor approach. A video of an entire indoor scene being reconstructed in real time appeared on YouTube shortly after the 2011 ICCV talk and received tens of thousands of views. The limitations were equally visible: the system required a GPU, was fragile under lighting changes, and could not scale to large outdoor environments. --- ## 2. Tracking the Edges: LSD-SLAM [Engel, Schöps & Cremers 2014. LSD-SLAM](https://doi.org/10.1007/978-3-319-10605-2_54) replaced DTAM's dense formulation with a semi-dense one and eliminated its dependence on a GPU. "Large-Scale Direct Monocular SLAM" tracks only pixels whose gradient magnitude exceeds a threshold. It ignores flat wall regions and retains pixels near edges with sufficient gradient. Rather than using a corner detector, it selects pixels solely by the strength of their intensity gradient. The tracking stage performs direct image alignment in SE(3). It warps the current frame onto a keyframe and minimizes the photometric residual with Gauss-Newton optimization. The map is keyframe-based, and each keyframe carries its own semi-dense depth map. A pose graph maintains connections between keyframes. Loop closure finds candidates through appearance-based relocalization and verifies them with a depth-consistency check. > 🔗 **Borrowed.** Gauss-Newton photometric registration is a classic of the image alignment field. The [Lucas & Kanade 1981](https://www.ijcai.org/Proceedings/81-2/Papers/017.pdf) tracker and its inverse compositional reformulation ([Baker & Matthews 2004](https://doi.org/10.1023/B:VISI.0000011205.11775.fd)) are the direct ancestors of LSD-SLAM's frontend. The Cremers group transplanted the language of the variational image-processing community into the entire SLAM pipeline. LSD-SLAM ran in real time on a CPU. Its use of keyframes and pose-graph optimization superficially resembled PTAM's tracking-and-mapping split, but the underlying measurement differed: LSD-SLAM used pixel intensity rather than binary descriptors such as ORB or BRIEF. LSD-SLAM also released footage of operation in large outdoor environments. A demo in which a semi-dense map was built while riding a bicycle for tens of meters showed that the direct approach could scale. On the KITTI benchmark, it was competitive with the leading feature-based methods of the time. Lighting changes remained a serious problem. Entering a tunnel, facing backlight through a window, or encountering a sudden flash violated photometric consistency and could destabilize the system immediately. --- ## 3. Sparse Direct Perfected: DSO [Engel, Koltun & Cremers 2018. DSO (PAMI)](https://doi.org/10.1109/TPAMI.2017.2658577) first appeared on arXiv in 2016. "Direct Sparse Odometry" is sparser than LSD-SLAM and uses far fewer pixels than DTAM, while applying a more thorough photometric calibration. The system selects roughly 2,000 high-gradient pixels in each keyframe. This is more than ORB-SLAM2's default setting (nFeatures=1000) and far fewer than LSD-SLAM's semi-dense set of all pixels with a gradient. DSO performs sliding-window bundle adjustment over these pixels, optimizing camera pose, inverse depth, and affine brightness parameters $(a_i, b_i)$. It marginalizes frames that leave the window and uses the Schur complement to eliminate variables and limit the size of the remaining window problem. Total computation depends on window size, pixel count, and iteration count. DSO separates the camera's photometric model into three layers. First, prior calibration corrects vignetting, the falloff in brightness toward the edges of the lens. Second, it inverts the camera response function, or gamma curve, in advance to convert the sensor's nonlinear light measurements into a linear intensity domain. Third, it uses exposure time $t_i$ as input metadata and estimates the remaining affine brightness changes $(a_i, b_i)$ online: $$E_{pj} = \sum_{\mathbf{p} \in \mathcal{N}_p} w_{\mathbf{p}} \left\| \left( I_j\!\left[\mathbf{p}'\right] - \frac{t_j e^{a_j}}{t_i e^{a_i}} I_i[\mathbf{p}] - \left(b_j - \frac{t_j e^{a_j}}{t_i e^{a_i}} b_i\right) \right) \right\|_\gamma$$ Here $t_i, t_j$ are exposure times, $(a_i, b_i)$ and $(a_j, b_j)$ are the affine brightness parameters of each frame (gain and bias), and $\|\cdot\|_\gamma$ is the Huber loss. Photometric calibration corrects vignetting during preprocessing, and the residual above is applied to the corrected intensities. DSO was the first direct SLAM system to divide exposure variation, vignetting, and the response curve between a separate calibration stage and real-time optimization variables. > 🔗 **Borrowed.** The formal basis of photometric camera calibration traces to the HDR-recovery work of [Debevec & Malik 1997](https://doi.org/10.1145/258734.258884). DSO likewise uses response-corrected intensities, but calibrates the response function in advance. The per-frame affine brightness parameters are the quantities adjusted online. On the TUM monocular dataset, DSO was reported to outperform ORB-SLAM2 across several sequences. In feature-poor environments, such as indoor corridors with large flat walls, DSO achieved a lower ATE than ORB-SLAM2. These results supported the claim that photometric methods retain information discarded by feature extraction. > 📜 **Prediction vs. outcome.** DSO required prior photometric calibration, and follow-up work soon addressed that dependency. In 2018, Bergmann, Wang, and Cremers proposed [online photometric calibration](https://doi.org/10.1109/LRA.2017.2777002), jointly estimating exposure time, the response function, and vignetting attenuation while the system ran. On the evaluated auto-exposure videos, the paper reported real-time calibration and VO accuracy on par with pre-calibrated input. This follow-up directly answered DSO's dependence on prior photometric calibration. > 📜 **Prediction vs. outcome.** DTAM achieved real-time dense SLAM on a single GPU, leaving broader access to dense reconstruction as a natural next step. Pure monocular dense reconstruction did not become deployable in real time until NeRF and 3DGS emerged in the 2020s. KinectFusion, also led by Newcombe, instead achieved GPU-based dense reconstruction with an RGB-D depth sensor in 2011 by changing the sensor rather than solving the monocular problem. --- ## 4. VI-DSO and the Lineage Extended At ICRA 2018, von Stumberg, Usenko, and Cremers presented [VI-DSO](https://doi.org/10.1109/ICRA.2018.8462905), which combined DSO with an IMU. Inertial measurements could support pose tracking during rapid lighting changes, a major failure mode for direct photometric methods. The IMU could also resolve the scale ambiguity of a monocular camera. VI-DSO adds an IMU preintegration factor to DSO's windowed photometric bundle adjustment. The IMU preintegration scheme was borrowed from [Forster et al.'s 2017 paper](https://doi.org/10.1109/TRO.2016.2597321). This recovered scale and improved robustness under extreme lighting. Follow-up work from the Cremers group, [Basalt](https://arxiv.org/abs/1904.06504) (2019) and [DM-VIO](https://doi.org/10.1109/LRA.2021.3140129) (2022), continued in the same direction. Both paired a direct photometric frontend with a tightly coupled inertial backend. This lineage developed in parallel with feature-based VIO systems such as VINS-Mono and OpenVINS, each with its own ecosystem. > 🔗 **Borrowed.** VI-DSO directly uses the manifold preintegration formulation of [Forster et al. 2017. On-Manifold Preintegration (IEEE TRO)](https://doi.org/10.1109/TRO.2016.2597321), placing Forster's inertial layer above DSO's photometric layer. --- ## 5. Limits of the Direct Method Direct methods retain more image information. They use pixels that feature detectors discard, including regions whose gradients are low but consistent. The photometric residual also provides a continuous optimization landscape without the discretization imposed by descriptor matching. As of 2026, however, most deployed systems remain feature-based for several reasons. Direct methods depend on photometric calibration. The vignetting correction, response-curve correction, and exposure control assumed by DSO are not readily available from a consumer camera. Smartphone cameras apply HDR fusion, auto-exposure, and real-time white balance internally without exposing that pipeline to the user. These processes violate DSO's photometric assumptions. Lighting changes remain difficult as well. Under auto-exposure or backlight, inter-frame brightness can change sharply and violate the direct method's assumption of photometric consistency. DSO's affine brightness model can absorb only gradual drift, so a passing cloud outdoors or a flickering fluorescent tube indoors remains a leading cause of tracking failure. Although DSO beat ORB-SLAM2 on sequences from controlled datasets, engineers deploying systems on robots often chose ORB-SLAM. It runs on many camera models without separate photometric calibration and can continue working after a camera change. DSO requires vignetting and response-curve calibration for each camera. Learned features such as [SuperPoint](https://arxiv.org/abs/1712.07629) (2018) and [LightGlue](https://arxiv.org/abs/2306.13643) (2023) also weakened the direct method's central criticism that features discard information. They retain more information than handcrafted descriptors while preserving the practical advantages of descriptor matching. --- ## 🧭 Still open **Direct tracking under rapid lighting change.** The direct method assumes that a scene's brightness distribution remains stable across frames, an assumption violated by auto-exposure, strong backlight, and tunnel-to-outdoor transitions. VI-DSO's IMU assistance partially mitigates the problem, but no complete solution dynamically estimates the lighting model itself. Learning-based photometric correction is being explored as an alternative but is not yet deployable in real time. **The shared weakness on textureless surfaces.** Feature-based methods fail on walls without corners, while the photometric residual in direct methods becomes largely insensitive to pose changes on surfaces without gradients. Both approaches are weak in indoor corridors, large warehouses, and homogeneous outdoor terrain. Semi-dense LSD-SLAM retained only pixels with a gradient, but it did not resolve the degeneracy that arises when those pixels are too sparsely distributed. **A possible transition to a learned photometric model.** Current direct SLAM systems represent the photometric model with a simple affine brightness correction or a fixed camera response function. Neural radiance field research instead represents scene appearance with a neural network. Whether such a model can become part of real-time direct SLAM, and where that would place the boundary between direct and learned methods, remain open questions as of 2026. A parallel approach had already appeared in 2011 through Newcombe's KinectFusion: replace the monocular camera with a different sensor. By measuring depth directly, an RGB-D camera enabled dense reconstruction without relying on brightness consistency. Direct methods tried to model photometric consistency, whereas RGB-D systems removed that assumption by changing the measurement source. --- # Ch.9 — Dense/RGB-D: From KinectFusion to BundleFusion When Richard Newcombe of Imperial College London presented KinectFusion at ISMAR in November 2011, its demo video drew more attention than the paper. A single handheld Kinect reconstructed an entire room as a 3D mesh in real time. DTAM, which Newcombe had released earlier that year, had pursued the same goal with a monocular camera; KinectFusion achieved it with an RGB-D sensor. The system combined the TSDF representation devised by Curless and Levoy for graphics in 1996, the ICP tracker introduced to robotics by Besl and McKay in 1992, and Microsoft's $150 Kinect sensor from 2010. Their combination began a brief, intense period of dense SLAM research. Davison's MonoSLAM (Ch.5) had used real-time, CPU-only processing to track sparse landmarks from a monocular camera. Newcombe's DTAM (Ch.8) then used a GPU for dense reconstruction through direct photometric optimization, while KinectFusion used the GPU with an RGB-D depth stream. --- ## 9.1 Dense reconstruction before Kinect Dense 3D reconstruction was possible before 2011, but not in *real time*. Offline pipelines could merge point clouds acquired by stereo or structured-light scanners, but indoor scanning rigs cost hundreds of thousands of dollars and saw little use outside laboratories. The SLAM community was already obtaining practical results from sparse landmarks, while dense reconstruction remained largely a graphics problem. [Curless and Levoy's 1996 SIGGRAPH paper, "A Volumetric Method for Building Complex Models from Range Images"](https://graphics.stanford.edu/papers/volrange/volrange.pdf) introduced the **TSDF (Truncated Signed Distance Function)** to this graphics pipeline. The method partitions 3D space into a uniform voxel grid and accumulates at each voxel the signed distance to the nearest surface. Moving from the sensor toward the surface, the sign convention assigns a positive value to the free space in front of the surface and a negative value to the solid behind it. Truncation clips the absolute value at a threshold $t$, giving $\text{TSDF}(x) = \text{clip}(d(x), -t, +t)$. Each incoming depth frame updates the value by a weighted average, reducing noise and sharpening the surface over time. Marching cubes then extracts the surface at the TSDF's zero-crossing. The method was accurate, but the voxel grid consumed substantial memory, and the available hardware could not update it in real time. For the next fifteen years, Curless and Levoy's paper remained primarily part of the graphics literature. During those fifteen years, GPUs entered the GPGPU era and Kinect appeared. --- ## 9.2 KinectFusion and TSDF Microsoft released Kinect for the Xbox 360 at about $150 in 2010. The sensor measured depth through structured light and streamed VGA-resolution depth maps at 30 Hz. Its precision was below that of research-grade ToF (Time-of-Flight) cameras, but its price was also far lower. Hackers responded first: open-source drivers appeared within weeks of the launch, followed by research applications. Newcombe had by then moved to Microsoft Research Cambridge, where he was developing GPU-based dense SLAM with Shahram Izadi's team. They already had the outline of a pipeline when Kinect launched, and its depth stream supplied the remaining input. The result, presented at ISMAR 2011, was [Newcombe et al. 2011. KinectFusion](https://doi.org/10.1109/ISMAR.2011.6092378). > 🔗 **Borrowed.** KinectFusion's core representation, the TSDF, was devised by Curless & Levoy (1996) for offline 3D scanning. Newcombe's team made it real-time via GPU parallel voxel updates. The pipeline has four stages. First, depth preprocessing denoises the raw depth map with a bilateral filter and computes surface normals. Second, ICP tracking aligns the current frame's point cloud with the virtual surface ray-cast from the previous TSDF. A point-to-plane variant of **ICP (Iterative Closest Point)** from [Besl & McKay (1992)](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf) runs thousands of iterations on the GPU. Its output is the camera's 6-DoF pose. The point-to-plane ICP objective transforms the current frame's point $\mathbf{p}_i$ by $T = (R, \mathbf{t})$ and matches it to $\hat{\mathbf{p}}_i$ on the ray-cast surface, whose normal is $\hat{\mathbf{n}}_i$. It then minimizes $$E(R, \mathbf{t}) = \sum_i \bigl(\hat{\mathbf{n}}_i^\top (R\,\mathbf{p}_i + \mathbf{t} - \hat{\mathbf{p}}_i)\bigr)^2$$ Unlike the original Besl-McKay point-to-point cost ($\|R\mathbf{p}_i + \mathbf{t} - \hat{\mathbf{p}}_i\|^2$), this objective measures only the error along the normal and is therefore less sensitive to motion along the surface. With the small-rotation approximation $R \approx I + [\boldsymbol{\omega}]_\times$, $E$ becomes a linear least-squares problem in the 6-DoF vector $(\boldsymbol{\omega}, \mathbf{t})$. The GPU solves it through parallel reduction. > 🔗 **Borrowed.** KinectFusion's tracking stage directly inherits ICP from Besl & McKay (1992), applying a classical robotics technique at GPU-scale density. Third, TSDF integration projects the depth map into the voxel grid at the estimated pose and updates the TSDF values. The paper's main configuration uses a 512³ voxel grid covering a room-scale volume about 3 m on a side (§4.2, Fig. 13). Fourth, surface rendering finds the TSDF's zero-crossing by ray marching to produce vertex and normal maps of the surface. The result becomes the reference surface for the next ICP step. Newcombe presented the monocular dense SLAM system DTAM in the same year. The two projects were closely related: DTAM used the GPU to optimize monocular photometric consistency, while KinectFusion used it for depth integration. Several researchers worked on both. > 🔗 **Borrowed.** The same researchers presented KinectFusion and DTAM in the same year. They share the use of a GPU for dense processing, but DTAM's photometric optimization and KinectFusion's depth alignment and TSDF fusion also differ in map representation and objective. The 512³ TSDF updated at 30 Hz, and a single indoor room could be reconstructed as a dense mesh within minutes. Within a fixed room-scale volume, dense model-to-frame ICP used many surface measurements and produced low tracking drift. The model was still built from earlier pose estimates, however, and was not an absolute reference surface. A 512³ voxel grid covers only a fixed spatial extent; once the camera leaves the room, voxels saturate or overwrite one another. The system had no loop closure, and Kinect's IR structured light did not work in sunlight. Outdoor use was therefore outside its scope from the outset. --- ## 9.3 Kintinuous — rolling volume Soon after KinectFusion appeared, Whelan at Imperial College addressed this limitation by moving the fixed-size TSDF volume with the camera. In July 2012, at the RSS workshop (RGB-D: Advanced Reasoning with Depth Cameras, Sydney), [Whelan et al. presented Kintinuous](https://www.cs.cmu.edu/~kaess/pub/Whelan12rssw.pdf), which introduced a "rolling TSDF volume." As the camera approached the boundary of the volume, slices on the far side were extracted as mesh and released, while new slices were attached in front. Memory stayed constant while the camera could move indefinitely. A demo that traversed an entire indoor corridor exceeded KinectFusion's fixed spatial range, but loop closure was still missing. After a long corridor traversal returned to its origin, the mismatch between the two ends remained unresolved. This misalignment limited global map consistency independently of the level of surface detail. --- ## 9.4 ElasticFusion: Surfels and non-rigid deformation After Kintinuous, Whelan replaced TSDF voxels with surfels. A **surfel (surface element)** is a point with a position, normal, radius, and color. In computer graphics, [Pfister et al. (2000)](https://www.merl.com/publications/docs/TR2000-10.pdf) proposed the concept as a rendering representation. Unlike a regular voxel grid, a surfel structure follows the observed surface. > 🔗 **Borrowed.** ElasticFusion adapted the surfel rendering technique of Pfister et al. (2000) as a SLAM map representation. [Whelan et al. 2016. ElasticFusion](https://doi.org/10.1177/0278364916669237) combined a surfel-based dense map with loop closure through *non-rigid deformation*. Loop closure had been difficult in earlier dense SLAM systems because updating a global mesh or voxel grid to satisfy a loop-closure constraint was expensive. ElasticFusion connected the surfel set to a deformation graph. When it detected a loop closure, it deformed the graph to distribute the error across the entire map, producing a non-rigid correction at the surfel-map level. Concretely, each node $g_k$ of the deformation graph carries a position $\mathbf{v}_k$ and a rotation $R_k$ and translation $\mathbf{t}_k$. A surfel $s$ lies within the influence of its $K$ nearest nodes, and the surfel's deformed position is computed as $$\tilde{\mathbf{p}}_s = \sum_{k \in \mathcal{N}(s)} w_k \bigl(R_k (\mathbf{p}_s - \mathbf{v}_k) + \mathbf{v}_k + \mathbf{t}_k\bigr)$$ The weight $w_k$ decreases with distance. When a loop-closure constraint is added, Gauss-Newton optimizes the graph nodes' $(R_k, \mathbf{t}_k)$ to distribute the error globally. The method can therefore correct the entire dense map consistently without rebuilding a TSDF from scratch. The ElasticFusion paper reported strong indoor reconstruction results on the ICL-NUIM synthetic dataset. Sequences kt0·kt1·kt2 recorded ATE RMSE below 1.4 cm, with kt0·kt1 at 0.9 cm; kt3, in which global loop closure activates, was an unusually large exception. These figures apply to the paper's ICL-NUIM setting rather than to KITTI or TUM RGB-D. --- ## 9.5 BundleFusion: offline-SfM quality, online In 2017, Dai, Nießner, Zollhöfer, Izadi, and Theobalt published [Dai et al. 2017. BundleFusion](https://doi.org/10.1145/3072959.3054739) in ACM Transactions on Graphics. KinectFusion and its successors had sought higher quality while preserving real-time operation. BundleFusion instead devoted extensive GPU computation to running SfM-grade bundle adjustment within an online system. BundleFusion used hierarchical optimization. At the fastest layer, dense depth alignment between the current and previous frames provides an initial pose. The next layer corrects it through sparse frame-to-frame alignment with SIFT features. At the third layer, hierarchical global bundle adjustment re-optimizes the poses of accumulated frames, including earlier poses. This "retroactive pose correction" sought to approach online the result that an offline SfM pipeline obtains by aligning all data after collection. The system back-projects updated pose sequences into the TSDF for reintegration, preventing tracking errors from remaining embedded in the map. Dai's team reported better results than ElasticFusion on TUM RGB-D. Its visual reconstruction quality approached that of the offline COLMAP pipeline by the standards of the time. > 📜 **Prediction vs. outcome.** BundleFusion claimed real-time online global bundle adjustment at "unprecedented speed," proposing a way to bring offline SfM quality into online operation. GPU performance continued to increase, but from 2021 one major research path instead used COLMAP for camera poses and NeRF for scene representation. This was less a direct successor to online TSDF mapping than a turn toward novel-view synthesis and neural scene representation. --- ## 9.6 Co-evolution of hardware and algorithm The six years from KinectFusion to BundleFusion reflect changes in both algorithms and hardware. The first-generation Kinect used structured light. Its depth precision was a few millimeters at meter-scale range, but sunlight obscured the IR pattern. Kinect 2, released in 2013, switched to ToF, improving both precision and dynamic range. Intel's RealSense series followed. As sensor options expanded, algorithms could make different assumptions about depth quality, and researchers explored methods that either exploited lower noise or tolerated higher noise. The CUDA ecosystem also matured. Between KinectFusion in 2011 and BundleFusion in 2017, GPU throughput and memory performance also improved. Any numerical comparison requires specifying the GPU models and arithmetic precision. The increasingly expensive real-time optimization in Whelan's ElasticFusion and Dai's BundleFusion depended on this hardware progress as well as on algorithm design. Had Kinect been priced like research instrumentation rather than a consumer device, this line of work would probably have spread more slowly. A mass-market sensor helped set its pace. > 📜 **Prediction vs. outcome.** The spatial range, drift, and outdoor limitations of KinectFusion's fixed 512³ volume shaped the research that followed. Kintinuous, ElasticFusion, and BundleFusion addressed volume extension in turn. Outdoor operation followed a different path because sunlight obscures IR structured-light patterns. Early Kinect-based dense SLAM remained largely indoors, while LiDAR became a main sensor for outdoor dense mapping. This limitation does not apply uniformly to every RGB-D sensing technology. --- ## 9.7 Why dense-only systems receded Between 2011 and 2017, dense RGB-D SLAM appeared likely to become the main direction of Visual SLAM, but it did not. Sparse backends continued to dominate. Practical SLAM systems after 2015, represented by [ORB-SLAM2](https://arxiv.org/abs/1610.06475) and [VINS-Mono](https://arxiv.org/abs/1708.03852), did not use dense maps by default. Several constraints reinforced this choice. A 512³ TSDF requires more than 512 MB, a substantial cost for mobile and embedded systems. [Voxblox](https://arxiv.org/abs/1611.03631), which stores TSDF values in hashed blocks, and [OctoMap](https://www.hrl.uni-bonn.de/papers/wurm10octomap.pdf), which stores occupancy probabilities in an octree, reduced memory use with different map representations but did not match sparse representations in efficiency. Real-time dense processing also required a GPU, making a KinectFusion-grade pipeline difficult to run on an autonomous vehicle's embedded processor or a lightweight drone platform. Kinect's IR depth sensing also failed outdoors, while commercially important applications such as autonomous driving and drones operated largely in outdoor environments. During the same period, dense map data structures extended KinectFusion's fixed 512³ volume in several directions. [Museth's VDB (2013)](https://doi.org/10.1145/2487228.2487235) combined block hashing with an internal tree, leaving sparse regions empty while refining only the neighborhood of a surface. Released as OpenVDB, it became a major infrastructure for large sparse volumetric data and provides a useful comparison with the nvblox lineage in Ch.17. [Reijgwart et al. (2023)'s wavemap](https://arxiv.org/abs/2306.08125) compressed occupancy with a wavelet transform to adjust the trade-off between resolution and memory. Ramos and Ott pursued continuous-function representations. [O'Callaghan and Ramos (2012)'s GPOM (Gaussian Process Occupancy Map)](https://doi.org/10.1177/0278364911435991) used Gaussian Process regression to relate depth measurements and assign probabilities even to unmeasured voxels. [Ramos and Ott (2016)'s Hilbert Map](https://doi.org/10.1177/0278364916684382) learned Hilbert-space features with logistic regression to provide streamable probabilistic occupancy. [Behley and Stachniss (2018)'s SuMa](https://www.ipb.uni-bonn.de/wp-content/papercite-data/pdf/behley2018rss.pdf) adapted ElasticFusion's indoor RGB-D surfel representation to outdoor LiDAR, producing a surfel-based SLAM system that operated on KITTI (→ Ch.17). These approaches extended dense mapping from a single room to outdoor and city-scale environments while incorporating probabilistic uncertainty. After NeRF appeared around 2020, high-quality dense reconstruction shifted toward NeRF and 3D Gaussian Splatting. The map representation changed, but measured RGB-D depth continued to constrain both tracking and the learned map geometry. The period of dense RGB-D SLAM was brief, but its components persisted. TSDF representations entered occupancy mapping for autonomous driving, and ICP became a standard tracking method in LiDAR SLAM. Dense-only systems receded while their techniques spread into other architectures. --- ## 🧭 Still open Large-scale outdoor dense reconstruction. Sunlight interference with IR structured light is a general limitation of active depth sensors. LiDAR operates at longer range but captures color and fine surface detail poorly. As of 2026, RGB-D still cannot densely process large outdoor environments. Learned stereo depth estimation has improved, but limitations in dark regions, on reflective surfaces, and at long range remain unresolved. Dense reconstruction of dynamic scenes. Every system from KinectFusion to BundleFusion was designed for static scenes. Dense reconstruction in spaces occupied by moving people requires separating dynamic objects, using semantic segmentation, geometric residuals, or other motion-separation methods. [DynaSLAM](https://arxiv.org/abs/1806.05620) and [MaskFusion](https://arxiv.org/abs/1804.09194) attempted this, but their computational cost and robustness remained unsuitable for practical deployment. Memory efficiency of voxel maps. Voxblox stores TSDF values in hashed blocks, while OctoMap stores occupancy probabilities in an octree; the two reduce memory use for different map representations. Dense representations at the scale of a building floor or city block nevertheless require tens of gigabytes. No general adaptive-resolution method automatically determines the appropriate resolution for each region. Implicit neural representations such as [Instant-NGP](https://arxiv.org/abs/2201.05989) address this problem, but real-time updates still trade off against query speed. While dense SLAM reconstructed individual rooms as meshes, another lineage addressed the problem of recognizing a return to the same room. Place recognition, which asks whether a system has seen a place before, had developed at Oxford since 2003 independently of dense mapping. KinectFusion had no loop closure; the researchers working on that problem focused on recognizing locations rather than constructing denser maps. --- # Ch.10 — The Parallel Line of Place Recognition: From FAB-MAP to NetVLAD, and on to AnyLoc Around 2003, while Davison was demonstrating real-time 3D tracking with a single webcam, Mark Cummins and Paul Newman at the Oxford Mobile Robotics Group were asking a different question: "How does a robot recognize a place it has visited before?" Because visual odometry (VO) accumulated drift, no SLAM system could close a loop without answering it. Place recognition developed through the 2000s alongside other components of Visual SLAM but followed its own lineage. FAB-MAP adapted Josef Sivic's bag-of-words (BoW) approach to robotics, DBoW2 made it practical, and NetVLAD introduced learning. In 2023, AnyLoc used foundation-model features without fine-tuning. The ORB-SLAM, DSO, and KinectFusion lineages developed different approaches to tracking and mapping, but each required loop closure. Place recognition supplied the decision behind it: where has the system seen this scene before? --- ## 10.1 Place recognition before BoW To close a loop without GPS, whether indoors, in a tunnel, or in an urban canyon, a robot must quickly find among thousands of candidates the image most similar to its current observation. Pixel-level comparison requires a linear scan, O(N), and becomes impractical in real time once the collection reaches tens of thousands of images. Sivic and Zisserman were among the first computer vision researchers to address this problem in the early 2000s. Their ICCV 2003 paper ["Video Google"](https://www.robots.ox.ac.uk/~vgg/publications/2003/Sivic03/sivic03.pdf) applied TF-IDF from document retrieval to images. They clustered SIFT descriptors with k-means to form "visual words" and represented each image as a frequency vector over those words. An inverted index avoided scanning the full image collection by retrieving only the posting lists for visual words in the query, and place-recognition researchers quickly adopted the idea. --- ## 10.2 FAB-MAP — probabilistic BoW and the Chow-Liu tree (2008) Mark Cummins and Paul Newman, at the Oxford Mobile Robotics Group, published [Cummins & Newman. FAB-MAP: Probabilistic Localization and Mapping in the Space of Appearance](https://doi.org/10.1177/0278364908090961) in 2008. FAB-MAP (**Fast Appearance-Based Mapping**) asks whether a scene corresponds to a place already in the database or to an entirely new location. A simple similarity score cannot resolve this distinction: if dozens of corridors look alike, the highest score does not guarantee a correct match. Cummins and Newman framed this as a Bayesian inference problem. Given an observation $z_t$ (the set of visual-word occurrences), they computed the probability that the current location is each database place $\ell_i$: $$P(\ell_i \mid z_t) \propto P(z_t \mid \ell_i) P(\ell_i)$$ The hard part is $P(z_t \mid \ell_i)$. Assuming visual words are independent gives a naïve Bayes model, but in practice visual words are correlated. If the word "door" appears, the word "doorknob" tends to appear along with it. The independence assumption distorts the probability. FAB-MAP modeled this correlation with a **Chow-Liu tree**. A Chow-Liu tree is a tree-structured graphical model that maximizes pairwise mutual information among words. The mutual information between two words $e_i, e_j$ is defined as $$I(e_i; e_j) = \sum_{e_i, e_j} P(e_i, e_j) \log \frac{P(e_i, e_j)}{P(e_i)P(e_j)}$$ and the Chow-Liu algorithm uses this as an edge weight to build a maximum spanning tree. Factorizing the joint likelihood through this tree gives $$P(z_t \mid \ell_i) = \prod_k P(z_t^k \mid z_t^{\text{pa}(k)}, \ell_i)$$ where $z_t^k \in \{0,1\}$ indicates the occurrence of the $k$-th word and $\text{pa}(k)$ is its parent in the tree. Unlike naïve Bayes, this factorization captures co-occurrence patterns across words and lowers false positives in visually similar places such as corridors. During training, the vocabulary and the tree are learned together from a large image set. FAB-MAP also explicitly modeled the possibility that the current location was absent from the database. Adding this "new place" hypothesis reduced false positives, which can cause catastrophic failure in loop closure. > 🔗 **Borrowed.** FAB-MAP's visual-word approach was transplanted directly from Sivic & Zisserman's "Video Google" (2003). The inverted-index logic of document retrieval was applied to a robot's memory of places. In 2011, Cummins and Newman published [FAB-MAP 2.0](https://www.robots.ox.ac.uk/~mjc/Papers/cummins_newman_ijrr_fabmap2_2010_preprint.pdf). They sought to extend the processable map scale to around 1,000 km and demonstrated the system on a city-scale dataset. --- ## 10.3 DBoW2 — binary descriptors and the vocabulary tree (2012) FAB-MAP used floating-point descriptors such as SIFT. Around 2012, the SLAM community was moving toward faster binary descriptors, particularly BRIEF, ORB, and BRISK. Retaining a SIFT vocabulary imposed a substantial computational cost. In 2012, Dorian Gálvez-López and Juan D. Tardós of Universidad de Zaragoza published [Gálvez-López & Tardós. Bags of Binary Words for Fast Place Recognition in Image Sequences](https://doi.org/10.1109/TRO.2012.2197158). **DBoW2** uses a vocabulary tree of binary descriptors. Hamming-distance comparisons made word assignment tens of times faster than with SIFT. DBoW2's structure is a vocabulary tree built by hierarchical k-means. The BoW vector representing an image is a TF-IDF–weighted binary-word frequency vector. Each leaf node $w_i$ of a tree with branching factor $k$ and depth $d$ carries the TF-IDF weight $$\eta_i = \frac{n_i}{n} \cdot \log \frac{N}{N_i}$$ where $n_i$ is the word count of $w_i$ in the image, $n$ is the total word count, $N$ is the number of database images, and $N_i$ is the number of images containing $w_i$. The similarity between two images $a$, $b$ is given by the L1-norm $$s(\mathbf{v}_a, \mathbf{v}_b) = 1 - \frac{1}{2} \left\| \frac{\mathbf{v}_a}{|\mathbf{v}_a|} - \frac{\mathbf{v}_b}{|\mathbf{v}_b|} \right\|_1$$ Lookup runs in O(log N) through an inverted index. > 🔗 **Borrowed.** DBoW2's vocabulary-tree concept traces its lineage to Nistér & Stewénius's 2006 ["Scalable Recognition with a Vocabulary Tree"](https://people.eecs.berkeley.edu/~yang/courses/cs294-6/papers/nister_stewenius_cvpr2006.pdf) (CVPR). DBoW2 transplanted that structure into the binary-descriptor world and tuned the weighting scheme for SLAM. DBoW2's influence came as much from deployment as from its algorithm. Released as an open-source library, it became the loop-closure module in ORB-SLAM (2015), ORB-SLAM2, and ORB-SLAM3. From 2015 through the mid-2020s, it served as the standard place-recognition component in the ORB-SLAM family and many systems built on it. The Gálvez-López–Tardós collaboration also connected directly to later SLAM systems. Tardós subsequently led the ORB-SLAM trilogy with Mur-Artal and Campos, and DBoW2 supplied its place-recognition layer. --- ## 10.4 NetVLAD — CNN-based VPR (2016) The BoW family had a fundamental limitation: its vocabulary was trained for a specific descriptor and environment. Large changes in lighting, season, or viewpoint shifted the distribution of visual words and could make a pretrained vocabulary fail. At CVPR 2016, Relja Arandjelović, Petr Gronat, Akihiko Torii, Tomáš Pajdla, and Josef Sivic published [NetVLAD: CNN Architecture for Weakly Supervised Place Recognition](https://doi.org/10.1109/CVPR.2016.572). Sivic had co-authored "Video Google" in 2003, introducing BoW to image retrieval. Thirteen years later, he co-authored a method designed to move beyond its limitations. NetVLAD made **VLAD (Vector of Locally Aggregated Descriptors)** aggregation differentiable. VLAD is an aggregation scheme proposed in 2010 by [Jégou et al.](https://inria.hal.science/inria-00548637/file/jegou_compactimagerepresentation.pdf). It represents an entire image by accumulating the residual between each local descriptor and its nearest cluster center, or visual word. The VLAD sub-vector for cluster center $k$ is $$\mathbf{V}(k) = \sum_{\mathbf{x}_i : \text{NN}(\mathbf{x}_i)=k} (\mathbf{x}_i - \boldsymbol{\mu}_k)$$ and the full VLAD vector $\mathbf{V} = [\mathbf{V}(1)^\top, \ldots, \mathbf{V}(K)^\top]^\top$ is the concatenation across all clusters, L2-normalized. With $K$ clusters and $D$-dimensional descriptors, the final vector is $KD$-dimensional. The VLAD vector carries much richer information than BoW's binary assignment. > 🔗 **Borrowed.** NetVLAD's aggregation design directly inherits VLAD from Jégou et al.'s "Aggregating Local Descriptors into a Compact Image Representation" (CVPR 2010). NetVLAD replaced VLAD's hard assignment with a soft assignment and made the full pipeline trainable end to end. The NetVLAD layer softens the nearest-neighbor assignment of classical VLAD into a softmax: $$\bar{a}_k(\mathbf{x}_i) = \frac{e^{\mathbf{w}_k^\top \mathbf{x}_i + b_k}}{\sum_{k'} e^{\mathbf{w}_{k'}^\top \mathbf{x}_i + b_{k'}}}$$ Here $\mathbf{x}_i$ is a local feature extracted by the CNN, and $\mathbf{w}_k$ and $b_k$ are learnable parameters. Accumulating the NetVLAD vector with this soft assignment gives $$\mathbf{V}(k) = \sum_i \bar{a}_k(\mathbf{x}_i)\,(\mathbf{x}_i - \boldsymbol{\mu}_k)$$ After intra-normalization (L2 on each sub-vector) and a final L2 normalization, the full vector $\mathbf{V} = [\mathbf{V}(1)^\top, \ldots, \mathbf{V}(K)^\top]^\top$ becomes the VPR descriptor. Unlike hard-assignment VLAD, the soft assignment permits gradients to propagate through the layer, so the CNN backbone can be trained end to end. The training procedure also differed from earlier methods. Using Google Street View Time Machine data, the authors treated images of the same place at different times as positive pairs and images of different places as negatives under a weakly supervised triplet loss. GPS positions provided the supervision without manual labels. On the Pittsburgh 250k and Tokyo 24/7 benchmarks, NetVLAD substantially outperformed the DBoW family and earlier VLAD-based methods. It was more robust to lighting and seasonal changes and tolerated some viewpoint differences. NetVLAD was not immediately integrated into practical SLAM pipelines, however, because its inference and memory costs exceeded those of DBoW2 and the ORB-SLAM ecosystem was already built around DBoW2. --- ## 10.5 Patch-NetVLAD, MixVPR, AnyLoc (2020–2023) After NetVLAD, Visual Place Recognition (VPR) research increasingly focused on generalization. In 2021, Hausler et al. introduced [Patch-NetVLAD](https://arxiv.org/abs/2103.01486). Rather than representing a place with a single global descriptor, it divides the image into patches and spatially combines their NetVLAD representations. On Tokyo 24/7, it improved Recall@1 by about 10 percentage points over NetVLAD, at the cost of more expensive inference. In 2023, Ali-bey et al.'s [MixVPR](https://arxiv.org/abs/2303.02190) produced global features through Transformer-style feature mixing, seeking a balance between a lightweight design and high performance. VPR papers from this period commonly evaluated on Mapillary Street Level Sequences (MSLS) and seasonal-change datasets such as Nordland. Extreme lighting and seasonal changes remained common obstacles. In 2023, Keetha et al.'s [AnyLoc: Towards Universal Visual Place Recognition](https://arxiv.org/abs/2308.00688) used self-supervised DINOv2 features for place recognition without fine-tuning. > 🔗 **Borrowed.** AnyLoc uses pretrained ViT representations from Oquab et al.'s [DINOv2](https://arxiv.org/abs/2304.07193) (Meta AI, 2023) and applies VLAD aggregation to them. This connects the BoW–VLAD lineage that began with FAB-MAP to foundation-model features. DINOv2 is a Vision Transformer (ViT) trained on large-scale internet imagery. It produces general-purpose features applicable across environments, although this does not eliminate biases in the pretraining data. AnyLoc emphasized DINOv2's **facets**. Each ViT attention head produces query (Q), key (K), and value (V) matrices as well as final patch features. Keetha et al. found experimentally that the value (V) facet provided the most semantically stable representation for place recognition. Q and K facets favored structural and geometric information, whereas the V facet emphasized semantics, supporting consistent place representations across seasonal and lighting changes. By applying VLAD aggregation to the V-facet representation, they produced a single model that operated across diverse indoor, outdoor, underground, and aerial environments. Across seven or more settings, including Pittsburgh, Tokyo, indoor factories, underground parking garages, and libraries, the model matched or surpassed earlier specialized methods. Research then extended generalization across sensor modalities. Lee et al.'s [(LC)²](https://arxiv.org/abs/2304.08660) (RA-L 2023) projected camera imagery and LiDAR point clouds into a shared 2.5D depth image, enabling a 2D query to retrieve places from a LiDAR map. Datasets such as Lee et al.'s [ViViD++](https://arxiv.org/abs/2204.06183) (RA-L 2022) support cross-modal evaluation by synchronizing visible, thermal, event, LiDAR, inertial, and depth streams across indoor, outdoor, and underground settings. --- ## 10.6 Toward integrating place recognition and metric localization (2024–2025) Place-recognition research has developed alongside other SLAM components since the early 2000s. ORB-SLAM embedded DBoW2, but kept place recognition as a black box separate from mapping and tracking: an image entered, and a loop-candidate ID emerged. By 2024–2025, this boundary had begun to blur. Berton et al.'s [EigenPlaces](https://arxiv.org/abs/2308.10832) (2023) and Izquierdo & Civera's [SALAD](https://arxiv.org/abs/2311.15937) (2023 arXiv / CVPR 2024) explored using place-recognition descriptors directly for metric localization. They sought to estimate a 6-DoF pose from the place representation itself, rather than stopping after identifying a previously seen location. Around 2024, researchers also began combining Gaussian-map representations with place recognition, following the rise of 3DGS (3D Gaussian Splatting) as a map representation. > 📜 **Prediction vs. outcome.** In their 2011 FAB-MAP 2.0 paper, Cummins and Newman extended the scale of place recognition by demonstrating appearance-only loop closure on a 1,000 km trajectory. Compared with early FAB-MAP experiments on the Oxford campus and parts of the city, this was an increase by a factor in the tens. Later city-scale experiments with DBoW2 and large vocabularies reproduced the scale in practical SLAM. These methods increased the scale that could be handled, while deep learning added tools for reducing sensitivity to seasonal and lighting changes. Generalization across scale and appearance remains dependent on the environment. > 📜 **Prediction vs. outcome.** In the introduction to the 2016 NetVLAD paper, Arandjelović et al. identified three challenges for place recognition: a CNN architecture, sufficient training data, and an end-to-end training procedure. NetVLAD directly addressed the architecture and training procedure, while VPR research over the next seven years focused on generalization across season, lighting, and viewpoint. In 2023, AnyLoc demonstrated a single multi-environment model using foundation-model features without fine-tuning. This marked a shift from specialized models toward general-purpose ones rather than a complete solution. --- ## 10.7 🧭 Still open **Extreme seasonal and lighting change.** The Nordland dataset (a Norwegian railway in summer and winter) and Oxford RobotCar dataset (a year of seasonal change) have exposed the same limitation for more than a decade. DINOv2-based methods have narrowed the gap, but one model still does not maintain consistent precision and recall across conditions such as snow-covered winter and dense summer foliage. Place recognition under severe appearance changes remains an open problem as of 2026. **Integration of place recognition and metric localization.** In most current SLAM pipelines, place recognition only identifies a previously seen location; subsequent descriptor matching or another matching method establishes geometric correspondences, from which a method such as PnP estimates the pose. Methods introduced between 2023 and 2025 attempted to merge both processes into one representation, but whether they meet accuracy and speed requirements depends on the sensors, scenes, and deployment conditions. *Privacy of recognizable place representations.* Reconstruction attacks can use stored VPR representations to recover original images or 3D structure. This poses a practical concern for commercial robots that map homes, hospitals, and offices. No place-representation scheme yet guarantees privacy without sacrificing performance. --- By this point, three SLAM lineages had matured. ORB-SLAM standardized the feature-based pipeline, DSO developed the photometric formulation, and KinectFusion and its successors established the capabilities and limits of dense mapping. Place recognition differed because it grew from image retrieval in computer vision rather than from SLAM itself, later supplying the loop-closure component that SLAM required. This separation let place recognition adopt deep-learning methods faster than the rest of the SLAM pipeline. When AnyLoc appeared in 2023, it cited Sivic's work. He had introduced BoW to image retrieval in 2003 and co-authored NetVLAD in 2016, when learned aggregation moved beyond some of BoW's limitations. AnyLoc extended that lineage by applying foundation-model features to place recognition. A different challenge came from a graduate student at NYU. His depth-estimation CNN tested whether geometry had to be recovered only through geometric methods. --- # Ch.11 — The Return of Depth Estimation: From Eigen to Depth Anything The feature-based, direct, RGB-D, and place-recognition lineages each matured around geometric methods. ORB-SLAM reconstructed the world through epipolar geometry, DSO relied on photometric consistency, KinectFusion aligned surfaces with ICP, and RGB-D fusion pipelines closed loops with geometric features. Learning had little role in these systems. That boundary began to shift with a computer vision paper by a graduate student at NYU rather than with a SLAM paper. Monocular depth estimation was one of the oldest ill-posed problems in computer vision. In principle, a single image cannot determine depth because projection discards the third dimension. Humans nevertheless infer depth with one eye from perspective, occlusion, texture gradients, and surface shading. In 2014, David Eigen at NYU tested whether a CNN could learn those statistical cues. The experiment began a lineage that would alter the SLAM pipeline ten years later. --- ## 1. Eigen 2014 — early CNN depth Monocular depth estimation research predated 2014. In 2005, Ashutosh Saxena at Stanford published [Make3D, a system that combined support vector machines (SVMs) with a Markov Random Field (MRF) to predict a depth map from a single image](https://papers.nips.cc/paper/2921-learning-depth-from-single-monocular-images). Make3D predicted coarse, piecewise-planar 3D structure primarily for outdoor scenes and became a representative learned monocular-depth system before CNNs. [Eigen et al. 2014](https://arxiv.org/abs/1406.2283), by Eigen, Puhrsch, and Fergus, introduced a different approach. Its two-stage CNN used a coarse network to predict global structure and a fine network to refine local detail. Training used roughly 120,000 frames from 464 indoor scenes in NYU Depth v2. The method improved on Make3D by the standards of the time and demonstrated that a network could learn to estimate depth. One weakness remained: **scale ambiguity**. The network's absolute scale depended on the training distribution and camera setting. A model trained on indoor NYU data could therefore produce incorrect scale on an outdoor scene. The issue remained central to monocular metric depth in 2024. > 🔗 **Borrowed.** Eigen 2014 inherited the depth-estimation task from Make3D (Saxena 2005). It replaced the SVM and MRF with a CNN while retaining the task definition and evaluation metrics, including RMSE and threshold accuracy. --- ## 2. Garg → Godard — self-supervised depth The bottleneck in supervised depth learning was data. The Kinect works well indoors, but outdoors, especially in sunlight, the infrared pattern washes out. Building a large-scale outdoor RGB-D dataset is expensive. In 2016, [Ravi Garg at UCL introduced another approach](https://arxiv.org/abs/1603.04992): using stereo image pairs as the training signal. The system predicts depth from the left image, then uses that depth and the camera baseline to reconstruct the right image. Because the right image is observed, a photometric loss provides supervision without labels. Clément Godard at UCL developed the idea into **MonoDepth** in [Godard et al. 2017](https://doi.org/10.1109/CVPR.2017.699). Its left-right consistency term imposes a two-way constraint between depths predicted from the left and right images. Adding structural similarity (SSIM) to the photometric loss improved stability in textureless regions. Stereo pairs were required only during training; inference used a single image. Under the paper's contemporary KITTI protocol, it reported lower error than the self-supervised methods listed in its comparison table. > 🔗 **Borrowed.** Garg and Godard derived their photometric loss from the stereo-matching literature. They repurposed the intensity-consistency constraint from disparity estimation, [as organized by Scharstein and Szeliski (2002)](https://vision.middlebury.edu/stereo/taxonomy-IJCV.pdf), as the training signal for a depth network. In 2019, Godard's *MonoDepth2* ([Godard et al. 2019, ICCV](https://arxiv.org/abs/1806.01260)) extended self-supervision to monocular video instead of stereo pairs. A depth network and a pose network train jointly: the pose network predicts camera motion between consecutive frames, and the depth estimate warps the previous frame into the current one. Both networks optimize the resulting warping error. **Minimum reprojection loss** selects the source frame with the lowest photometric error, reducing errors in occluded regions. **Auto-masking** excludes pixels that move at the same speed as the camera, including the case of a stationary camera observing stationary objects. The design still had several limitations. Moving objects and reflective surfaces violated photometric consistency, while the textureless sky provided little signal. Scale also remained ambiguous because video supervision recovers only relative scale between frames. --- ## 3. MiDaS — mixing datasets [Ranftl et al. 2020](https://doi.org/10.1109/TPAMI.2020.3019967), **MiDaS** (Mixing Datasets for Zero-shot Cross-dataset Transfer), led by René Ranftl at Intel, trained one model on multiple datasets rather than a single dataset. Depth units and scales differ across these datasets. NYU contains indoor metric depth, KITTI provides outdoor LiDAR points, ReDWeb derives stereo data from movies, and MegaDepth uses SfM reconstructions. Combining them without normalization would give the network inconsistent targets. Ranftl addressed the mismatch with an **affine-invariant loss**. Before comparison, the method normalizes each image's predicted and ground-truth depth through an affine transformation consisting of scale and shift. It subtracts the median to remove the shift and divides by the mean absolute deviation from that median to remove scale. This scale-and-shift-invariant normalization eliminates unit mismatches across datasets, so the network learns relative depth ordering rather than absolute distance. The original MiDaS mixed several datasets and demonstrated zero-shot transfer to datasets excluded from training. Later MiDaS releases expanded the training mixture to as many as 12 datasets. The family estimated relative depth across distributions as different as indoor and outdoor imagery, but it did not recover absolute scale. Ranftl's team separately released [**DPT** (Dense Prediction Transformer)](https://arxiv.org/abs/2103.13413) in 2021, replacing the MiDaS backbone with a ViT-based architecture. DPT became the default backbone from MiDaS v3 onward, followed by a refinement in v3.1 (2022). The change substantially improved performance. > 🔗 **Borrowed.** DPT in MiDaS v3 used a ViT-based encoder, while the later Depth Anything used an encoder pretrained with DINOv2. Replacing the backbone became a common source of performance gains in the foundation-model era, and DPT (Ranftl 2021) provided an early large-scale demonstration in depth estimation. --- ## 4. Depth Anything — foundation scale In January 2024, **Depth Anything**, developed by Lihe Yang's team at TikTok Research ([Yang et al. 2024](https://arxiv.org/abs/2401.10891)), approached depth estimation through training-data scale. It used 1.5M labeled images, merged from existing datasets, and 62M unlabeled images. The method generated pseudo-labels for the unlabeled set and included them in training, using semantic-segmentation features as auxiliary supervision to improve their quality. The model surpassed MiDaS and earlier methods across the major KITTI, NYU, ScanNet, and DIODE benchmarks. Its ViT-L backbone contained 335M parameters. Inference was far from real-time, reflecting the method's emphasis on prediction quality. [**Depth Anything v2**](https://arxiv.org/abs/2406.09414), released later the same year, trained a teacher on synthetic datasets including Virtual KITTI and Hypersim, then trained student models on real images pseudo-labeled by that teacher. Such data covers surfaces that are difficult to annotate in real imagery, including reflective and transparent materials. Version 2 visibly improved edge detail and thin structures over v1. Depth Anything still produces relative depth without scale. [**ZoeDepth** (Shariq Farooq Bhat et al. 2023)](https://arxiv.org/abs/2302.12288) combined relative-depth pretraining with domain-specific metric-bin heads and automatic routing. [**Depth Anything v2** (2024)](https://arxiv.org/abs/2406.09414) offered metric models obtained by fine-tuning a strong relative-depth backbone on metric labels. [**Metric3D v2** (2024)](https://arxiv.org/abs/2404.15506) handled ambiguity across focal lengths through a canonical-camera-space transformation and converted predictions back with camera parameters at inference. All three extended the relative-depth lineage toward metric prediction, but generalization across training domains and camera settings remained a separate problem. --- ## 5. Re-entry into SLAM CNN-SLAM incorporated monocular depth predictions into SLAM in 2017. Later systems also used learned depth for initialization. Monocular SLAM requires a sufficient baseline to triangulate from two frames and has ambiguous scale from the outset. Adding a depth prior to the first frame accelerates initialization and supplies an initial scene structure. A metric model, or a prior calibrated against a known reference, can also initialize metric scale. [DROID-SLAM, released in 2021 by Teed and Deng](https://arxiv.org/abs/2108.10869), combines recurrent optical flow with BA. Follow-up work in that lineage incorporated monocular depth priors into geometric initialization. Metric-calibrated depth predictions can also act as periodic scale anchors for monocular visual odometry and suppress scale drift. A relative-depth output such as MiDaS, however, cannot determine absolute scale without an external metric reference. > 📜 **Prediction vs. outcome.** Eigen's 2014 paper identified integration with 3D geometric information, such as surface normals, as a natural extension. PAD-Net, VPD, and other systems later implemented parts of this joint multi-task approach. By 2024, however, shared ViT backbones arguably had greater impact than explicit task combination. The field developed along a different path from the one proposed. > 📜 **Prediction vs. outcome.** MiDaS (2020) used a scale-and-shift-invariant loss and focused on relative depth. ZoeDepth and the metric models of Depth Anything v2 later used fine-tuning on metric labels, while Metric3D v2 corrected camera-model variation through a canonical space. Metric prediction became more broadly useful, but scale generalization to unseen cameras and domains remained in progress. --- ## 🧭 Still open **Depth on reflective and transparent surfaces.** For glass, water, and metallic reflections, the image does not directly represent the physical surface. This optical ambiguity remains difficult even with additional synthetic training data, and generalization to real reflective scenes is unstable. Specialized approaches such as [ClearGrasp (Sajjan et al. 2020)](https://arxiv.org/abs/1910.02550) exist, but no general solution does. Even foundation-scale models produce large structural errors in this regime. **Separating ego-depth and object-depth in dynamic scenes.** Moving cars and pedestrians violate photometric consistency. Self-supervised methods often mask moving objects, avoiding rather than solving the problem. Several later studies, including [Ranjan et al. (2019)](https://arxiv.org/abs/1805.09806), have attempted to estimate moving-object depth separately from ego-motion, but a practical solution remains difficult. **Generalization of metric scale.** ZoeDepth and the metric variants of Depth Anything v2 learn scale from metric labels, while Metric3D v2 explicitly handles camera-parameter variation. CCTV and archival imagery may still lack reliable intrinsics or metadata. Metric depth that generalizes independently of training domain and camera model remains difficult even for foundation models and was a central open question in 2025. --- By 2024, while Depth Anything was advancing depth benchmarks, a Cambridge paper had posed another unresolved problem to the SLAM community for nine years: estimating absolute pose directly from a single image without feature extraction, optimization, or an initial map. The paper was [PoseNet](https://arxiv.org/abs/1505.07427). --- # Ch.12 — The End-to-End Frustration Eigen's network estimated metric depth from pixels, and SfMLearner obtained geometric supervision without labels. Once learning could infer shape, researchers asked whether one network could perform the entire SLAM task rather than only pose estimation or loop closure. Work from 2015 to 2018 did not produce a successful solution. In 2015, Alex Kendall, a PhD student working under Roberto Cipolla at the Cambridge Computer Laboratory, trained a neural network on Google Street View images to estimate a 6-DoF pose from a single photograph. [Kendall et al. 2015. PoseNet](https://doi.org/10.1109/ICCV.2015.336) drew immediate attention at ICCV in Santiago de Chile. It suggested that a single CNN might replace the feature extraction, matching, optimization, and map management developed over thirty years of SLAM research. Dozens of papers explored that possibility between 2015 and 2018, and nearly all encountered the same limitations. --- ## 12.1 PoseNet PoseNet followed [AlexNet (Krizhevsky et al. 2012)](https://papers.nips.cc/paper/4824-imagenet-classification-with-deep-convolutional-neural-networks), adapting the high-level visual representations learned by ImageNet classification networks to pose estimation. > 🔗 **Borrowed.** PoseNet uses the [GoogLeNet (Inception, Szegedy et al. 2014)](https://arxiv.org/abs/1409.4842) architecture as its backbone. It replaces the classification head with a 7-dimensional regression head for x, y, z, and four quaternion components, directly adapting an ImageNet feature hierarchy to localization. On the Cambridge Landmarks dataset collected by Kendall, which included King's College Chapel, streets, a former hospital, and several other outdoor scenes, PoseNet achieved position errors of around 2 m and orientation errors of 5–8°, depending on the scene (§5 of the original paper). A single GPU produced a pose within 5 ms without feature extraction, RANSAC, or map lookup. The paper triggered immediate follow-ups. [Bayesian PoseNet (Kendall & Cipolla 2016)](https://arxiv.org/abs/1509.05909) tried to estimate pose uncertainty via Monte Carlo Dropout. LSTM PoseNet integrated sequence information. Variants with added geometric loss appeared. Kendall himself released a 2017 version combining a recurrent structure with photometric loss. As benchmark performance improved, however, the gap became visible. On the same scenes, [Active Search (Sattler et al. 2012)](https://www.graphics.rwth-aachen.de/media/papers/sattler_eccv12_preprint_1.pdf) and DenseVLAD achieved position errors of roughly 0.2 m. The PoseNet family rarely improved beyond errors of several meters, revealing a fundamental limitation of regressing an absolute pose from one image. --- ## 12.2 DeepVO To address PoseNet's reliance on a single image, Sen Wang of Heriot-Watt University in Edinburgh and his co-authors used image sequences. They presented [Wang et al. 2017. DeepVO](https://arxiv.org/abs/1709.08429) at ICRA in 2017. A CNN influenced by FlowNet extracted optical-flow features from consecutive frame pairs, and an LSTM accumulated temporal context to estimate visual odometry directly. > 🔗 **Borrowed.** DeepVO's training labels are KITTI's GPS/IMU ground truth, and its feature extraction design is borrowed directly from the optical flow CNN architecture of [FlowNet (Dosovitskiy et al. 2015)](https://arxiv.org/abs/1504.06852). The LSTM was intended to use temporal context to suppress drift. DeepVO showed lower drift than DVO-SLAM or VISO2-M on parts of the KITTI sequences, but only under conditions resembling the training data in driving patterns, lighting, urban appearance, and speed profiles. When conditions differed, the accumulated context became a source of bias. [Zhou et al. 2017. SfMLearner](https://arxiv.org/abs/1704.07813), released the same year by Tinghui Zhou at UC Berkeley, took a different approach. It jointly estimated depth and ego-motion through self-supervised learning, using photometric reprojection loss as the training signal and requiring no labels. > 🔗 **Borrowed.** SfMLearner's photometric loss is mathematically identical to the intensity residual in classical direct SLAM. It placed the photometric principle of [DSO (Engel et al. 2018)](https://arxiv.org/abs/1607.02565) in a differentiable learning framework, and its self-supervision later appeared in MonoDepth2 and DROID-SLAM. SfMLearner's evaluation on short VO snippets must be distinguished from ORB-SLAM's system performance over longer trajectories with loop closure. --- ## 12.3 Three causes of failure Between 2019 and 2020, researchers increasingly examined the limitations of this work. In a 2019 talk, Sudeep Pillai of MIT, later TRI, organized the structural problems of end-to-end methods into three categories. **First: absence of inductive bias.** Classical SLAM explicitly encoded geometric constraints accumulated over decades, including the epipolar constraint, rigid-body motion, scale invariance, and spatial continuity. A CNN had to infer them from data, while ImageNet classification offered little supervision for the metric geometry of 3D space. Even when a regression network estimated the correct pose, it was difficult to determine whether it had learned 3D structure or memorized combinations of lighting, color, and texture. **Second: generalization failure.** Performance collapsed outside the training distribution. A PoseNet trained on Cambridge Landmarks was unusable on Oxford streets, and a DeepVO trained on KITTI accumulated exponentially growing drift on other vehicle datasets without radar. Classical ORB-SLAM also failed when feature detection broke down or lighting changed drastically, but its failure was predictable and the system could reinitialize. End-to-end models could instead return an incorrect estimate without indicating the magnitude of the error. **Third: absence of uncertainty quantification.** SLAM cannot function only as a pose estimator because downstream systems, including path planning and obstacle avoidance, require the covariance of the localization estimate. EKFs and factor graphs propagate covariance naturally. Bayesian PoseNet attempted to estimate variance through dropout, but the calibration between that variance and actual position error was difficult to verify. On inputs outside the training distribution, it could return a confident but incorrect estimate, which is more dangerous for a robotic system than a visible failure. --- ## 12.4 Reassessment After completing his PhD in 2019, Kendall moved to Wayve and shifted toward imitation learning and world-model research for autonomous driving. He continued to study learning-based localization but rejected absolute pose regression from a single image as the appropriate problem formulation. Earlier, in 2017, Federico Tombari's group at TU Munich, later Google, developed [CNN-SLAM (Tateno et al. 2017)](https://arxiv.org/abs/1704.03489). It fused dense depth predicted by a CNN with the depth estimate from direct monocular SLAM. Because learning was confined to dense depth, the method was not fully end to end, but it tested whether a CNN could address scale ambiguity and low-texture regions in monocular SLAM. Results varied across scenes, and the method did not consistently improve accuracy. > 📜 **Prediction vs. outcome.** In the PoseNet paper (2015), Kendall identified uncertainty estimation, temporal integration, and extension to larger scenes as the next tasks. Bayesian PoseNet (2016), LSTM PoseNet (2016), and multiple outdoor experiments pursued all three directions. Each encountered further limitations, and researchers ultimately abandoned the broader absolute-pose-regression approach. The proposed extensions could not overcome the weakness of the underlying formulation. Some components remained useful in other settings. SfMLearner's photometric self-supervision continued in monocular-depth methods such as MonoDepth2 (Godard 2019). DROID-SLAM (Teed & Deng 2021) also uses differentiable geometry, but trains with pose and optical-flow supervision. DeepVO's LSTM-based temporal modeling also reappeared in modified form in visual-inertial learning research. The methods changed even as their individual ideas persisted. > 📜 **Prediction vs. outcome.** In the SfMLearner paper (2017), Zhou identified dynamic-object handling and robustness to photometric noise as remaining tasks. Later self-supervised work, including [GeoNet (Yin & Shi 2018)](https://arxiv.org/abs/1803.02276), made partial progress. Self-supervised VO did not replace SLAM in the mainstream, but photometric self-supervision persisted after the field rejected end-to-end VO as the larger objective. --- ## 12.5 The lesson that remained Around 2020, a common design rule was to keep geometry explicit and use learning for features and priors. > 🔗 **Borrowed.** CodeSLAM (Bloesch 2018) and DROID-SLAM (Teed & Deng 2021) implement this principle. Both retain the geometric structure of a factor graph or bundle adjustment and use learning for depth representation in CodeSLAM and for dense correspondence and recurrent updates in DROID-SLAM, retaining constraints that PoseNet had removed. The classical pipeline was not uniformly superior to learning-based alternatives. ORB-SLAM also failed in textureless environments, at night, and in rain. The distinction was that errors from end-to-end models were less interpretable and predictable than failures in classical SLAM. Neither a particular dataset nor a particular architecture fully explained the failure. A direct mapping from image to pose omitted the geometric constraints accumulated over thirty years. --- ## 🧭 Still open **Which inductive bias to inject, and how.** Even if geometry remains part of the algorithm, it is unclear which constraints should be encoded and at what level. Candidates include rigid-body motion and the epipolar constraint. Foundation models are again blurring this boundary, while GaussianSLAM and 3DGS-based systems explore how much geometry can reside within a learned representation. **Calibration of learned uncertainty.** Bayesian PoseNet did not resolve this problem. Whether uncertainty estimates from deep learning remain calibrated to actual error, especially for out-of-distribution inputs, was still open as of 2026. Autonomous-driving applications make the question practically urgent. **Redefining "end-to-end."** PoseNet's definition of end to end, learning a direct image-to-pose mapping, failed. Since foundation models appeared in 2023, however, the term has begun to change. Researchers are again deciding which SLAM modules should be learned and which should remain explicit algorithms. CodeSLAM, released by Andrew Davison's lab at Imperial College London in 2018, embodied this division between explicit geometry and learned representation. --- # Ch.13 — The Hybrid Victory: From CodeSLAM to DROID-SLAM When Michael Bloesch presented CodeSLAM at CVPR 2018, he was affiliated with the Dyson Robotics Lab at Imperial College London and advised by Andrew Davison. Richard Newcombe had built DTAM in the same lab in 2011. Jan Czarnowski would release DeepFactors there in 2020, and Edgar Sucar and Tristan Laidlow would continue related work. Bloesch combined Davison's view of SLAM as probabilistic inference, developed since 2002, with the mid-2010s premise that representations could be learned. --- ## 13.1 CodeSLAM — latent code and the map Traditional monocular SLAM estimated depth as an optimization variable, whether for a few hundred sparse landmarks or, as in [DTAM](https://www.doc.ic.ac.uk/~ajd/Publications/newcombe_etal_iccv2011.pdf) (Newcombe et al. 2011), for every pixel. The dimension of this variable space grew with image resolution. A 640×480 dense depth map for one keyframe contains 307,200 independent variables, making optimization expensive, initialization sensitive, and prior information difficult to incorporate. [Bloesch et al. 2018. CodeSLAM](https://doi.org/10.1109/CVPR.2018.00271) optimized a low-dimensional **latent code** that generated the depth map rather than optimizing each depth value directly. A variational autoencoder (VAE) trained on real depth distributions learned a bottleneck space that approximated the manifold of realistic depth maps. Restricting optimization to this manifold reduced the number of variables from hundreds of thousands to hundreds. > 🔗 **Borrowed.** CodeSLAM's latent depth representation uses the encoder-decoder latent-space structure established in [Kingma & Welling 2013. VAE](https://arxiv.org/abs/1312.6114). Training follows the VAE framework, but during SLAM inference, **z** becomes a MAP optimization variable without stochastic sampling. A representation developed for generative image models became, a decade later, a low-dimensional space for SLAM optimization. For each keyframe, a VAE encoder extracts a latent code **z** from the image, and a decoder reconstructs a dense depth map from that code. The system jointly optimizes the camera pose and **z**. A photometric loss enforces consistency, while a latent prior regularizes **z** toward its prior distribution. Written as an equation, the objective is: $$E(\mathbf{z}, T) = \sum_{i,j} \rho\bigl(I_j(\pi(T_{ij}, D_\mathbf{z}(u_i), u_i)) - I_i(u_i)\bigr) + \lambda \|\mathbf{z}\|^2$$ $D_\mathbf{z}$ is the decoder, $\pi$ is the projection, $\rho$ is a robust cost, and $T_{ij}$ is the relative pose between keyframes. The latent-prior term $\lambda\|\mathbf{z}\|^2$ corresponds to the negative log-likelihood of the standard normal prior $p(\mathbf{z}) = \mathcal{N}(0, I)$ and arises directly from MAP inference under a Gaussian prior. > 🔗 **Borrowed.** Dellaert and Kaess's [GTSAM](https://gtsam.org/tutorials/intro.html) factor graph provided the backend structure for DeepFactors. Czarnowski reformulated CodeSLAM's joint pose-latent optimization as an explicit factor graph in which a learned latent variable appeared alongside conventional pose nodes. Graph edges expressed the coupling between learned and geometric variables. CodeSLAM surpassed earlier methods in reconstructing geometry from sparse input, but it did not operate in real time. Both VAE inference and the optimization loop were slow, a limitation the paper stated explicitly. > 📜 **Prediction vs. outcome.** CodeSLAM showed that dense SLAM could incorporate a compact learned representation, but remained limited in speed and scale. DeepFactors (2020), from the same Imperial group, moved closer to real-time operation without reaching deployment-grade performance. Teed and Deng at Princeton later achieved broader support for monocular, stereo, and RGB-D inputs through a different design: a learned frontend combined with Dense Bundle Adjustment (DBA). --- ## 13.2 DeepFactors — Imperial Dyson Lab, factor graph integration In 2020, Jan Czarnowski, also supervised by Davison at the Imperial Dyson Robotics Lab, released [Czarnowski et al. 2020. DeepFactors](https://doi.org/10.1109/LRA.2020.2969036). He sought to integrate CodeSLAM's approach into a complete SLAM pipeline. DeepFactors extended CodeSLAM's latent-depth representation and joint optimization into an explicit factor graph, explicitly separated tracking from mapping, and introduced a keyframe-selection criterion. On an NVIDIA GTX 1080, tracking against keyframes ran at about 250 Hz, but computing the network Jacobian took several hundred milliseconds per keyframe and became the pipeline's bottleneck. The system demonstrated the approach without achieving deployment-grade real-time performance. DeepFactors placed a learned representation in a factor graph as a node and applied geometric optimization in its latent space. Czarnowski's result favored replacing selected pipeline components with learnable modules rather than replacing the full pipeline end to end. Around the same time, Daniel Cremers's group at TU Munich arrived at the same design principle from a different starting point. The Davison group paired CodeSLAM's VAE latent variables with a factor graph, whereas the Cremers group inserted neural predictions into its 2016 direct sparse odometry system, [DSO](https://arxiv.org/abs/1607.02565). [Yang, Wang, Stückler, Cremers 2018. DVSO](https://arxiv.org/abs/1807.02570) added neural depth to monocular DSO as "virtual stereo," synthesizing a second camera observation in a monocular setting. [Yang, von Stumberg, Wang, Cremers 2020. D3VO](https://arxiv.org/abs/2003.01060) incorporated depth, pose, and uncertainty as three self-supervised predictions and additional factors in DSO's factor graph. [Wimbauer et al. 2021. MonoRec](https://arxiv.org/abs/2011.11814) and [Wimbauer et al. 2023. Behind the Scenes](https://arxiv.org/abs/2301.07668) extended the approach to dense reconstruction of dynamic scenes and single-view density fields. Although the Cremers and Imperial groups were institutionally separate, both integrated neural predictions into classical optimization. Princeton produced another form of the same hybrid design in 2021. --- ## 13.3 RAFT — recurrent optical flow Zachary Teed and Jia Deng at Princeton presented [Recurrent All-Pairs Field Transforms (RAFT)](https://arxiv.org/abs/2003.12039) at ECCV 2020. RAFT addressed optical-flow estimation rather than SLAM. RAFT's design later became the core of DROID-SLAM and consists of three parts. 1. Feature encoder: a CNN extracts feature maps from two images. 2. Correlation volume: the method stores similarities between all pixel pairs in a 4D volume with a four-level pyramid. 3. Update operator: a Gated Recurrent Unit (GRU) iteratively queries the correlation volume and refines the flow field. The phrase "all-pairs" describes the method's distinguishing feature. Rather than examining only selected neighboring pixels, RAFT considers every candidate location and progressively refines the flow field at a fixed resolution. Unlike earlier coarse-to-fine methods such as PWC-Net, it keeps the flow field at one-eighth of the input width and height while querying the correlation pyramid, then upsamples the output. Against the best published results available at the time, the RAFT paper reported a 16% relative reduction in KITTI F1-all error and a 30% reduction in Sintel final-pass endpoint error. RAFT did not originate in SLAM, but Teed recognized a structural similarity between its update operator and iterative bundle adjustment. A GRU refining a flow field could play a role analogous to an optimization step that refines pose and depth. --- ## 13.4 DROID-SLAM — the update operator and BA At NeurIPS 2021, Teed and Deng presented [DROID-SLAM](https://arxiv.org/abs/2108.10869). DROID stands for "Differentiable Recurrent Optimization-Inspired Design." Its architecture makes the hybrid design explicit. The frontend has the same structure as RAFT. A CNN encoder extracts feature maps, an all-pairs correlation volume is built, and a GRU update operator iteratively estimates dense flow. The difference is that flow is estimated not between a single pair of images but simultaneously on every edge of a keyframe graph. The backend performs DBA with pose and inverse depth as optimization variables. It uses the 2D correspondences from flow estimation as constraints for their joint optimization and solves the linear system efficiently with the Schur complement. The **DBA layer** connects the two components. Flow and uncertainty estimates from the GRU enter DBA; the updated pose and depth then provide the reference for the next GRU iteration, forming a recurrent loop. > 🔗 **Borrowed.** The idea behind DBA reaches back ten years to Newcombe's DTAM (2011), a precursor of photometric bundle adjustment over every pixel. DROID-SLAM combined that approach with the more robust correspondences supplied by learned flow, despite the separate institutional lineages of Newcombe and Teed. > 🔗 **Borrowed.** Teed and Deng adapted DROID-SLAM's update operator directly from their work on RAFT. All-pairs recurrent refinement, originally designed for optical flow, proved structurally compatible with the iterative optimization of bundle adjustment. On the EuRoC MAV dataset, DROID-SLAM recorded a lower RMSE ATE than ORB-SLAM3, then the state of the art. Evaluation covered synthetic TartanAir data as well as real indoor and outdoor sequences. The method was particularly robust to lighting changes and scarce texture compared with feature-based systems. In his later Handbook retrospective, Teed reported that on EuRoC V1_02, global optimization reduced the frontend-only ATE from 16.5 cm to 1.2 cm. Classical BA reached single-digit-centimeter accuracy using constraints supplied by learned correspondences. DROID-SLAM differed from the end-to-end methods in Ch.12 by separating the roles of learning and geometry. PoseNet regressed pose directly without geometric constraints and did not generalize. Teed and Deng instead used learning for dense correspondence estimation and BA for enforcing geometric constraints. Neural networks handled feature extraction and dense matching, while geometric optimization maintained consistency and propagated uncertainty. The system replaced hand-designed features but retained the optimization structure, distinguishing the 2021 hybrid from the 2015 end-to-end approach. --- ## 13.5 The Imperial Dyson Lab lineage The Imperial Dyson Robotics Lab connected CodeSLAM with several later hybrid systems. Andrew Davison led SLAM research at Imperial for twenty years after MonoSLAM in 2002. His students and collaborators developed several related lines of work. - **Richard Newcombe** (Davison advisee, Imperial): DTAM (2011), [KinectFusion](https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/ismar2011.pdf) (2011); later Oculus → Meta Reality Labs. - **Michael Bloesch** (Davison advisee, Imperial): CodeSLAM (2018), touch and inertial SLAM research. - **Jan Czarnowski** (Davison advisee, Imperial): DeepFactors (2020). - **Edgar Sucar** (Davison group, Imperial): [iMAP](https://arxiv.org/abs/2103.12352) (2021), later extended to the NeRF-SLAM lineage. - **Tristan Laidlow** (Davison group, Imperial): dense 3D reconstruction, later extended to the neural implicit SLAM lineage. The group retained its emphasis on factor graphs and uncertainty as representations changed from sparse monocular landmarks to dense latent variables and then to implicit fields. In [FutureMapping](https://arxiv.org/abs/1803.11288) (2018) and [FutureMapping 2](https://arxiv.org/abs/1910.14139) (2019, with Ortiz), Davison proposed that a Spatial AI map should include both computational structure and multiple representations. The papers argued for combining diverse geometric and semantic representations in one probabilistic graph. CodeSLAM and DeepFactors were early experiments in that program. Teed and Deng developed DROID-SLAM independently at Princeton. Nevertheless, DTAM by Newcombe and Davison was a conceptual predecessor of DROID-SLAM's DBA, showing that technical lineages need not follow direct institutional connections. --- ## 13.6 2023–2025 — extensions after DROID After DROID-SLAM, researchers both extended its design and pursued alternatives. [GO-SLAM](https://arxiv.org/abs/2309.02436) (Zhang et al. 2023) extended DROID-SLAM's tracking with online loop closure and full bundle adjustment, while mapping with an Instant-NGP-style neural implicit representation based on multi-resolution hash encoding. It combined DROID-family dense flow and BA for tracking with an implicit map, adding another hybrid layer. [NICER-SLAM](https://arxiv.org/abs/2302.03594) (Zhu et al. 2023) followed a different design. Rather than building on DROID, it solved tracking and mapping jointly over a single hierarchical neural implicit representation. It pursued the same goal of RGB-only dense SLAM from outside the DROID lineage. [SplaTAM](https://arxiv.org/abs/2312.02126) (Keetha et al. 2024) replaced the map representation with 3D Gaussian Splatting and based tracking on silhouette-guided differentiable rendering rather than DROID-style dense flow. It belongs more directly to the 3DGS lineage than to the DROID lineage. [DPV-SLAM](https://arxiv.org/abs/2408.01654) (Lipson, Teed, Deng 2024) came from the same Princeton group as DROID-SLAM. Built on [DPVO](https://github.com/princeton-vl/DPVO) (Deep Patch Visual Odometry) rather than DROID, it added proximity-based loop closure and CUDA block-sparse BA. The resulting system was roughly 2.5× faster and used less memory than DROID-SLAM. Its main change was a patch-based sparse representation combined with efficient loop closure rather than a simple feature replacement. Outside the DROID lineage, work in 2024–2025 followed Naver Labs' [DUSt3R](https://arxiv.org/abs/2312.14132) (Wang et al. 2023). DUSt3R changed the conventional SfM procedure by directly predicting pointmaps from two images, as Ch.16 discusses in detail. The same Revaud group then added symmetric multi-view processing and working memory in [Cabon et al. 2025. MUSt3R](https://arxiv.org/abs/2503.01661), extending the image-pair formulation to many frames in an attempt to support both offline SfM and online VO/SLAM with one network. DROID-family components also appeared in this ecosystem. [Li et al. 2024. MegaSAM](https://arxiv.org/abs/2412.04463) extended DROID-SLAM's differentiable DBA to dynamic scenes and uncalibrated video, jointly optimizing camera intrinsics at inference time. NVIDIA's [Huang et al. 2025. ViPE](https://arxiv.org/abs/2508.10934) combined DROID-SLAM's dense-flow network, cuvslam's sparse points, and a monocular depth network as three constraints in one DBA, producing a large-scale annotation pipeline for unconstrained video at YouTube scale. These systems applied DROID's learned-frontend, classical-backend design to the harder conditions of calibration-free and dynamic scenes. Several approaches developed in parallel: GO-SLAM placed a neural map above DROID tracking; DPV-SLAM redesigned the system around lightweight patch odometry; NICER-SLAM and SplaTAM based tracking on implicit or splatting representations; and MegaSAM and ViPE extended DROID's DBA to uncalibrated and dynamic settings. Teed and Deng's 2021 combination of a learned frontend and classical backend supplied a starting point for several successors, while NICER-SLAM and SplaTAM pursued the same broader problem through separate representations. > 📜 **Prediction vs. outcome.** DROID-SLAM established a reference architecture that combined differentiable DBA with end-to-end learning. Three years later, the same group improved efficiency with DPV-SLAM. GO-SLAM, NICER-SLAM, and SplaTAM instead replaced the map with implicit or Gaussian-splatting representations. As of 2026, these variations on DROID's "learned frontend plus classical backend" design had not converged on a general-purpose solution. --- ## 🧭 Still open Generalization of learned priors beyond the training distribution remains a primary problem. The VAEs in CodeSLAM and DeepFactors learn the depth distribution of their training data. In substantially different environments, including open outdoor scenes, non-uniform textures, and nighttime conditions, a learned prior can direct optimization toward an incorrect solution. DROID-SLAM's flow estimator also performs worse outside its training domain. As of 2026, no learned SLAM system operated reliably in every environment. Training on diverse synthetic data such as TartanAir helps, but a sim-to-real gap remains. Real-time operation remains another constraint. DROID-SLAM averages about 10–15 fps on an NVIDIA RTX 2080Ti and slows as the keyframe graph grows, with DBA as the bottleneck. As of 2026, it was still impractical for low-power applications requiring at least 30 Hz, such as mobile robots and AR/VR. Attempts to reduce computation through fewer keyframes or approximate BA incur performance trade-offs. Integrating learned loop closure also remains unresolved. DROID-SLAM does not handle loop closure explicitly. In his later Handbook retrospective, Teed wrote that "DROID-SLAM doesn't include any relocalization module, so large loops with lots of drift cannot be closed." Unlike the local frontend, the backend performs global BA over the keyframes. Without a separate relocalization module, however, it can fail to recover the correspondences needed to close large loops after substantial drift. Some efforts have attempted to integrate learned loop closure, including the place-recognition work in Ch.10, into DROID's factor graph, but they have not converged on a single system. The connection between the NetVLAD lineage in Ch.10 and the DROID lineage in Ch.13 remains open. --- DROID-SLAM used an inverse-depth map, a strong dense representation in 2021, but [NeRF](https://arxiv.org/abs/2003.08934) (Neural Radiance Field) had proposed a different option in 2020: representing a scene as a continuous function rather than as points, lines, planes, or meshes. Differentiable rendering offered another way to enforce photometric consistency. At Imperial College in 2021, Edgar Sucar treated this as a SLAM problem rather than a rendering task: could an MLP replace the TSDF voxel grid entirely? His answer, iMAP, followed after fourteen months of work. --- # Ch.14 — NeRF Enters SLAM: iMAP → NICE-SLAM In 2021, Edgar Sucar at Imperial College used a learned representation not for tracking but for *the map itself*. His iMAP system adapted NeRF, a method developed outside SLAM. In March 2020, Ben Mildenhall and his colleagues posted [Mildenhall et al. 2020. NeRF](https://arxiv.org/abs/2003.08934) to arXiv. The paper synthesized photorealistic novel views from calibrated images, and the SLAM community initially treated it as a rendering method rather than a map-building technique. Fourteen months later, Sucar's ICCV 2021 presentation of iMAP demonstrated that NeRF could serve as the map representation itself. Extending the KinectFusion lineage from Ch.9, iMAP was an early system testing whether an implicit neural field could replace a TSDF voxel grid. NeRF followed several coordinate-based MLP representations of 3D introduced around 2019. [Park et al.'s DeepSDF](https://arxiv.org/abs/1901.05103) represented object surfaces with an MLP that mapped coordinates to signed distance. [Mescheder et al.'s Occupancy Networks](https://arxiv.org/abs/1812.03828) mapped the same input to occupancy probability. [Sitzmann et al.'s SRN](https://arxiv.org/abs/1906.01618) stored a scene feature vector at each coordinate and formed images through differentiable ray marching. All three mapped coordinates to a field value. Mildenhall et al.'s 2020 NeRF added volume-rendering integration and positional encoding, applying the formulation to view synthesis. iMAP inherited this entire one-year lineage rather than one paper alone. --- ## NeRF: MLP-based spatial representation NeRF represents an entire 3D space implicitly in one MLP. Its inputs are a spatial coordinate $(x, y, z)$ and viewing direction $(\theta, \phi)$; its outputs are the color $(r, g, b)$ and density $\sigma$ at that position. Volume rendering combines these local values into an image of the full scene. Rendering uses the volume rendering equation. A ray leaving camera origin $\mathbf{o}$ in direction $\mathbf{d}$ is sampled along parameter $t$: $$\hat{C}(\mathbf{r}) = \int_{t_n}^{t_f} T(t)\,\sigma\!\left(\mathbf{r}(t)\right) \mathbf{c}\!\left(\mathbf{r}(t), \mathbf{d}\right)\, dt$$ Here $T(t) = \exp\!\left(-\int_{t_n}^{t} \sigma(\mathbf{r}(s))\, ds\right)$ is the accumulated transmittance, or the probability that the ray reaches $t$ without being blocked. A piecewise Riemann sum approximates the integral. > 🔗 **Borrowed.** The volume-rendering equation comes from the classical graphics paper [Kajiya & Von Herzen (1984)](https://courses.cs.duke.edu/cps296.8/spring03/papers/RayTracingVolumeDensities.pdf). For nearly forty years, it served as a physics-based tool for offline rendering. Mildenhall used differentiable rendering to turn it into an optimization objective for reconstructing a scene from images. Mildenhall et al. (2020) used positional encoding to address the difficulty MLPs have in learning high-frequency spatial signals. Projecting coordinates $(x, y, z)$ through sine and cosine functions at multiple frequencies allows the network to learn fine textures and sharp boundaries: $$\gamma(p) = \left(\sin(2^0 \pi p),\, \cos(2^0 \pi p),\, \ldots,\, \sin(2^{L-1} \pi p),\, \cos(2^{L-1} \pi p)\right)$$ > 🔗 **Borrowed.** NeRF's positional encoding appeared in the original Mildenhall et al. (2020) paper. The same year, [Tancik et al. (2020)](https://arxiv.org/abs/2006.10739) "Fourier Features Let Networks Learn High Frequency Functions" explained why this technique works through neural tangent kernel (NTK) theory. NeRF trains by comparing images captured from known camera poses with rendered outputs and minimizing a per-pixel L2 loss. After optimization, the MLP weights encode the scene's geometry and appearance without explicit voxels or a mesh. The original NeRF also had clear limitations. Training took hours, and each model represented one scene whose camera poses had already been estimated by an external SfM system such as COLMAP. A SLAM system instead had to estimate poses and learn the map jointly at close to real-time speed. --- ## iMAP: early neural implicit SLAM At ICCV 2021, Edgar Sucar of Imperial College's Dyson Robotics Lab presented [Sucar et al. 2021. iMAP](https://doi.org/10.1109/ICCV48922.2021.00612). **iMAP** (Implicit MAP) took RGB-D input, optimized camera poses, and represented the map with a single MLP. The system alternates between two optimization loops. The *mapping* loop samples rays from the current keyframe and a random set of past keyframes to update the MLP. The *tracking* loop freezes the MLP and optimizes the current frame's pose against the rendering loss. Both loops use the same MLP. The objective has two terms: a color loss $\mathcal{L}_{\text{color}} = \|\hat{C} - C\|_2^2$ and a depth loss $\mathcal{L}_{\text{depth}} = \|\hat{D} - D\|_2^2$. Because iMAP uses RGB-D input, direct depth supervision stabilizes geometry learning. iMAP was a proof of concept that operated on small indoor scenes but had two structural limitations. First, the single MLP forgot earlier regions as it learned new ones, exhibiting catastrophic forgetting. Keyframe replay partially mitigated this effect without removing its cause. Second, the MLP's representational capacity became insufficient as scenes grew because every forward pass treated the entire space as one function. > 📜 **Prediction vs. outcome.** In iMAP's conclusion, Sucar wrote that "future directions for iMAP include how to make more structured and compositional representations that reason explicitly about the self similarity in scenes." Structured and compositional representations became central to later work. Five months later, an ETH Zürich pre-release of NICE-SLAM partitioned space hierarchically with a multi-resolution voxel feature grid. Wang et al.'s Co-SLAM (2023) later combined hash-grid and coordinate encodings and reported 10–17 Hz on an RTX 3090 Ti. Explicit reasoning about self-similarity received less attention in mainstream NeRF-SLAM, and approaches that refined a single MLP also became less central. --- ## NICE-SLAM: hierarchical grid and scalability Zihan Zhu and Songyou Peng at ETH Zürich addressed iMAP's single-MLP limitation in [Zhu et al. 2022. NICE-SLAM](https://arxiv.org/abs/2112.12130), presented at CVPR 2022. **NICE-SLAM** (Neural Implicit Scalable Coding for SLAM) replaced the single MLP with a multi-resolution voxel feature grid and a small MLP decoder. NICE-SLAM partitions space into an explicit voxel grid and stores a learnable feature vector at each voxel. During rendering, trilinear interpolation combines features from the voxels surrounding a sample coordinate, and a small MLP decodes them into color and occupancy. Most spatial information resides in the grid, reducing the required MLP size. NICE-SLAM organized three grid resolutions hierarchically. The coarse-to-fine grids represented geometry at multiple levels, while a separate color feature grid and decoder represented appearance. Adding a new region required updating only its corresponding voxel features, substantially reducing catastrophic forgetting elsewhere. Like iMAP, NICE-SLAM froze the MLP and grid features while optimizing pose during tracking. During mapping, it updated the grid features. On Replica and ScanNet, the system handled larger spaces than iMAP and reconstructed finer details. The grid introduced its own limitations because memory grew with the cube of resolution. The system could handle one or two indoor rooms but did not scale to multi-story buildings or outdoor environments. It also remained far from real-time operation. At SIGGRAPH 2022, Thomas Müller's [Müller et al. 2022. Instant-NGP](https://nvlabs.github.io/instant-ngp/) addressed this bottleneck with a hash-table-based feature encoding. The representation reduced the memory growth of voxel grids and cut training time from minutes to seconds. Although Instant-NGP was not a SLAM paper, most subsequent NeRF-SLAM systems adopted hash encoding. > 🔗 **Borrowed.** NICE-SLAM's multi-resolution feature grid was designed independently at roughly the same time as Instant-NGP's hash encoding, but Instant-NGP's hash grid soon replaced it in many NeRF-SLAM implementations. The feature grid also extends the grid-based representation of KinectFusion (Ch.9), replacing stored TSDF values with learned features. --- ## Co-SLAM and NeRF-SLAM: two integration directions From late 2022 onward, systems following iMAP and NICE-SLAM developed in two directions: more efficient implicit representations and combinations of a NeRF map with a classical SLAM backend. UCL's [Wang et al. (2023) **Co-SLAM**](https://arxiv.org/abs/2304.14377) followed the first direction. It combined a multi-resolution hash grid with one-blob encoding in a joint coordinate and parametric representation. The hash grid represented densely observed regions efficiently, while the coordinate encoding supplied a smooth prior over unobserved areas, balancing convergence speed with surface completeness. The paper reported 15–17 Hz on Replica with an RTX 3090 Ti, bringing NeRF-based SLAM close to real-time operation. At the same CVPR, [Johari et al.'s **ESLAM**](https://arxiv.org/abs/2211.11704) from Idiap and EPFL addressed the same problem differently. It replaced the 3D feature grid with multi-scale axis-aligned feature planes, reducing memory growth from $O(n^3)$ to $O(n^2)$, and decoded TSDF rather than volume density to accelerate convergence. Antoni Rosinol at MIT released [**NeRF-SLAM**](https://arxiv.org/abs/2210.13641) in 2023. It retained classical SLAM tracking and factor-graph optimization while replacing only the map representation with NeRF. A DROID-SLAM frontend supplied poses and dense depth; NeRF-SLAM used those poses, depths, and uncertainties to construct an Instant-NGP-based map in parallel. > 🔗 **Borrowed.** The NeRF-SLAM backend uses Dellaert's factor-graph optimization (Ch.6). NeRF changed the map representation while leaving the graph-based mathematics of pose estimation established after 2005 intact. Rosinol used a modular design, replacing only the map representation rather than introducing NeRF throughout the pipeline. Classical SLAM functions such as loop closure remained in place. --- ## The structural limitation of iMAP iMAP showed that a single MLP could represent an entire scene and be updated near real time while jointly optimizing pose, although its measured performance remained limited. A single MLP lacks locality: rendering any region requires evaluating the entire network. Learning a new region changes all weights and can degrade older regions through catastrophic forgetting. A growing scene also increases the spatial variation that one MLP must represent, requiring a larger network and more iterations. Representation capacity grows with parameter count, whereas scene complexity grows with spatial volume, making a nonlocal representation increasingly inefficient at larger scales. NICE-SLAM's grid, Instant-NGP's hash encoding, and Co-SLAM's dual encoding all addressed locality. Partitioning space allows each component to represent its own region, reducing interference with previously learned areas and decoupling the rendering cost of one region from the total scene size. --- ## 🧭 Still open **Real-time NeRF-SLAM.** As of 2023, iMAP and NICE-SLAM were far from real time. Co-SLAM reported 10–17 Hz on an RTX 3090 Ti but remained too slow for mobile and embedded robotic hardware. Gaussian Splatting (Ch.15) later addressed speed by returning to an explicit representation, while implicit neural fields still did not support conventional real-time SLAM at 30 fps or more without a consumer GPU. Instant-NGP greatly accelerated rendering, but the combined tracking-and-mapping loop remained constrained. **Large-scale outdoor environments.** Methods such as [Block-NeRF](https://arxiv.org/abs/2202.05263) (2022, Tancik et al.) partition space into many local NeRFs, but do not yet integrate cleanly with SLAM's requirements for loop closure and global consistency. City-scale NeRF-SLAM remains an open problem. **Semantic and editable implicit maps.** Because a NeRF map is optimized for rendering, inserting semantic labels and editing the map afterward are difficult. Removing an object or reclassifying a region is substantially harder than in a TSDF or point cloud. [LERF](https://arxiv.org/abs/2303.09553) aligns language features with space for semantic queries; it does not itself provide object deletion or editing. Semantic representations and editing are explored using tools such as [Nerfstudio](https://arxiv.org/abs/2302.04264), but real-time integration with SLAM remained a research problem as of 2026. --- As iMAP and NICE-SLAM developed implicit fields, other researchers reconsidered explicit map representations. Millions of small ellipsoids placed throughout space promised faster rendering and more intuitive editing than a map encoded in MLP weights or a feature grid. Splatting already had predecessors such as EWA splatting in 2001. Bernhard Kerbl's SIGGRAPH 2023 paper demonstrated the combination of trainable 3D Gaussians and a fast differentiable renderer. --- # Ch.15 — The Gaussian Splatting Era: From 3DGS to GS-SLAM iMAP and NICE-SLAM represented space with an MLP, but their representations were difficult to edit directly. Updating the global MLP in iMAP could affect other regions; NICE-SLAM reduced that problem with local feature grids. NICE-SLAM also ran below 1 fps on an RTX 3090, far from real-time SLAM. Its scene representation remained difficult to inspect within the network parameters. At SIGGRAPH in August 2023, Bernhard Kerbl of INRIA, Georgios Kopanas, Thomas Leimkuhler, and George Drettakis presented [their paper](https://arxiv.org/abs/2308.04079). Kerbl retained the differentiable scene optimization that NeRF had developed over the preceding three years but changed the representation's form. Instead of encoding scenes in MLPs or voxel feature grids, as iMAP, NICE-SLAM, and Co-SLAM did, 3DGS represented them with millions of explicit ellipsoidal Gaussian primitives. The SLAM community adopted the representation within six months. Its rasterization drew on Matthias Zwicker's twenty-year-old EWA splatting technique (2001), while its optimization retained the differentiable-rendering framework associated with NeRF. --- ## The structure of 3DGS Kerbl represented scenes as an explicit set of Gaussians. Each Gaussian has a position (mean) $\boldsymbol{\mu} \in \mathbb{R}^3$, a covariance matrix $\boldsymbol{\Sigma} \in \mathbb{R}^{3 \times 3}$, an opacity $\alpha \in (0,1]$, and a color expressed in spherical harmonics coefficients. For training stability the covariance is factored into a scale vector $\mathbf{s}$ and a unit quaternion $\mathbf{q}$: $$\boldsymbol{\Sigma} = \mathbf{R}\mathbf{S}\mathbf{S}^\top\mathbf{R}^\top$$ Rendering alpha-blends the projected 2D Gaussians in depth order. Each Gaussian's effective opacity $\alpha_i$ is the product of the learnable opacity $\sigma_i$ and the 2D Gaussian density $G_i(\mathbf{x})$ evaluated at the pixel location. The pixel color $C$ is $$C = \sum_{i \in N} c_i \alpha_i \prod_{j
🔗 **Borrowed.** 3DGS's rasterization-based splatting descends directly from Zwicker et al.'s [EWA splatting (2001)](https://www.cs.umd.edu/~zwicker/publications/EWAVolumeSplatting-VIS01.pdf). Zwicker wrapped each point in an elliptical weighted-average kernel to render point clouds. Kerbl replaced that kernel with a learnable Gaussian and accelerated it with a GPU tile rasterizer. --- ## Structural fit between 3DGS and SLAM Implicit representations posed several difficulties for SLAM. An MLP-based NeRF had to update the entire network for each new observation, while catastrophic forgetting made incremental learning difficult. Expanding the map required increasing the network's capacity. NICE-SLAM's voxel grid mitigated these problems but retained a trade-off between resolution and memory. 3DGS addressed these problems with an explicit representation. When a new keyframe arrives, the system can add Gaussians only to the corresponding region. Densification aligns naturally with keyframe insertion, while rendering retains NeRF-level quality at real-time speed. GS-SLAM papers began appearing in quick succession in late 2023. --- ## GS-SLAM: an early attempt Chi Yan (HKU) and collaborators posted [Yan et al. 2023. GS-SLAM](https://arxiv.org/abs/2311.11700) on arXiv in November 2023. Several 3DGS-SLAM manuscripts appeared around the same time; GS-SLAM was among the early systems to combine 3DGS with tracking and mapping. GS-SLAM followed the classical SLAM framework: tracking estimated the pose of the current frame, and mapping updated the Gaussian map. Yan introduced two mechanisms. Adaptive Gaussian expansion inserted Gaussians into low-coverage regions when a new keyframe was added. Geometry-aware Gaussian selection optimized only Gaussians that contributed substantially to the rendering loss, reducing computation during backpropagation. Tracking optimizes the pose against a rendered photometric loss. GS-SLAM's tracking loss is an L1 color loss over sampled pixels: $$\mathcal{L}_{track} = \sum_m \|\mathbf{C}_m - \hat{\mathbf{C}}_m\|_1$$ During mapping, Yan used a weighted sum of color L1 and depth L1 losses. Training the Gaussian map also inherited the original 3DGS objective, $(1-\lambda)\mathcal{L}_1 + \lambda\mathcal{L}_{D\text{-}SSIM}$ with $\lambda=0.2$, as its default form. The differentiable rasterizer makes this loss differentiable with respect to pose. On the Replica dataset, GS-SLAM matched NICE-SLAM's PSNR at higher throughput. It nevertheless required an RGB-D camera and was not validated in large-scale outdoor environments. --- ## SplaTAM: silhouette-based densification [Keetha et al. 2024. SplaTAM (CVPR)](https://arxiv.org/abs/2312.02126), by Nikhil Keetha at Carnegie Mellon and his colleagues, used a simpler densification method than GS-SLAM based on a silhouette mask. The **silhouette mask** identifies regions in the current view that the existing Gaussians do not explain. SplaTAM adds new Gaussians to these empty areas of the rendered mask, using absence of coverage as a direct densification criterion. Tracking optimizes the pose, while mapping optimizes the Gaussian parameters. Tracking holds the map fixed, and mapping updates the Gaussians. GS-SLAM also separates pose and map variables, so separation alone does not distinguish SplaTAM. > 🔗 **Borrowed.** SplaTAM applies PTAM's keyframe-based map-management principle (Klein & Murray, 2007) to a new representation. Selective keyframe insertion, which PTAM used to maintain its map, becomes the trigger for Gaussian densification in SplaTAM. On the Replica dataset SplaTAM recorded PSNR 34.11 dB. In the same paper's table, NICE-SLAM recorded 24.42 dB. The rendering-quality gap was clear. In the Limitations & Future Work section of the 2024 CVPR paper, Keetha listed sensitivity to motion blur, depth noise, and aggressive rotation. He also identified removal of the dependence on known intrinsics and dense depth, along with improved scalability, as future work. > 📜 **Prediction vs. outcome.** In SplaTAM's Limitations section (2024), Keetha identified sensitivity to motion blur, depth noise, and aggressive rotation, as well as dependence on known intrinsics and dense depth. Matsuki's MonoGS (CVPR 2024) removed the depth requirement with a monocular RGB system in the same year. Work on unknown intrinsics and large-scale operation remained in progress in 2024–2025. --- ## MonoGS: monocular RGB [Matsuki et al. 2024. MonoGS (CVPR)](https://arxiv.org/abs/2312.06741), by Hidenobu Matsuki at Imperial College's Dyson Robotics Lab and his colleagues, removed the depth-sensor requirement. It performs 3DGS SLAM from a single monocular RGB camera. Scale is the central difficulty in the monocular setting because metric recovery without depth remains unsolved even in SfM. Matsuki optimized the Gaussian geometry directly and added a geometric-consistency loss between rendered depth and neighboring Gaussians. $$\mathcal{L}_{iso} = \sum_k \| \mathbf{s}_k - \bar{s}_k \mathbf{1} \|_1$$ Here $\mathbf{s}_k \in \mathbb{R}^3$ is the scale vector of the $k$-th Gaussian, and $\bar{s}_k = \frac{1}{3}\sum_j s_{k,j}$ is the mean scale across its three axes. This isotropy regularization prevents Gaussians from degenerating into excessively thin plates. Without depth supervision, monocular Gaussians tend to remain near the camera plane, and the regularizer suppresses that failure mode. For tracking, Matsuki directly optimizes pose with the photometric loss from Gaussian rendering. During first-frame initialization, a monocular depth prior provides the initial Gaussian positions. Later frames begin from the previous pose and undergo rendering-based refinement. Without a depth sensor, isotropic regularization and geometry-consistency losses across keyframes jointly maintain scale. > 🔗 **Borrowed.** MonoGS's monocular depth prior follows the Godard MonoDepth2 lineage covered in Ch.11. MonoGS applies the premise that self-supervised monocular depth can recover structure without direct depth supervision to Gaussian initialization. Matsuki worked at Imperial's Dyson Robotics Lab, as did Sucar on iMAP and Bloesch on CodeSLAM under Davison's supervision. MonoGS reflects the lab's transition from implicit MLPs to explicit Gaussian representations. On the TUM-RGBD dataset, MonoGS recorded an average ATE RMSE of 4.44 cm monocular and 1.58 cm RGB-D. Rendering quality in the RGB-D setting on Replica reached an average PSNR of 37.50 dB, comparable to other same-generation GS-SLAM systems. --- ## RTG-SLAM and real-time processing Speed remained a problem for GS-SLAM systems because neither GS-SLAM nor SplaTAM operated convincingly in real time. Peng Zhexi's group at Zhejiang University published [RTG-SLAM (SIGGRAPH 2024)](https://arxiv.org/abs/2404.19706) in 2024 with real-time performance as an explicit objective. RTG-SLAM controls the number of Gaussians by optimizing only those that contribute substantially to the current camera view. It initializes Gaussians from surfels (surface elements), preserving geometry with fewer primitives. On the Replica dataset, it approached real-time throughput. --- ## Representation choices after 3DGS The rapid growth of 3DGS papers after 2024 broadened the representation choices in SLAM mapping. TSDF and occupancy grids remain in use for embedded, planning, and safety applications, while NeRF and 3DGS are selected for different objectives. 3DGS papers emphasize rendering speed and explicit primitive updates, but that trend alone does not establish that the other representations left the mainstream. The shift reflected both representation design and compatibility with available hardware. GPU rasterizers are more highly optimized than GPU ray marchers, and 3DGS's ability to use the existing graphics pipeline accelerated its adoption relative to NeRF. > 🔗 **Borrowed.** 3DGS directly inherits differentiable scene optimization from NeRF. Mildenhall et al. (2020) established the use of gradients and a photometric loss to connect observations with rendering. Kerbl retained that framework while replacing the implicit MLP with explicit Gaussians. > 📜 **Prediction vs. outcome.** In §7.4, Limitations, of the 3DGS paper (2023), Kerbl et al. identified elongated artifacts and popping in sparsely observed regions, the absence of regularization, and memory consumption of more than 20 GB during training and several hundred MB when rendering large scenes. They proposed antialiasing, more principled culling, and point-cloud compression as future work. The [Compact 3DGS](https://arxiv.org/abs/2311.13681) line and [Niedermayr et al.](https://arxiv.org/abs/2401.02436) directly addressed compression in 2024. The original paper did not explicitly propose dynamic-scene extensions such as [4DGS](https://arxiv.org/abs/2310.08528) and [Deformable 3DGS](https://arxiv.org/abs/2309.13101), or generation and editing systems such as [DreamGaussian](https://arxiv.org/abs/2309.16653) and [GaussianEditor](https://arxiv.org/abs/2311.14521), but these became separate research directions around 2024. --- ## 🧭 Still open Memory scaling. The number of Gaussians grows linearly with scene size. A few hundred thousand primitives may suffice for the indoor Replica dataset, but an outdoor city block can require tens of millions. Researchers are studying Gaussian pruning and level-of-detail hierarchies, but no consensus exists on managing the trade-off between memory and rendering quality at large scale. The Compact 3DGS line (Lee et al. 2024, Niedermayr et al. 2024) explores compression. Semantic integration. In 2023, [LERF](https://arxiv.org/abs/2303.09553) combined language features with NeRF, while [LangSplat](https://arxiv.org/abs/2312.16084) combined them with a Gaussian representation, and later work coupled semantic Gaussians to SLAM. A common protocol that compares real-time updates, tracking quality, and semantic accuracy across scenes and hardware has not settled. Interference between jointly optimized semantic and geometric variables remains a central evaluation target. Dynamic scenes. 4DGS and Deformable 3DGS added a time dimension to Gaussians. In SLAM, dynamic objects move independently of the background and require separate treatment. GS-SLAM (Yan et al. 2023), SplaTAM (Keetha et al. 2024), and MonoGS (Matsuki et al. 2024) all retain a static-world assumption. Ch.15b separately traces SLAM's treatment of moving objects, from mask-based outlier rejection and multi-object factor graphs to deformable reconstruction. 3DGS also left its initialization unresolved. Gaussians could originate from an SfM point cloud or a depth sensor, but placing them required a known pose, while estimating a pose required an existing map. Systems handled this dependency differently: RGB-D methods used measured depth, while MonoGS formed initial depth hypotheses internally without an external depth predictor. DUSt3R and its successors, covered in Ch.16, instead learned geometry directly rather than initializing it from an existing representation. --- # Ch.15b — Where the Static-World Assumption Breaks: Dynamic and Deformable SLAM GS-SLAM, SplaTAM, and MonoGS all retained the static-world assumption. A separate line of research had addressed moving and deforming environments from the outset. In 2015, Javier Fuentes-Pacheco, Ruiz-Ascencio, and Rendón-Mancha published [*Visual simultaneous localization and mapping: a survey*](https://link.springer.com/article/10.1007/s10462-012-9365-8) in *Artificial Intelligence Review*. Its final section addressed "Dynamic and Deformable Environments." Earlier papers had studied moving objects, but most treated them as outliers for RANSAC to reject. The survey was an early account that treated dynamic environments as a distinct topic. Ten years later, the 2025 *SLAM Handbook* devoted 37 pages to it. Its six authors were Lukas Schmid, José María Martínez Montiel, Shoudong Huang, Daniel Cremers, José Neira, and Javier Civera. Although static scenes had been SLAM's starting point, self-driving cars, domestic service robots, and endoscopes all had to operate in changing environments. --- ## 15b.1 Three axes In Handbook Ch.15 §15.1, Schmid et al. revise the earlier definition of "dynamic SLAM." They define an environment as dynamic or static relative to *the observation*, not as an intrinsic property of the environment. The same physical motion may appear as a short-term change to one robot and a long-term change to another, depending on the ratio between observation rate $\text{Obs}$ and change rate $\text{Dyn}$. When $\text{Dyn} \ll \text{Obs}$, motion is visible between frames; when $\text{Dyn} \gg \text{Obs}$, the scene changes between visits. This perspective defines three axes. The observation axis distinguishes short-term from long-term change. The reconstruction axis distinguishes pose-only estimation, joint scene geometry, and full 4D spatio-temporal understanding. The time axis separates online from offline methods. Earlier accounts often reduced the field to removing dynamic objects, but that task occupies only one region of this three-axis space. Researchers working in different regions of the taxonomy had used the same terminology for different problems. --- ## 15b.2 Short-term: from masking to multi-object SLAM The earliest solution removed moving regions from the measurements. Berta Bescos, then a doctoral student at Zaragoza, published [DynaSLAM](https://arxiv.org/abs/1806.05620) in RA-L in 2018. The system added Mask R-CNN to the ORB-SLAM2 frontend, masking people and cars before excluding those regions from keypoint extraction. On the TUM-RGBD walking sequence, this direct approach reduced ATE to single-digit centimeters. During the same period, Martin Rünz at UCL tracked moving objects separately rather than removing them. Under Lourdes Agapito's supervision, he released [Co-Fusion (Rünz & Agapito, 2017)](https://arxiv.org/abs/1706.06629) and [MaskFusion (Rünz et al., 2018)](https://arxiv.org/abs/1804.09194) in consecutive years. Each object received its own surfel model, and the system jointly estimated camera and object trajectories. At ICRA 2018, Raluca Scona of Edinburgh and Stefan Leutenegger of Imperial presented another approach in [StaticFusion](https://arxiv.org/abs/1806.05628). It separated dynamic regions through residual clustering without semantic segmentation, avoiding errors from that component. The next approach included moving objects in the estimated state. Jun Zhang at QUT led [VDO-SLAM (Zhang et al., 2020)](https://arxiv.org/abs/2005.11052), which represented each dynamic object as a variable in the factor graph. The camera pose $T_i^w \in SE(3)$ and the pose of object $k$, $T_{k,i}^w \in SE(3)$, appeared in the same graph. A constant-velocity factor enforced continuity in each object's linear and angular velocities, and joint optimization operated over the product manifold of the camera and object SE(3) states. In 2021, Bescos at Zaragoza implemented the same principle on ORB-SLAM2 in [DynaSLAM II (Bescos et al., 2021)](https://arxiv.org/abs/2010.07820). Yuheng Qiu at CMU extended it to articulated objects such as human bodies in [AirDOS](https://arxiv.org/abs/2109.09903), published in RA-L in 2022. > 🔗 **Borrowed.** VDO-SLAM's factor-graph extension directly follows the iSAM tradition established by Dellaert and Kaess in the graph-SLAM work discussed in Ch.6. Dynamic SLAM represented a moving car by adding its state variables and measurement factors to the map. A third approach used inertial information. Song, Lim, Lee, and Myung at KAIST URL published [DynaVINS](https://arxiv.org/abs/2208.11500) in RA-L in 2022 without semantic masks or multi-object tracking. During bundle adjustment, the method reduced the factor weights of observations that disagreed with the pose prior from IMU preintegration, limiting the influence of dynamic features on the joint state. The same group's [DynaVINS++](https://arxiv.org/abs/2410.15373), published in RA-L in 2024, reformulated the method as adaptive truncated least squares. It also addressed the failure mode in which dynamic features affected IMU-bias estimation and caused divergence. The Handbook groups this work under §15.2.3, "Dense Dynamic SLAM," and presents Schmid's [Dynablox (Schmid et al., 2023)](https://arxiv.org/abs/2304.10049) as a current LiDAR MOS system. [AnyCam](https://arxiv.org/abs/2503.23282) (2025) uses a transformer backbone to recover 4D structure directly from ordinary video, extending the simultaneous tracking-and-reconstruction approach introduced by Rünz in 2017. --- ## 15b.3 Long-term: maps across time Short-term dynamics describe motion between frames, whereas long-term dynamics describe changes between visits, such as a chair moved overnight. Research on this problem followed a different lineage. Under Michaud's supervision at Sherbrooke, Mathieu Labbé developed [RTAB-Map](https://introlab.github.io/rtabmap/) beginning in 2013, drawing directly on models of human memory. It organized short-term, working, and long-term memory hierarchically and moved nodes according to time and observation frequency. A node remained in working memory during a session, moved to long-term memory if it was not revisited often, and was discarded if it no longer carried useful information. In a 2019 JFR paper, Labbé described how this structure scaled to multi-session SLAM. Hyungtae Lim, working under Ayoung Kim at KAIST, took a different approach in [ERASOR](https://arxiv.org/abs/2103.04316) (2021). He formulated map cleaning as scene differencing, identifying points that disappeared between two traversals of the same location. Handbook §15.3 repeatedly distinguishes **absence of evidence from evidence of absence**. A mapping system must determine whether an object has disappeared or was simply not observed. Without this distinction, map cleaning can erase valid objects and change detection can misclassify occluded regions. Schmid's [Panoptic Multi-TSDF](https://arxiv.org/abs/2109.10165), published in RA-L in 2022, addressed the problem with independent submaps for each object and a local-consistency rule for active and inactive states. The same group's [Khronos](https://arxiv.org/abs/2402.13817) (2024) further used graduated non-convexity for robust association. After loop closure, it performed deformable geometric change detection and estimated when each object changed, extending a metric-semantic map into a 4D spatio-temporal representation. > 🔗 **Borrowed.** Panoptic Multi-TSDF adapts the multi-map management introduced by Atlas in the ORB-SLAM work of Ch.7. It replaces keyframe submaps with panoptic object submaps while retaining the principle of partitioning a map that has become too large or heterogeneous. LiDAR research addressed the same problem separately. Jang, Lee, Nahrendra, and Myung at KAIST URL released [Chamelion](https://arxiv.org/abs/2602.08189) in 2026. It combined scene-mixing augmentation with a dual-head network to perform change detection without ground truth in transient environments such as construction sites and frequently rearranged indoor spaces. Khronos built a 4D representation from RGB-D and panoptic inputs, whereas Chamelion addressed long-term maintenance of point-cloud maps. Another line of work modeled recurring changes. Beginning in 2014, Tomáš Krajník and Achim Lilienthal at Örebro in Sweden developed **frequency maps**, representing periodic events such as commuter traffic and day-night lighting changes with a Fourier basis. In 2019, Martin Magnusson's group at the Stockholm Royal Institute of Technology consolidated this work as Maps of Dynamics (MoD), encoding *typical motion patterns* directly in the map. A statement such as "people usually walk to the left in this corridor" became part of the representation. [Changing-SLAM (Schmid et al., 2023)](https://arxiv.org/abs/2301.09479) combined a Kalman filter for short-term changes with semantic class matching for long-term changes in an ORB-SLAM extension. --- ## 15b.4 Deformable: when the shape itself changes Deformable SLAM addresses scenes in which even the background changes shape, a problem studied extensively by Civera and Montiel in Zaragoza. An earlier starting point was [DynamicFusion](https://grail.cs.washington.edu/projects/dynamicfusion/), the 2015 CVPR best paper by Newcombe, Fox, and Seitz at Microsoft Research. It placed an embedded deformation graph over KinectFusion's canonical TSDF to reconstruct non-rigid objects such as faces and torsos in real time. Each graph node carried a rotation and translation that the system optimized in every frame. In related work, Matthias Innmann at TU München added color information in [VolumeDeform](https://arxiv.org/abs/1603.08161) (2016). In 2017, Miroslava Slavcheva introduced [KillingFusion](https://campar.in.tum.de/pub/slavcheva2017cvpr/slavcheva2017cvpr.pdf), using Killing-vector-field regularization to permit topological changes such as a hand separating from the torso. Under Tedrake's supervision at MIT, Wei Gao's [SurfelWarp](https://arxiv.org/abs/1904.13073) (2019) used surfels instead of a TSDF to support exploration more readily. > 🔗 **Borrowed.** DynamicFusion directly adapted the embedded deformation graph published by Sumner, Schmid, and Pauly in computer graphics in 2007. A sparse control graph for mesh deformation became the variable representation for real-time non-rigid SLAM. Monocular deformable SLAM developed in Zaragoza. Juan Lamarca, who completed his doctorate under Montiel, published [DefSLAM](https://arxiv.org/abs/1908.08918) in RA-L in 2021. The system recomputed a template at each keyframe with isometric NRSfM and combined an ORB frontend with Lucas-Kanade optical flow to maintain tracks. It assumed planar topology. In 2023, Juan J. Gómez Rodríguez from the same group removed this limitation with [NR-SLAM](https://arxiv.org/abs/2308.04036), using a dynamic deformation graph for arbitrary topology and a visco-elastic model for temporal regularization. Handbook §15.4.2 groups this work as the "monocular line of deformable SLAM." Many applications are medical. Song at Tsinghua released [MIS-SLAM](https://ieeexplore.ieee.org/document/8458232) in 2018 to track intraoperative organ deformation with stereo endoscopy. The Jayender group at Children's National developed EMDQ (Expectation Maximization + Dual Quaternion), which estimated a smooth deformation field over SURF features. Both systems targeted intraoperative navigation for minimally invasive surgery. Handbook §15.4.1 highlights a fundamental problem: **Floating Map Ambiguity**. Without a prior, observations cannot distinguish the rigid motion of a non-rigid object from the camera's rigid motion. The image alone cannot determine whether a hand moved 30 cm or the camera moved 30 cm. This differs from the conventional scale ambiguity of monocular SLAM because scale, trajectory, and deformation become coupled in one ill-posed problem. DefSLAM and NR-SLAM partially constrain the ambiguity with isometric and visco-elastic priors, but no principled solution existed as of 2026. > 📜 **Prediction vs. outcome.** In §7, Future Work, of DynamicFusion (2015), Newcombe identified extension to larger scenes and topological changes, along with integration with loop closure, as the next challenges. KillingFusion addressed topological change in 2017, and the surfel-based SurfelWarp (2019) partially addressed larger scenes. Loop closure did not appear until Khronos introduced deformable geometric change detection in 2024, nine years later. --- ## 15b.5 One way to read the connections These works do not divide cleanly into three fixed schools. The papers, collaborations, and movement of researchers nevertheless reveal three overlapping lines. **The Zaragoza line** (Montiel, Neira, Civera, Lamarca, Rodríguez) pushed monocular geometry from MonoSLAM (Ch.5) and ORB-SLAM (Ch.7) into DynaSLAM, DefSLAM, and NR-SLAM. **The dense dynamic-reconstruction line** ran from KinectFusion to DynamicFusion and from SLAM++ to Co-Fusion and MaskFusion, linking work by Davison, Newcombe, Agapito, Rünz, and Cremers across several institutions. **Schmid's path** from the Cremers group at TUM through the Carlone group at MIT to JPL connects Dynablox, Panoptic Multi-TSDF, and Khronos. This is the chapter's interpretation of the lineage, not a taxonomy explicitly declared by the Handbook. --- ## 🧭 Still open **Absence vs. evidence of absence.** Determining whether an object has disappeared from the map or was merely occluded remains a foundational problem in long-term SLAM. Schmid's Panoptic Multi-TSDF provided a partial answer through active submaps. The Handbook treats occlusion above 70% as an extreme-environment case that remains difficult; it does not report this as a universal error threshold for Panoptic Multi-TSDF. As of 2026, no paper had claimed a principled solution. **Floating Map Ambiguity.** Deformable SLAM still relies on isometric and visco-elastic priors to separate the camera's rigid motion from an object's rigid motion. The conditions under which observations alone can identify both motions remain unknown. Lamarca's [2023 IJRR paper](https://arxiv.org/abs/2302.03710) described some observation conditions, but no general theory exists. **Online deformable SLAM.** DefSLAM and NR-SLAM approach real-time operation, but no system performs Khronos-level change-aware integration online from monocular RGB input. Its optimization cost exceeds real-time limits. GPU acceleration and learned priors may help, but no validated pipeline has yet appeared. **The real-world gap in medical MIS.** MIS-SLAM and NR-SLAM operate on phantoms and ex vivo data, but robustness declines in surgical environments containing blood, smoke, tool occlusions, and abrupt lighting changes. Gaussian-based methods such as EndoGS (2024) are emerging, but no system has been reported at deployment level. --- The question running through this chapter is how to represent a world that changes. GS-SLAM and NeRF-SLAM improved the speed or compactness of scene representations while retaining a static-world assumption. Ch.16 follows DUSt3R and its successors along another route: learning the geometric prior rather than refining the representation pipeline alone. --- # Ch.16 — Foundation 3D: From DUSt3R to VGGT Philippe Weinzaepfel and Jerome Revaud of Naver Labs Europe released CroCo in 2022, proposing a cross-view self-supervised pretraining scheme that learned visual representations from two images of the same scene. The paper concerned feature learning. A year later, the same team used CroCo's architecture in DUSt3R to output pointmaps directly without calibration, turning the model into a reconstruction system. The work later reached the VGG group at Oxford, and by 2026 it had changed how SfM pipelines divided matching, calibration, and reconstruction. --- ## 16.1 DUSt3R — learned pointmap For roughly ten years after 2013, 3D reconstruction followed the same procedure: find feature points, match them, estimate intrinsic and extrinsic camera parameters, build a point cloud through triangulation, and refine it with bundle adjustment. [COLMAP (Schönberger & Frahm, 2016)](https://openaccess.thecvf.com/content_cvpr_2016/html/Schonberger_Structure-From-Motion_Revisited_CVPR_2016_paper.html) was the most complete form of this pipeline. Errors dropped, but the structure of the procedure did not change. [Shuzhe Wang et al. 2023. DUSt3R: Geometric 3D Vision Made Easy](https://arxiv.org/abs/2312.14132) bypasses this procedure. It takes two images as input and outputs 3D coordinates for each pixel directly. It does not require intrinsic parameters such as focal length or principal point. The output, called a pointmap, gives coordinates in a common 3D space rather than in image coordinates, without prior knowledge of the camera lens. DUSt3R's transformer uses the encoder-decoder structure inherited from CroCo. Each image is encoded independently, then the decoder reads the other image's encoder output through cross-attention. Self-attention handles relations among pixels within one image; cross-attention learns correspondences between images implicitly from large-scale data instead of applying a coded matching rule. The paper assembled 8.5 million pairs from Habitat, MegaDepth, ARKitScenes, Static Scenes 3D, BlendedMVS, ScanNet++, CO3D-v2, and Waymo. Their annotations came from three sources: synthetic generation, reconstructions by SfM software, and dedicated sensors. Classical SfM therefore supplied part, not all, of the supervision for its learned successor. > 🔗 **Borrowed.** DUSt3R's backbone comes from ViT ([Dosovitskiy et al. 2020](https://arxiv.org/abs/2010.11929)), and its direct precursor is CroCo ([Weinzaepfel et al. 2022](https://arxiv.org/abs/2210.10716), also from Naver Labs Europe). CroCo proposed cross-view self-supervised pretraining in which information from one image reconstructs masked regions in the other. DUSt3R retained CroCo's encoder-decoder structure and changed the task to pointmap prediction. Both pointmaps already use the first camera's coordinate frame. Relative camera pose can be recovered with PnP-RANSAC from predicted 3D points and their pixel correspondences in the second image. For multiple image pairs, a separate global alignment adjusts the pointmaps and cameras. Pose estimation is derived from the pointmaps. When extending to three or ten images, DUSt3R solves a global alignment. It is an optimization problem that registers the pointmaps of all image pairs into one common coordinate frame. Only at this stage does something resembling bundle adjustment appear, but it proceeds without feature matching or camera models. --- ## 16.2 Making matching explicit: MASt3R DUSt3R's output is closer to reconstruction than to novel view synthesis. Yet it handles an important reconstruction subtask, finding precise pixel correspondences between two images, only implicitly. Replacing the explicit matching performed by SuperPoint+SuperGlue or LightGlue needed additional machinery. [Vincent Leroy et al. 2024. Grounding Image Matching in 3D with MASt3R (ECCV)](https://arxiv.org/abs/2406.09756) adds a matching head to DUSt3R. It is trained to output a feature descriptor for each pixel along with the pointmap, with joint learning that keeps the 3D position and the feature consistent. The resulting features are anchored in 3D space rather than in the image plane. Matching simplifies into nearest-neighbor search over these feature descriptors. > 🔗 **Borrowed.** MASt3R addresses the ambiguity in 2D descriptors that SuperGlue ([Sarlin et al. 2020](https://arxiv.org/abs/1911.11763)) handled with contextual reasoning. SuperGlue used a graph neural network to reduce ambiguity in 2D matching; MASt3R instead learns 3D structure directly. Within months of MASt3R's release, multiple groups in the SLAM community reported experiments that replaced the SuperPoint+SuperGlue combination with MASt3R. In late 2024, [Riku Murai, Eric Dexheimer, Andrew Davison](https://arxiv.org/abs/2412.12392) at Imperial College London released MASt3R-SLAM, using MASt3R's matching as the frontend and graph-based global optimization as the backend. The system retained the classical SLAM architecture while replacing most of its internal components. MASt3R's strength is that dense matching is possible without ground-truth calibration. As of 2026, researchers are testing DUSt3R or MASt3R in the initialization and matching stages of COLMAP-based SfM pipelines. > 📜 **Prediction vs. outcome.** The DUSt3R paper itself did not include a dedicated "Future Work" section, but the structure of pairwise processing plus global alignment implies sequence processing and real-time operation as the next tasks. Spann3R arrived in August 2024 and MASt3R-SLAM at the end of 2024. The two follow-up works addressed sequential extension and SLAM integration within 6–12 months. --- ## 16.3 Spann3R — sequential processing Batch processing has a practical constraint: in SLAM, the images are not all available in advance. DUSt3R and MASt3R take a complete image set as input and register it in a batch. SLAM receives images in temporal order and must update the map at each frame. [Hengyi Wang & Lourdes Agapito 2024. 3D Reconstruction with Spatial Memory (Spann3R)](https://arxiv.org/abs/2408.16061) reshapes DUSt3R's structure for sequential processing. It stores information from processed frames in a spatial-memory bank and applies cross-attention to that memory when a new frame arrives. Attention selects the past-frame information associated with each pixel of the new image. > 🔗 **Borrowed.** Spann3R's spatial memory resembles other cross-attention memory mechanisms. It retains DUSt3R's pretrained ViT encoder-decoder and constructs memory keys from decoder outputs (geometric features) and image features, so lookup reflects both appearance and distance. DUSt3R's geometric representation becomes the index for sequential memory. Spann3R retains DUSt3R's ability to work without a calibrated camera and updates the map incrementally as each image arrives. It is not fully real-time, but it moves the method from batch reconstruction toward SLAM. --- ## 16.4 VGGT — multi-view joint inference Spann3R enabled sequential processing but retained DUSt3R's pairwise pointmaps and global alignment. At the start of 2025, Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny at Oxford's VGG group took an arbitrary number of images as simultaneous input and produced camera pose, depth, and a point cloud in one forward pass. [Jianyuan Wang et al. 2025. VGGT: Visual Geometry Grounded Transformer](https://arxiv.org/abs/2503.11651) turned DUSt3R's pairwise processing into multi-view joint inference. Exhaustively pairing N images for DUSt3R produces N(N-1)/2 pointmap pairs followed by global alignment. VGGT passes all N images through the transformer at once, and attention processes the relations among every image pair simultaneously. > 🔗 **Borrowed.** There is an inversion in the fact that SfM reconstructions supplied part of DUSt3R's training supervision. VGGT then performs, within one model, functions that classical SfM separated into pairwise geometry estimation, graph construction, and global optimization. The stages of the classical pipeline have been absorbed in a different form. In quantitative comparisons with DUSt3R, VGGT showed consistently better camera-pose accuracy and point-cloud quality. Processing was also faster because it required no global-alignment optimization. The change made the boundary between pose estimation and reconstruction less distinct. --- ## 16.5 Pose estimation and reconstruction converge Traditional computer vision distinguished the two problems. Map-based localization finds the current position in an already-known map, and 3D reconstruction recovers the geometry of an unknown environment. SLAM was hard because it solved both at the same time. Systems from DUSt3R to VGGT use shared learned representations for geometry and camera estimation. DUSt3R recovers pose from pointmaps through separate operations and joins multiple views with global alignment. VGGT outputs cameras and geometry in one forward pass. Their common direction does not imply identical postprocessing requirements. DUSt3R, MASt3R, and VGGT have not discarded multi-view geometry. Their transformer weights encode principles implemented explicitly by the epipolar constraint, triangulation, and bundle adjustment. The change lies in how those principles are implemented: implicitly in model weights rather than as separate algorithms. This implementation is harder to debug than Schönberger's COLMAP code. The causes of a DUSt3R failure are buried inside attention weights, returning interpretability to the list of unresolved problems. > 📜 **Prediction vs. outcome.** The MASt3R paper closed briefly, suggesting that matching without ground-truth calibration was open to several downstream tasks. It was not an explicit prediction of pipeline reshaping. As of 2026 several photogrammetry software packages are evaluating DUSt3R/MASt3R as an initialization stage, and the pattern looks more like hybrid insertion than full replacement. Within two years, one team at Naver Labs Europe released CroCo (2022), DUSt3R (2023), and MASt3R (2024), covering the path from pretraining method to matching system. The group centered on Weinzaepfel, Revaud, and Leroy was small compared with Google Brain, DeepMind, or Meta AI. Davison's group at Imperial College London then carried the work into SLAM with MASt3R-SLAM. --- ## 16.6 Another branch — semantic foundation enters the map DUSt3R, MASt3R, and VGGT form the geometric branch of foundation 3D: they deal with pointmaps, camera poses, and geometric structure. Around 2022, the phrase also came to include a semantic branch that brought CLIP, DINO, and SAM into the map. The geometric branch reduced dependence on calibration; the semantic branch reduced dependence on a fixed label dictionary. The semantic branch began in Luca Carlone's group at MIT. [Nathan Hughes et al. 2022. Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization](https://arxiv.org/abs/2201.13360) placed an online hierarchy of objects → places → rooms → buildings on top of Kimera's (Rosinol 2020) metric-semantic mesh. Its closed-set classifier remained limited to a predefined dictionary of roughly 100–1000 labels, but it showed that a hierarchical map could run in real time. Foundation models removed the fixed-dictionary constraint. [Songyou Peng et al. 2023. OpenScene: 3D Scene Understanding with Open Vocabularies (CVPR)](https://arxiv.org/abs/2211.15654) came from the ETH/Pollefeys group, followed by [Qiao Gu et al. 2024. ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning (ICRA)](https://arxiv.org/abs/2309.16650) from a Montréal-MIT collaboration. OpenScene distilled CLIP features onto 3D point clouds, allowing natural-language queries such as "how close is this point to a chair." In ConceptGraphs, a VLM generated language descriptions as node attributes and an LLM described relations between objects. These methods connected concepts outside a predefined class dictionary to 3D representations. ConceptGraphs is distinct from methods that directly extend Hydra's hierarchy. [Dominic Maggio et al. 2024. Clio: Real-time Task-Driven Open-Set 3D Scene Graphs](https://arxiv.org/abs/2404.13696) turned this lineage toward tasks. Clio treats a natural-language task as an information bottleneck and retains only the level of abstraction that task needs in the scene graph. For an instruction such as "clean near the coffee machine," it preserves the coffee machine and surrounding objects while grouping unrelated details. The exposed layer of the hierarchy varies by task. > 🔗 **Borrowed.** Clio is a direct successor to Hydra in the Carlone group, adding task-driven abstraction. ConceptGraphs is a separate open-vocabulary object-graph approach and should not be described as inheriting Hydra's objects-places-rooms hierarchy. Ch.18 §18.4 traces the contraction of the object-as-landmark lineage in 2017–2019. Semantic SLAM later returned in hierarchical scene graphs, but that semantic branch and the geometric DUSt3R branch had not yet converged as of 2026. No reported end-to-end system attached CLIP features to VGGT's pointmap or combined Clio's scene graph with DUSt3R's calibration-free geometry. Their possible point of contact remains an open problem. --- ## 16.7 What classical SLAM still supplies MASt3R-SLAM borrows the architecture of classical SLAM. Keyframe selection, loop closure, and map management remain necessary on top of the new representation. The DUSt3R family replaced the internal machinery of feature matching and reconstruction but retained these system-level decisions in their classical form. The same pattern appears in the other recent lineages. NeRF-SLAM adopted NeRF as a map representation yet kept keyframe-based tracking. 3DGS-SLAM adopted Gaussians but performed loop closure in the classical way (Ch.15). Dynamic SLAM in Ch.15b changed the frontend for mask removal while retaining the backend. The representation changed more often than the system structure. Foundation 3D followed this pattern as well. In 2025, Dominic Maggio, Hyungtae Lim, and Luca Carlone at MIT released [VGGT-SLAM](https://arxiv.org/abs/2505.12549). VGGT reconstructs local submaps; the system then optimizes 15-degree-of-freedom projective transforms between sequential submaps on \(\mathrm{SL}(4)\), including loop-closure constraints. The transformer supplies local geometry, but global graph optimization remains. Revaud wrote in Handbook Ch.13 that "a form of factor graph is still necessary." Real-time large-scale sequences and integration with the semantic branch in §16.6 remained unresolved in 2026–2027. --- ## 🧭 Still open **Large-scale sequence processing.** The transformers in DUSt3R and VGGT require memory that scales quadratically with the number of images. Up to 100 images is realistic, but 1,000 or 10,000 is another matter. Spann3R's incremental approach is a partial answer, but smooth handling of large outdoor environments is unresolved. Sparse-attention and hierarchical-global-alignment approaches have not yet produced a consensus method. **Loop closure for pointmaps.** In classical SLAM, loop closure recognizes a previously visited place and corrects accumulated error. The DUSt3R family must represent that place and propagate the correction through a pointmap-based map. MASt3R-SLAM uses the existing approach, but whether it is the best or a principled solution is unknown. **Metric-scale generalization.** DUSt3R's pointmaps have relative scale. The depth ratio between two images is recovered, but absolute scale is unknown. As with Metric3D or Depth Anything v2, metric scale remains a problem for foundation 3D. The physical constraint of determining absolute scale without GPS or an IMU remains regardless of data scale. **SLAM lineage or separate branch?** MASt3R-SLAM and VGGT-SLAM brought foundation 3D into SLAM systems in 2024–2025. Real-time operation on large sequences and integration with the semantic branch in §16.6 (Hydra → Clio and the separate ConceptGraphs branch) remain unclear. No common architecture yet combines geometric and semantic foundation models in one system. --- NeRF and foundation models changed the internal machinery of reconstruction and localization, while keyframes, loop closure, and map management remained above the new representations. LiDAR-based SLAM developed in parallel over the same period. Its sensors and research culture differed, but it faced many of the same engineering problems with a rotating laser rather than a lens. --- # Ch.17 — The LiDAR Parallel Universe: From LOAM to FAST-LIO The lineage from Ch.1 photogrammetry through Ch.16 Foundation 3D shares one premise: the sensor is a camera. MonoSLAM, PTAM, ORB-SLAM, DSO, and DUSt3R all read the world through pixels. Over the same period, LiDAR SLAM developed separately around ICP and point-cloud registration rather than keypoints, photometric consistency, or feature descriptors. The two communities rarely cited each other, and they used different benchmarks and conferences. LOAM combined three earlier elements: point-to-plane matching from Besl & McKay's 1992 ICP, the network-of-poses formulation from Lu & Milios's 1997 work on globally consistent scan alignment, and the demand for real-time outdoor operation after the 2007 DARPA Urban Challenge. Ji Zhang's 2014 contribution was to split high-frequency odometry from low-frequency mapping on a spinning Velodyne. That division became a standard LiDAR-system architecture. When Ji Zhang presented LOAM at RSS 2014, it drew little attention from the Visual SLAM community, which was occupied with ElasticFusion and LSD-SLAM. LOAM shared no code with camera-based methods, and the researcher populations barely overlapped. Both lineages belonged to robotics but developed apart for roughly a decade. LOAM built on [ICP (Besl·McKay, 1992)](https://graphics.stanford.edu/courses/cs164-09-spring/Handouts/paper_icp.pdf), while the factor graph already standard in Graph SLAM crossed into LiDAR systems only later. --- ## 17.1 LOAM: edges and planes, and the capture of KITTI In 2014, Google's Waymo predecessor program was already driving on roads, and the influence of the DARPA Urban Challenge remained strong. A Velodyne HDL-64E cost $75,000 per unit. LiDAR research was therefore concentrated in large, well-funded groups able to afford the hardware. CMU Robotics Institute's Autonomous Mobile Robot Lab, led by Professor Sanjiv Singh, was one of them. Attempts to build maps with LiDAR existed before LOAM. [Lu & Milios 1997. "Globally Consistent Range Scan Alignment for Environment Mapping" (Autonomous Robots)](https://doi.org/10.1023/A:1008854305733) placed 2D range scans as nodes, tied them together with relative scan-to-scan constraints as edges, and jointly optimized the full trajectory. This "network of poses" idea became an origin point of pose-graph SLAM (see Ch.6). Alongside Besl and McKay's ICP, [Biber and Straßer 2003. "The Normal Distributions Transform" (IROS)](https://doi.org/10.1109/IROS.2003.1249285) proposed NDT, a distribution-based method that aligns per-cell Gaussian distributions; Magnusson later extended it to 3D. These methods were either 2D or offline 3D. LOAM brought real-time 3D operation. Ji Zhang, under Singh's supervision, released [Zhang & Singh 2014. "LOAM: Lidar Odometry and Mapping in Real-time" (RSS)](https://www.roboticsproceedings.org/rss10/p07.pdf). He classified LiDAR points into two kinds of features. An **edge point** is a point with high smoothness $c$ (high curvature); a **planar point** is one with low $c$ (low curvature). Rather than registering the whole point set like ICP, LOAM matches only these two feature sets. Edge points are constrained point-to-line against edge lines in the neighboring scan, and planar points are constrained point-to-plane against local planes. This selection lowers computational cost enough for real-time operation. The algorithm is split into two stages. Lidar Odometry estimates the 6-DoF transform between scans at 10 Hz. Lidar Mapping, at a lower frequency (1 Hz), registers against the full map to correct the error. Separating high-frequency odometry from low-frequency mapping suppresses drift while retaining real-time performance. Later LiDAR SLAM systems widely adopted this two-tier structure. LOAM entered the leading group in public KITTI comparisons after its release. The widely reported relative translation error is 0.78% on sequence 00 and 0.84% averaged over the listed sequences. These figures show its competitiveness, but they do not establish a universal advantage over visual odometry under different sensor and evaluation conditions. > 🔗 **Borrowed.** LOAM's feature-based point registration starts from Besl·McKay's (1992) ICP. The difference is that it selectively matches only edge and planar features rather than all points. Selective reuse of classical registration bought both speed and precision. --- ## 17.2 LeGO-LOAM: cut the ground first LOAM did not treat the ground plane explicitly. In outdoor driving environments, a significant share of the point cloud is road surface. Grouping it with other edge and planar features produces matching noise. At Stevens Institute of Technology's Robust Field Autonomy Lab, Tixiao Shan and his advisor Brendan Englot separated ground segmentation as the first stage in [Shan & Englot 2018. LeGO-LOAM](https://doi.org/10.1109/IROS.2018.8594299). The point cloud is projected onto a range image, the ground points are separated first, and the non-ground points are then re-clustered. Ground is used for roll and pitch estimation, and clusters are used for yaw and translation. This is a two-stage optimization. The result required less computation than LOAM. Where the original LOAM struggled to run in real time on a Velodyne VLP-16, LeGO-LOAM runs on the same sensor even on embedded NVIDIA Jetson platforms. The reduction came with a cost: segmentation can fail and odometry can degrade in sparse scans or environments with irregular ground, including occluded sections, rough off-road terrain, and building interiors. Beyond its lower computational cost, LeGO-LOAM established a design pattern: preprocess the sensor input into structured components before running odometry. FAST-LIO and LIO-SAM later used related modular preprocessing. Around the same time as LeGO-LOAM, Jens Behley and Cyrill Stachniss at the University of Bonn brought **surfels** (surface elements) to outdoor LiDAR instead of using edge and plane features. Their **SuMa**, described in [Behley & Stachniss 2018. "Efficient Surfel-Based SLAM using 3D Laser Range Data in Urban Environments" (RSS)](http://www.roboticsproceedings.org/rss14/p16.pdf), summarized each point's neighborhood as a disk-shaped surfel and performed scan-to-model registration. The follow-up [Chen et al. 2019. "SuMa++" (IROS)](https://doi.org/10.1109/IROS40897.2019.8967704) used semantic segmentation to filter moving objects at the surfel level. The surfel representation used by ElasticFusion in the indoor RGB-D lineage (Ch.9) had crossed into outdoor Velodyne systems. By 2018, feature selection (LOAM), segmentation-first processing (LeGO-LOAM), and surfel accumulation (SuMa) were competing approaches. --- ## 17.3 FAST-LIO — tightly coupled LiDAR-IMU LiDAR scan frequency sits at 10–20 Hz. Fast motion between scans produces motion distortion in the point cloud. The sensor position at the end of a scan differs from its position at the start, which degrades the LOAM family on high-speed platforms. An IMU runs at 100–400 Hz and can fill the gaps between LiDAR scans. Performance depends on how the two sensors are combined. A loosely coupled system estimates each independently and fuses the results; a tightly coupled system handles both inside one state estimator. The latter can use their cross-correlation but is harder to implement. At Hong Kong University (HKU)'s MaRS Lab, Wei Xu and his advisor Fu Zhang presented [**FAST-LIO**](https://arxiv.org/abs/2010.08196) in RA-L 2021. Their drone-control work supplied a concrete field requirement: LiDAR odometry had to withstand heavy rotor vibration and fast UAV maneuvers. They used an **iterated Extended Kalman Filter (iEKF)**, which repeatedly re-linearizes at the current estimate during the measurement update. This can reduce measurement-model linearization error relative to a basic EKF that linearizes once, but the improvement depends on the initial estimate and the motion and observation conditions. The following year they published **FAST-LIO2** ([Xu et al. 2022](https://doi.org/10.1109/TRO.2022.3141876)) in TRO, adding the ikd-Tree. Conventional kd-Trees carry heavy reconstruction cost every time a point is added. The ikd-Tree is an incremental variant that performs only partial reconstruction. Real-time nearest-neighbor search stays feasible even with millions of map points. Experiments showed consistent performance on UAVs, handheld rigs, and autonomous cars. Drift stayed low even in drone environments. The next FAST-LIO system addressed motion distortion at the point level. [He et al. 2023. "Point-LIO: Robust High-Bandwidth Light Detection and Ranging Inertial Odometry" (Advanced Intelligent Systems)](https://doi.org/10.1002/aisy.202200459), also from the MaRS Lab, updates the state whenever a LiDAR point arrives instead of collecting a full scan before updating. Each point is fused at its own timestamp rather than correcting intra-scan distortion with a constant-velocity model or IMU interpolation. The authors reported lower drift than FAST-LIO2 on high-agility platforms. > 🔗 **Borrowed.** FAST-LIO brings to LiDAR a tightly coupled IMU framework developed in Visual-Inertial SLAM. [Forster et al. 2016. "On-Manifold Preintegration" (TRO)](https://doi.org/10.1109/TRO.2016.2597321) established the preintegration formulation, while FAST-LIO implemented tightly coupled inertial estimation in iEKF form. --- ## 17.4 LIO-SAM: the factor graph crosses over to LiDAR Factor graphs were already standard in Visual SLAM. [GTSAM (Dellaert·Kaess, 2012)](https://gtsam.org/) had become a common backend for Visual-Inertial systems, while LiDAR systems still relied mainly on EKF variants or scan matching. LiDAR pipelines therefore made less use of graph optimization's ability to correct the full trajectory after loop closure. Tixiao Shan, after LeGO-LOAM, released [Shan et al. 2020. LIO-SAM](https://doi.org/10.1109/IROS45743.2020.9341176), which explicitly adopted GTSAM's factor graph as the backend of a LiDAR-IMU system. The [public implementation](https://github.com/TixiaoShan/LIO-SAM) maintains two graphs. The long-term mapping graph accumulates LiDAR odometry, GPS, and loop-closure constraints between keyframes. A separate IMU-preintegration graph estimates state and bias from inertial measurements and LiDAR odometry and is reset periodically to bound computation. > 🔗 **Borrowed.** LIO-SAM imports directly into a LiDAR system the GTSAM factor graph backend that Dellaert (from 2006 onward) had standardized on the Visual SLAM side. LIO-SAM handles accumulated drift better than FAST-LIO2 because it includes loop closure, but its computational cost is higher. Without GPS or another sensor, the factor graph offers less advantage. The two systems serve different design goals: FAST-LIO2 prioritizes speed and precision in a real-time single-sensor configuration, while LIO-SAM prioritizes consistency in multi-sensor long-term mapping. For nearly ten years after LOAM, LiDAR odometry added feature selection, surfels, and neural descriptors. In 2023, [Vizzo et al. 2023. "KISS-ICP: In Defense of Point-to-Point ICP" (RA-L)](https://doi.org/10.1109/LRA.2023.3236571) at Bonn took the opposite direction. An adaptive threshold and a point-to-point ICP, with no feature extraction or learned descriptors and little tuning, produced competitive odometry on KITTI. The name stood for Keep It Small and Simple. The result showed that classical registration remained competitive when motion compensation, sampling, and robust correspondence handling were combined effectively. It does not establish why LOAM arose historically. --- ## 17.5 Falling sensor prices and wider use: 2007–2024 Sensor price shaped LiDAR SLAM alongside its technical papers. At the 2007 DARPA Urban Challenge, the Velodyne HDL-64E used by leading teams cost $75,000 per unit, putting it beyond most groups outside autonomous-driving and defense research. In 2012 the HDL-32E was still around $30,000. By 2014, when LOAM appeared, the VLP-16 had dropped to $7,999, still a significant share of a research budget. The LiDAR market spread across a much wider range of prices over the following decade. In 2019, Livox (part of DJI) [announced a US retail price of $599 for the Mid-40](https://www.livoxtech.com/news/1). Ouster's 128-channel OS1-128 cost $18,000 that year; in 2020, Ouster [announced a $600 target price for the solid-state ES2 in 2024 series-production programs](https://investors.ouster.com/news-releases/news-release-details/ouster-announces-first-high-performance-true-solid-state-digital). Products and production targets in the hundreds of dollars marked a real shift, but they do not establish a uniform 100-fold decline across LiDAR: channel count, field of view, range, and retail-versus-volume pricing differ. Wider availability did not eliminate algorithmic constraints. Solid-state LiDARs generally have a more limited field of view (FoV) than spinning sensors, and some use non-repetitive scan patterns. The original LOAM, designed around a 360° rotating scan, does not transfer unchanged. FAST-LIO and FAST-LIO2, by contrast, were designed to handle both mechanical and solid-state LiDARs and reported operation with small FoV and irregular sampling. Lower prices therefore changed the algorithmic questions rather than removing them. --- ## 17.6 Why the Visual and LiDAR lineages split Visual SLAM and LiDAR SLAM developed in the same period, yet the two communities exchanged little for years. Several differences reinforced the split. The sensors produced different measurements. Cameras capture texture and color; LiDAR measures range and geometry. Camera-based methods developed around keypoints, descriptors, and photometric consistency, while LiDAR methods used edges, planes, and range images. Their problem formulations differed accordingly. The conferences differed as well. Camera-based methods appeared mainly at CVPR and ICCV, while LiDAR SLAM appeared mostly at ICRA, IROS, and RSS. The researcher populations overlapped little. During the early-to-mid 2010s, as Velodyne supplied Google and the autonomous-driving industry, LiDAR SLAM concentrated in self-driving robotics groups. Place recognition methods diverged too. Cameras use visual appearance, as in DBoW2 and NetVLAD. LiDAR uses the structural features of a 3D point cloud, as in [Scan Context (Kim·Kim, 2018)](https://gisbi-kim.github.io/publications/gkim-2018-iros.pdf) or [PointNetVLAD](https://arxiv.org/abs/1804.03492). Even for the same location, the signal being recognized is different. The first signs of convergence appeared in the early 2020s, when LiDAR-camera fusion papers began reaching CVPR. Tixiao Shan's [LVI-SAM (2021)](https://arxiv.org/abs/2104.10831) added a visual-inertial subsystem to LIO-SAM. The authors presented a tightly coupled factor graph, but the LIS and VIS subsystems operate largely independently and support each other during failures. A fully unified state estimate remained open. --- ## 17.7 Visual-LiDAR convergence attempts: 2024–2025 From 2024, more work attempted to handle camera and LiDAR data in one representation as foundation models became less tied to a single sensor. Two approaches emerged. One approach uses multimodal pretrained features to align LiDAR and camera data in the same embedding space. It adapts the contrastive-learning principle of [CLIP (Radford et al., 2021)](https://arxiv.org/abs/2103.00020) from image-text alignment to LiDAR-image pairs. In 2023–2024, this work remained experimental. The other approach converts sensor outputs into shared geometric primitives or neural fields and processes them in one backend. This work also remained at the research-paper stage, with few demonstrations of real-time operation. Neither approach had produced a common system lineage by 2026. FAST-LIO2 and ORB-SLAM3 were still used independently. --- ## 17.8 Radar is outside this book's scope Radar SLAM developed as another independent subfield, divided between spinning radar such as the Navtech CIR family and SoC-based 4D mmWave radar. Direct Doppler radial-velocity measurements enable correspondence-free odometry, while radio-specific noise models address speckle, multipath, and receiver saturation. The lineage runs from Oxford's radar localisation in [Cen & Newman 2018](https://doi.org/10.1109/ICRA.2018.8460687), through Adolfsson and Magnusson's **CFEAR** and its successor **TBV-SLAM**, to Burnett and Barfoot's continuous-time ICP. Oxford Radar RobotCar, Boreas, and MulRan supply dedicated benchmarks. Radar's operation in bad weather and smoke offers a clear practical advantage, but it has little historical overlap with the photogrammetry → SfM → Visual SLAM → learning → 3D foundation lineage. Its separate history falls outside this book; Handbook of SLAM (2026) Ch.9 provides a technical account. --- ## 📜 Prediction vs. outcome > Zhang and Singh named two items as explicit future work in the conclusion of the 2014 LOAM paper: loop closure to correct drift and fusion with IMU output through a Kalman filter. Both appeared within the next ten years. FAST-LIO (2021) and FAST-LIO2 (2022) integrated the IMU with a tightly coupled iEKF, while LIO-SAM (2020) added loop closure through a factor-graph backend. One field problem absent from the paper's conclusion remained: dynamic-object handling. As of 2026, real-time separation of moving pedestrians and vehicles from LiDAR points often used deep-learning segmentation. Geometric methods within SLAM also existed, but no general solution covered the range of dynamic environments. --- ## 🧭 Still open **Full Visual+LiDAR fusion.** Even after LVI-SAM, no tightly coupled design that handles both sensors inside one state estimator has become a broadly accepted common architecture. Autonomous-driving systems need LiDAR to compensate when a camera weakens in fog or rain, but algorithm design and sensor calibration remain barriers. Transformer-based fusion in 2024–2025 remained at the research-prototype stage. **Algorithms optimized for solid-state LiDAR.** The original LOAM assumed the scan-line structure of a 360° spinning LiDAR. Non-repetitive scans in some Livox products and the limited FoV of solid-state sensors change observability and motion distortion. FAST-LIO2's direct point-to-map formulation and Livox LOAM already address parts of this setting, but no single configuration covers the differing fields of view and scan patterns across sensor families. **Dynamic object handling.** The static-world assumption remained in LiDAR SLAM from Zhang's 2014 work through 2026. Systems commonly hand real-time separation of moving objects to segmentation networks. Geometric approaches inside SLAM still face computational and stability costs. Production pipelines are mostly proprietary, while public research has not converged on one generally accepted solution. --- The LiDAR and visual lineages matured with separate technical vocabularies, and their integration remained incomplete. Other approaches developed beside both of them, including biologically inspired, event-based, and semantic SLAM, without becoming the main line of either community. --- # Ch.18 — Failed Cases and Lost Lineages While the camera and LiDAR communities developed their own methods, other robotics researchers followed different directions. Some of these lines began with RatSLAM in 2004, well before LOAM; others ran through the 2010s. They did not become mainstream, but they remained part of SLAM's history. These approaches had identifiable sources. Milford and Wyeth's RatSLAM (2004) drew on O'Keefe and Dostrovsky's 1971 place-cell work through cognitive-map theory. Event SLAM inherited the silicon-retina line through Lichtsteiner, Posch, and Delbruck's 2008 DVS at ETH Zürich INI. Salas-Moreno et al.'s SLAM++ (2013) extended 1990s object-level scene understanding into the SLAM state. Each encountered a different engineering limit. Some approaches accumulated papers and promising early results without entering the mainstream. They reached scaling limits or lost adoption to a more practical alternative, outcomes distinct from a failure of the underlying technique. --- ## 18.1 RatSLAM — place cell-based topological map RatSLAM, presented by [Milford et al. 2004](https://doi.org/10.1109/ROBOT.2004.1302555) at ICRA 2004, approached place recognition through a biological model. It imitated the firing patterns of **place cells** and **head direction cells** in the rat hippocampus to form a place representation during exploration. The computational model was a **Continuous Attractor Network (CAN)**. Its neurons form a continuous activation 'bump' on a 2D grid, which moves according to the robot's velocity and rotation input (path integration). Visual input is compared with stored place representations and corrects the bump's position. RatSLAM alternates between propagation from motion and correction from visual matching. > 🔗 **Borrowed.** The place-cell discovery by [O'Keefe and Dostrovsky (1971)](https://pubmed.ncbi.nlm.nih.gov/5124915/) began in neuroscience and led to the theory of cognitive maps. RatSLAM was an early complete implementation of that biological mechanism in an engineering system, but few later SLAM systems adopted it directly. Milford and Gordon Wyeth, based at the Queensland University of Technology (QUT) robotics lab, repeatedly tested RatSLAM on suburban roads in Brisbane between 2004 and 2008. A roof-mounted camera supplied the image stream as the system recognized previously traveled routes and closed loops. The [Milford & Wyeth 2008](https://doi.org/10.1109/TRO.2008.2004520) IEEE T-RO paper reported tens of thousands of images over a 66 km route. Contemporary geometric SLAM systems often operated over only a few hundred meters, so RatSLAM had a much larger demonstrated range. The system did not scale much further. CAN grew in computational complexity with the number of places, and its topological map could recognize a return without reliably producing meter-level metric positions. Autonomous driving and manipulation needed precise coordinates that this cognitive-map formulation did not supply. > 📜 **Prediction vs. outcome.** In the conclusion of the 2008 T-RO paper, Milford and Wyeth described RatSLAM as "an alternative approach to vision-only SLAM" and reported repeatable, reliable loop closure on long routes with large accumulated error and visual ambiguity. They presented it as an alternative, not a replacement. RatSLAM remained competitive on specific benchmarks, but after 2012 graph-based SLAM and visual odometry moved ahead in accuracy and speed. Topological maps persisted in some place-recognition work, while RatSLAM's metric-topological integration did not continue as a major system lineage. RatSLAM's algorithm saw limited adoption, but its geometry-free place representation entered the place-recognition literature. In 2012, [SeqSLAM](https://doi.org/10.1109/ICRA.2012.6224623) emerged from the same Milford group, and image-sequence-based recognition became one line of visual place-recognition benchmarks. --- ## 18.2 The engineering limits of biologically-inspired SLAM RatSLAM was the most complete case of biologically inspired SLAM, but not the only one. From the mid-2000s to the early 2010s, researchers proposed SLAM variants based on cognitive maps, entorhinal grid cells, and hippocampal replay. They encountered similar problems. Biological models describe *how* the brain represents space. Whether the same representation fits an engineering objective is a separate question. Evolution shaped the rat hippocampus for particular environments and behavior, which differ from a robot's operating conditions. Engineering SLAM requires sub-meter position accuracy, real-time processing, rapid adaptation to new environments, and verifiable error bounds. Cognitive models did not guarantee these properties, leaving a substantial gap between neuroscience and robotics. The question reopened in the 2020s because representations learned by foundation models invited comparison with cognitive maps. Whether the resemblance can be made precise, or amounts only to analogy, remains unknown. --- ## 18.3 Event SLAM — the gap between hardware and algorithm maturity The [Dynamic Vision Sensor (DVS)](https://doi.org/10.1109/JSSC.2007.914337), developed by Patrick Lichtsteiner, Christoph Posch, and Tobi Delbruck at the ETH Zürich Institute of Neuroinformatics (INI), was first disclosed at ISSCC 2008. Each pixel independently compares logarithmic light-intensity change against a threshold and asynchronously outputs an event of positive (ON) or negative (OFF) polarity. With no global shutter, it records each pixel's firing time at microsecond resolution, producing a camera without frames. > 🔗 **Borrowed.** The DVS event sensor (Lichtsteiner et al. 2008) was hardware inspired by the change-detection mechanism of the biological retina. Event SLAM started with this sensor in hand. Hardware ran ahead of algorithms, and closing that gap took ten years. Event cameras offered μs-level temporal resolution, little blur during high-speed motion, high dynamic range (HDR) across tunnels and direct sunlight, and a fraction of the power consumption of conventional cameras. At ICRA 2014, [Weikersdorfer et al. 2014](https://doi.org/10.1109/ICRA.2014.6906882) presented event-based 3D SLAM. The same year, other groups released event-based optical-flow and depth-estimation methods. Between 2016 and 2018, Henri Rebecq in Davide Scaramuzza's RPG lab at the University of Zurich released [EVO](https://doi.org/10.1109/LRA.2016.2645143) (RA-L 2017) and [ESIM](https://proceedings.mlr.press/v87/rebecq18a.html) (CoRL 2018), completing more of the event-SLAM pipeline. Results in real environments remained limited. Early DVS sensors had 128×128 pixels rather than VGA resolution, a serious constraint for feature matching and map building. Existing frame-based algorithms also did not apply directly to event streams, so the field needed new methods as well as better hardware. From 2014 to 2018, event SLAM produced good results in controlled environments and low-texture conditions, but did not outperform existing visual-inertial odometry in general environments. The approach also spread beyond odometry. [EventVLAD](https://ieeexplore.ieee.org/document/9635907/) (Lee & Kim, IROS 2021) combined edge images reconstructed from event streams with NetVLAD descriptors, demonstrating place recognition under sudden illumination changes and motion blur, conditions that challenged frame-based VPR. --- ## 18.4 Semantic SLAM — the shrinking of the object-as-landmark path From 2017 to 2019, semantic methods appeared throughout the programs of CVPR, ECCV, and IROS. Deep learning was improving instance segmentation and object detection, and researchers widely explored their integration with SLAM. Implementations progressed more slowly than the surrounding claims. [Salas-Moreno et al. 2013](https://doi.org/10.1109/CVPR.2013.178)'s **SLAM++** was an early large-scale system in this lineage. Salas-Moreno and his advisor Andrew Davison at Imperial College used *objects* rather than points or patches as the map's basic unit. They stored predefined 3D models of chairs, desks, and monitors in a database, recognized them in RGB-D input through ICP (Iterative Closest Point) alignment, and placed them on the map. Tens of objects could replace thousands of points, reducing map size and grounding place recognition and loop closure in semantic entities. > 🔗 **Borrowed.** SLAM++'s object-level representation combined scene graphs from graphics with model-based recognition from computer vision. The approach later reappeared in LERF (Language Embedded Radiance Field) and LangSplat in the 2020s, with language features replacing objects as the representation unit. After SLAM++, [SemanticFusion](https://arxiv.org/abs/1609.05130) (McCormac et al., 2017, ICRA) and [MaskFusion](https://arxiv.org/abs/1804.09194) (Rünz et al., 2018, ISMAR) used semantic information in mapping and dynamic-object separation. [SuperPoint](https://arxiv.org/abs/1712.07629) (DeTone et al., 2018)-based systems appeared in the same period, but SuperPoint is not a semantic feature: it jointly learns a keypoint detector and local descriptor. Both lines used learning, but they addressed different problems and representation units. Semantic-SLAM papers argued that integrating semantic understanding with geometric pipelines could make maps more robust to environmental change. Through 2019, traditional geometric pipelines such as ORB-SLAM2, VINS-Mono, and LIO-SAM produced the main gains on autonomous-driving benchmarks. Systems with deep semantic features remained competitive only in specific indoor environments and with fixed object classes. On new categories or unseen environments, semantic priors sometimes increased drift. > 📜 **Prediction vs. outcome.** In the Conclusion of the SLAM++ paper, Salas-Moreno described his method as "a first step toward a more generic SLAM method," hoping it would extend to objects with low-dimensional shape variation, and ultimately to systems that segment and define object classes on their own. The paper's introduction added that an object-unit representation would bring "large map compression" and "gains in efficiency and robustness." The actual development partially hit the mark. Object-level maps found a place in AR and certain manipulation applications, and the compression and efficiency advantages were confirmed again in indoor environments with repeated objects. But mainstream geometric SLAM still retains sparse points and keyframe-based graphs as of 2026, and the stage where objects are segmented and defined autonomously has not been reached. Object-as-landmark adoption remained limited, while semantics found other roles in dynamic-region separation within SLAM and in downstream semantic mapping and task planning. Semantic-first SLAM depended on accurate segmentation; a segmentation error could corrupt the map, whereas robust estimation let a geometric pipeline survive some incorrect matches. Generalization posed a second problem. Semantic priors trained on particular object classes did not transfer beyond those classes, while SLAM systems had to operate in a much wider range of environments. The object-as-landmark path contracted, but semantic information continued at a different map layer. [SuMa++](https://doi.org/10.1109/IROS40897.2019.8967704) (Chen et al., IROS 2019) overlaid semantic classes on LiDAR point clouds to filter dynamic objects, and [Kimera](https://doi.org/10.1109/ICRA40945.2020.9196885) (Rosinol et al., ICRA 2020) combined a metric-semantic mesh with a 3D scene graph. [Hydra](https://doi.org/10.15607/RSS.2022.XVIII.050) (Hughes et al., RSS 2022) extended the graph into a real-time hierarchy. [ConceptGraphs](https://doi.org/10.1109/ICRA57147.2024.10610243) (Gu et al., ICRA 2024) and [Clio](https://doi.org/10.1109/LRA.2024.3451395) (Maggio et al., RA-L 2024) later added open-vocabulary foundation features. This lineage remained active in 2026 and reappears in [Ch.15b](chapter_15b_dynamic.md), [Ch.16](chapter_16_foundation_3d.md), and [Ch.19 §19.7](chapter_19_open_problems.md#197-the-return-of-semantic-representation-and-open-world). --- ## 18.5 The Manhattan-world assumption — scope and continued use as an auxiliary constraint Another line of research used the Manhattan-world assumption and later receded. Following the Manhattan-world concept of [Coughlan & Yuille 1999](https://doi.org/10.1109/ICCV.1999.790349), these methods treated indoor walls, floors, and ceilings as aligned with three orthogonal axes (x, y, z) of the world coordinate frame. Parallel image lines converge to vanishing points, each described by the relation `v = K R d` between the camera rotation matrix R and direction vector d (K is the camera intrinsic matrix). Three orthogonal vanishing points allow direct recovery of the three columns of R. Visual-odometry systems could use this geometric constraint to suppress drift without an IMU or feature matching. The constraint reduced drift in long corridors and rectangular rooms but failed outdoors, around curved structures, and in irregular industrial environments where the assumption did not hold. A prior fitted tightly to one environment became a liability elsewhere. As general-purpose visual-inertial odometry matured after 2015, Manhattan-world methods receded. The assumption remains an auxiliary constraint in some indoor mapping tools rather than an independent research lineage. --- ## 18.6 Rediscovery patterns of extinct lineages The discontinued lineages left different kinds of descendants. RatSLAM's topological-map idea carried into SeqSLAM and visual place recognition. SLAM++'s object-level map returned in another form after 2022 when NeRF and Gaussian splatting were combined with language features, as in [LERF](https://arxiv.org/abs/2303.09553) (Kerr et al., 2023) and [LangSplat](https://arxiv.org/abs/2312.16084) (Qin et al., 2023). Event-camera SLAM followed a different path because early hardware remained limited. After 2022, event cameras above 640×480 reached the market, and high-speed drones and HDR environments supplied clear applications. Between 2020 and 2024, the event-vision community around [Guillermo Gallego](https://arxiv.org/abs/1904.08405) at TU Berlin reported competitive results in event-based depth and ego-motion estimation. A promising idea can still wait years for suitable hardware and algorithms. Adoption also depends on whether a more practical alternative becomes established before both mature. --- ## 🧭 Still open **Biologically inspired SLAM.** Spatial representations formed by foundation models through large-scale unsupervised learning have structural similarities to cognitive maps. Whether a transformer's internal representation implements anything comparable to a place cell remains unverified. A RatSLAM-like lineage returning through foundation models is therefore a hypothesis, not an observed convergence. **Event-camera SLAM adoption.** Commercial high-resolution event cameras widened the research base after 2022, but event processing had not settled on a stable common framework by 2026. Integration with frame-based pipelines, event representations, real-world benchmarks, and evaluation standards were all still developing. Broad adoption remained uncertain. **The direction of the semantic-map concept.** As interest in semantic SLAM cooled after 2017, semantic representation developed other roles in dynamic-region separation and downstream tasks. From 2023, LERF and language-based Gaussian-splatting systems combined language features with dense scene representations. It remained unclear whether semantics would become part of SLAM itself or stay downstream, and whether these representations could relax the usual requirement that geometry be reliable first. Full visual-LiDAR fusion, solid-state sensor algorithms, dynamic-object handling, event-camera maturity, and the return of semantic maps all connect with unresolved problems from earlier chapters; Ch.19 brings those threads together. --- # Ch.19 — Today's Map and Tomorrow's Open Questions By 2026, AR layers remained fixed to walls, indoor delivery robots distinguished kitchens from conference rooms without a supplied map, and DUSt3R-family models recovered 3D structure from a few photographs in seconds. These systems solved many of the problems that defined SLAM in 2003, but only under particular assumptions. The 2003 formulation assumed a static scene, stable lighting, bounded space, and the geometry of a single camera. Under those conditions, the EKF tracked state, graph SLAM closed loops, and ORB-SLAM managed keyframes. The solutions remain valid within the simplifying assumptions on which they were built. Across the preceding chapters, the unresolved problems recur beside the conditions under which individual methods succeed. Collected together, they reveal where apparently separate lineages face the same constraints. --- ## 19.1 Lighting and environmental change: reality the camera cannot handle Visual SLAM has struggled with environmental change since its first outdoor deployments. Field conditions repeatedly expose the limits of a camera's photometric model. Learned descriptors beat ORB inside the training domain but lose consistency on underwater, thermal, and low-light imagery; as of 2026 there is still no consensus on which is more robust (see Ch.2 §2.7). The low-light and dynamic-tracking failures recorded in Ch.5 explain why the 2007 PTAM paper bounded itself as "Small AR Workspaces." Most feature-based SLAM still assumes more stable conditions (see Ch.5 §🧭). The direct-method lineage faces a structural version of the problem. Brightness preservation fails under auto-exposure, strong backlight, and tunnel-to-outdoor transitions, and no complete method dynamically estimates the lighting model (see Ch.8 §🧭). Place recognition has faced a related limit for more than ten years. Even with [DINOv2](https://arxiv.org/abs/2304.07193)-based methods narrowing the gap, a single model does not yet maintain consistent precision and recall across the extreme seasonal and lighting conditions in [Nordland](https://nikosuenderhauf.github.io/projects/placerecognition/) and [Oxford RobotCar](https://robotcar-dataset.robots.ox.ac.uk/) (see Ch.10 §10.7). ORB-SLAM's long-term map reuse has the same limitation. Atlas made multi-map maintenance possible, but lighting changes can prevent recognition of the same place between morning and evening (see Ch.7 §🧭). Ch.2, 5, 7, 8, and 10 each encounter this problem through a different component of the pipeline. --- ## 19.2 Static-world assumption: the oldest simplification hits its limits The static-world assumption is one of SLAM's oldest simplifications and recurs across more lineages than any other. In the SfM lineage dynamic objects are a shared weak point of every current system, COLMAP included, and as of 2026 no Dynamic SfM implementation has COLMAP-level generality (see Ch.3 §3.7). Everything in Ch.9 from KinectFusion through BundleFusion assumed a static scene, and while DynaSLAM, MaskFusion, and others coupled real-time segmentation into dense SLAM, neither cost nor robustness reached practical deployment (see Ch.9 §🧭). In monocular depth, self-supervised methods mask moving objects, avoiding rather than solving the problem (see Ch.11 §🧭). 3DGS SLAM still assumed a static world in 2025. [4DGS](https://arxiv.org/abs/2310.08528) and [Deformable 3DGS](https://arxiv.org/abs/2309.13101) add a time dimension, but no integrated SLAM system both represents and tracks dynamic objects (see Ch.15 §🧭). LiDAR SLAM is not exempt: dynamic-object handling remains a challenge distinct from LOAM's stated future work, while production autonomous-driving stacks are largely proprietary and difficult to compare directly with public research systems (see Ch.17 §🧭). Five chapters reach the same unresolved problem through different representations. The long-term dynamic and deformable problems in [Ch.15b](chapter_15b_dynamic.md) are related. **Absence vs evidence of absence**, whether an object vanished or was occluded, received a partial answer through the active submaps of [Schmid's Panoptic Multi-TSDF](https://doi.org/10.1109/LRA.2022.3148854) (2022). The Handbook treats occlusion above 70% as an extreme-environment case that remains difficult, but does not report it as a universal error threshold for Panoptic Multi-TSDF. **Floating Map Ambiguity**, separating rigid camera motion from rigid object motion, is constrained only through isometric and visco-elastic priors; identification without a prior remains unresolved. No system performs Khronos-level change-aware integration online from monocular RGB, and medical MIS systems lose robustness when moving from phantom and ex vivo data to surgical conditions. All four items from Ch.15b remain open. --- ## 19.3 Scale and representational memory: the problem changes when the size does Scaling a SLAM system from one room to a building, and from a building to a city, exposes related limits in several representations. Monocular scale is a geometric ambiguity established in 1980s SfM theory. Without an IMU, depth sensor, or another metric cue, projective geometry in monocular imagery alone cannot determine the metric scale of the trajectory and scene (see Ch.5 §🧭). [Metric3D v2](https://arxiv.org/abs/2404.15506) and [Depth Anything v2](https://arxiv.org/abs/2406.09414) produce metric depth when camera intrinsics are known, yet smartphones, CCTV footage, archives, and satellite imagery often lack them. Foundation-scale training has not removed this constraint (see Ch.11 §🧭). In the TSDF lineage, memory became a limit of the representation. [Voxblox](https://arxiv.org/abs/1611.03631) and [OctoMap](https://octomap.github.io/) reduced the cost, but memory for a dense representation still rises rapidly with mapped volume and voxel resolution, and no general adaptive-resolution policy has been established (see Ch.9 §🧭). NeRF-SLAM also remains open at city scale (see Ch.14 §🧭). In Gaussian Splatting, the required number of Gaussians rises with scene extent and detail; the [Compact 3DGS](https://arxiv.org/abs/2311.13681) (Lee et al. 2024) family explores compression, but no approach has become standard (see Ch.15 §🧭). Foundation 3D moves the limit into the transformer: attention across all image tokens has quadratic memory cost in the token count, making long sequences difficult, and Spann3R provides only a partial incremental solution (see Ch.16 §🧭). The representations differ, but each reaches a scale-dependent memory limit. Scale also raises **data movement cost**, the energy needed to move bits between processor and memory rather than the capacity of the representation alone. In Handbook Ch.18 §18.8, Davison proposes "on-device data movement, measured in bits × millimetres" as the 12th SLAM performance metric. [Hughes et al.](https://doi.org/10.15607/RSS.2022.XVIII.050) likewise describe how a hierarchical scene graph reduces memory from $O(L \cdot V/\delta^3)$ to $O(N_\text{sub} + N_\text{obj} + N_\text{rooms})$ (Handbook Ch.16 Eq. 16.34-16.36). It remains unclear whether mainstream evaluations will adopt data movement as a metric. --- ## 19.4 Uncertainty calibration for learning-based systems Since Julier and Uhlmann established the EKF's inconsistency in Ch.4, SLAM researchers have had to ask whether a system's reported uncertainty matches its actual localization error. Non-Gaussian uncertainty violates a central EKF assumption. Real sensor errors are often multimodal or heavy-tailed; Stein particles, normalizing flows, and learned uncertainty have been tested, but their real-time validation remains limited (see Ch.4 §4.8). In graph SLAM, robust-cost selection still relies heavily on judgment because no principled method determines in advance whether Huber, Cauchy, or Geman-McClure fits a particular sensor and environment (see Ch.6 §🧭). [Ch.6b](chapter_06b_certifiable.md) raises a related issue of tightness bounds. SE-Sync's exact-recovery theorem gives the sufficient condition "noise below $\beta$" without a way to compute $\beta$ in advance for an actual instance. Extending certifiable methods to Visual SLAM and VIO, and re-solving the SDP online as measurements arrive, also remain open. Learning-based methods make the calibration problem harder to observe. After Bayesian PoseNet, uncertainty under out-of-distribution input remained unresolved (see Ch.12 §🧭). DROID-SLAM and related systems showed that learned priors can degrade outside the training domain without an explicit failure signal. [TartanAir](https://arxiv.org/abs/2003.14338)-style synthetic training still leaves a sim-to-real gap (see Ch.13 §🧭). Foundation 3D adds the question of how to propagate a loop-closure correction through a pointmap. MASt3R-SLAM uses existing methods, but whether this is a principled solution is unknown (see Ch.16 §🧭). Autonomous driving and medical robotics need calibrated uncertainty, yet few systems address it at that level. Davison frames the issue with a question: *"If a network has built a 3D model from 100 images, does adding one more image require running the whole thing again"* (Handbook Ch.18, p.528). Long-term representation and fusion bring probabilistic state estimation and modular scene representations back into the system. The [GBP Learning](https://arxiv.org/abs/2312.14294) lineage (Nabarro et al.) represents network weights as random variables in a factor graph, reducing the distinction between *"training time"* and *"test time"* (p.543). Whether this formulation resolves the problem or moves it into another set of assumptions remains unclear. --- ## 19.5 Sensor fusion and new modalities: integration unfinished Visual SLAM and LiDAR SLAM addressed localization and mapping with different measurements and algorithms. Their lineages have not yet merged into a common architecture. LVI-SAM coupled visual odometry with LIO-SAM, but largely at a loosely coupled level. Autonomous-driving systems need LiDAR to take over when cameras fail in fog or rain, yet tightly coupled fusion remains difficult both algorithmically and in calibration (see Ch.17 §🧭). Solid-state LiDAR poses a related problem. Limited fields of view and non-repetitive patterns differ from the original LOAM's 360° scan-line assumptions. FAST-LIO2 and Livox LOAM address parts of this setting, but generalization across sensor families remains limited (see Ch.17 §🧭). Wide-baseline matching presents a related integration problem. Beyond 45 degrees of viewpoint change, Harris- and ORB-based matching drops sharply. DUSt3R bypasses explicit matching, but it is too early to know whether this resolves the descriptor problem or avoids it temporarily (see Ch.2 §2.7). Place recognition and metric localization also remain separate pipeline stages. Attempts from 2023–2025 to unify them in one representation achieved neither the required precision nor speed (see Ch.10 §10.7). Event cameras show the lag between new hardware and mature algorithms. Commercial high-resolution event cameras spread after 2022, while integration with frame-based pipelines, event representations, and real-world benchmarks all remained under development (see Ch.18 §🧭). Kinect followed a similar sequence: the sensor launched in 2010 and KinectFusion arrived a year later. Two modalities fall outside this account: **4D imaging radar** and **legged/proprioceptive SLAM**. Radar can complement cameras and LiDAR in conditions such as fog and rain, where both optical modalities may degrade. Oxford Radar RobotCar (2019), the radar channel in NuScenes, and 4D imaging radar development by companies including Arbe and Mobileye broadened this modality's role in autonomous-driving research and products. Legged SLAM formed a separate lineage that fused kinematic and contact priors for outdoor deployment of ANYmal, Spot, and Unitree in the 2020s. Both have distinct origins and benchmarks from the visual, LiDAR, and foundation-3D lineages and warrant separate histories. --- ## 19.6 Recoupling compute structure and hardware Davison's Handbook Ch.18 emphasizes a topic rarely covered in SLAM histories: matching the graph structure of an algorithm to the graph structure of the silicon. Dennard scaling broke, and single-core CPU clock speed stalled near 4 GHz in the mid-2000s; *"this has stopped being true"* (Handbook Ch.18, p.528). Wearable Spatial AI still has to fit into glasses weighing 65 g and consuming less than 1 W. The gap favors heterogeneous, specialized, parallel hardware. Several hardware examples appeared by the mid-2020s. [Apple Vision Pro R1](https://www.apple.com/apple-vision-pro/specs/) (2023) is a dedicated chip for 12 ms sensor processing; [Meta ARIA Gen 2](https://www.projectaria.com/ariagen2/) (2024) carries custom silicon for "ultra low power and on-device machine perception." The [Graphcore IPU](https://www.graphcore.ai/products/ipu) has thousands of independent cores with local memory communicating by message passing. Manchester's [SCAMP5](https://personalpages.manchester.ac.uk/staff/p.dudek/papers/carey-iscas2013.pdf) implements 256×256 per-pixel in-plane processing at 1.2 W, while [SpiNNaker](https://apt.cs.manchester.ac.uk/projects/SpiNNaker/) connects up to one million ARM cores in a neuromorphic structure. Each supports a different graph topology, and no systematic theory maps Spatial AI algorithms to these architectures. Davison's later work on **Gaussian Belief Propagation** addresses this hardware structure. [Ortiz et al.](https://arxiv.org/abs/2203.11618) (2022) accelerated bundle adjustment on the IPU with GBP by 30× over a CPU, and [Murai et al. Robot Web](https://arxiv.org/abs/2306.04620) (2024) demonstrated multi-robot SLAM in which robots shared factor-graph fragments over Wi-Fi and converged through asynchronous message passing. The motivation was that *"we must get away from the idea that a 'god's eye view' of the whole structure of the graph will ever be available"* (Handbook Ch.18, p.541). The factor graph becomes the main representation, and local messages replace full-posterior computation. Whether this approach will combine with transformer-based systems such as MASt3R-SLAM remains unanswered. Among Davison's twelve metrics, number 11 is "power usage" and number 12 is "on-device data movement." Both extend evaluation beyond accuracy to power consumption and the bit volume and physical distance of data movement within the device. TUM, KITTI, and EuRoC do not yet include these measures, and no consensus exists on how to add them to mainstream benchmarks. --- ## 19.7 The return of semantic representation and Open-World Semantic objects did recede from SLAM's landmark representation, as [Ch.18 §18.4](chapter_18_dead_ends.md#184-semantic-slam--the-shrinking-of-the-object-as-landmark-path) describes. Neither ORB-SLAM3 nor MASt3R-SLAM uses object-level primitives. Over the same period, however, semantics moved to an upper layer of the map and produced practical systems. [Kimera](https://doi.org/10.1109/ICRA40945.2020.9196885) (2020) combined a metric-semantic mesh with a 3D scene graph, and [Hydra](https://doi.org/10.15607/RSS.2022.XVIII.050) (2022) extended it into the *"first online system to produce fully hierarchical scene graphs that included objects, places, and rooms"* (Handbook Ch.16, §16.4.2). Foundation features were then added to this layer. [ConceptFusion](https://arxiv.org/abs/2302.07241) and [VLMaps](https://arxiv.org/abs/2210.05714) (2023) placed CLIP features in dense maps; [ConceptGraphs](https://doi.org/10.1109/ICRA57147.2024.10610243) (2024) used open-vocabulary object nodes; [Clio](https://doi.org/10.1109/LRA.2024.3451395) (2024) built task-driven hierarchies; and [LERF](https://arxiv.org/abs/2303.09553) and [LangSplat](https://arxiv.org/abs/2312.16084) attached language to radiance fields and Gaussian splatting. Semantic representation persisted above the geometric map rather than as its landmarks. This work introduced unresolved questions of its own. Hughes and Carlone identify one directly: *"performing uncertainty quantification in hierarchical representations mixing discrete and continuous variables is still a largely unexplored problem"* (p.488). Uncertainty propagation remains insufficiently explored when discrete variables such as object category and room ID share a graph with continuous variables such as pose and surface. Scene graphs also remain difficult to extend into outdoor and unstructured environments, while Clio's task-driven hierarchy (Handbook Ch.17 Eq. 17.8) has not generalized broadly. A broader question is whether a system still needs an explicit map. Paull and the editors address it in Handbook Ch.17 §17.4.2, "Revisiting the Question of the Need for Maps." A long-context VLM might plan from past frames without an explicit scene graph. [OpenEQA](https://open-eqa.github.io/) and [Mobility VLA](https://arxiv.org/abs/2407.07775) (2024) show that map-free methods work on short, simple tasks but degrade as spatial and temporal horizons lengthen. *"the need for an explicit map representation ... largely depend[s] on the spatial and temporal horizons of the considered tasks and remains an active area of research"* (p.515). The evidence supports neither universal map use nor universal map-free operation. The relation between SLAM and generative robot policies raises the same question. VLA models such as [RT-2](https://robotics-transformer2.github.io/) (2023), [OpenVLA](https://arxiv.org/abs/2406.09246) (2024), and [π₀](https://www.physicalintelligence.company/blog/pi0) (2024) might replace SLAM or operate above it. The Handbook's final sentence argues that *"true generalization and scalability to compositional tasks ... could be achieved through some form of explicit structure that is learned through a process such as SLAM. ... these two paradigms ... are entirely complementary"* (Paull/Carlone, Handbook Ch.17, p.520). The architecture implied by "complementary" remains open. --- ## 19.8 The shape of the open questions The open problems differ in both age and kind. The monocular scale ambiguity of Ch.5 is a geometric fact established in SfM theory and retains the same formulation in 2026. The limits of the static-world assumption, by contrast, have returned in changing forms over twenty years: SfM in Ch.3, dense SLAM in Ch.9, Gaussian maps in Ch.15, and LiDAR in Ch.17. Loop closure for foundation 3D and calibration of learned uncertainty are newer formulations that emerged only in the preceding few years. Ch.0 described a period in which SLAM is often treated as solved. The five editors of the 2026 SLAM Handbook offer the internal counterpoint in their epilogue: *"If someone tells you 'SLAM is solved,' don't listen to them."* New methods repeatedly relax one assumption and expose another problem. Particle filters addressed limits in the EKF's local linearization and single-Gaussian approximation, while dense methods retained image information discarded by sparse features. Neither transition invalidated the earlier method; each changed the assumptions under which the system operated. Methods regarded as solved in 2026 remain conditional in the same way. Their next open problem will appear when one of those conditions no longer holds. --- ## 19.9 Lineage map ```mermaid graph TD PM[사진측량 1858] BA[Bundle Adjustment
Brown 1958] SfM[Photo Tourism 2006] COLMAP[COLMAP 2016] SC[Smith-Cheeseman 1986] Mono[MonoSLAM 2003] PTAM[PTAM 2007] ORB[ORB-SLAM 2015] ORB3[ORB-SLAM3 2020] LSD[LSD-SLAM 2014] DSO[DSO 2016] VIDSO[VI-DSO 2018] LM[Lu-Milios 1997] FG[Factor Graph
Dellaert 2000s] iSAM[iSAM 2008] iSAM2[iSAM2 2012] g2o[g2o 2011] Forster[Preintegration
Forster 2015] VINS[VINS-Mono 2018] Kinect[KinectFusion 2011] Elastic[ElasticFusion 2015] SESync[SE-Sync 2019] TEASER[TEASER 2020] LOAM[LOAM 2014] FAST[FAST-LIO 2021] NeRF[NeRF 2020] iMAP[iMAP 2021] NICE[NICE-SLAM 2021] GS3D[3DGS 2023] Spla[SplaTAM 2023] MonoGS[MonoGS 2023] DROID[DROID-SLAM 2021] DPV[DPV-SLAM 2024] DUSt3R[DUSt3R 2023] MASt[MASt3R 2024] VGGT[VGGT 2025] MASlam[MASt3R-SLAM 2024] Hydra[Hydra 2022] Clio[Clio 2024] PM --> BA --> SfM --> COLMAP SC --> Mono --> PTAM --> ORB --> ORB3 PTAM -.-> LSD --> DSO --> VIDSO LM --> FG --> iSAM --> iSAM2 FG --> g2o iSAM2 -.-> SESync --> TEASER Forster --> VINS --> ORB3 VIDSO --> Forster Kinect --> Elastic Elastic -.-> LOAM --> FAST NeRF --> iMAP --> NICE NICE -.-> GS3D --> Spla --> MonoGS PTAM -.-> DROID --> DPV COLMAP -.-> DUSt3R --> MASt --> VGGT MASt --> MASlam ORB3 -.-> Hydra --> Clio click PM "#chapter-1" "Ch.1 Prehistory — Photogrammetry" click BA "#chapter-1" "Ch.1 Prehistory — Bundle Adjustment" click SfM "#chapter-3" "Ch.3 Structure from Motion" click COLMAP "#chapter-3" "Ch.3 SfM — COLMAP" click SC "#chapter-4" "Ch.4 EKF-SLAM — Smith-Cheeseman" click Mono "#chapter-5" "Ch.5 MonoSLAM·PTAM" click PTAM "#chapter-5" "Ch.5 MonoSLAM·PTAM" click ORB "#chapter-7" "Ch.7 ORB-SLAM family" click ORB3 "#chapter-7" "Ch.7 ORB-SLAM3" click LSD "#chapter-8" "Ch.8 Direct Methods — LSD-SLAM" click DSO "#chapter-8" "Ch.8 Direct Methods — DSO" click VIDSO "#chapter-8" "Ch.8 Direct Methods — VI-DSO" click LM "#chapter-6" "Ch.6 Graph SLAM — Lu-Milios" click FG "#chapter-6" "Ch.6 Graph SLAM — Factor Graph" click iSAM "#chapter-6" "Ch.6 Graph SLAM — iSAM" click iSAM2 "#chapter-6" "Ch.6 Graph SLAM — iSAM2" click g2o "#chapter-6" "Ch.6 Graph SLAM — g2o" click Forster "#chapter-7" "Ch.7b IMU Preintegration (after Ch.7)" click VINS "#chapter-7" "Ch.7 — VINS-Mono" click Kinect "#chapter-9" "Ch.9 RGB-D — KinectFusion" click Elastic "#chapter-9" "Ch.9 RGB-D — ElasticFusion" click SESync "#chapter-6" "Ch.6b Certifiable (after Ch.6)" click TEASER "#chapter-6" "Ch.6b Certifiable — TEASER" click LOAM "#chapter-17" "Ch.17 LiDAR — LOAM" click FAST "#chapter-17" "Ch.17 LiDAR — FAST-LIO" click NeRF "#chapter-14" "Ch.14 NeRF-SLAM" click iMAP "#chapter-14" "Ch.14 NeRF-SLAM — iMAP" click NICE "#chapter-14" "Ch.14 NeRF-SLAM — NICE-SLAM" click GS3D "#chapter-15" "Ch.15 Gaussian Splatting" click Spla "#chapter-15" "Ch.15 — SplaTAM" click MonoGS "#chapter-15" "Ch.15 — MonoGS" click DROID "#chapter-13" "Ch.13 Hybrid — DROID-SLAM" click DPV "#chapter-13" "Ch.13 — DPV-SLAM" click DUSt3R "#chapter-16" "Ch.16 Foundation 3D — DUSt3R" click MASt "#chapter-16" "Ch.16 — MASt3R" click VGGT "#chapter-16" "Ch.16 — VGGT" click MASlam "#chapter-16" "Ch.16 — MASt3R-SLAM" click Hydra "#chapter-16" "Ch.16 §16.6 Semantic Foundation" click Clio "#chapter-16" "Ch.16 §16.6 — Clio" ```