← Insights AI-frontier & modellen 14 July 2026 5 min Written with AI assistance

The Map Is Becoming a Copy of the World

Spatial capture just collapsed from lidar rigs to the phone in your pocket — and that changes what a map is.

Ruben Horbach Ruben Horbach Co-founder

In short

  • A single phone camera can now rebuild a room in 3D live at 20 fps, no lidar or rig needed.
  • Spatial capture cost has collapsed in ~18 months, moving mapping from centralized projects to crowd-sourced side effects.
  • Apple bringing Gaussian splatting to Maps signals mapping's shift from pictures of places to the places themselves.
  • Cheap single-camera 3D reconstruction could replace lidar in robot perception, lowering costs for embodied AI.
  • As representations get dense enough, the distinction between a map and a copy of the world stops mattering.

A single phone camera can now rebuild a room in 3D at 20 frames per second. Live, over long sequences, end to end. No lidar, no rig, one camera, streaming.

That sentence would have sounded absurd two years ago. Spatial capture meant expensive lidar-plus-RGB scanners, careful sweeps and hours of post-processing before you had anything explorable. As recently as this spring the honest read in the field was that distribution of 3D Gaussian splats was more or less solved, since you could stream a big indoor scan into a browser and walk through it in real time, while capture stayed the blocker because good results still needed hardware most people will never touch.

That blocker is dissolving, and it is worth slowing down to look at what happens once it is gone.

A single phone camera rebuilding a room at 20 frames per second

The capture curve, compressed

Here is the trajectory, all of it inside roughly the last eighteen months.

Niantic put collaborative splat capture on phones and a shared map, so anyone could capture anything in 3D and pin it to a place. World Labs' Marble went from generating 3D scenes you could stitch into larger environments to reconstructing real-world locations from a handful of images, with sixteen iPhone photos of a Shinjuku alley becoming a fully explorable 3D world in minutes. NVIDIA's Lyra removed the manual stitching step for large-scale explorable worlds entirely. An open-source repo now takes any city name and returns a 3D model with buildings and streets pulled from OpenStreetMap, MIT-licensed and with no login. SuperSplat added automatic voxel collision, turning any splat into solid walkable geometry at 5 cm resolution, more compressed than a mesh but solid enough to move through.

And in June, Apple announced at WWDC that Gaussian splatting is coming to Apple Maps.

Each of those is a modest engineering step on its own. Together they trace one curve: the cost of turning a real place into an explorable digital one is falling toward zero, and the device doing it is the one already in a few billion pockets.

the cost of turning a real place into an explorable digital one is falling toward zero, and the device doing it is the one already in a few billion pockets.

What people do with free capture

You do not have to speculate about the demand, because people are already crowd-sourcing real places into explorable splats. Cathedrals, alleys, plazas, shared in tools like SuperSplat. You can walk into St. Stephan's from your desk. It is uneven and hobbyist right now, in roughly the way early Wikipedia was uneven and hobbyist, and I would not overread the current quality. The shape is what is visible: a next-generation map you can step inside, assembled bit by bit by whoever happens to be standing there with a phone.

A Gaussian splat you can walk through rather than orbit

The historical rhyme is Street View. Google spent years driving camera cars down every road on earth to build a photographic index of the planet, and it had to be a centralised, capital-intensive project because capture was expensive. When capture costs one camera and a few seconds, the same index gets built as a side effect of people existing in places. Apple putting splats into Maps is the institutional signal about where mapping goes next: from pictures of places to the places themselves, in geometry.

The second customer: machines

There is a quieter implication underneath the consumer one. A live single-camera 3D reconstruction of a space is exactly what a robot needs in order to understand the room it is moving through. Perception has been one of the stubborn bottlenecks in embodied AI and the hardware answer has usually been lidar, which is precise and expensive and power-hungry. If a commodity camera stream can produce dense persistent 3D structure in real time, the perception stack for household and warehouse robots gets radically cheaper.

The crowd-sourced map feeds back into that. A robot entering a building that has already been captured does not start from zero; it loads the space. Machine-readable 3D models of cities are the connective layer between the world of bits and the world of atoms, and I wrote back in 2024 that I could no longer walk around a city without seeing that model in my head. What has changed is that the model is starting to actually exist, and ordinary phones are the thing building it.

Volumetric capture as a machine-readable layer over the city
Volumetric capture as a machine-readable layer over the city

From picture to copy

A useful way to hold all this is that maps have always been compressions. A road atlas compresses geography into lines. Street View compresses it into photographs. A splat map barely compresses at all, because it keeps the geometry, the light, and increasingly the physics. At some point the representation gets dense enough that the distinction between a map of the place and a copy of the place stops mattering for most uses, and tourism, real estate, insurance, robotics, games and memory itself all run on the copy once it exists.

We have watched this pattern before. First we map, then we predict, then we act. Spatial computing spent a decade stuck at expensive partial mapping, the capture cost just collapsed, and the mapping phase looks likely to finish itself one phone at a time.

So the question I find interesting is no longer whether a walk-in copy of the world gets built. It is who ends up holding it, and what gets permission to move through it.

Ruben Horbach

Ruben Horbach

Co-founder · Back From the Future

Ruben researches how organisations adopt AI meaningfully — not as technology, but as a change in work and people. He builds the agent infrastructure behind BFF and speaks about the near future of work.

Translate this to your situation?

Book a conversation — we're happy to think along about what this means for you.