
Spatial Intelligence
A picker in a large distribution centre spends more of the shift walking than picking. That part is well understood and it is why slotting exists. The part that gets less attention is that a meaningful share of the walking is not travel at all. It is search: standing in the right aisle, scanning racking, checking a label, walking back because the bin was two bays down.
Travel time can be optimised with better slotting and better routing. Search time cannot, because it is not a routing problem. It is a positioning problem. The picker does not know precisely where they are, and the system does not know either, so nothing in the stack can point at the bin.
That is the gap AR indoor navigation closes in a warehouse, and it only closes if the position is accurate to the bin rather than to the aisle.
Every warehouse positioning technology works in a demo. They separate at scale, in a building that changes weekly, with a budget that has to survive a second year.
| Approach | Accuracy | What you install | Where it breaks |
|---|---|---|---|
| BLE and Wi-Fi beacons | 1 to 5 m | Hundreds of transmitters, mounted and powered | That is the aisle, not the bin. Accuracy degrades as racking fills, and batteries are a standing maintenance line |
| Floor markers and QR | Exact at the marker | Thousands of printed targets, surveyed | Drift between markers. They get scuffed, painted over, and driven across by forklifts |
| Voice and RF picking | No position at all | Headsets or handhelds | Solves sequencing, not location. Still assumes the picker knows where they are standing |
| Ultra-wideband | 10 to 30 cm | Anchors with power and survey, plus a tag on everything tracked | Genuinely accurate and genuinely an infrastructure project. Untagged things stay invisible |
| Visual positioning | Sub-5 cm | Nothing. One capture of the building | Needs a scan, and the map has to be kept current as the layout changes |
The last row is the only one where cost does not scale with floor area in hardware, and the only one where the honest objection is about map maintenance rather than about capital. That objection is dealt with below, because it is the one that actually decides these projects.
A visual positioning system uses the camera already in the picker's hand. The device sends a frame, the system matches it against a prebuilt 3D map of the building, and it returns 6-DoF pose: exactly where the camera is and exactly which way it is facing, in the coordinate frame of that map.
Two properties matter more than the accuracy number.
It is absolute, not relative. The on-device tracking in every modern phone and headset is relative. It knows how far you have moved since the session started, in a frame that begins wherever you happened to be. Two devices never agree, and nothing survives closing the app. A VPS localizes against a map that already exists and is shared, so every device gets a pose in the same frame, and it is the same frame next week. That is the requirement for putting a label on a bin and having it still be on that bin tomorrow.
It is shared. One map, many devices. A phone, a tablet on a forklift, a pair of smart glasses and an autonomous mobile robot can all query it and all get answers that agree with each other.
In production the two work together rather than competing. On-device tracking carries the user smoothly between fixes, and the VPS supplies the absolute position and corrects the drift.
Anyone who has run a warehouse pilot asks this within the first ten minutes, and it is the right question. Warehouses are not static buildings. Racking moves, zones get reconfigured, seasonal peaks change the whole floor plan, and inventory profile changes what the aisles look like week to week. If keeping the map current means recapturing the whole site and re-authoring every anchor, the system will quietly fall out of use inside a year.
Two things have to be true for it to survive.
Partial updates. Recapture only the zones that changed and merge them into the existing map. A reslotted aisle should cost an aisle's worth of capture, not a building's.
Anchors that carry forward. The content authored against the map (labels, routes, bin markers, work instructions) has to survive a map update rather than being re-placed by hand. Otherwise the authoring cost is the real cost and it recurs every quarter.
Worth being precise about which changes visual positioning absorbs and which it does not, because the claims in this category get loose. A moving pallet, a changed inventory profile, a new stack in an aisle: fine, because the system matches against the structural geometry of the space rather than against what is sitting in it. A wall moved, an entire racking run relocated, a zone rebuilt: that needs a partial recapture. Ask any vendor for that distinction explicitly, and treat a claim of complete immunity to change as a red flag.
The map is the expensive part, and it is a fixed cost. Everything below runs on the same one.
The obvious use and the least valuable. Route a picker to a bin or a zone with a path drawn on the floor in front of them. Worth having, and it mostly helps new starters and agency staff during peak. Experienced pickers already know the building.
This is where the return is. Wayfinding assumes the destination is fixed and known. Asset finding is the opposite problem: locating the things that move. Cages, totes, pallet trucks, cleaning equipment, returns trolleys, the specific piece of MHE that is due a service.
These are the items nobody can find, that get walked past, that get bought twice because the first one is somewhere in the building. Because the position is absolute and shared, the last confirmed location of an asset is a lookup rather than a search, and confirming a new location is a glance rather than a scan-in at a fixed station.
A Fortune 100 industrial customer running this in a private cloud measured 2.5x technician productivity on asset finding and a 4x reduction in mean time to repair. Those are maintenance numbers rather than picking numbers, but the mechanism is identical: the time went into finding the thing, not into working on it.
Costs nothing extra, because the map already exists. New starters learn on the real floor, with guidance anchored to the real racking, rather than in a classroom against a floor plan they then have to translate. The instruction sits on the actual bay, at the actual height, in the actual aisle.
The reason this matters commercially is churn. Warehouse onboarding is a recurring cost that spikes exactly when the building is busiest, and a training mode that runs off infrastructure already paid for is close to free capacity at peak. More on the authoring side of this on the work instructions and asset navigation page.
Most large warehouses now run autonomous mobile robots alongside people. The robots already localize, competently, using their own SLAM stack. The problem is that they do it in their own coordinate frame, which nothing else in the building can read.
So the fleet manager knows where its robots are, the WMS knows where its inventory is meant to be, and the person on the floor knows where they are, and none of those three agree in a common frame. Every handoff between them becomes an integration project.
Resolving people, robots and content into one shared map turns that into a lookup. A robot can be directed to a position a person marked. A person can be routed to a robot that has stopped. The two can be kept apart in a shared zone because both positions are expressed the same way. That argument is set out in full on the ground truth for AMRs page and on the robotics page.
Most distribution centres have already been scanned, usually more than once, for BIM, facilities management, insurance or a fit-out. That scan is normally treated as a document rather than as infrastructure, and it sits on a drive.
Scan-agnostic ingestion means it does not have to be recaptured to become a localization map. MultiSet accepts E57 point clouds from Matterport, Leica, NavVis, Faro and XGRIDS, meshes from Polycam and Scaniverse, metric-scaled Gaussian splats, plain 360 video, and native phone LiDAR.
Two consequences. A warehouse scanned last year can be localizing this week, without anyone booking a capture crew into a live building. And the capture decision stops being locked to the positioning vendor, which matters across a deployment that will outlast at least one generation of scanning hardware. The full input list and what each one produces is on the 3D mapping page.
Where the compute runs is a separate decision and worth settling early. Public cloud is the fast path. Private cloud, self-hosted and fully on-device exist for sites where a third-party API call is not acceptable, which in this sector usually means a customer contract rather than a regulation.
Warehouse pilots fail for predictable reasons, most of them about where the pilot was run rather than what was piloted.
Slotting and routing have already taken most of the travel time out of the pick cycle. What is left is search, and search is a positioning problem. Beacons put you in the aisle. Markers put you at the marker. UWB puts you within thirty centimetres of a tag, on a tag you had to buy and fit.
Visual positioning puts a phone within a few centimetres of the bin, on a scan of the building you probably already own, with nothing installed on the racking. The map then pays for itself three times: wayfinding, asset finding and training all run off the same one.
See how the VPS works, or send us one scan file and we will process it.