
Spatial Intelligence
Indoor navigation is not one market. A shopping mall wants to route a shopper to a store in under a minute and does not care about ten centimetres. A warehouse wants a picker to find bin 14-C-08 among nine identical aisles, and ten centimetres is the difference between the right bin and the wrong one. A refinery wants a technician to identify one valve among forty that look the same, and getting it wrong is a safety event.
Most comparison articles treat these as the same problem. They are not, and the technology that wins each is different.
This guide compares the platforms on the four things that actually decide a deployment: how accurate the position is, what capture data the system will accept, where it can be deployed, and what it costs to keep running.
GPS does not work indoors. Satellite signals do not survive a roof and a steel frame, and even where a phone reports a position inside a building it is typically interpolating from the last known fix and the Wi-Fi environment. The error is measured in tens of metres. That is enough to know which building someone is in and useless for anything else.
Everything called indoor navigation is a way of replacing that missing signal. There are four families, and the choice between them is mostly a question of what you are willing to install.
Wi-Fi and BLE beacons. Physical transmitters placed through the building. The device measures signal strength from several and trilaterates. Accuracy is typically one to five metres, degrading as people and inventory move around. The cost is not the hardware, it is the installation and the batteries: a large distribution centre can need hundreds of beacons, each needing replacement on a cycle.
Ultra-wideband. Purpose-built anchors and tags exchanging precisely timed radio pulses. Accuracy of ten to thirty centimetres is realistic, which is genuinely good. The cost is that it is an infrastructure project. Anchors need power, mounting, surveying and commissioning, and every tracked thing needs a tag.
Markers and QR codes. Printed targets at known positions. Cheap and accurate at the marker. The problem is between markers, where the device dead-reckons and drifts, and the fact that someone has to place, survey and maintain thousands of stickers in an environment that scuffs and repaints them.
Visual positioning. The camera already on the device compares what it sees against a prebuilt 3D map of the space and returns a full position and orientation. No transmitters, no tags, no stickers. The cost moves from hardware installation to capture: someone has to scan the building once, and the map has to be kept current as the space changes.
The last of these is what a visual positioning system does, and it is the only one of the four where the cost does not scale with floor area in hardware.
Worth being precise, because the term gets used loosely.
A VPS returns 6-DoF pose: three numbers for position and three for orientation, expressed in the coordinate frame of a map. Not "you are near aisle 14". The exact point in space the camera occupies and the exact direction it is facing.
That distinction is what makes AR overlay possible. Knowing a technician is somewhere in a two-metre circle is enough to show a floor arrow. Knowing precisely where the camera is and which way it points is what lets you draw a label on the correct valve and have it stay there as the person walks around it.
The sequence is always the same three steps. Capture the space. Process the capture into a map. Localize against that map at runtime, usually many times a second, each query returning a pose and a confidence value the application is expected to check before trusting it.
| Platform | Best for | Accuracy | Capture input | Deployment | Watch for |
|---|---|---|---|---|---|
| MultiSet AI | Industrial, warehouse, multi-floor, regulated sites | Sub-5 cm to sub-10 cm depending on capture | Scan-agnostic: E57, Matterport, NavVis, Leica, Faro, XGRIDS, Polycam, Gaussian splats, 360 video, phone LiDAR | Public cloud, private cloud, self-hosted, on-device | Smaller consumer footprint than the mapping giants |
| Google ARCore Geospatial | Outdoor consumer AR where Street View exists | Roughly one to three metres outdoors | Street View, no capture needed | Cloud only | Not built for indoors or private facilities |
| Vuforia Area Targets | Existing PTC estates | Centimetre-level in the target area | Specific supported scanners | Cloud or on-device | Capture is constrained to approved devices. Licensing |
| UWB vendors | Asset tracking where tags are acceptable | 10 to 30 cm | No capture, hardware survey | On-premises | Anchor and tag installation across the whole area |
| BLE and Wi-Fi vendors | Malls, airports, large public venues | 1 to 5 m | No capture, site survey | Cloud or on-premises | Battery maintenance. Accuracy drifts as the space fills |
| Marker systems | Fixed stations, low budget | Exact at the marker | None | Anything | Drift between markers. Physical maintenance |
Absolute or relative. This is the distinction most comparisons miss and the one that decides whether a deployment works.
SLAM, visual or LiDAR, builds a map as it goes and tracks against it. Excellent for a robot navigating a space. But the coordinate frame starts wherever the device started, so two devices do not agree with each other, and nothing persists once the session ends. Put a label on a valve today and it will not be there tomorrow.
A VPS localizes against a map that already exists and is shared. Every device that queries it gets a pose in the same frame, and it is the same frame next week. That is the requirement for persistent AR content, for multi-user work, and for a robot and a technician to agree on where something is.
The two are complementary rather than competing. In production, SLAM handles frame-to-frame tracking and the VPS supplies the absolute fix and corrects the drift.
How accurate do you actually need to be? If the job is routing a person to a room, one to five metres is fine and beacons are cheaper. If the job is putting a label on the right object among identical objects, you need centimetres and only visual positioning or UWB will do it. Be honest about which job you have, because paying for centimetres you do not need is the most common way these projects get expensive.
What capture data do you already have? Most industrial sites have been scanned already, for BIM, for facilities management, for insurance. If a platform will ingest that scan, the capture cost is zero and the project starts this week. If it requires its own capture device and its own workflow, add weeks and a rescan to the plan. This is the question buyers underweight most.
Where is the compute allowed to run? A public cloud API is the easy path and it is a non-starter in a defence facility, a pharmaceutical plant under GMP, or anywhere with data-residency obligations. If private cloud, on-premises or fully on-device are likely to be required, that constraint should be in the first conversation, not the security review.
What happens when the space changes? This is the question that separates a pilot from a deployment. Warehouses reslot. Plants get new equipment. If updating the map means recapturing the whole site and re-authoring every anchor, the system will quietly fall out of use within a year. Ask specifically what a partial update costs and whether authored content survives it.
Shopping centres, airports and stadiums show up constantly in indoor navigation searches, so it is worth saying plainly: for those venues, this guide is probably the wrong one.
Public-venue wayfinding is a different product. The accuracy requirement is metres. The user is a member of the public on their own phone who will not install anything. The economics are driven by footfall analytics and retail media, not by technician productivity. Beacon and Wi-Fi platforms built for that market handle it well and are cheaper.
Visual positioning earns its place where accuracy is operationally load-bearing: finding one asset among many identical ones, overlaying a model on the thing it describes, or giving a robot and a person the same coordinate frame.
Most platforms tie the map to their own capture. That single constraint drives most of the cost and most of the delay in an indoor navigation project, because it means the scan you already paid for cannot be used.
Scan-agnostic ingestion means the platform accepts whatever the site already has. In practice that is E57 point clouds from Matterport, Leica, NavVis, Faro or XGRIDS, meshes from Polycam or Scaniverse, metric-scaled Gaussian splats, plain 360 video, or native phone LiDAR.
Two consequences follow. The first is speed: a site scanned last year for BIM can become a localization map without anyone going back. The second is that the capture decision stops being locked to the positioning vendor, which matters over a multi-year deployment where scanning hardware will change at least once.
One capture then produces two things from the same pipeline: a map machines localize against, and a human-readable reconstruction for everyone else. More on the formats and what each one produces on the 3D mapping page.
Pilots fail for predictable reasons. A short protocol that avoids most of them:
Two things are changing the requirement.
Smart glasses are becoming a real deployment target rather than a demo. Once the display is on someone's face rather than in their hand, position accuracy stops being a nice-to-have, because a label that sits ten degrees off is worse than no label.
And robots and people increasingly need the same map. An autonomous mobile robot and a technician working the same floor have historically run on separate positioning stacks that disagree with each other. Resolving both into one coordinate frame is becoming the requirement rather than the ambition, and it is the reason positioning is now discussed as infrastructure rather than as a feature of an AR app. That case is set out on the robotics page.
If you need metres and your users are the public, use a beacon or Wi-Fi platform. If you need centimetres on tagged assets and can install hardware, UWB works. If you need centimetres on anything a camera can see, without installing hardware, and you already have a scan of the building, visual positioning is the only one of the four that fits.
MultiSet returns sub-5 cm pose from the scan you already own, and deploys in public cloud, private cloud, self-hosted or fully on-device. See how the VPS works, or book a pilot.