Spatial Intelligence

·
September 9, 2025

AR Indoor Navigation Platforms Compared: 2026 Buyer's Guide

Eight AR indoor navigation platforms compared on accuracy, scan input, deployment mode and cost. Built for warehouse and industrial sites, not shopping malls.
Shadnam Khan
Shadnam Khan
MultiSet AI

Indoor navigation is not one market. A shopping mall wants to route a shopper to a store in under a minute and does not care about ten centimetres. A warehouse wants a picker to find bin 14-C-08 among nine identical aisles, and ten centimetres is the difference between the right bin and the wrong one. A refinery wants a technician to identify one valve among forty that look the same, and getting it wrong is a safety event.

Most comparison articles treat these as the same problem. They are not, and the technology that wins each is different.

This guide compares the platforms on the four things that actually decide a deployment: how accurate the position is, what capture data the system will accept, where it can be deployed, and what it costs to keep running.

What indoor navigation has to solve

GPS does not work indoors. Satellite signals do not survive a roof and a steel frame, and even where a phone reports a position inside a building it is typically interpolating from the last known fix and the Wi-Fi environment. The error is measured in tens of metres. That is enough to know which building someone is in and useless for anything else.

Everything called indoor navigation is a way of replacing that missing signal. There are four families, and the choice between them is mostly a question of what you are willing to install.

Wi-Fi and BLE beacons. Physical transmitters placed through the building. The device measures signal strength from several and trilaterates. Accuracy is typically one to five metres, degrading as people and inventory move around. The cost is not the hardware, it is the installation and the batteries: a large distribution centre can need hundreds of beacons, each needing replacement on a cycle.

Ultra-wideband. Purpose-built anchors and tags exchanging precisely timed radio pulses. Accuracy of ten to thirty centimetres is realistic, which is genuinely good. The cost is that it is an infrastructure project. Anchors need power, mounting, surveying and commissioning, and every tracked thing needs a tag.

Markers and QR codes. Printed targets at known positions. Cheap and accurate at the marker. The problem is between markers, where the device dead-reckons and drifts, and the fact that someone has to place, survey and maintain thousands of stickers in an environment that scuffs and repaints them.

Visual positioning. The camera already on the device compares what it sees against a prebuilt 3D map of the space and returns a full position and orientation. No transmitters, no tags, no stickers. The cost moves from hardware installation to capture: someone has to scan the building once, and the map has to be kept current as the space changes.

The last of these is what a visual positioning system does, and it is the only one of the four where the cost does not scale with floor area in hardware.

What a visual positioning system actually returns

Worth being precise, because the term gets used loosely.

A VPS returns 6-DoF pose: three numbers for position and three for orientation, expressed in the coordinate frame of a map. Not "you are near aisle 14". The exact point in space the camera occupies and the exact direction it is facing.

That distinction is what makes AR overlay possible. Knowing a technician is somewhere in a two-metre circle is enough to show a floor arrow. Knowing precisely where the camera is and which way it points is what lets you draw a label on the correct valve and have it stay there as the person walks around it.

The sequence is always the same three steps. Capture the space. Process the capture into a map. Localize against that map at runtime, usually many times a second, each query returning a pose and a confidence value the application is expected to check before trusting it.

The comparison

Platform Best for Accuracy Capture input Deployment Watch for
MultiSet AI Industrial, warehouse, multi-floor, regulated sites Sub-5 cm to sub-10 cm depending on capture Scan-agnostic: E57, Matterport, NavVis, Leica, Faro, XGRIDS, Polycam, Gaussian splats, 360 video, phone LiDAR Public cloud, private cloud, self-hosted, on-device Smaller consumer footprint than the mapping giants
Google ARCore Geospatial Outdoor consumer AR where Street View exists Roughly one to three metres outdoors Street View, no capture needed Cloud only Not built for indoors or private facilities
Vuforia Area Targets Existing PTC estates Centimetre-level in the target area Specific supported scanners Cloud or on-device Capture is constrained to approved devices. Licensing
UWB vendors Asset tracking where tags are acceptable 10 to 30 cm No capture, hardware survey On-premises Anchor and tag installation across the whole area
BLE and Wi-Fi vendors Malls, airports, large public venues 1 to 5 m No capture, site survey Cloud or on-premises Battery maintenance. Accuracy drifts as the space fills
Marker systems Fixed stations, low budget Exact at the marker None Anything Drift between markers. Physical maintenance

The row that matters most

Absolute or relative. This is the distinction most comparisons miss and the one that decides whether a deployment works.

SLAM, visual or LiDAR, builds a map as it goes and tracks against it. Excellent for a robot navigating a space. But the coordinate frame starts wherever the device started, so two devices do not agree with each other, and nothing persists once the session ends. Put a label on a valve today and it will not be there tomorrow.

A VPS localizes against a map that already exists and is shared. Every device that queries it gets a pose in the same frame, and it is the same frame next week. That is the requirement for persistent AR content, for multi-user work, and for a robot and a technician to agree on where something is.

The two are complementary rather than competing. In production, SLAM handles frame-to-frame tracking and the VPS supplies the absolute fix and corrects the drift.

Four questions that decide it

How accurate do you actually need to be? If the job is routing a person to a room, one to five metres is fine and beacons are cheaper. If the job is putting a label on the right object among identical objects, you need centimetres and only visual positioning or UWB will do it. Be honest about which job you have, because paying for centimetres you do not need is the most common way these projects get expensive.

What capture data do you already have? Most industrial sites have been scanned already, for BIM, for facilities management, for insurance. If a platform will ingest that scan, the capture cost is zero and the project starts this week. If it requires its own capture device and its own workflow, add weeks and a rescan to the plan. This is the question buyers underweight most.

Where is the compute allowed to run? A public cloud API is the easy path and it is a non-starter in a defence facility, a pharmaceutical plant under GMP, or anywhere with data-residency obligations. If private cloud, on-premises or fully on-device are likely to be required, that constraint should be in the first conversation, not the security review.

What happens when the space changes? This is the question that separates a pilot from a deployment. Warehouses reslot. Plants get new equipment. If updating the map means recapturing the whole site and re-authoring every anchor, the system will quietly fall out of use within a year. Ask specifically what a partial update costs and whether authored content survives it.

Where the mall answer is different

Shopping centres, airports and stadiums show up constantly in indoor navigation searches, so it is worth saying plainly: for those venues, this guide is probably the wrong one.

Public-venue wayfinding is a different product. The accuracy requirement is metres. The user is a member of the public on their own phone who will not install anything. The economics are driven by footfall analytics and retail media, not by technician productivity. Beacon and Wi-Fi platforms built for that market handle it well and are cheaper.

Visual positioning earns its place where accuracy is operationally load-bearing: finding one asset among many identical ones, overlaying a model on the thing it describes, or giving a robot and a person the same coordinate frame.

What scan-agnostic changes

Most platforms tie the map to their own capture. That single constraint drives most of the cost and most of the delay in an indoor navigation project, because it means the scan you already paid for cannot be used.

Scan-agnostic ingestion means the platform accepts whatever the site already has. In practice that is E57 point clouds from Matterport, Leica, NavVis, Faro or XGRIDS, meshes from Polycam or Scaniverse, metric-scaled Gaussian splats, plain 360 video, or native phone LiDAR.

Two consequences follow. The first is speed: a site scanned last year for BIM can become a localization map without anyone going back. The second is that the capture decision stops being locked to the positioning vendor, which matters over a multi-year deployment where scanning hardware will change at least once.

One capture then produces two things from the same pipeline: a map machines localize against, and a human-readable reconstruction for everyone else. More on the formats and what each one produces on the 3D mapping page.

How to run an evaluation that tells you something

Pilots fail for predictable reasons. A short protocol that avoids most of them:

  1. Test in your worst space, not your best. A long repetitive aisle, a featureless corridor, a room that changes daily. Any platform demos well in a furnished office.
  2. Measure first-lock time. How long from opening the app to a trusted position. Anything over a few seconds gets abandoned by the people who have to use it.
  3. Measure drift over a realistic session. Walk the route for as long as the actual job takes, not thirty seconds.
  4. Change something and re-test. Move a pallet, open a door, change the lighting. Then re-test without remapping.
  5. Check the confidence value. Any serious system returns one. Confirm your application gates on it rather than trusting every pose.
  6. Price the second year. Rescan cadence, map updates, per-call or per-area cost. The first year is never the expensive one.

Where this is heading

Two things are changing the requirement.

Smart glasses are becoming a real deployment target rather than a demo. Once the display is on someone's face rather than in their hand, position accuracy stops being a nice-to-have, because a label that sits ten degrees off is worse than no label.

And robots and people increasingly need the same map. An autonomous mobile robot and a technician working the same floor have historically run on separate positioning stacks that disagree with each other. Resolving both into one coordinate frame is becoming the requirement rather than the ambition, and it is the reason positioning is now discussed as infrastructure rather than as a feature of an AR app. That case is set out on the robotics page.

The short version

If you need metres and your users are the public, use a beacon or Wi-Fi platform. If you need centimetres on tagged assets and can install hardware, UWB works. If you need centimetres on anything a camera can see, without installing hardware, and you already have a scan of the building, visual positioning is the only one of the four that fits.

MultiSet returns sub-5 cm pose from the scan you already own, and deploys in public cloud, private cloud, self-hosted or fully on-device. See how the VPS works, or book a pilot.