Where Visual Positioning Works: 12 Environments and What Each One Takes

A technical walk-through of the 12 failure modes every autonomy team faces when moving from controlled environments to the field.
The argument

Localization is not one problem

In most built environments the geometry pins down four or five degrees of freedom and leaves one or two almost free. That free axis is nearly always the one the building labelled, because somebody had to number the repeating modules, and MultiSet VPS reads those labels.What a VPS is and how it works

The part most teams do not realise

If your map was built with a survey-grade scanner, the visual layer is almost certainly already in the capture. Leica RTC360 and BLK360, NavVis, Faro, XGRIDS and Matterport Pro2 and Pro3 all record panoramic imagery alongside the point cloud. MultiSet ingests the structured E57 with panoramas included and builds a VPS map from the same file. One capture, both layers, one coordinate frame. Check the export setting, not the scanner.

Scan requirements
Twelve environments

Twelve environments, and what each one takes

Pick the one that looks like your site. Each opens to the mechanics, the mechanism MultiSet uses, and what to send us if you want it tested on your own building.

01Indoor environmentsWarehouses, hospitals, data centres

Bounded, lit, predictable, and still the hardest class of site to localize in.

Plan of a corridor lined with eight identical repeating door modules. A dimension line marks lateral and vertical position as locked by the walls, while an arrow along the corridor shows position along it is quantised to the door pitch and otherwise unconstrained.
Why this environment is hard
  • Identical repeating modules, bay to bay
  • Cubicle curtains: the largest surfaces, gone by lunchtime
  • Mesh cabinet fronts, returns straddling two depths
What MultiSet VPS does here

Bay numbers, zone signage, cabinet asset tags and shutter numbers sit on exactly the axis the geometry leaves free, and unlike curtains they are still there next week. Nothing gets installed, so the scan is the deployment.

Bring us this: a scan of the floor you think is hardest. Book a demo
02Moving objectsPeople, vehicles, equipment

The static-world assumption fails on day one. Multi-image fusion is how MultiSet VPS absorbs it.

A room whose fixed structure is unchanged, with one large object drawn in two positions, was and now. An arrow between them marks the axis normal to the moved surface as the one that goes free.
Why this environment is hard
  • Ceiling booms, somewhere new every case
  • Belt stanchions: least permanent, most heavily weighted
  • Drapes over fixed plant, changing every surface
What MultiSet VPS does here

Cloud processing filters transients at map build time, then a multi-image query sends four to six frames over a short trajectory and requires them to agree. A trolley dominating one frame is outvoted by the other five.

Bring us this: the log where the pose jumped. Book a demo
03Changing environmentsRetail, events, fit-outs

Every map goes stale. Refreshing it should cost a scan, not a re-authoring project.

Two plans side by side, the stored map and the world today, sharing almost no common shapes. A question mark between them shows that geometric registration has no basis on which to relate the two.
Why this environment is hard
  • The same hall, bare on Monday and built by Wednesday
  • A fixture reset, moving every dominant plane
  • A crop closing the aisle a little more each day
What MultiSet VPS does here

Map Versioning computes the alignment between an old scan and a new one with the VPS itself rather than ICP, so it survives a change of layout, scanner or point density. The base scan stays the coordinate frame of record, so your content layer never moves.

Bring us this: ask what a map update costs in your current stack. Map Versioning
04GPS-denied environmentsBasements, tunnels, vessels

MultiSet VPS never touches a satellite, and still reports a global position from below deck.

A section below grade or below the waterline with the satellite reference crossed out, showing a run of repeated plant bays. An arrow along the run marks position along it as the unconstrained axis.
Why this environment is hard
  • No signal below the waterline, on a hull that is moving
  • Utility vaults, reached through a hatch
  • Platform screen doors: glazing and aliasing at once
What MultiSet VPS does here

The pose comes from matching camera imagery against a pre-scanned map, so there is no degraded mode to design for and the denied case is the default case. Georeference once and the result still returns as GeoPose in WGS 84 with the antenna disconnected.

Bring us this: your denial mode. Absent, degraded, or contested. Book a demo
05Changing conditionsLight, fog, washdown

Light, heat and airborne water degrade every sensor. The multi-image query is how MultiSet VPS absorbs it.

A sensor with a short band of usable range before a curtain of fog, frost and washdown aerosol. Beyond it the line is dashed, showing returns coming from the medium rather than the surface behind it.
Why this environment is hard
  • Cold store doors, fogging the lens on schedule
  • Rolling mill heat, bending sightlines
  • Washdown aerosol, returning before any surface
What MultiSet VPS does here

A multi-image query fuses four to six frames and selects the most reliable viewport rather than trusting whichever one arrived. Escalate only when needed: the standard query at about two seconds, Deep Search when a frame returns no pose or low confidence, then multi-image.

Bring us this: captures from your worst hour, not your best. Book a demo
06Unstructured terrainSolar, agriculture, mining

No walls, no corners, no reliable planes. MultiSet VPS localizes on appearance instead.

Identical rows on sloping graded ground. An upward arrow marks the surface normal as the only constrained direction, while an arrow along the slope shows both ground tangent directions sliding free.
Why this environment is hard
  • Tracker panels rotating through the day
  • Vineyard rows pruned to one profile
  • Water at the toe, returning nothing usable
What MultiSet VPS does here

Panel serial plates, row-end markers, bark texture and spoil arrangement stay stable on ground that is geometrically uniform, and appearance matching has no equivalent of geometric degeneracy. Sites this size can now be mapped from the air: 360 drone capture builds the map, and either drones or ground devices localize against it.

Bring us this: the two hundred metres of route you trust least. 360 to VPS
07Featureless and low-textureCleanrooms, studios, vessels

Featureless describes two different problems. MultiSet VPS solves the more common one.

A cylindrical vessel interior in section. An arrow along the axis and a curved arrow around the end face mark translation along the axis and rotation about it as two degrees of freedom with no geometric support, while height stays constrained.
Why this environment is hard
  • Cleanrooms, stripped of cues by protocol
  • Green-screen stages, stripped of them on purpose
  • Silo interiors, symmetric about the axis
What MultiSet VPS does here

MultiSet VPS matches on attention: it weights the few genuinely distinctive pixels in a view and down-weights the generic repeated ones. Where a space is poor in both channels, Object Tracking on a fixed asset is a cheaper first move than a marker survey, provided the asset is rigid, textured and asymmetric.

Bring us this: send a scan and we will tell you which kind of low-feature site you have. Object Tracking
08Multi-level mapsHotels, stadia, car parks

Two levels look identical from inside. MultiSet VPS lets you rule one out before matching begins.

Four stacked building levels with identical section, door pitch and fittings. A dimension line marks the plan as identical on every level, and a vertical arrow shows vertical position, and with it level identity, as the ambiguous axis.
Why this environment is hard
  • Guest floors differing only in printed numbers
  • Passenger decks, stacked and in motion
  • Concourse rings, one directly above the other
What MultiSet VPS does here

Room numbering, door graphics and wear differ level to level even where the footprint does not, and that is the channel a geometric matcher cannot read. hintFloorHeight and hintMapCodes let you pass in a level you already know from a lift button, a beacon or a barometer.

Bring us this: let us look at what your stack already knows about level. Book a demo
09Large-scale sitesTerminals, plants, campuses

MultiSet VPS has no drift budget, because there is no trajectory in the estimate.

A chart of error against distance travelled over ten kilometres. A widening wedge shows error growing when position is integrated from motion, against a constant band showing error staying flat when each pose is computed against the map.
Why this environment is hard
  • Steel that corrupts the fix you would use to search
  • Session joins, where pose discontinuities live
  • Rolling rebuild, so part of the map is always stale
What MultiSet VPS does here

Each pose is computed directly against a server-side map from one observation, so there is no trajectory-length term in the error and accuracy at ten kilometres is the accuracy of one query. MapSet aligns each added map against the whole set automatically, and versioning one map leaves the inter-map transforms untouched.

Bring us this: ask what share of your accuracy engineering exists purely to contain drift. Book a demo
10Repetitive environmentsCorridors, aisles, cabins

The structure repeats, the labelling does not, and MultiSet VPS reads the labelling.

Nine identical bays above a match-cost curve whose minima recur at the repeat pitch. The true position and the selected position sit on adjacent minima exactly one pitch apart, showing perceptual aliasing.
Why this environment is hard
  • Seat pitch, identical for forty metres
  • Glasshouse bays, glazing as the least trustworthy cue
  • Archive shelving that moves on rails
What MultiSet VPS does here

MultiSet VPS weights the row numbers and bay markers over the repeated module, and recall on exactly these scenes has risen with every engine generation while false positives have fallen. A tight hintPosition and hintRadius is the documented single most effective fix, and the SDK cross-checks each pose against device motion and raises a false-positive callback instead of a success when the two disagree.

Bring us this: the log where the pose jumped by exactly one repeat unit. Book a demo
11Glass, mirrors and reflective surfacesAtria, curtain wall, retail

The surfaces that define the space are the ones a sensor cannot trust.

A glazed pane with the real room on one side and a dashed ghost copy of it appearing behind the glass. The mullion and etched signage are marked as the stable cues, while the extent of the space stays uncertain.
Why this environment is hard
  • Glazing passing straight through, or reflecting twice
  • Mirrors presenting a plausible copy of the room
  • Specular runs, moving with the observer
What MultiSet VPS does here

Keep glass and mirrors out of the centre of frame and capture the mullions, etched signage, structure and soffit instead, which is what panoramas carry well. Heavily reflective areas are one of the three cases we name as genuinely hard, alongside the dark and drastic structural change, so it is worth identifying before a pilot rather than during one.

Bring us this: the atrium, the mirrored lift lobby or the fitting-room wall. Book a demo
12Crowded and occluded spacesStations, concourses, events

When half the field of view is people, what is left has to be enough.

A concourse at peak with the lower field of view filled by a crowd. A highlighted band across the upper third marks the signage gantry, soffit, fascia and column heads that crowds never occlude, captioned capture high.
Why this environment is hard
  • Bodies across the lower field of view at peak
  • Density moving with the timetable
  • A years-old reference scan underneath it all
What MultiSet VPS does here

Crowded scenes queried against a reference scan years older than the query are part of MultiSet's standard benchmark set, and recall on them has improved with every engine generation. Capture high, because the signage gantry and soffit are the part crowds never occlude.

Bring us this: your peak-hour footage, not your six-in-the-morning survey walk. Overnight AR event navigation
What makes it work

Six mechanisms, and which environments each one covers

Appearance matching

Reads signage, labels, wear and texture, the cues sitting on the axis geometry leaves free. A 6-DoF pose from one camera frame, with no drift term.

Covers 01, 04, 06, 08, 10
Future proofed releases

Attention-based matching weights distinctive pixels and down-weights the repeated ones. Every engine generation raises recall and cuts false positives on the hardest, most repetitive scenes, without changing the API you build against.

Covers 07, 10, 11, 12
Multi-image query

Four to six frames over a short trajectory, each tagged with the device tracking pose, fused into one result. Selects the most reliable viewport. About five seconds against two.

Covers 02, 05, 08, 10, 12
Localization hints

hintPosition, hintRadius, hintFloorHeight, hintMapCodes, geoHint, use2DFiltering. Each shrinks the candidate set before the aliasing-prone step.

Covers 04, 08, 09, 10
Map Versioning and MapSet

Alignment computed by the VPS rather than ICP, so it survives drastic change between captures. Multiple scans active in one frame, merged into a single coordinate system.

Covers 03, 06, 09
Capture discipline

Panoramas mandatory. Corridors both directions. Glass and mirrors out of centre frame. Capture high in crowds. Fixes more complaints than tuning does.

Covers 01, 07, 10, 11, 12
Getting started

From scan to production in three steps

01
Capture

Use the survey you already own, or take one walk with a 360 camera. Scan formats and scanners

02
Map and localize

Cloud processing builds the VPS map and filters transient objects. Query from Unity, native iOS, native Android, WebXR, Meta Quest or ROS 2. Developer docs

03
Keep it true

Map Versioning for sites that change, MapSet for sites too big for one capture. Map Versioning

Questions we get

Frequently asked questions

Why does a robot report a confident position that is wrong by exactly one aisle?

Because the space is auto-similar. Every aisle, bay or row was built the same way, so an observation is genuinely consistent with the map in many positions at once. This is perceptual aliasing, and the error is not random noise: it is one row, one bay, one module. Two fixes: capture every repeating run in both directions, which removes most of the symmetry ambiguity, and use an engine that weights the distinctive cues in a view over the repeated module. MultiSet VPS does the second, with reject-gating aimed squarely at the confident-wrong pose, and the SDK also cross-checks each pose against device motion and flags a false positive when the two disagree.

Which environments is visual positioning hardest in, and which is it easiest in?

Easiest where a space is visually distinctive and lit: signage, labels, asset tags, painted markings, texture, fixed equipment. Hardest where a space has been deliberately stripped of visual cues, such as a cleanroom under protocol or a green-screen stage, or where there is no light at all. The useful thing is that visual difficulty and geometric difficulty rarely coincide. Repeating structures are exactly the ones humans had to label, so the spaces where geometry runs out are often the ones where appearance is richest.

What happens to localization when a warehouse or a store gets rearranged?

Consistency between what the system expects and what it observes degrades, and the failure is gradual rather than announced. The fix is a re-scan, and the cost that matters is not the scan, it is the merge. MultiSet computes the alignment between an old scan and a new one with the VPS itself rather than with ICP, so it survives fixture resets, refits and a change of scanner. The original scan stays the coordinate frame of record, so anchors and content authored against it never move. More than one scan can be active at once, which means an empty hall and a built-out hall can both be live in the same frame.

Does localization accuracy get worse the further a device travels?

Not with MultiSet. Systems that integrate motion accumulate error along the trajectory: a tenth of a per cent of linear drift is fine over two hundred metres, produces a metre over a kilometre, and produces tens of metres across a site-wide route. MultiSet computes each pose directly against the map from a single observation, so there is no trajectory-length term in the error. Ten kilometres into a route, accuracy is the accuracy of one query. What grows with site size is search cost, and a rough position or geo prior bounds that.

How does a localizer tell the fourth floor from the eleventh when they look identical?

Two ways. Appearance, because room numbering, door graphics, notices, wayfinding and wear differ level to level even where the footprint does not. And explicit level knowledge, which you usually already have: someone pressed a lift button, a beacon fired in the lobby, the barometer registered a step, or the app has a level selector. hintFloorHeight restricts the search to a vertical band, hintMapCodes restricts it to specific maps in a set. Both shrink the candidate set before the ambiguous step runs. The post-lift pattern is to clear the position hint, set the new band, and fire a multi-image query as the doors open.

What does featureless actually mean, and does it stop visual positioning working?

It describes two opposite problems that get filed under one word. Geometric poverty with visual richness is the common one: a large uniform floor plate constrains almost nothing geometrically, while its painted zone markings, equipment labels and patch repairs identify the location precisely. Visual poverty with geometric richness is the reverse, an unlit plant room or bare concrete in the dark. Teams frequently report the first and specify for the second, and that single misclassification explains a lot of deployments that underperform for no visible reason. Visual positioning is built for the first case and needs a geometric layer for the second.

Does visual positioning work in spaces full of glass and mirrors?

Usually yes, and mostly by capturing correctly. Glass either passes light through or reflects it, so ghost geometry appears behind the pane, and a mirror presents a fully detailed copy of the room in the wrong place. The rule is to keep mirrors, glass and blank walls out of the centre of frame and to frame the mullions, the etched signage, the structure behind the glass, the floor and the soffit instead. Those are fixed, distinctive and unaffected by reflection. Panoramic imagery carries them well, which is one reason panoramas are mandatory for third-party scans. A fully mirrored volume with no fixed graphics anywhere in frame is hard for any modality, and it is worth identifying that before a pilot rather than during one.

Does visual positioning still work when a space is full of people?

Yes, and it is measured. Crowded scenes queried against a reference scan years older than the query are part of MultiSet's standard benchmark set, and recall on them has improved with every engine generation. Two practical things help further. Capture high, because the upper third of a concourse is the part crowds never occlude, and use a multi-image query so a crowd has to occlude four to six frames from different viewpoints to be decisive.

Does visual positioning work with no satellite signal at all, underground or below deck?

Yes, because no satellite is involved at any point. The pose comes from matching a camera frame against a pre-scanned map, so there is no degraded GNSS-denied mode to design for: the denied case is the default case. That covers all three denial types, outright absence, multipath degradation inside steel-framed buildings, and contested signal where jamming or spoofing is assumed. Georeference the map once and the result still returns as GeoPose in WGS 84, so a device in a sub-basement with its antenna disconnected reports a valid latitude, longitude and altitude. On a vessel it is better than that: the map frame travels with the hull, which is the frame you actually wanted.

Do I have to re-scan a site before visual positioning will work on it?

Usually not. If the site was captured with a survey-grade scanner, the visual layer is probably already in the file. Leica RTC360 and BLK360, NavVis, Faro, XGRIDS and Matterport Pro2 and Pro3 all record panoramic imagery alongside the point cloud, and a geometry-only pipeline discards it. So the thing to check is the export setting, not the scanner. Panoramic images must be included in the export, they are the part that carries the localization, and they can be cubic or equirectangular because the format is detected automatically. Density of both the scan and the accompanying imagery should be high. If a capture genuinely has no imagery, one walk with an Insta360 X4 or X5 adds a visual layer over the existing map at roughly a thousand square feet per sixty seconds of footage, which is usually cheaper than re-surveying.