In most built environments the geometry pins down four or five degrees of freedom and leaves one or two almost free. That free axis is nearly always the one the building labelled, because somebody had to number the repeating modules, and MultiSet VPS reads those labels.What a VPS is and how it works
If your map was built with a survey-grade scanner, the visual layer is almost certainly already in the capture. Leica RTC360 and BLK360, NavVis, Faro, XGRIDS and Matterport Pro2 and Pro3 all record panoramic imagery alongside the point cloud. MultiSet ingests the structured E57 with panoramas included and builds a VPS map from the same file. One capture, both layers, one coordinate frame. Check the export setting, not the scanner.
Scan requirementsPick the one that looks like your site. Each opens to the mechanics, the mechanism MultiSet uses, and what to send us if you want it tested on your own building.
Bounded, lit, predictable, and still the hardest class of site to localize in.
Bay numbers, zone signage, cabinet asset tags and shutter numbers sit on exactly the axis the geometry leaves free, and unlike curtains they are still there next week. Nothing gets installed, so the scan is the deployment.
The static-world assumption fails on day one. Multi-image fusion is how MultiSet VPS absorbs it.
Cloud processing filters transients at map build time, then a multi-image query sends four to six frames over a short trajectory and requires them to agree. A trolley dominating one frame is outvoted by the other five.
Every map goes stale. Refreshing it should cost a scan, not a re-authoring project.
Map Versioning computes the alignment between an old scan and a new one with the VPS itself rather than ICP, so it survives a change of layout, scanner or point density. The base scan stays the coordinate frame of record, so your content layer never moves.
MultiSet VPS never touches a satellite, and still reports a global position from below deck.
The pose comes from matching camera imagery against a pre-scanned map, so there is no degraded mode to design for and the denied case is the default case. Georeference once and the result still returns as GeoPose in WGS 84 with the antenna disconnected.
Light, heat and airborne water degrade every sensor. The multi-image query is how MultiSet VPS absorbs it.
A multi-image query fuses four to six frames and selects the most reliable viewport rather than trusting whichever one arrived. Escalate only when needed: the standard query at about two seconds, Deep Search when a frame returns no pose or low confidence, then multi-image.
No walls, no corners, no reliable planes. MultiSet VPS localizes on appearance instead.
Panel serial plates, row-end markers, bark texture and spoil arrangement stay stable on ground that is geometrically uniform, and appearance matching has no equivalent of geometric degeneracy. Sites this size can now be mapped from the air: 360 drone capture builds the map, and either drones or ground devices localize against it.
Featureless describes two different problems. MultiSet VPS solves the more common one.
MultiSet VPS matches on attention: it weights the few genuinely distinctive pixels in a view and down-weights the generic repeated ones. Where a space is poor in both channels, Object Tracking on a fixed asset is a cheaper first move than a marker survey, provided the asset is rigid, textured and asymmetric.
Two levels look identical from inside. MultiSet VPS lets you rule one out before matching begins.
Room numbering, door graphics and wear differ level to level even where the footprint does not, and that is the channel a geometric matcher cannot read. hintFloorHeight and hintMapCodes let you pass in a level you already know from a lift button, a beacon or a barometer.
MultiSet VPS has no drift budget, because there is no trajectory in the estimate.
Each pose is computed directly against a server-side map from one observation, so there is no trajectory-length term in the error and accuracy at ten kilometres is the accuracy of one query. MapSet aligns each added map against the whole set automatically, and versioning one map leaves the inter-map transforms untouched.
The structure repeats, the labelling does not, and MultiSet VPS reads the labelling.
MultiSet VPS weights the row numbers and bay markers over the repeated module, and recall on exactly these scenes has risen with every engine generation while false positives have fallen. A tight hintPosition and hintRadius is the documented single most effective fix, and the SDK cross-checks each pose against device motion and raises a false-positive callback instead of a success when the two disagree.
The surfaces that define the space are the ones a sensor cannot trust.
Keep glass and mirrors out of the centre of frame and capture the mullions, etched signage, structure and soffit instead, which is what panoramas carry well. Heavily reflective areas are one of the three cases we name as genuinely hard, alongside the dark and drastic structural change, so it is worth identifying before a pilot rather than during one.
When half the field of view is people, what is left has to be enough.
Crowded scenes queried against a reference scan years older than the query are part of MultiSet's standard benchmark set, and recall on them has improved with every engine generation. Capture high, because the signage gantry and soffit are the part crowds never occlude.
Reads signage, labels, wear and texture, the cues sitting on the axis geometry leaves free. A 6-DoF pose from one camera frame, with no drift term.
Attention-based matching weights distinctive pixels and down-weights the repeated ones. Every engine generation raises recall and cuts false positives on the hardest, most repetitive scenes, without changing the API you build against.
Four to six frames over a short trajectory, each tagged with the device tracking pose, fused into one result. Selects the most reliable viewport. About five seconds against two.
hintPosition, hintRadius, hintFloorHeight, hintMapCodes, geoHint, use2DFiltering. Each shrinks the candidate set before the aliasing-prone step.
Alignment computed by the VPS rather than ICP, so it survives drastic change between captures. Multiple scans active in one frame, merged into a single coordinate system.
Panoramas mandatory. Corridors both directions. Glass and mirrors out of centre frame. Capture high in crowds. Fixes more complaints than tuning does.
Use the survey you already own, or take one walk with a 360 camera. Scan formats and scanners
Cloud processing builds the VPS map and filters transient objects. Query from Unity, native iOS, native Android, WebXR, Meta Quest or ROS 2. Developer docs
Map Versioning for sites that change, MapSet for sites too big for one capture. Map Versioning
Because the space is auto-similar. Every aisle, bay or row was built the same way, so an observation is genuinely consistent with the map in many positions at once. This is perceptual aliasing, and the error is not random noise: it is one row, one bay, one module. Two fixes: capture every repeating run in both directions, which removes most of the symmetry ambiguity, and use an engine that weights the distinctive cues in a view over the repeated module. MultiSet VPS does the second, with reject-gating aimed squarely at the confident-wrong pose, and the SDK also cross-checks each pose against device motion and flags a false positive when the two disagree.
Easiest where a space is visually distinctive and lit: signage, labels, asset tags, painted markings, texture, fixed equipment. Hardest where a space has been deliberately stripped of visual cues, such as a cleanroom under protocol or a green-screen stage, or where there is no light at all. The useful thing is that visual difficulty and geometric difficulty rarely coincide. Repeating structures are exactly the ones humans had to label, so the spaces where geometry runs out are often the ones where appearance is richest.
Consistency between what the system expects and what it observes degrades, and the failure is gradual rather than announced. The fix is a re-scan, and the cost that matters is not the scan, it is the merge. MultiSet computes the alignment between an old scan and a new one with the VPS itself rather than with ICP, so it survives fixture resets, refits and a change of scanner. The original scan stays the coordinate frame of record, so anchors and content authored against it never move. More than one scan can be active at once, which means an empty hall and a built-out hall can both be live in the same frame.
Not with MultiSet. Systems that integrate motion accumulate error along the trajectory: a tenth of a per cent of linear drift is fine over two hundred metres, produces a metre over a kilometre, and produces tens of metres across a site-wide route. MultiSet computes each pose directly against the map from a single observation, so there is no trajectory-length term in the error. Ten kilometres into a route, accuracy is the accuracy of one query. What grows with site size is search cost, and a rough position or geo prior bounds that.
Two ways. Appearance, because room numbering, door graphics, notices, wayfinding and wear differ level to level even where the footprint does not. And explicit level knowledge, which you usually already have: someone pressed a lift button, a beacon fired in the lobby, the barometer registered a step, or the app has a level selector. hintFloorHeight restricts the search to a vertical band, hintMapCodes restricts it to specific maps in a set. Both shrink the candidate set before the ambiguous step runs. The post-lift pattern is to clear the position hint, set the new band, and fire a multi-image query as the doors open.
It describes two opposite problems that get filed under one word. Geometric poverty with visual richness is the common one: a large uniform floor plate constrains almost nothing geometrically, while its painted zone markings, equipment labels and patch repairs identify the location precisely. Visual poverty with geometric richness is the reverse, an unlit plant room or bare concrete in the dark. Teams frequently report the first and specify for the second, and that single misclassification explains a lot of deployments that underperform for no visible reason. Visual positioning is built for the first case and needs a geometric layer for the second.
Usually yes, and mostly by capturing correctly. Glass either passes light through or reflects it, so ghost geometry appears behind the pane, and a mirror presents a fully detailed copy of the room in the wrong place. The rule is to keep mirrors, glass and blank walls out of the centre of frame and to frame the mullions, the etched signage, the structure behind the glass, the floor and the soffit instead. Those are fixed, distinctive and unaffected by reflection. Panoramic imagery carries them well, which is one reason panoramas are mandatory for third-party scans. A fully mirrored volume with no fixed graphics anywhere in frame is hard for any modality, and it is worth identifying that before a pilot rather than during one.
Yes, and it is measured. Crowded scenes queried against a reference scan years older than the query are part of MultiSet's standard benchmark set, and recall on them has improved with every engine generation. Two practical things help further. Capture high, because the upper third of a concourse is the part crowds never occlude, and use a multi-image query so a crowd has to occlude four to six frames from different viewpoints to be decisive.
Yes, because no satellite is involved at any point. The pose comes from matching a camera frame against a pre-scanned map, so there is no degraded GNSS-denied mode to design for: the denied case is the default case. That covers all three denial types, outright absence, multipath degradation inside steel-framed buildings, and contested signal where jamming or spoofing is assumed. Georeference the map once and the result still returns as GeoPose in WGS 84, so a device in a sub-basement with its antenna disconnected reports a valid latitude, longitude and altitude. On a vessel it is better than that: the map frame travels with the hull, which is the frame you actually wanted.
Usually not. If the site was captured with a survey-grade scanner, the visual layer is probably already in the file. Leica RTC360 and BLK360, NavVis, Faro, XGRIDS and Matterport Pro2 and Pro3 all record panoramic imagery alongside the point cloud, and a geometry-only pipeline discards it. So the thing to check is the export setting, not the scanner. Panoramic images must be included in the export, they are the part that carries the localization, and they can be cubic or equirectangular because the format is detected automatically. Density of both the scan and the accompanying imagery should be high. If a capture genuinely has no imagery, one walk with an Insta360 X4 or X5 adds a visual layer over the existing map at roughly a thousand square feet per sixty seconds of footage, which is usually cheaper than re-surveying.