Case Study
Six Environments. One Localization Layer.
VPS Navigation
6
environments
distinct environment classes, one system
5 cm
median accuracy
from Gaussian splat ingest to production map
<1,500
ms
pose queries in production

Most visual positioning works somewhere. The hard problem is working everywhere, because every environment fails differently. Repetitive geometry defeats retrieval. Crowds occlude the features you mapped against. Daylight and windowless corridors sit forty seconds apart on the same route. A system tuned for one of these is not a system tuned for the next. That is why most deployments stay inside the environment they were built for.

360Maps by 360 Channel did not. Since October 2025, 360Channel has deployed its AR wayfinding and anchored-content product across six distinct environment classes in Tokyo and Yokohama: a vertical mixed-use tower, open streets, an indoor-to-outdoor route spanning separate buildings, an exhibition gallery, a major arena, and a working government building. MultiSet provides the localization layer underneath.

Here is why each environment breaks most visual positioning, and why MultiSet VPS holds.

The six environments

Where Why most VPS breaks Why MultiSet VPS holds
Toranomon Hills Station Tower / TOKYO NODE, Tokyo
236,640 m², 49 floors
Vertical interior, extreme lighting range
  • Floors near-identical by design
  • Confident match on the wrong floor: a silent failure
  • One route: atrium daylight to windowless corridor to dim gallery
  • Multi-floor localization, no GPS, no beacons
  • Retrieval tuned for feature-sparse interiors
  • Tolerant of low light and direct glare
  • Map footprints beyond 10,000 m²
K Arena Yokohama
20,033 seats
Large repetitive volume
  • Bowl sections and concourse rings are near-duplicates
  • Right features matched in the wrong place
  • The canonical visual-positioning failure case
  • Features weighted by how distinctive they are
  • Generic, repeated detail down-weighted
  • 15 to 22 recall points gained in the most repetitive environments (MultiSet internal benchmark)
Akasaka, Tokyo
Street-scale, TBS Sacas district
Open street, far-field anchoring
  • Nothing controlled: crowds, vehicles, weather, day to night
  • Anchors at ~100 m: small angular error, large positional error
  • Rotational accuracy becomes the whole problem
  • Outdoor and indoor in one system
  • Tolerant of crowds, vehicles and moved equipment
  • 2° rotational accuracy on initial lock
Toranomon Hills Oval Plaza and T-Deck
Indoor to outdoor, across buildings
  • The seam at the threshold
  • Illumination swings by orders of magnitude
  • No usable GPS either side: urban canyon outside, nothing inside
  • One coordinate frame, inside and out
  • No GPS dependency at the boundary
TOKYO NODE Gallery, 45F
Dark, crowded, changing space
  • Lighting designed for attention, not feature extraction
  • Crowds occlude the mapped geometry
  • The space itself changes across a ten-week run
  • Low-light and occlusion tolerance
  • Built for dynamic environments
  • Map versioning: anchors survive a re-scan
Tokyo Metropolitan Government Building No. 1, Nishi-Shinjuku
243 m, 48 floors
Live public building, no construction
  • No authority to install anything
  • No ability to close floors for mapping
  • Window too short to amortize a fit-out
  • Camera-only localization
  • Nothing installed in the building

Three of these deserve a closer look

Repetitive geometry is where visual positioning has historically broken

Plain walls, uniform surfaces, corridors that resemble other corridors. A retrieval system trained on feature-rich scenes finds a strong match and places you a section away, confidently. An arena is this problem in its purest form, since a seating bowl is repetitive by design. The response is to stop treating all features as equal evidence: an attention-based mechanism up-weights the details that genuinely distinguish one location from another and discounts the ones that recur everywhere. On MultiSet's internal benchmark of more than 70 datasets and 16,068 test queries, this raised overall recall from 54.6% to 61.8%. On the hardest transit scene in that benchmark, median position error fell from 6.0 m to 1.8 m and the false-positive rate dropped from 73% to 13%.

The indoor-outdoor threshold is where most systems reveal they are actually two systems

One handles interiors, another leans on GNSS outdoors, and the handoff between them is visible to the user as a jump. A route that leaves a tower, crosses an open plaza and enters a second building offers no clean place to make that switch: GPS is unreliable between towers and unavailable inside them. Holding a single coordinate frame across the threshold removes the seam rather than managing it.

Far-field anchoring is a different regime from near-field

Placing a marker three metres ahead forgives a degree or two of rotational error. Placing content a hundred metres away does not: at that distance, angular accuracy dominates the result, and a lock that is precise in position but loose in orientation puts the content in the wrong part of the sky.

How one system covers all six

Capture: Gaussian splats into VPS maps

360Channel surveys its venues with 3D scanners that produce 3D Gaussian Splatting data, the same survey data behind its digital twin work. MultiSet ingests metric-scaled Gaussian splats directly and converts them into production localization maps: a .ply and a poses.json in, a queryable VPS map out, at 5 cm median localization accuracy with pose queries under 1,500 ms. One survey of a building yields both a machine-readable map and a photorealistic viewable asset.

That is a deliberate bet on where capture is heading. Gaussian splatting is becoming the JPEG of 3D, with Khronos standardization targeted for Q2 2026, and treating it as a first-class localization input means a venue's existing capture pipeline feeds positioning without a separate mapping exercise.

Models built for the difficult cases

Retrieval and matching are trained specifically on repetitive, feature-sparse interiors like malls, transit stations, and campuses, across varying lighting and crowd densities, rather than on the well-lit, feature-rich scenes that make benchmarks look good.

Maps that outlive the space

Venues change: tenants move, exhibits are installed and struck, layouts shift. Anchors and annotations persist across re-scans, with rollback and rate-of-change analytics. For a venue running several activations a year against one map, that is the difference between infrastructure and rebuilding from scratch each time.

Also worth knowing: 360 video into VPS

Gaussian splats are not the only route in. A single consumer 360-camera recording converts into a VPS map at roughly 1,000 sq ft per 60 seconds of footage: a $500 camera in place of $50,000-plus survey equipment. The same pass produces a viewable 3D Gaussian Splat alongside the machine-readable map. Where a survey crew is impractical, the cost of localizing a building drops to the cost of walking through it.

Map your venue

Six environments, one localization layer, nothing installed in the building.

Explore MultiSet VPS · Talk to us

Join industry leaders in changing the way you build spatial computing solutions and scale your deployments.