Back to all roles

Senior Research Scientist, Vision Transformers and Feed-Forward Networks

Push the transformer and MLP side of our localization stack on accuracy, latency and cost.
Research
Bengaluru, India
Full-time
Hybrid

About MultiSet

MultiSet builds the Visual Positioning System for Physical AI. Our platform gives robots, smart glasses, phones and facility workers a single shared coordinate frame, indoors and in GPS-denied environments. MultiSet is an Auggie Award winner for Best Developer Tool.

Scan-agnostic is the part that matters. Most positioning vendors require their own capture hardware or their own app. We take what a customer already owns. E57 point clouds, LiDAR, Matterport and MatterPak, Leica, NavVis, FARO, XGRIDS, Insta360 walkthroughs, Gaussian splats, textured meshes, phone scans. Vision Fusion normalizes scale, lighting and noise into one map. MapSet stitches captures from different scanners into a single coordinate frame across a whole site.

The result runs today across 5,000+ locations in 150+ countries, covering more than 30M sq ft, with 1M+ device interactions. Every map is private and consented. Your data never becomes our model.

About the Role

You will lead model research for visual place recognition and localization, reporting to our CTO. The work is hands-on. You will design architectures, run the training, and validate results against our production baseline on our own data.

We train our own models on data we collected ourselves. Your job is to push the transformer and feed-forward parts of that stack on accuracy, latency and cost at the same time.

What You Will Work On

  • Design and train transformer backbones for visual place recognition and global retrieval.
  • Improve the feed-forward and MLP components of our architectures: attention-free blocks, token mixing, MLP head design for descriptor projection and pose regression.
  • Push descriptor quality and compactness so retrieval stays accurate as map size grows.
  • Cut inference cost per query without giving up recall, using distillation, pruning, quantization and architecture search.
  • Prepare models for on-device deployment through CoreML and LiteRT, working with our SDK engineers.
  • Build honest evaluation. Design benchmarks on our internal datasets that predict real customer performance, not leaderboard performance.
  • Run failure analysis on live customer maps and turn what you find into model changes.
  • Write up results clearly for the team, and publish or patent where it makes sense.
  • Set research direction with the CTO, and mentor junior researchers and engineers on the team.

What We Are Looking For

  • MS or PhD in computer vision, machine learning or a related field, or equivalent research experience.
  • 5+ years of applied research experience, including at least one model you took from idea to production.
  • Strong grasp of vision transformer architectures and what the feed-forward blocks actually do inside them.
  • Real experience training vision models from scratch, not only fine-tuning. You have debugged a run that would not converge and found the reason.
  • Fluent in PyTorch. Comfortable with multi-GPU training and cloud GPU workflows.
  • Working knowledge of visual localization, SLAM, structure from motion, or image retrieval.
  • Ability to read a new paper, judge whether it is worth the effort, and reproduce it fast.
  • Care about latency and memory, not only accuracy.

Nice to Have

  • Published work in visual place recognition, feature matching, or self-supervised representation learning.
  • Experience with self-supervised or foundation model vision backbones.
  • Model compression and edge deployment experience on mobile or embedded targets.
  • ROS 2, robotics, or XR development experience.
  • Experience with 3D Gaussian Splatting or COLMAP reconstruction pipelines.
  • Contributions to open source vision or localization projects.

Location and Work Model

Bengaluru, India. Hybrid. You will work closely with our CTO and with the SDK engineers who take your models to device.

Compensation

Competitive salary plus meaningful equity. We talk about the range on the first call, before you invest time in a process. We are angel-backed and raising our seed now.

How to Apply

Apply through the form on the careers page. Choose this role from the dropdown, attach your CV, and include links to papers, code or demos. Come ready to talk about one model you built, and what you learned when it did not work. That conversation matters more to us than a list of benchmarks.

Inclusive Application

The strongest candidates rarely match every line. If this role fits what you want to do next and you can do the work, apply. A partial match is still worth our time and yours.

Equal Opportunity

MultiSet AI is an equal opportunity employer. We hire on ability and evidence. We consider all applicants without regard to race, colour, ancestry, national origin, religion, age, sex, gender identity or expression, sexual orientation, marital or family status, disability, veteran status, or any other characteristic protected by applicable law. Accommodations are available on request at any stage of the process.

Apply for this role
Name this role in the application form. A resume and one paragraph is enough to start.
Apply