Forward deployed researchers

The deployment layer for embodied AI.

Embodied AI isn't limited by models. It's limited by deployment. Synjuku builds the forward deployed researcher, the missing layer between the lab and the real world, and the infrastructure they run.

$5M+ contracted with three world-class model companies, and a network reaching 200+ real work sites.

The vision

In-house data doesn't reach the real world.

Labs that want real-world applications need what no in-house team can create: the diversity of environments and tasks that captures how work is actually done. That takes infrastructure: forward deployed researchers connected to real sites, and the tooling to work with them. Synjuku is that layer.

image · candidates pending
The problem

The forward deployed researcher doesn't exist.

Every robot policy improves through the same deployment cycle. The middle of that cycle happens at a real work site, not in the lab.

01

Train

Lab pretrains the policy.

02

Deploy

Robot enters a real site.

03

Experiment

Observe what fails, and why.

04

Capture

Collect targeted data on site.

05

Retrain

Model improves. Repeat.

Steps 02 through 04 need a researcher standing on site: a forward deployed researcher. That role barely exists, so the cycle stalls at the lab door.
The solution

Deployment infrastructure that drives adoption at scale.

We turn workers at real work sites into research technicians, and turn their sites into supervised learning environments for robot fleets. The deployment loop is the product.

01
Upskill

Workers become research technicians

We train the people at real sites to run capture sessions, triage failures, and feed the loop: a new class of talent between the research scientist and the on-site operator.

02
Capture

The long tail, at research grade

Rare scenarios, failure modes, and real environments, captured where work happens. We optimize novelty per hour of collection, not hours collected, and grade every delivery.

03
Deploy

Sites become deployment-ready

Each trained site is an environment a lab can send robots into, with people who know how to run supervised learning sessions and a loop that turns failures into training data.

The capture standard

Four synchronized views. One timeline.

Every episode is recorded on a synchronized four-camera rig: egocentric, exocentric, and both wrists, cross-synced onto a single timeline. Quality gates run at the edge, so bad captures are caught on site, not after delivery.

The rig is built so a trained site worker, not an engineer, runs a session end to end. Task cards and a manifest travel with every episode, so provenance starts at the moment of capture.

Ego + exo + both wrists Cross-camera sync QC gates at the edge Run by trained workers
On-site capture4-cam · synced
Ego29.97 fps
Exo29.97 fps
Wrist · L29.97 fps
Wrist · R29.97 fps
SYNC  shared timeline LABELS  verified spans QC  pass ✓
The data

One capture. Every layer a policy needs.

Each episode is delivered as aligned layers, not raw video. Vision, motion, language, and geometry, time-synced and grounded in the same task, in your lab's native training format.

Multi-viewMulti-camera calibration and cross-view feature matching
LAYER 01

Multi-view capture

Egocentric, exocentric, and both wrists, cross-calibrated and synced onto one timeline.

Motion21-keypoint hand-pose skeleton overlay across an ego sequence
LAYER 02

Hand pose & motion

Per-frame 21-keypoint hand skeletons for left, right, and grip center, tracked through the episode.

LanguageManipulation primitives labeled with action language and timestamps
LAYER 03

Verified labels

Task → action → primitive segmentation with natural-language spans, verified against calibrated checkers and removed if they can't be confirmed.

SpatialReconstructed 3D hand mesh beside the source frame
LAYER 04

3D hand & geometry

Reconstructed 3D hand mesh, camera pose from SLAM, and exocentric body pose for spatial grounding.

CoverageEnvironment relationship chord diagram and task distribution charts
LAYER 05

Coverage & diversity

Every batch reports its spread across environments and tasks: coverage you can audit, not just count.

Provenancetask card → site → operator → episode
LAYER 06

Provenance

Every delivered episode traces back through its task card, capture site, and operator. 1:1 verified, versioned end to end, with a QC report in every delivery.

Layer previews are internal pipeline outputs, pending clearance for public use.
Your data is top tier for robot learning.
A leading robotics-model lab · active customer
Traction

Stage one is live.

$5M+
revenue. Teach capture is live, stage one of the FDR timeline.
2
deployment design projects, co-designed with customers.
200+
real work sites within network reach for collection and deployment.

Already collecting long-tail, diverse data that improves model performance today.

Why Synjuku

Everyone else sells training. We run deployment.

Real environments Demo environments Training Deployment Volume data vendors Egocentric capture, gig collectors Studio & teleop factories Staged scenes, repeated tasks Labs in-house Marquee demos run as rehearsal Synjuku FDRs deploying in real environments
Demo vs real environments, training vs deployment. The upper-right quadrant is empty: no one else puts researchers on site to run deployment.
Real, not staged

Real environments, real work

No staged scenes, no repeated studio tasks. Capture happens at working sites, from cafés and kitchens to plants and warehouses, where the long tail actually lives.

Graded, not claimed

A QC report with every delivery

Two-layer quality gates: signal integrity, then label truth against calibrated human gold sets. Labels that can't be confirmed are removed, not shipped.

Your format

Data your policy trains on directly

Delivered in your lab's ingestion format, LeRobot v3 and MCAP shipped today, with primitive-level annotation. No translation step between our delivery and your training run.

The site network

Deployment infrastructure, not a labor pool

Trained research technicians at real sites, with reach to 200+ environments. Each certified site lowers the cost of the next capture, and is a place robots can actually go to work.

Work with us

One loop, two sides. We sit in the middle.

For model & robotics labs

Reach the real world without building a field org.

  • Long-tail training data from real work sites, in your ingestion format
  • Co-designed collection, targeted at what your model gets wrong, not generic volume
  • Deployment environments: trained sites your robots can enter, with people who know how to run the loop
For site owners

Put your site on the map of embodied AI.

  • Revenue from your environment: your site's work becomes training data labs pay for
  • Upskilled workers: we train your people into certified research technicians
  • First in line for robots: deployment-ready sites are where working robots land first
Team

Operators who already run the loop.

Frank Jiang

Frank Jiang

Co-Founder · CEO

Ops, data, sites. Scaled a VLM training-data library from 7K to 6M+ hours at Troveo.

Gary Zhang

Gary Zhang

Co-Founder · Research

UC Berkeley Robotics PhD; built the full capture-to-delivery pipeline.

LN

Lucas Negritto

Co-Founder · Product

Ex-OpenAI, previously YC founder. Data and product across the pipeline.

CH

Cancy Han

Founding Partnerships · BD

Partnerships and business development across labs and sites.

S

Seth

Founding Engineer

Core infrastructure and tooling across the capture and delivery stack.

Plus a 20+ person field operations team running collection across Southeast Asia.

Advisors

Advised by people who built the field.

Robotics-lab researchers, VLA founding contributors, world-model builders, and robotics investors and founders.

Industry & labs
Physical Intelligence Google Robotics NVIDIA Meta Dyna Robotics AMD
Research & universities
UC Berkeley Harvard Tsinghua NTU Singapore

The loop starts with a conversation.

Whether you're a lab that needs the real world or a site that's ready for robots, tell us what you're working on and we'll scope where to begin.