RoboForce Inc. · Robot Learning Data
Teleoperation
at Scale
20+ operators on real-time robotic control. My job is volume that a model can actually learn from.
Data Operations Shift Lead · Milpitas, CA · Apr 2026 – Present
- Role
- Data Operations Shift Lead
- Platforms
- RoboForce Titan · Universal Robots UR5e
- Controllers
- GELLO · UMI · Vive Trackers · VR
- Output
- Demonstration data for VLA training
Where This Started
I was RoboForce's first data collector.
Before the team, the shifts, or the stations, there was one person and one open question, worked out with the AI engineers: what makes a demonstration good? I ran the rigs, tested what models could and couldn't learn from, and set the quality bar with the people training on it.
The SOPs, the training, and the standard the team runs on all trace back to that. Benchmark first, scale second.
The Challenge
Vision-Language-Action models train on human demonstrations. That puts people on the critical path — for two things at once, pulling opposite ways.
Quantity
Models need volume.
An idle station collects nothing. A minute spent fighting the rig is a minute lost across 20+ people. Scale comes from the setup, not from asking anyone to work faster.
Quality
Models need consistency.
A demonstration the model can't learn from is worse than none — it costs time, then teaches the wrong thing. Twenty operators solving a task twenty ways is a dataset with no signal.
Every task has two answers: the way a human can repeat a hundred times without drifting, and the way a model can learn. The job is the overlap.
Where We Sit
Data collection is the first stage of the project pipeline. Everything after it inherits the ceiling we set.
- Data Collectionour team
- Annotation
- AI Model
- Model Eval
Inside our stage
- 01
Requirements
AI teamWhat the models need next, in the AI team's language.
- 02
Task Design
MeTurned into an SOP the floor can execute identically every time.
- 03
Collection
Me + 20 operators10+ stations, 2 shifts, rigs kept running.
- 04
Episodes
→ training1,000+ a day, shipped as demonstration data.
Collection statistics and model feedback set the next week's targets — which is where stage 01 comes from.
What I Own
Four surfaces, one outcome: more demonstrations, more learnable.
Teleoperation Systems
The rig is the ceiling on everything downstream.
I improve the setup in hardware and software. Every gain multiplies across 20+ operators at once.
Task Design & SOPs
I write the SOPs from doing the task, not from watching it.
Not just the procedure — the specific way of performing it that a human repeats without drifting and a model can still learn.
Team Operations
Hiring through performance review, built as systems that run themselves.
Interviewing, onboarding, training, daily summaries, scheduling, reviews, continuous feedback — automated wherever it would otherwise become recurring manual work.
Station Uptime
A down station is data that never gets collected.
I triage failures directly — terminal, logs, ROS 2 graph — and work with engineering on the rest.
Systems I Built
Two examples, and the tradeoff behind each.
Shift Output Tracking
A terminal UI, not a web dashboard.
Stakeholders wanted per-operator output per shift. Operators already work in a Linux terminal, so a TUI meant no context switch and nothing new to log into. QR tracking keeps entry fast on the floor; the same numbers roll up to leadership.
Structured Technical Interview
Rubrics, so the interview measures aptitude instead of confidence.
Collectors come from mostly non-technical backgrounds, so an unstructured interview selects for vocabulary and self-assurance. Explicit scoring plus a guide that walks the interviewer to an objective write-up — same bar for everyone, whoever runs it.
Between the AI Team and the Floor
The engineers writing requirements have mostly never run a controller. The operators executing them mostly aren't technical. I sit between the two, because I've done both jobs.
EngineersCollectors
An instruction that can't survive contact with the floor isn't a requirement yet.
Requirements arrive in the AI team's language. I turn them into something a non-technical operator can execute identically every time — Linux terminal included.
CollectorsEngineers
What's actually collectable, in terms engineers can design against.
At what quality, at what rate, which tasks are realistic at volume. That is why targets are reachable and the data comes back as asked.
Weekly targets come out of that exchange — collection statistics on one side, the AI team's requirements on the other. Reviews measure against that, not raw volume.
I was the first person collecting this data, and I still run a controller. Every system here was built by someone who has to use it.
Let's Connect
Building something? Let's talk.
jonathangoenadibrata@gmail.com