Back

RoboForce Inc. · Robot Learning Data

Teleoperation
at Scale

20+ operators on real-time robotic control. My job is volume that a model can actually learn from.

1,000+episodes per day
10+collection stations
20+operators supported
2daily shifts

Data Operations Shift Lead · Milpitas, CA · Apr 2026 – Present

Role
Data Operations Shift Lead
Platforms
RoboForce Titan · Universal Robots UR5e
Controllers
GELLO · UMI · Vive Trackers · VR
Output
Demonstration data for VLA training

I was RoboForce's first data collector.

Before the team, the shifts, or the stations, there was one person and one open question, worked out with the AI engineers: what makes a demonstration good? I ran the rigs, tested what models could and couldn't learn from, and set the quality bar with the people training on it.

The SOPs, the training, and the standard the team runs on all trace back to that. Benchmark first, scale second.

Vision-Language-Action models train on human demonstrations. That puts people on the critical path — for two things at once, pulling opposite ways.

Quantity

Models need volume.

An idle station collects nothing. A minute spent fighting the rig is a minute lost across 20+ people. Scale comes from the setup, not from asking anyone to work faster.

Quality

Models need consistency.

A demonstration the model can't learn from is worse than none — it costs time, then teaches the wrong thing. Twenty operators solving a task twenty ways is a dataset with no signal.

Every task has two answers: the way a human can repeat a hundred times without drifting, and the way a model can learn. The job is the overlap.

Data collection is the first stage of the project pipeline. Everything after it inherits the ceiling we set.

  1. Data Collectionour team
  2. Annotation
  3. AI Model
  4. Model Eval

Inside our stage

  1. 01

    Requirements

    AI team

    What the models need next, in the AI team's language.

  2. 02

    Task Design

    Me

    Turned into an SOP the floor can execute identically every time.

  3. 03

    Collection

    Me + 20 operators

    10+ stations, 2 shifts, rigs kept running.

  4. 04

    Episodes

    → training

    1,000+ a day, shipped as demonstration data.

Collection statistics and model feedback set the next week's targets — which is where stage 01 comes from.

Four surfaces, one outcome: more demonstrations, more learnable.

01 GELLO · UMI · Vive Trackers · VR

Teleoperation Systems

The rig is the ceiling on everything downstream.

I improve the setup in hardware and software. Every gain multiplies across 20+ operators at once.

02 Methodology

Task Design & SOPs

I write the SOPs from doing the task, not from watching it.

Not just the procedure — the specific way of performing it that a human repeats without drifting and a model can still learn.

03 20+ operators

Team Operations

Hiring through performance review, built as systems that run themselves.

Interviewing, onboarding, training, daily summaries, scheduling, reviews, continuous feedback — automated wherever it would otherwise become recurring manual work.

04 Linux · Python · ROS 2

Station Uptime

A down station is data that never gets collected.

I triage failures directly — terminal, logs, ROS 2 graph — and work with engineering on the rest.

RoboForce TitanUniversal Robots UR5eGELLOUMIVive TrackersVR ControllersROS 2PythonLinux

Two examples, and the tradeoff behind each.

Python · curses · QR

Shift Output Tracking

A terminal UI, not a web dashboard.

Stakeholders wanted per-operator output per shift. Operators already work in a Linux terminal, so a TUI meant no context switch and nothing new to log into. QR tracking keeps entry fast on the floor; the same numbers roll up to leadership.

Rubrics · scoring

Structured Technical Interview

Rubrics, so the interview measures aptitude instead of confidence.

Collectors come from mostly non-technical backgrounds, so an unstructured interview selects for vocabulary and self-assurance. Explicit scoring plus a guide that walks the interviewer to an objective write-up — same bar for everyone, whoever runs it.

The engineers writing requirements have mostly never run a controller. The operators executing them mostly aren't technical. I sit between the two, because I've done both jobs.

EngineersCollectors

An instruction that can't survive contact with the floor isn't a requirement yet.

Requirements arrive in the AI team's language. I turn them into something a non-technical operator can execute identically every time — Linux terminal included.

CollectorsEngineers

What's actually collectable, in terms engineers can design against.

At what quality, at what rate, which tasks are realistic at volume. That is why targets are reachable and the data comes back as asked.

Weekly targets come out of that exchange — collection statistics on one side, the AI team's requirements on the other. Reviews measure against that, not raw volume.

I was the first person collecting this data, and I still run a controller. Every system here was built by someone who has to use it.

Building something? Let's talk.

jonathangoenadibrata@gmail.com