APIs, integration & security — in depth
Training DataLong read

Annotation Pipelines for Robot Demonstration Data

Labeling raw robot video and sensor data is now the bottleneck in training effective policies.

Staff Writer · · 11 min read
Cover illustration for “Annotation Pipelines for Robot Demonstration Data”
Training Data · October 7, 2026 · 11 min read · 2,499 words

A teleoperation session ends after several hours, and what's left behind is a file, or more precisely a stack of files: RGB video, depth frames, LiDAR point clouds, joint-state streams, end-effector trajectories, force and haptic readings, a log of action commands. None of it is labeled. None of it is segmented. A neural network pointed at that raw stack has no way to tell where one task ends and the next begins, no way to know which frames capture a successful grasp and which capture a near-miss, no way to connect a spike in the force sensor to the moment a gripper lost contact with an object. The data is dense with behavior, but it is not structured in a way any learning algorithm can use.

This is the actual starting condition of most robot demonstration data today, and it's a harder problem than it sounds. Demonstrations come from teleoperation, kinesthetic teaching, virtual reality control, joystick input, motion capture, or straight recordings of a human doing the task by hand. No matter the source, you end up with one continuous, unsegmented record of several sensor modalities running in parallel. Without annotation, a learning system has no way to determine task boundaries, intent, temporal dependencies, or the surrounding environmental context. Models trained on that raw, unannotated signal tend to produce brittle policies that work in the lab and fall apart the moment conditions shift even slightly outside it.

In-the-wild robot data collection is accelerating, and the rate at which demonstrations are being recorded is now outpacing the rate at which they can be labeled by hand. So annotation, and not data collection, is becoming the limiting step when you build capable robot policies. The rest of this piece is about what that annotation step actually involves, why the choices made inside it matter as much as the data itself, and how the field's answers to those choices are starting to change.

Annotation pipeline layers

You can't turn a raw demonstration into training data with just one labeling pass. You need several, stacked on top of each other, and each one answers a separate question about the same recording. An annotation pipeline has to classify actions, label object states, estimate poses, track interactions, map the environment, and tag temporal events, and none of those jobs substitutes for another.

In manipulation tasks, this gets specific fast. Annotators mark grip points, trace end-effector trajectories, track object contact states, and divide the task into manipulation phases, and each of those requires looking across multiple sensor modalities at once. One layer answers what happened (the action), another answers to what (the object). Another answers in what phase of the task (temporal structure). Another answers toward what goal, the task itself, and another answers with what outcome: success or failure.

No single schema covers all of this, because the tasks themselves are too different from one another. A pipeline built for a pick-and-place task on a tabletop looks nothing like one built for a long-horizon kitchen task, where a dozen subgoals and several tools come into play. The diversity of robotic applications becomes, directly and unavoidably, a diversity of annotation types, conventions, and edge cases. Every section that follows examines one or more of these layers in detail, because the layer being annotated determines both the method you use and the stakes if that layer is handled poorly.

How temporal segmentation shapes what a policy can learn

Of all these layers, temporal segmentation is the foundation, because it decides the basic unit of behavior the model is being asked to learn. Robot trajectory segmentation breaks a continuous motion into structured pieces: motion phases, manipulation steps, navigation transitions, recovery actions, moments where the robot is simply waiting. Done well, this gives the model a hierarchy of behavior instead of a flat, undifferentiated stream of frames.

Picture a robot arm that reaches into a cluttered bin to pick out a single object. That one episode contains an approach phase, a repositioning adjustment as the gripper nears the target, the grasp itself, a lift, and a retreat. Segment the whole thing as one undivided "pick" action, and the model learns to complete the task end to end but has nothing to fall back on if something goes wrong mid-reach: no internal representation of the sub-phases means no way to recognize which phase failed or what recovery from that point should look like. But if you segment too finely, cutting the motion into small slices tied to the quirks of one specific teleoperator's movements, the model starts fitting to that person's particular style rather than the task, and it loses the ability to generalize to a different operator or a different grip.

Choosing the right granularity is harder than it sounds, because a demonstration is a continuous physical motion, not a static image with clean edges. Identifying the exact moment an action starts, stops, or transitions into another phase takes real expertise, and human demonstrations are full of ambiguous, borderline moments where even a skilled annotator could draw the line in more than one place. Task decomposition labeling pushes this further still: real tasks involve subtask boundaries, intermediate goals, prerequisites, dependencies between steps, and sometimes more than one valid path to the same outcome. Annotating that structure is what lets a model plan hierarchically, rather than simply replaying a sequence of actions it has memorized.

Why multimodal synchronization is a prerequisite, not an option

When the sensor streams being segmented don't line up with each other in time, temporal segmentation can't give you a usable training signal. If the visual feed lags behind the tactile signal by even a fraction of a second, what the robot saw and what it felt get scrambled, and the model learns the wrong relationship between cause and effect.

A full-fidelity pipeline has to align RGB video, depth, LiDAR point clouds, joint states, end-effector trajectories, force and haptic signals, action commands, and environmental context to one shared timeline. Get that wrong, and no amount of careful segmentation downstream will fix it, because the labels will be attached to the wrong moment in the robot's actual experience. AgiBotWorld 2026 treats this as an infrastructure requirement rather than something to patch after the fact: its G2 hardware platform captures RGB(D), tactile signals, force, LiDAR point clouds, IMU data, and full-body joint states within one unified, synchronized pipeline from the moment of collection.

This matters most in manipulation, where the event that actually determines success or failure, a slip, the instant a grasp either holds or fails, often appears in the force or tactile channel before it appears anywhere in the video. Force-controlled data collection captures contact dynamics alongside motion trajectories, so it can catch that kind of event. If you skip it, the model is trained on a visual record of a task that looks fine right up until the object hits the floor, with no signal anywhere in the data explaining why.

How discarding failed demonstrations has limited robot learning

Once a demonstration has been segmented and synchronized, a pipeline still has to decide what to do with the ones that end badly. The standard practice across the field has been to keep the successful runs and throw out everything else, treating failure as a signal the model never gets to see. The logic behind this is simple enough on its face: train a model only on clean, correct examples, and it should learn to behave correctly.

Real deployment does not cooperate with that logic. Lighting changes, objects sit at angles no one planned for, and grippers meet friction that never showed up in any simulation, and because a model is trained exclusively on success, it has never seen a single example of what to do when something starts to go wrong. It has no learned response to failure, because failure was deliberately removed from everything it was shown.

The trajectories being thrown away aren't noise. Most of them contain high-quality motion right up to the moment things fall apart: the approach was correct, the positioning was correct, the force applied was reasonable, and then one thing, a slip, a misjudged angle, broke the sequence. Discarding the whole trajectory because of that last moment throws away the most instructive part of it: what a near-correct attempt looks like just before it fails. The scarcity of natural-language and structured annotation for this kind of diverse robot data is a recognized constraint on the field, and what labeled data does exist tends to rely on post-hoc crowdsourced labeling, which brings its own costs: high price, inconsistent quality, reduced diversity because of templated instructions, and uneven granularity from one annotator to the next. What remains open is what a pipeline would look like if it annotated failure directly.

Structured failure annotation

Diagram: Three Equally Valid States: AGIBOT's Redefinition of Clean Data. Visualizes: Visualize the shift from a binary (success/discard) annotation model to a three-state model where success, failure, and recovery are all labeled training inputs.

AgiBotWorld 2026, released by AGIBOT in April 2026, answers that question concretely. Rather than discarding failed trajectories, the dataset keeps them and attaches two additional structured fields: error_cause, which records why the failure happened (gripper slip, position error, and similar categories), and restorable, which records whether recovery from that specific failure was possible. A trajectory that ends badly is no longer dropped from the dataset. It becomes a labeled event the model can learn from directly, in the same way a successful trajectory is.

AGIBOT's free-form collection approach has teleoperators responding to real conditions as they happen rather than following a fixed script, which naturally produces a wide range of error types alongside a range of recovery attempts, the raw material from which a model can begin to learn how to correct itself mid-task. The annotation hierarchy built on top of this spans task-level descriptions, action sequences, atomic skill labels, object annotations (2D bounding boxes with attributes including name and color), and error-recovery trajectories, covering the full span from high-level task framing down to individual object properties.

The underlying shift here is in what counts as clean data. Clean used to mean success-only: filter out anything that didn't go right, and keep the rest. AgiBotWorld 2026's annotation philosophy treats success, failure, and recovery as three equally valid labeled states, and each one tells the model something the other two cannot. That is a real redefinition of data quality, and it is a measurable one: annotation philosophy, not scale alone, is what determines how much a model trained on this data can actually do when conditions stop cooperating.

The quality-coverage problem that automated annotation pipelines have not yet solved

Hand annotation at this level of detail cannot grow at the rate data collection does, which limits the scale automated pipelines are meant to close. Automated annotation pipelines exist to close that gap, but they introduce a different one: without a reliable way to measure whether an automatically generated label is actually correct, a pipeline designer is stuck choosing between two bad options, accepting noisy labels or throwing away samples that might otherwise be useful.

The RoboAnnotatorX paper, presented at ICCV 2025, names this limitation directly: existing automated methods struggle to maintain temporal coherence and semantic richness across long, extended demonstrations, the exact kind of multi-step tasks where annotation matters most. A related failure mode occurs when you use vision-language models for zero-shot labeling. These systems perform well when the target concept lines up closely with whatever the model saw during pre-training, and the performance doesn't degrade gracefully when it doesn't, it collapses. A model asked to label a concept well outside its training distribution will often produce a confident, specific, and wrong answer.

That's the deeper issue: a high-confidence detection from an automated system doesn't mean the annotation is correct for the purposes of training a policy. Detector confidence and annotation correctness measure two different things, and if you conflate them, you get forced into the binary choice described above. Automated pipelines haven't solved the problem by producing bad labels. It's that automation, on its own, gives no reliable way to tell a good label from a bad one, which leaves no principled way to set a quality threshold without losing large fractions of otherwise usable data in the process.

How SPARC's reliability scoring replaces binary accept/reject filtering

SPARC, short for Spatial Annotations from Robot Demonstrations with Reliability Calibration, is a direct response to that gap. Built by researchers at KIT, NVIDIA, and Robotics Institute Germany and published in June 2026, SPARC automatically labels robot demonstrations with structured spatial annotations and attaches a reliability score to each one.

The structure of the task itself produces the score. SPARC draws on the spatio-temporal structure already present in how robots perform tasks, including phase-aware motion and gripper proximity, and it applies a robot-overlap filter, so the reliability signal it produces is grounded in interaction evidence rather than in how certain a vision model feels about its own output. Tested against 1,700 human-annotated demonstrations spanning a range of embodiments and scenarios, SPARC outperformed detection-only baselines on object localization accuracy by a wide margin, while retaining three times more samples at high-precision operating points, a separate measure of how much usable data survives at a given quality bar.

That improvement in annotation quality carries through to downstream policy performance. Policies trained on SPARC-annotated reasoning data achieved more than three times the success rate of a no-reasoning baseline across cluttered manipulation tasks, a result that measures something distinct from the localization accuracy figure above: it measures what the robot can actually do once trained on the resulting labels. Together, the two results separate a claim about annotation quality from a claim about downstream capability, and SPARC's contribution is that it replaces a hard accept or reject line with a reliability score, so pipeline designers can set it wherever their task demands and trade coverage for precision as needed rather than being locked into one fixed operating point.

Diagram: SPARC: Reliability Scoring vs. Binary Accept/Reject. Visualizes: Visualize SPARC's key measured outcomes against the detection-only baseline it replaces.

Zero-shot language labeling and the scale problem for long-horizon data

SPARC solves spatial annotation but leaves natural-language, semantic annotation unaddressed. A policy trained on well-localized objects and reliable spatial labels still needs natural-language descriptions of what the task is and why each step happens, and generating that kind of annotation at scale is a separate challenge from the one SPARC addresses.

Most robot interaction datasets available today carry no language annotation at all, because it costs too much to produce by hand. Where language labels do exist, they typically come from post-hoc crowdsourced labeling, the same method responsible for the cost, quality, and granularity problems already described in the failure-annotation discussion above. Zero-shot language labeling generates these descriptions directly from a vision-language model without human intervention at each step, and it's the approach the field now turns to for long-horizon tasks, where a demonstration can span a dozen subgoals and hours of recording. You have to reason over extended temporal context, understanding how an early action in a sequence relates to a goal reached several minutes later, rather than just pattern-matching against isolated frames. That requirement is what separates this problem from spatial annotation, and it marks automation's advance along a second, distinct axis of the pipeline.

Sources

  1. ICCV 2025 Open Access Repository
  2. SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale
  3. Towards Generalist Robot Learning from Internet Video: A Survey
  4. How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
Filed underTraining Data

More in Training Data