In our last post, deterministic checks found large amounts of corrupted data in open-source UMI and teleoperation datasets. Many errors, however, are too context-dependent or nuanced for deterministic checks to catch. Beyond the video, misrepresented labels, extraneous footage, and missing annotations can also all poison a model's understanding of well-formed data.

To remedy this, we built Argus, an open-source annotation and quality pipeline to audit everything that a model learns from an episode. Argus reads every color camera stream in a recording, along with the recorded robot state and actions, the task instructions, and the publisher's annotations. It produces dense timeline annotations, flags faults in both operator execution and the recording itself, and runs deterministic checks to verify dataset metadata. We have run the pipeline on OpenAI's GPT-6 Astra, GPT-6 and 6.1 Sol, Claude Opus 5.5, and DeepSeek's v4.1 Flash, with varying results (see How Argus performs across models).

We share results from 3,546 of the episodes we have audited with Argus (66.5 hours). These span nine datasets of teleoperated robots, UMI and human ego video. For all nine datasets, we are releasing every annotation from our pipeline.

Process

PipelineFrom a recorded episode to a published annotation
  1. 01Episode as recorded

    We take every camera, the recorded motion and the dataset's own instruction exactly as they were recorded.

  2. 02Exact frames

    We cut each episode's frames by their exact timestamps and send Astra one frame every 1 to 1.5 seconds, two a second on human ego video.

  3. 03Astra

    Astra sees the cameras side by side in grids.

  4. 04Dense annotation

    Astra returns a timeline, key events, the outcome and goal frame, mistakes and recoveries, and an instruction check.

  5. 05Checks

    Our deterministic checks are stored next to the labels.

  6. 06Dashboard

    Every episode can be browsed, downloaded as JSON and exported as JSON Lines.

Constructing a faithful annotation pipeline is tricky because of the enormous diversity that exists across datasets. For example, one tele-op dataset features arms mounted on a mobile base while another has three robot arms and four cameras; ego datasets contain human hands while UMI also features grippers. The lack of normalization in the field means that a pipeline must reconcile different data formats and metadata fields. Our design strives for accuracy while remaining cost-effective, and is structured to take advantage of Astra's ability to interpret its given frames. Below were some of our key design considerations:

  • Reading recordings in their original format. Robotics datasets rarely share the same format, file structure, camera layout, and more. Argus reads LeRobot datasets (v2.0, v2.1 and v3.0), MCAP files, plain video (mp4, mov, mkv, webm and avi) and zip or tar archives of any of these, whether each episode has its own files or many episodes share one.
  • Checking that the number of frames is consistent for every episode. LeRobot v3, for example, is a major format structured such that an mp4 from one camera contains back-to-back episodes, separated by timestamps. In order to avoid incorrect frame parsing for each episode, and rounding errors in those timestamps, which are floats stored to the microsecond, we divide the frames into episodes with PyAV by integer PTS. Argus keeps only the frame whose PTS matches the expected value on the video's time base, and raises an error on any mismatch, so an episode is never labelled from the wrong frames.
  • Adapting resolution to task specifications. We first filter tele-op episodes with a smaller model, using GPT-6 Sol to read the task specification to determine whether the task needs high resolution to display fine-grained details, such as lettering or displays. The harness then sends the footage from each camera at 224 pixels only if the task is a confirmed low resolution task, and sends footage at 448 pixels if the task is uncertain or confirmed to require high resolution. To ensure that we have precise evaluation on the success of a task, we also send at high resolution any frames where the gripper closes and opens on an object, along with the first and last frames for each episode.
  • Accounting for moving camera angles. In order to prevent mis-annotations due to variable perspectives from mobile wrist cameras, Argus describes how each camera is mounted (fixed, or mobile on a wrist or head) and asks Astra to infer each camera's direction based on scene geometry (see Astra out of the box).
  • Reconstructing raw timelines. Argus infers the events between frames based on concrete observations rather than assuming optimal movements. For example, if a shirt is grasped in one frame and on the table in the next, our harness prompts Astra to determine whether it was dropped or placed, based on observations such as the position of its landing and how crumpled the shirt is.
  • Checks based on physical observations. A number of errors can be caught from the physics of the recording. Since the cameras are fixed onto the grippers, we can verify the accuracy of UMI pose annotations. Astra is given each gripper's recorded motion, checks it against the recorded video and flags clear contradictions, such as when the camera turns while the recorded pose remains fixed. Deterministic checks detect any sped-up sessions from the delay between leader and follower arms, and any swapped camera streams where the left and right sides are flipped.

For every episode we input, Argus outputs the following:

  • A timeline of each action phase of each arm, gripper or hand, with the object it acts on, where the object goes, whether the step advances the task, is wasted or is idle, and how much of the task is done at that point
  • Every change in an object's state that the cameras show (closed to open, unstacked to stacked), and on robot footage the camera views that show the state before and after
  • Key events, the moments a reviewer would mark to judge progress (every fold of a T-shirt, every cup stacked), each with its outcome
  • The goal frame, the first frame where the full goal holds, and for a goal that was later undone, when and how it was undone
  • The completion outcome (success, success then undone, failure or unclear, with a task left partly done counted as a failure), the full end state it was judged against, and the reason when it is not a clean success
  • On human ego video, each separate activity with its own outcome and goal frame, and for every step whether the hands are in view and what covers them
  • An instruction check, which says whether the footage shows the task the instruction asks for, more than it, only part of it, or a different task, along with the model's own one-sentence description of what was done
  • Every operator mistake that would teach a model a bad habit (a drop, a failed grasp, a long struggle), with when it happened, its severity, and the camera and time that show it
  • Every failed attempt, with whether and when the operator recovered, what they actually did, and the ideal fix
  • Every problem in the recording or its labels, found by the model or by our deterministic checks, such as an instruction that describes a different task, a person reaching into the scene, a task object out of every view, an episode cut off or left idle, swapped camera streams, a sped-up recording, or a recorded gripper that never moves while the video shows it opening and closing
  • A severity rating for each issue (low, medium or high), assigned by Argus based on how much of the episode it affects and how directly it corrupts what a model learns

An example of a complete output is shown below.

What one annotation containsFlip three blocks to show different faces (MolmoAct2 episode 1346)
Failure

Data issue, high severityThe instruction describes flipping three blocks, while the recorded demonstration collects three blocks into a row and aligns that row.

Left armRight armBoth armsProgress0 to 100% of the goalKey eventssubgoalState changesOutcomethe instructed goal is never reached0s5s10s15s20s25s30s35s40s45s50s55s60s65s70s75s80s85s90s95sLeft armRight armBoth armsProgress0 to 100% of the goalKey eventssubgoalState changesOutcomegoal never reached0s20s40s60s80s
Actionsadvancing the taskwasted effortidle
Eventskey eventsubgoal completestate change
Outcomeinstructed goal never reached
At 1:18the instructed goal is never reachedprogress 0%move along the timeline to read any second

What the arms are doing

Both armsWithdraw and open empty grippersto clear of the row1:18 to 1:21

Scene, as of 1:18

  • P block next-to W block
  • W block next-to Y block
  • P block nearer-table-edge-than W block
  • W block nearer-table-edge-than Y block
  • three-block row separate-from remaining wooden blocks
1:18three-block row roughly side-to-side across the foreground table to aligned roughly front-to-back, with P nearest and Y farthest
1:18subgoal complete Three-block row finished in its final orientation

Astra out of the box

To assess Astra's out-of-the-box performance, we gave it just the task and one 448-pixel-wide frame per second from each camera. We asked Astra to generate a timeline, judge the outcome, and flag any problems.

MolmoAct2 episode 1346, Flip three blocks to show different faces.

The left gripper holds the P block. The wrist camera looks down on its top face, the letter P, and past it onto a side face with a 3.Left wrist camera, 0:28
Astra, given only the task
“The left gripper picks up and reorients the blue P block, then places it beside the other two. Its 3 face is visible on top, completing a three-block row.”
What the footage shows

The wrist camera looks down on the block, so the P facing it is the top, and the 3 above it is a side seen at an angle. Astra mistook the 3 for the top because it sits higher in the picture. No block is turned over, and the letters were on top from the start.

Given only the task, Astra consistently misjudged episode 1346 (shown above), incorrectly reporting that the block was turned over in all nine runs. This was because the wrist cameras move with the grippers, so the view of a block changes even when the block does not. Our harness overcomes this problem by describing how each camera is mounted and instructing Astra to treat an object as changed only when confirmed by the fixed camera or a comparable view.

Incorrect task instructions can also bias Astra towards “seeing” erroneously named objects, especially when the footage itself is visually ambiguous. In FastUMI Prepare_tableware, the task instruction asks for chopsticks even though the gripper holds a fork. Astra reports “chopsticks picked up from the tray”, likely mistaking the black lines on the gripper's two claws as chopsticks. In MolmoAct2 episode 8276, the instruction asks for black pants even though the garment is a black polo shirt; Astra accepts the instruction's word instead of reconciling the garment's shape across frames and angles, and judges that the “pants” are correctly folded. Our harness therefore presents the objects an instruction names as claims to verify against the frames, and requires that Astra describe each object's identifying features before naming it.

FastUMI‑100K Prepare_tableware episode 001768, Put the chopsticks and spoon from the left-hand tray into the right-hand bowl.

The gripper holds a fork upright over the table; the label below the video says chopsticks.Gripper camera, 0:03
Astra, taking the object's name from the instruction
“Chopsticks picked up from the tray”
What the footage shows

The gripper holds a fork. There are no chopsticks anywhere in the episode.

MolmoAct2 episode 8276, Grasp the black pants, fold them in half lengthwise, then fold again widthwise. Place the folded pants neatly on the table.

A wrist camera close to a black knit garment; a short sleeve with its hemmed opening lies across the table.Left wrist camera, 0:33
Astra, taking the object's name from the instruction
“The black pants are folded in half lengthwise, folded crosswise into a compact bundle, and placed neatly on the table.”
What the footage shows

A short sleeve and its hemmed opening. The garment is a short-sleeved knit polo shirt, and there are no pants in the episode.

How Argus performs across models

Argus is model agnostic, and can run on any sufficiently capable VLM. We ran Argus on GPT-6 Astra, Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol (all at medium reasoning), anchoring to Astra as a benchmark. Each model was given the same 193 episodes to annotate, keeping the prompt, frame, and output budget consistent. These episodes cover about one hour each of tele-op, UMI and human ego footage, and all answers parsed successfully.

Outcome agreementHow often two models give an episode the same outcome
AstraAstraClaude Opus 5.5Opus 5.5GPT-6 SolGPT-6 SolDeepSeek v4.1 flashDeepSeekGPT-6.1 SolGPT-6.1 SolAstraAstra96%84%73%95%Claude Opus 5.5Opus 5.596%88%75%95%GPT-6 SolGPT-6 Sol84%88%70%86%DeepSeek v4.1 flashDeepSeek73%75%70%73%GPT-6.1 SolGPT-6.1 Sol95%95%86%73%

Each share is over the 171 teleoperation and UMI episodes that all five models answered. A human ego clip holds several tasks, so it has no single outcome to compare.

The models usually converged on the same outcome. Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol agreed with Astra on 73% to 96% of the tele-op and UMI episodes. However, Astra labels much more densely than the first three models we compared, marking 44 events per minute of footage on these episodes, 1.9 times as many as the densest of them, GPT-6 Sol. The difference in density is largest on human ego video, where Astra marks 26 key events per episode while none of the three marks more than 15. GPT-6.1 Sol is the exception, at 43 events per minute and 28 key events per human ego episode.

In-context learning from Astra's traces does not close the gap. We added one fully annotated Astra example to the prompts of Claude Opus 5.5, GPT-6 Sol and DeepSeek v4.1 flash, then asked each one to label 64 of the episodes again. Outcome agreement shifted only marginally and annotation density increased by only 4 to 9%, at most reaching 59% of the 43.5 events Astra marked per minute on those episodes. We expect Argus-style pipelines to improve as models become increasingly accurate, fast, and cheap, until every episode of every dataset can be checked this way. GPT-6.1 Sol, the newest model we tried, already labels 96% as many events per minute as Astra on the same episodes and gives the same outcome on 95% of the tele-op and UMI ones, for $5 per hour of footage where Astra costs $27.

In-context learningThe same models, each given one complete Astra annotation of another episode in its prompt

All footage64 episodesAstra labels 43.5 events per minute on these

Without an Astra trace
With one Astra trace in context
Astra
ModelLabelled events per minuteEvents per minSame outcome as Astra*Outcome*Subgoals per episodeSubgoals
  1. Claude Opus 5.517.8 → 19.496% → 96%3.9 → 2.8
  2. GPT-6 Sol24.7 → 25.777% → 82%3.7 → 2.4
  3. DeepSeek v4.1 flash16.2 → 17.074% → 70%2.8 → 2.6

* Teleoperation and UMI episodes only.

Models by releaseTimeline events per minute of footage, in release order
Labelled events per minute01020304050Astra, 44.3 labelled events per minuteAstra44.3Sep 2026DeepSeek v4.1 flash, 15.3 labelled events per minuteDeepSeek v4.1 flash15.3Sep 2026Claude Opus 5.5, 17.9 labelled events per minuteClaude Opus 5.517.9Sep 2026GPT-6 Sol, 23.9 labelled events per minuteGPT-6 Sol23.9Sep 2026GPT-6.1 Sol, 42.6 labelled events per minuteGPT-6.1 Sol42.6Sep 2026Labelled events per minute02550AstraSep 202644.3DeepSeek v4.1 flashSep 202615.3Claude Opus 5.5Sep 202617.9GPT-6 SolSep 202623.9GPT-6.1 SolSep 202642.6

Each model labelled the same 193 episodes (186 minutes of teleoperation, UMI and human ego footage) through our harness, with identical prompts and frames. Each point counts the timeline events a model labels per minute of footage, which measures how finely it divides what happens, not whether each event is right, and the line follows the highest rate released so far.

Every annotation from this comparison can be found on the data dashboard, kept separately from Astra's: Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol, and the charts that compare them all.

Example Results

01Success then undone

Issue
The demonstration is successful, but the recording continues past when the goal is achieved and shows the operator knocking the finished arrangement apart. As such, the last frame does not show the goal.
Downstream Impact
Goal-conditioned models and value functions that read the last frame as the goal learn the wrong goal, and successful demonstrations are miscategorized as failures. These episodes can be converted into clean success demonstrations by trimming at the reported goal frame.
How often
  • 10%Galaxea Open‑World23 of 222 episodes
  • 1.9%MolmoAct2‑BimanualYAM25 of 1,284 episodes
  • 1.8%10Kh‑RealOmin5 of 280 episodes
Four blocks in one straight row, the task's goal.
0:18, goal reachedFour blocks in one straight row, the task's goal.
The row has been pushed apart and the arms have withdrawn.
0:37, end of recordingThe row has been pushed apart and the arms have withdrawn.

What to watchThe row is complete at 0:18. At 0:29, the grippers begin pushing the blocks apart, and the episode ends without the completed row.

Task given by the dataset“Push blocks to align them in a straight line.”

topframe 555

Statefour selected blocks loosely grouped with gaps and differing angles → compact straight row

bothlift and reposition at row ends → opposite ends of row

leftframe 555
rightframe 555

Statefour selected blocks loosely grouped with gaps and differing angles → compact straight row

bothlift and reposition at row ends → opposite ends of row

18.50 s
555 / 1,114
Flagged frames
The model's labels
left
right
both
state
recover
flags

02Task instructions do not match footage

Issue
The task instructions describe a task different from what is shown. The magnitude of this problem ranges from low severity cases with minor deviations, to medium and high severity cases that could teach a world model incorrect dynamics. 20% of MolmoAct2 episodes were medium or high severity cases, and 177 of its 202 “failed” demonstrations carry such an instruction.
Downstream Impact
Language-conditioned policies learn incorrect associations between words and behaviors, and datasets look worse than they are as successful demonstrations are misclassified. This is fixed by replacing the original instruction with Astra's description of the episode.
How often
  • 79%OpenAoE‑2000h85 of 107 episodes (32 low)
  • 34%Gen‑HumanEgo27 of 79 episodes (19 low)
  • 26%Galaxea Open‑World57 of 222 episodes (13 low)
  • 25%MolmoAct2‑BimanualYAM319 of 1,284 episodes (60 low)
  • 13%FastUMI‑100K128 of 964 episodes (53 low)
  • 13%10Kh‑RealOmin37 of 280 episodes (4 low)
  • 8.2%ABC‑130k15 of 183 episodes (11 low)
  • 7.3%HABIT23 of 315 episodes (14 low)
What the dataset saysWhat the footage shows

MolmoAct2, episode 17777 (the video below)

“Pick up each item, scan its barcode, and place it in the basket.”

The operator takes the four packages out of the basket, scans each one, and collects them on the right side of the table.

FastUMI-100K, Prepare_tableware (all 32 episodes)

“Put the chopsticks and spoon from the left-hand tray into the right-hand bowl.”

A fork and a spoon are moved from a plate into a rectangular wire basket. There are no chopsticks and no bowl.

What to watchFrom the first seconds the arms lift packages out of the basket and scan them onto the table, the opposite of the instruction.

Task given by the dataset“Pick up each item, scan its barcode, and place it in the basket.”

topframe 30

bothremain parked

leftframe 30
rightframe 30

bothremain parked

1.00 s
30 / 2,411
Flagged frames
The model's labels
left
right
both
state
flags

03Head camera taken off during recording

Issue
In ego data, the operator takes the head camera off during the recording, so part of the clip shows a sudden change in perspective and has the camera pointing at the operator's face or the ceiling.
Downstream Impact
The video should primarily be of hand movement; cutting the affected segment is sufficient to fix this.
How often
  • 12%Egocentric‑100K13 of 112 episodes (3 low)
factory 126, 1:12 to 1:15Lifted off the head at the sewing machine until it points up at the wearer's face.
factory 167, 0:01 to 0:04Taken off and turned round to face the wearer.
factory 102, 1:20 to 1:23Covered by a hand, then set down facing the ceiling.

What to watchEach clip is the three seconds around the moment the camera comes off. We blurred the faces ourselves.

04Recording plays faster than real time

Issue
The episode is sped up. This occurs when the control loop records at less than 30 Hz, but the episode is timestamped at 30 Hz. Timestamps, video timing and frame-rate metadata are all uniform, so the files do not reveal the problem. We found these sessions with a deterministic check: the delay between the follower and leader arms is fixed at about 125 ms, so if the delay spans too few frames, the session must be playing artificially quickly.
Downstream Impact
World models learn incorrect velocities and policies learn to move faster than they should. The compression factor in a sped-up episode can be measured by the frame lag between the follower and leader arms, so the session can be corrected to run at real time or dropped.
How often
  • 11%MolmoAct2‑BimanualYAM3,536 of all 32,246 episodes (294 of the 1,284 labelled)
3.75 framesfollower lag at a true 30 Hz (125 ms)
2.69 framesfollower lag in this session
≈ 22 Hzthe rate it was actually recorded at
≈ 1.39×faster than real time on playback

What to watchThe motion looks brisk and slightly stuttered.

Task given by the dataset“Place dirty dishes in dishwasher rack, sort waste into bins, and clean table.”

topframe 120

rightCarry and lower blue cup → dishwasher rack

rightRelease blue cup and withdraw → dishwasher rack

leftframe 120
rightframe 120

rightCarry and lower blue cup → dishwasher rack

rightRelease blue cup and withdraw → dishwasher rack

4.00 s by the file's timestamp, captured at 5.58 s
120 / 1,178
Timestamps against capture39.3 s of timestamps at 30 Hz, captured over 54.8 s at 22 Hz
The frame's timestamp in the file
1.6 s behind
4.0 s
When the frame was captured
5.6 s
The model's labels
left
right
both
state
recover
flags

05Camera files swapped

Issue
The left camera stream shows the right gripper's view and vice versa, but nothing in the files flags this. Cases can be identified with a deterministic check which compares each stream's motion change with each gripper's movement. In the example below, the left stream moves with the right gripper's recorded motion far more than with the left gripper (0.68 correlation versus 0.28), and the other gripper also sits on the wrong side of the frame. It's important to note that recorded motion alone cannot be trusted to find all these swaps, as the motion data itself is often flawed. For example, in all 14 of FastUMI‑100K's swapped episodes, the recorded gripper opening does not actually happen, and in 50 other episodes, the recorded pose remains still while the camera moves, or jumps while the camera does not. Argus instead uses Astra to detect the camera swap independently by examining the raw footage.
Downstream Impact
World models learn incorrect cause-and-effects: one gripper's motion moves the other gripper's view. Swapping the streams then fixes the mispairing.
How often
  • 1.5%FastUMI‑100K14 of 964 episodes
  • 0.36%10Kh‑RealOmin1 of 280 episodes

What to watchAs we can see from the example below, the left camera stream changes when the right gripper moves (purple), not when the left one moves (blue).

IPEC-COMMUNITY/FastUMI_100k_lerobotFold_the_T-shirt / episode_002711

Task given by the dataset“First align and fold the two sleeves of the T-shirt over the body, then fold the T-shirt vertically in half from top to bottom.”

leftframe 40
rightframe 40

leftalign fingers around first sleeve → shirt body

leftgrasp sleeve → held between fingers

2.00 s
40 / 268
Flagged frames
The model's labels
left
right
both
state
flags
Left camera change vs. each gripper's recorded motionindependent scales
left camera change 1.1220 luma MADleft gripper motion 0.4695 cmright gripper motion 0.7904 cm

06Privacy blurring hides task objects

Issue
Blurring intended to mask people's faces can mistakenly mask objects that the robot is handling.
Downstream Impact
Policies learn to grasp objects they cannot properly see. The frames can be masked as a partial remedy, but only the publisher is able to completely fix this issue by correctly masking the original footage.
How often
  • 3.8%HABIT12 of 315 episodes (11 low)

What to watchFrom 0:19 the donut in the left gripper's camera turns into a blurred rectangle. It is sharp again for a moment at 0:21, then blurred again until 0:23.

configinc/HABITepisode 5733

Task given by the dataset“Move the first donut from the left to the tray.”

left wrist viewframe 160

leftgrasp box edge → between fingers

rightremain parked

exo viewframe 160
right wrist viewframe 160

leftgrasp box edge → between fingers

rightremain parked

16.00 s
160 / 240
Flagged frames
The model's labels
left
right
both
state
flags

07Task object leaves every camera

Issue
Part of the task happens out of view for all cameras.
Downstream Impact
Since the recorded actions in those timestamps have no visual evidence, world models are asked to predict frames without the context on what the arms are doing, and the model cannot verify the correctness of the end state. These episodes can still be used once the period of no visual activity is masked.
How often
  • 0.78%MolmoAct2‑BimanualYAM10 of 1,284 episodes (2 low)
  • 0.45%Galaxea Open‑World1 of 222 episodes

What to watchAt 0:42, the red packet is placed at the far right edge of the table. From that point onwards, only a small section can be seen in the top view, and the right gripper's camera sees it only in passing, at 0:49 and 1:24.

Task given by the dataset“Pick up the red packet and scan its barcode.”

topframe 1,290

Statered packet held at scanner → resting at far-right edge

rightApproach and align with orange carton → staging area

leftframe 1,290
rightframe 1,290

Statered packet held at scanner → resting at far-right edge

rightApproach and align with orange carton → staging area

43.00 s
1,290 / 2,933
Flagged frames
The model's labels
left
right
both
state
recover
flags

08Labelling operator mistakes as clean demonstrations

Issue
The recording completes the task, but there are mistakes in the operation, including missing grasps, dropping objects, and meandering paths. These episodes are then released as good demonstrations despite containing iterations of suboptimal actions.
Downstream Impact
Policies trained by imitation will copy the suboptimal approach to the task. To improve behavior cloning, these episodes should either be filtered out or trimmed to remove the mistakes. However, for world models, these operator mistakes can become useful data points for learning how to recover from failed attempts, and should be labeled as such and kept instead.
How often
  • 6.6%ABC‑130k12 of 183 episodes
  • 5.5%MolmoAct2‑BimanualYAM71 of 1,284 episodes
  • 3.6%Galaxea Open‑World8 of 222 episodes
  • 2.1%10Kh‑RealOmin6 of 280 episodes
  • 0.62%FastUMI‑100K6 of 964 episodes

What to watchAt 1:59, the left gripper attempts to pick up the scissors, but they repeatedly slip out of the grasp, and from 2:06 the right gripper tries too, repeating the grabbing motion for over twenty seconds.

XDOF/ABC-130kepisode 3d75a76a

Task given by the dataset“organize the medicine kit”

leftframe 3,513

RecoveredInitial right-arm scissors grasp fails. After repeated attempts, both arms repositioned the scissors and the left arm secured the blade and packed them.

leftattempt grasp along blades

topframe 3,513
rightframe 3,513

RecoveredInitial right-arm scissors grasp fails. After repeated attempts, both arms repositioned the scissors and the left arm secured the blade and packed them.

leftattempt grasp along blades

117.10 s
3,513 / 6,397
Flagged frames
The model's labels
left
right
both
state
recover
flags

Summary of results

Of the 3,546 episodes we labelled across nine datasets, 27% have at least one issue that Argus flags at medium severity or above, either from the model or from a deterministic check. In these episodes, part or all of the footage is wrong or wasted for training as it was recorded.

27%of all 3,546 episodes have at least one issue that Argus flags at medium or high severity
46different problems found across the nine datasets, 25 in the data and 21 in how operators demonstrated
20%of episodes, across the eight datasets with instructions or labels, have an instruction or label that does not fully match the footage (14% at medium or high severity)
38labelled events per minute of footage, 153,213 in all, each with its start, end, action and object
Issue profileThe problems in each dataset, and how many of its episodes have them

MolmoAct2‑BimanualYAM1,284 labelled episodes25.1 hours of footage

High
The episode cannot be used as it is.
Medium
Part of it is wrong, and the rest is good once that part is trimmed or masked.
Low
Worth knowing, but a model trained on it would barely be affected.
ProblemShare of episodesShareEpisodesShare of hoursHours

Data issuesfaults in the recording, the scene or the label

  1. Instruction or labels do not match footage25%319 (60 low)
  2. Recording plays faster than real timeDeterministic checkmeasured over all 32,246 episodes of the dataset11%3,536
  3. Success then undone1.9%250.88%
  4. Long idle stretch1.0%13 (8 low)0.48%
  5. Task object leaves every camera0.78%10 (2 low)0.02%
  6. Scale or display unreadable0.62%80.14%
  7. Person changes the scene0.31%4 (1 low)0.03%
  1. Recorded motion disagrees with video0.16%20.01%
  2. Camera image frozen or lost0.08%10.00%

Operator mistakesthe recording is faithful, but the demonstration went wrong

  1. Repeated attempts12%155 (149 low)
  2. Missed grasp8.8%113 (112 low)
  3. Wasted motion7.6%98 (98 low)
  4. Missed placement4.4%57 (57 low)
  5. Dropped object3.8%49 (16 low)
  1. Knocked object2.4%31 (12 low)
  2. Task left unfinished2.3%30 (12 low)
  3. Missed insertion0.70%9 (9 low)
  4. Undid its own result0.39%5 (4 low)
  5. Missed rotation0.39%5 (5 low)
  6. Collision0.23%3 (3 low)
  7. Long pause0.23%3 (3 low)
  8. Spilled contents0.08%1
  9. Misaligned placement0.08%1 (1 low)
  10. Selection error0.08%1 (1 low)
  11. Failed activation0.08%1 (1 low)
  12. Spillage0.08%1 (1 low)
Labelled episodesThe episodes labelled from each dataset, every one with a full dense annotation
DatasetSetupEpisodesHours
MolmoAct2-BimanualYAMTeleoperated arms1,28425.1
ABC-130kTeleoperated arms1835.5
Galaxea Open-WorldTeleoperated mobile robot2225.7
HABITTeleoperated arms3155.0
FastUMI-100KUMI9644.8
10Kh-RealOminUMI2805.5
Egocentric-100KHuman ego1125.6
Gen-HumanEgoHuman ego793.5
OpenAoE-2000hHuman ego1075.7

Labelling costs about $26 per hour of footage on teleoperated arms, $30 on UMI and $19 on human ego video, measured on about 20 episodes of each setup drawn from these datasets, with median episodes of 71 s, 21 s and 3 min.

Every data issueEvery data issue found, at any severity, as a share of each dataset's labelled episodes
ProblemMolmoAct2ABC-130kGalaxeaHABITFastUMIRealOminEgocentric-100KGen-HumanEgoOpenAoE
Instruction or labels do not match footage25%8.2%26%7.3%13%13%·34%79%
Recording plays faster than real time11%········
Recorded motion disagrees with video0.16%·0.45%1.6%6.0%6.8%···
Long idle stretch1.0%·0.90%3.5%0.21%7.1%12%1.3%·
Success then undone1.9%·10%··1.8%···
Recorded pose jumps, video does not····3.0%4.3%···
Recorded gripper opening never changes····1.7%6.1%···
Person changes the scene0.31%·1.4%·0.10%2.5%4.5%·0.93%
Camera files swapped····1.5%0.36%···
Task object leaves every camera0.78%·0.45%···1.8%··
Video starts or ends mid-task·0.55%··0.21%0.71%0.89%8.9%·
Privacy blur hides task objects···3.8%···1.3%·
Camera turned away or covered······12%··
Scale or display unreadable0.62%········
Setup changes mid-episode·····1.8%0.89%··
Camera image frozen or lost0.08%···0.21%0.36%···
Duplicate adjacent frames·1.6%·0.32%·····
Large action with no camera motion····0.10%0.71%···
Recorded pose jumps away and back····0.31%····
Instruction missing··0.90%······
Frozen camera···0.32%0.10%····
Too dark or too bright to see······1.8%··
Another person handles the same object······0.89%··
Limited visual detail······0.89%··
Camera not where described········0.93%

A dot means none were found. Playback speed can only be checked on MolmoAct2, whose leader and follower arms make it measurable, and its cell covers every episode of the dataset.

Every operator mistakeEvery operator mistake found, at any severity, as a share of each dataset's labelled episodes
ProblemMolmoAct2ABC-130kGalaxeaHABITFastUMIRealOminEgocentric-100KGen-HumanEgoOpenAoE
Repeated attempts12%17%1.8%··6.4%··1.9%
Missed grasp8.8%7.1%9.5%0.32%0.10%2.1%···
Wasted motion7.6%3.3%1.8%0.63%·0.71%···
Dropped object3.8%4.4%2.3%·0.10%1.4%·1.3%·
Missed placement4.4%0.55%0.45%··0.36%···
Task left unfinished2.3%2.7%0.45%·0.83%1.4%···
Knocked object2.4%1.1%0.90%······
Missed insertion0.70%1.1%·······
Undid its own result0.39%········
Missed rotation0.39%········
Collision0.23%·0.45%······
Long pause0.23%········
Misaligned placement0.08%0.55%·······
Spilled contents0.08%·0.45%······
Selection error0.08%········
Failed activation0.08%········
Spillage0.08%········
Failed thread engagement·0.55%·······
Incomplete placement··0.45%······
Slipped grasp·····0.36%···
Food hygiene········0.93%

None of Egocentric-100K's 112 clips has an operator mistake.

The mix of errors varies significantly by dataset, and each error affects training in a specific way. Incorrect instructions teach a language-conditioned policy the wrong association and can make failures look like good demonstrations, an undone goal gives goal-conditioned models the wrong target, and a sped-up session teaches wrong dynamics. Most of these errors are easy to fix once found, but they are pernicious when left unnoticed.

Conclusion

Clean data is essential for effectively scaling model capabilities. As a research lab focused on training World Models, we benefit from as much high-quality data as possible. By building and open-sourcing Argus, we aim to improve the quality standard for public robotics data, to ensure that anyone can inspect these datasets with ease and use with confidence.

We are publishing all annotations on our data dashboard at pantheon.inc/data-board, including a JSON download for each episode and a JSON Lines export of any dataset and filter. You can run Argus on your own recordings with data.pantheon.inc/review, or from its open-source code on GitHub.

Credits

This project was undertaken by Eric Li. We thank XK Lu, Jenna Hong, and Janet Guo for their valued help with the writeup, and Jannik Schilling, Kunvar Thaman, Sambhav Gupta, and Calix Huang for suggestions and ideas.

If you are interested in a data-related collaboration or have any questions or comments, please contact us at data@pantheon.inc.

Dataset citations