In our last post, deterministic checks found large amounts of corrupted data in open-source UMI and teleoperation datasets. Many errors, however, are too context-dependent or nuanced for deterministic checks to catch. Beyond the video, misrepresented labels, extraneous footage, and missing annotations can also all poison a model's understanding of well-formed data.
To remedy this, we built Argus, an open-source annotation and quality pipeline to audit everything that a model learns from an episode. Argus reads every color camera stream in a recording, along with the recorded robot state and actions, the task instructions, and the publisher's annotations. It produces dense timeline annotations, flags faults in both operator execution and the recording itself, and runs deterministic checks to verify dataset metadata. We have run the pipeline on OpenAI's GPT-6 Astra, GPT-6 and 6.1 Sol, Claude Opus 5.5, and DeepSeek's v4.1 Flash, with varying results (see How Argus performs across models).
We share results from 3,546 of the episodes we have audited with Argus (66.5 hours). These span nine datasets of teleoperated robots, UMI and human ego video. For all nine datasets, we are releasing every annotation from our pipeline.
- MolmoAct2-BimanualYAM annotations
- ABC-130k annotations
- Galaxea Open-World annotations
- HABIT annotations
- FastUMI-100K annotations
- 10Kh-RealOmin annotations
- Egocentric-100K annotations
- Gen-HumanEgo annotations
- OpenAoE-2000h annotations
Process
- 01Episode as recorded
We take every camera, the recorded motion and the dataset's own instruction exactly as they were recorded.
- 02Exact frames
We cut each episode's frames by their exact timestamps and send Astra one frame every 1 to 1.5 seconds, two a second on human ego video.
- 03Astra
Astra sees the cameras side by side in grids.
- 04Dense annotation
Astra returns a timeline, key events, the outcome and goal frame, mistakes and recoveries, and an instruction check.
- 05Checks
Our deterministic checks are stored next to the labels.
- 06Dashboard
Every episode can be browsed, downloaded as JSON and exported as JSON Lines.
Constructing a faithful annotation pipeline is tricky because of the enormous diversity that exists across datasets. For example, one tele-op dataset features arms mounted on a mobile base while another has three robot arms and four cameras; ego datasets contain human hands while UMI also features grippers. The lack of normalization in the field means that a pipeline must reconcile different data formats and metadata fields. Our design strives for accuracy while remaining cost-effective, and is structured to take advantage of Astra's ability to interpret its given frames. Below were some of our key design considerations:
- Reading recordings in their original format. Robotics datasets rarely share the same format, file structure, camera layout, and more. Argus reads LeRobot datasets (v2.0, v2.1 and v3.0), MCAP files, plain video (mp4, mov, mkv, webm and avi) and zip or tar archives of any of these, whether each episode has its own files or many episodes share one.
- Checking that the number of frames is consistent for every episode. LeRobot v3, for example, is a major format structured such that an mp4 from one camera contains back-to-back episodes, separated by timestamps. In order to avoid incorrect frame parsing for each episode, and rounding errors in those timestamps, which are floats stored to the microsecond, we divide the frames into episodes with PyAV by integer PTS. Argus keeps only the frame whose PTS matches the expected value on the video's time base, and raises an error on any mismatch, so an episode is never labelled from the wrong frames.
- Adapting resolution to task specifications. We first filter tele-op episodes with a smaller model, using GPT-6 Sol to read the task specification to determine whether the task needs high resolution to display fine-grained details, such as lettering or displays. The harness then sends the footage from each camera at 224 pixels only if the task is a confirmed low resolution task, and sends footage at 448 pixels if the task is uncertain or confirmed to require high resolution. To ensure that we have precise evaluation on the success of a task, we also send at high resolution any frames where the gripper closes and opens on an object, along with the first and last frames for each episode.
- Accounting for moving camera angles. In order to prevent mis-annotations due to variable perspectives from mobile wrist cameras, Argus describes how each camera is mounted (fixed, or mobile on a wrist or head) and asks Astra to infer each camera's direction based on scene geometry (see Astra out of the box).
- Reconstructing raw timelines. Argus infers the events between frames based on concrete observations rather than assuming optimal movements. For example, if a shirt is grasped in one frame and on the table in the next, our harness prompts Astra to determine whether it was dropped or placed, based on observations such as the position of its landing and how crumpled the shirt is.
- Checks based on physical observations. A number of errors can be caught from the physics of the recording. Since the cameras are fixed onto the grippers, we can verify the accuracy of UMI pose annotations. Astra is given each gripper's recorded motion, checks it against the recorded video and flags clear contradictions, such as when the camera turns while the recorded pose remains fixed. Deterministic checks detect any sped-up sessions from the delay between leader and follower arms, and any swapped camera streams where the left and right sides are flipped.
For every episode we input, Argus outputs the following:
- A timeline of each action phase of each arm, gripper or hand, with the object it acts on, where the object goes, whether the step advances the task, is wasted or is idle, and how much of the task is done at that point
- Every change in an object's state that the cameras show (closed to open, unstacked to stacked), and on robot footage the camera views that show the state before and after
- Key events, the moments a reviewer would mark to judge progress (every fold of a T-shirt, every cup stacked), each with its outcome
- The goal frame, the first frame where the full goal holds, and for a goal that was later undone, when and how it was undone
- The completion outcome (success, success then undone, failure or unclear, with a task left partly done counted as a failure), the full end state it was judged against, and the reason when it is not a clean success
- On human ego video, each separate activity with its own outcome and goal frame, and for every step whether the hands are in view and what covers them
- An instruction check, which says whether the footage shows the task the instruction asks for, more than it, only part of it, or a different task, along with the model's own one-sentence description of what was done
- Every operator mistake that would teach a model a bad habit (a drop, a failed grasp, a long struggle), with when it happened, its severity, and the camera and time that show it
- Every failed attempt, with whether and when the operator recovered, what they actually did, and the ideal fix
- Every problem in the recording or its labels, found by the model or by our deterministic checks, such as an instruction that describes a different task, a person reaching into the scene, a task object out of every view, an episode cut off or left idle, swapped camera streams, a sped-up recording, or a recorded gripper that never moves while the video shows it opening and closing
- A severity rating for each issue (low, medium or high), assigned by Argus based on how much of the episode it affects and how directly it corrupts what a model learns
An example of a complete output is shown below.
Data issue, high severityThe instruction describes flipping three blocks, while the recorded demonstration collects three blocks into a row and aligns that row.
What the arms are doing
Scene, as of 1:18
- P block next-to W block
- W block next-to Y block
- P block nearer-table-edge-than W block
- W block nearer-table-edge-than Y block
- three-block row separate-from remaining wooden blocks
Astra out of the box
To assess Astra's out-of-the-box performance, we gave it just the task and one 448-pixel-wide frame per second from each camera. We asked Astra to generate a timeline, judge the outcome, and flag any problems.
MolmoAct2 episode 1346, Flip three blocks to show different faces.
Left wrist camera, 0:28“The left gripper picks up and reorients the blue P block, then places it beside the other two. Its 3 face is visible on top, completing a three-block row.”What the footage shows
The wrist camera looks down on the block, so the P facing it is the top, and the 3 above it is a side seen at an angle. Astra mistook the 3 for the top because it sits higher in the picture. No block is turned over, and the letters were on top from the start.
Given only the task, Astra consistently misjudged episode 1346 (shown above), incorrectly reporting that the block was turned over in all nine runs. This was because the wrist cameras move with the grippers, so the view of a block changes even when the block does not. Our harness overcomes this problem by describing how each camera is mounted and instructing Astra to treat an object as changed only when confirmed by the fixed camera or a comparable view.
Incorrect task instructions can also bias Astra towards “seeing” erroneously named objects, especially when the footage itself is visually ambiguous. In FastUMI Prepare_tableware, the task instruction asks for chopsticks even though the gripper holds a fork. Astra reports “chopsticks picked up from the tray”, likely mistaking the black lines on the gripper's two claws as chopsticks. In MolmoAct2 episode 8276, the instruction asks for black pants even though the garment is a black polo shirt; Astra accepts the instruction's word instead of reconciling the garment's shape across frames and angles, and judges that the “pants” are correctly folded. Our harness therefore presents the objects an instruction names as claims to verify against the frames, and requires that Astra describe each object's identifying features before naming it.
FastUMI‑100K Prepare_tableware episode 001768, Put the chopsticks and spoon from the left-hand tray into the right-hand bowl.
Gripper camera, 0:03“Chopsticks picked up from the tray”What the footage shows
The gripper holds a fork. There are no chopsticks anywhere in the episode.
MolmoAct2 episode 8276, Grasp the black pants, fold them in half lengthwise, then fold again widthwise. Place the folded pants neatly on the table.
Left wrist camera, 0:33“The black pants are folded in half lengthwise, folded crosswise into a compact bundle, and placed neatly on the table.”What the footage shows
A short sleeve and its hemmed opening. The garment is a short-sleeved knit polo shirt, and there are no pants in the episode.
How Argus performs across models
Argus is model agnostic, and can run on any sufficiently capable VLM. We ran Argus on GPT-6 Astra, Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol (all at medium reasoning), anchoring to Astra as a benchmark. Each model was given the same 193 episodes to annotate, keeping the prompt, frame, and output budget consistent. These episodes cover about one hour each of tele-op, UMI and human ego footage, and all answers parsed successfully.
Each share is over the 171 teleoperation and UMI episodes that all five models answered. A human ego clip holds several tasks, so it has no single outcome to compare.
The models usually converged on the same outcome. Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol agreed with Astra on 73% to 96% of the tele-op and UMI episodes. However, Astra labels much more densely than the first three models we compared, marking 44 events per minute of footage on these episodes, 1.9 times as many as the densest of them, GPT-6 Sol. The difference in density is largest on human ego video, where Astra marks 26 key events per episode while none of the three marks more than 15. GPT-6.1 Sol is the exception, at 43 events per minute and 28 key events per human ego episode.
In-context learning from Astra's traces does not close the gap. We added one fully annotated Astra example to the prompts of Claude Opus 5.5, GPT-6 Sol and DeepSeek v4.1 flash, then asked each one to label 64 of the episodes again. Outcome agreement shifted only marginally and annotation density increased by only 4 to 9%, at most reaching 59% of the 43.5 events Astra marked per minute on those episodes. We expect Argus-style pipelines to improve as models become increasingly accurate, fast, and cheap, until every episode of every dataset can be checked this way. GPT-6.1 Sol, the newest model we tried, already labels 96% as many events per minute as Astra on the same episodes and gives the same outcome on 95% of the tele-op and UMI ones, for $5 per hour of footage where Astra costs $27.
All footage64 episodesAstra labels 43.5 events per minute on these
- Without an Astra trace
- With one Astra trace in context
- Astra
- Claude Opus 5.517.8 → 19.496% → 96%3.9 → 2.8
- GPT-6 Sol24.7 → 25.777% → 82%3.7 → 2.4
- DeepSeek v4.1 flash16.2 → 17.074% → 70%2.8 → 2.6
* Teleoperation and UMI episodes only.
Each model labelled the same 193 episodes (186 minutes of teleoperation, UMI and human ego footage) through our harness, with identical prompts and frames. Each point counts the timeline events a model labels per minute of footage, which measures how finely it divides what happens, not whether each event is right, and the line follows the highest rate released so far.
Every annotation from this comparison can be found on the data dashboard, kept separately from Astra's: Claude Opus 5.5, GPT-6 Sol, DeepSeek v4.1 flash and GPT-6.1 Sol, and the charts that compare them all.
Example Results
01Success then undone
- Issue
- The demonstration is successful, but the recording continues past when the goal is achieved and shows the operator knocking the finished arrangement apart. As such, the last frame does not show the goal.
- Downstream Impact
- Goal-conditioned models and value functions that read the last frame as the goal learn the wrong goal, and successful demonstrations are miscategorized as failures. These episodes can be converted into clean success demonstrations by trimming at the reported goal frame.
- How often
- 10%Galaxea Open‑World23 of 222 episodes
- 1.9%MolmoAct2‑BimanualYAM25 of 1,284 episodes
- 1.8%10Kh‑RealOmin5 of 280 episodes


What to watchThe row is complete at 0:18. At 0:29, the grippers begin pushing the blocks apart, and the episode ends without the completed row.
Task given by the dataset“Push blocks to align them in a straight line.”
02Task instructions do not match footage
- Issue
- The task instructions describe a task different from what is shown. The magnitude of this problem ranges from low severity cases with minor deviations, to medium and high severity cases that could teach a world model incorrect dynamics. 20% of MolmoAct2 episodes were medium or high severity cases, and 177 of its 202 “failed” demonstrations carry such an instruction.
- Downstream Impact
- Language-conditioned policies learn incorrect associations between words and behaviors, and datasets look worse than they are as successful demonstrations are misclassified. This is fixed by replacing the original instruction with Astra's description of the episode.
- How often
- 79%OpenAoE‑2000h85 of 107 episodes (32 low)
- 34%Gen‑HumanEgo27 of 79 episodes (19 low)
- 26%Galaxea Open‑World57 of 222 episodes (13 low)
- 25%MolmoAct2‑BimanualYAM319 of 1,284 episodes (60 low)
- 13%FastUMI‑100K128 of 964 episodes (53 low)
- 13%10Kh‑RealOmin37 of 280 episodes (4 low)
- 8.2%ABC‑130k15 of 183 episodes (11 low)
- 7.3%HABIT23 of 315 episodes (14 low)
MolmoAct2, episode 17777 (the video below)
“Pick up each item, scan its barcode, and place it in the basket.”
The operator takes the four packages out of the basket, scans each one, and collects them on the right side of the table.
FastUMI-100K, Prepare_tableware (all 32 episodes)
“Put the chopsticks and spoon from the left-hand tray into the right-hand bowl.”
A fork and a spoon are moved from a plate into a rectangular wire basket. There are no chopsticks and no bowl.
What to watchFrom the first seconds the arms lift packages out of the basket and scan them onto the table, the opposite of the instruction.
Task given by the dataset“Pick up each item, scan its barcode, and place it in the basket.”
03Head camera taken off during recording
- Issue
- In ego data, the operator takes the head camera off during the recording, so part of the clip shows a sudden change in perspective and has the camera pointing at the operator's face or the ceiling.
- Downstream Impact
- The video should primarily be of hand movement; cutting the affected segment is sufficient to fix this.
- How often
- 12%Egocentric‑100K13 of 112 episodes (3 low)
What to watchEach clip is the three seconds around the moment the camera comes off. We blurred the faces ourselves.
04Recording plays faster than real time
- Issue
- The episode is sped up. This occurs when the control loop records at less than 30 Hz, but the episode is timestamped at 30 Hz. Timestamps, video timing and frame-rate metadata are all uniform, so the files do not reveal the problem. We found these sessions with a deterministic check: the delay between the follower and leader arms is fixed at about 125 ms, so if the delay spans too few frames, the session must be playing artificially quickly.
- Downstream Impact
- World models learn incorrect velocities and policies learn to move faster than they should. The compression factor in a sped-up episode can be measured by the frame lag between the follower and leader arms, so the session can be corrected to run at real time or dropped.
- How often
- 11%MolmoAct2‑BimanualYAM3,536 of all 32,246 episodes (294 of the 1,284 labelled)
What to watchThe motion looks brisk and slightly stuttered.
Task given by the dataset“Place dirty dishes in dishwasher rack, sort waste into bins, and clean table.”
05Camera files swapped
- Issue
- The left camera stream shows the right gripper's view and vice versa, but nothing in the files flags this. Cases can be identified with a deterministic check which compares each stream's motion change with each gripper's movement. In the example below, the left stream moves with the right gripper's recorded motion far more than with the left gripper (0.68 correlation versus 0.28), and the other gripper also sits on the wrong side of the frame. It's important to note that recorded motion alone cannot be trusted to find all these swaps, as the motion data itself is often flawed. For example, in all 14 of FastUMI‑100K's swapped episodes, the recorded gripper opening does not actually happen, and in 50 other episodes, the recorded pose remains still while the camera moves, or jumps while the camera does not. Argus instead uses Astra to detect the camera swap independently by examining the raw footage.
- Downstream Impact
- World models learn incorrect cause-and-effects: one gripper's motion moves the other gripper's view. Swapping the streams then fixes the mispairing.
- How often
- 1.5%FastUMI‑100K14 of 964 episodes
- 0.36%10Kh‑RealOmin1 of 280 episodes
What to watchAs we can see from the example below, the left camera stream changes when the right gripper moves (purple), not when the left one moves (blue).
Task given by the dataset“First align and fold the two sleeves of the T-shirt over the body, then fold the T-shirt vertically in half from top to bottom.”
06Privacy blurring hides task objects
- Issue
- Blurring intended to mask people's faces can mistakenly mask objects that the robot is handling.
- Downstream Impact
- Policies learn to grasp objects they cannot properly see. The frames can be masked as a partial remedy, but only the publisher is able to completely fix this issue by correctly masking the original footage.
- How often
- 3.8%HABIT12 of 315 episodes (11 low)
What to watchFrom 0:19 the donut in the left gripper's camera turns into a blurred rectangle. It is sharp again for a moment at 0:21, then blurred again until 0:23.
Task given by the dataset“Move the first donut from the left to the tray.”
07Task object leaves every camera
- Issue
- Part of the task happens out of view for all cameras.
- Downstream Impact
- Since the recorded actions in those timestamps have no visual evidence, world models are asked to predict frames without the context on what the arms are doing, and the model cannot verify the correctness of the end state. These episodes can still be used once the period of no visual activity is masked.
- How often
- 0.78%MolmoAct2‑BimanualYAM10 of 1,284 episodes (2 low)
- 0.45%Galaxea Open‑World1 of 222 episodes
What to watchAt 0:42, the red packet is placed at the far right edge of the table. From that point onwards, only a small section can be seen in the top view, and the right gripper's camera sees it only in passing, at 0:49 and 1:24.
Task given by the dataset“Pick up the red packet and scan its barcode.”
08Labelling operator mistakes as clean demonstrations
- Issue
- The recording completes the task, but there are mistakes in the operation, including missing grasps, dropping objects, and meandering paths. These episodes are then released as good demonstrations despite containing iterations of suboptimal actions.
- Downstream Impact
- Policies trained by imitation will copy the suboptimal approach to the task. To improve behavior cloning, these episodes should either be filtered out or trimmed to remove the mistakes. However, for world models, these operator mistakes can become useful data points for learning how to recover from failed attempts, and should be labeled as such and kept instead.
- How often
- 6.6%ABC‑130k12 of 183 episodes
- 5.5%MolmoAct2‑BimanualYAM71 of 1,284 episodes
- 3.6%Galaxea Open‑World8 of 222 episodes
- 2.1%10Kh‑RealOmin6 of 280 episodes
- 0.62%FastUMI‑100K6 of 964 episodes
What to watchAt 1:59, the left gripper attempts to pick up the scissors, but they repeatedly slip out of the grasp, and from 2:06 the right gripper tries too, repeating the grabbing motion for over twenty seconds.
Task given by the dataset“organize the medicine kit”
Summary of results
Of the 3,546 episodes we labelled across nine datasets, 27% have at least one issue that Argus flags at medium severity or above, either from the model or from a deterministic check. In these episodes, part or all of the footage is wrong or wasted for training as it was recorded.
MolmoAct2‑BimanualYAM1,284 labelled episodes25.1 hours of footage
- High
- The episode cannot be used as it is.
- Medium
- Part of it is wrong, and the rest is good once that part is trimmed or masked.
- Low
- Worth knowing, but a model trained on it would barely be affected.
Data issuesfaults in the recording, the scene or the label
- Instruction or labels do not match footage25%319 (60 low)
- Recording plays faster than real time11%3,536
- Success then undone1.9%250.88%
- Long idle stretch1.0%13 (8 low)0.48%
- Task object leaves every camera0.78%10 (2 low)0.02%
- Scale or display unreadable0.62%80.14%
- Person changes the scene0.31%4 (1 low)0.03%
- Recorded motion disagrees with video0.16%20.01%
- Camera image frozen or lost0.08%10.00%
Operator mistakesthe recording is faithful, but the demonstration went wrong
- Repeated attempts12%155 (149 low)
- Missed grasp8.8%113 (112 low)
- Wasted motion7.6%98 (98 low)
- Missed placement4.4%57 (57 low)
- Dropped object3.8%49 (16 low)
- Knocked object2.4%31 (12 low)
- Task left unfinished2.3%30 (12 low)
- Missed insertion0.70%9 (9 low)
- Undid its own result0.39%5 (4 low)
- Missed rotation0.39%5 (5 low)
- Collision0.23%3 (3 low)
- Long pause0.23%3 (3 low)
- Spilled contents0.08%1
- Misaligned placement0.08%1 (1 low)
- Selection error0.08%1 (1 low)
- Failed activation0.08%1 (1 low)
- Spillage0.08%1 (1 low)
The mix of errors varies significantly by dataset, and each error affects training in a specific way. Incorrect instructions teach a language-conditioned policy the wrong association and can make failures look like good demonstrations, an undone goal gives goal-conditioned models the wrong target, and a sped-up session teaches wrong dynamics. Most of these errors are easy to fix once found, but they are pernicious when left unnoticed.
Conclusion
Clean data is essential for effectively scaling model capabilities. As a research lab focused on training World Models, we benefit from as much high-quality data as possible. By building and open-sourcing Argus, we aim to improve the quality standard for public robotics data, to ensure that anyone can inspect these datasets with ease and use with confidence.
We are publishing all annotations on our data dashboard at pantheon.inc/data-board, including a JSON download for each episode and a JSON Lines export of any dataset and filter. You can run Argus on your own recordings with data.pantheon.inc/review, or from its open-source code on GitHub.
Credits
This project was undertaken by Eric Li. We thank XK Lu, Jenna Hong, and Janet Guo for their valued help with the writeup, and Jannik Schilling, Kunvar Thaman, Sambhav Gupta, and Calix Huang for suggestions and ideas.
If you are interested in a data-related collaboration or have any questions or comments, please contact us at data@pantheon.inc.
Dataset citations
- Allshire, Arthur, et al. “Scalable Behavior Cloning with Open Data, Training, and Evaluation.” arXiv, 2026. XDOF/ABC-130k, Apache-2.0
- Build AI. “Egocentric-100k.” Hugging Face, 2025. builddotai/Egocentric-100K, Apache-2.0
- Fang, Haoquan, et al. “MolmoAct2: Action Reasoning Models for Real-world Deployment.” arXiv, 2026. allenai/MolmoAct2-BimanualYAM-Dataset, Apache-2.0
- Galaxea Team. “Galaxea G0: Open-World Dataset and Dual-System VLA Model.” arXiv, 2025. OpenGalaxea/Galaxea-Open-World-Dataset, CC BY-NC-SA 4.0
- GenRobot. “10Kh-RealOmin-OpenData.” Hugging Face, 2025. genrobot2025/10Kh-RealOmin-OpenData, CC BY-SA 4.0
- GenRobot. “Gen-HumanEgo.” Hugging Face, 2026. genrobot2025/Gen-HumanEgo, CC BY-SA 4.0
- Li, Zishuo, et al. “Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning.” arXiv, 2026. inclusionAI/OpenAoE-2000h, Open-AoE Dataset License
- Liu, Kehui, et al. “FastUMI-100K: Advancing Data-driven Robotic Manipulation with a Large-scale UMI-style Dataset.” arXiv, 2025. IPEC-COMMUNITY/FastUMI_100k_lerobot, Apache-2.0
- Song, Jaehwi, et al. “HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation.” arXiv, 2026. configinc/HABIT, CC BY 4.0



















