In just eight weeks and with a core operations team of 5, we scaled an international data collection operation to a team of 90 operators. To support it, we built our own custom in-house Universal Manipulation Interface (UMI) hardware, multiple colocated storage and compute clusters, custom reverse-engineered camera firmware, and dedicated internet infrastructure. We also took several ambitious, nonconsensus bets on data distribution, and have collected what is to the best of our knowledge the most diverse UMI dataset in the world.
We have since produced over 1 million unique tasks at an average cost of $10 per hour of data, all while paying our operators three times the local living wage. In doing so, we have validated a collection process we can scale to millions of hours of UMI data. To push the frontier of dexterous manipulation forward, we are also releasing a 100-hour sample of our dataset, fully annotated with dense task labels from Argus, the most advanced open-source annotation and data quality pipeline in robotics.
What does good data look like?
Before UMI, robotics data was broadly either egocentric data (humans performing tasks with a head cam) or teleoperated data. UMI bridged the gap between the ease of ego data collection and the embodiment similarity of teleoperation. It enabled collection of 1 DoF gripper actions, end-effector deltas, and wrist cam views while leaving open an avenue to scaling past millions of hours of data collection.
We are betting big on UMI. As a data collection method, we believe UMI is the ideal point on the Pareto frontier between scalability and retaining embodiment transfer. However, many existing UMI datasets suffer from imitation learning-derived priors. Namely, most such datasets are generated as repetitive, task-sparse successful iterations of a very small set of tasks, with little consideration to diversity. This data is suitable for imitation learning on a single task at a time, but fails to model the real-world state distribution and is insufficient for our goal of true zero-shot task generalization in world models.
It is also prohibitively expensive to buy petabytes of UMI data from external suppliers. Existing UMI datasets cost around $60/hr, and custom datasets raise the cost to as much as $150/hr. These rates are for task-sparse data; suppliers generally don’t operationally support diverse, task-dense data, so the type of data we need from suppliers is unobtainable at any cost.
Under these conditions, it was clear to us very early on that we’d need our own in-house UMI data collection operation where we could redesign every aspect of collection from the ground up.
Collection approach
We knew we needed task-diverse and environment-diverse data, but we also knew this could be operationally challenging. In existing operations, the pattern is usually:
- Start recording an episode
- Perform a small task, usually ~20 seconds
- End the episode
- Reset the task to the initial state
Every step outside of performing the task adds operational overhead that significantly reduces the throughput of an operator. Worse, the structure of this collection style can only support VLA-style imitation learning, since the overhead of switching to a different task and recording it in the hardware is too high. To justify the overhead of switching tasks, each task would have to be repeated at least ten times in a row.
Freeform collection
Here, we made a big bet. Six months ago, frontier VLMs were not good enough at labeling video data. For any semantic-level data annotation work, like task verification, subtask annotation, or goal frame annotation, the only choice was human operators through expensive data labeling contracts with large commitments. However, we had high confidence that this was a temporary constraint and that VLMs would soon perform this work accurately and cost-effectively. Therefore, we decided to do the following:
- Introduce the collection of freeform data. In freeform data, an operator is given a set of 15 or so items scattered around the table and is asked to perform any random tasks on the objects with no guidance. In its raw form, this data cannot easily be used for training planners since there are no task labels or clear goal states. It also can't be used for any form of text-conditioned policy training.
- Wait for VLMs to reach our standards of data annotation and QA, and only use the data once this becomes the case. At the time we designed the operation, we knew the data we produced wouldn't immediately be useful since it wouldn't have clear task annotations, but expected VLMs to eventually catch up. That was sufficient if it meant we could collect at much higher throughput. We took on the risk that VLMs would never reach or surpass human level, and the data we collected in this scheme would be unusable.
It turned out that we only had to wait five months between our first VLM experiments and the point when Argus became good enough at annotation to fully serve both our annotation and QA needs, making all of our collected data usable.
Activity in this excerpt“Move and replace tiles on a shape-sorting board.”
Activity in this excerpt“Assemble and rearrange a multicolored brick structure.”
Freeform collection takes us a long way towards our goals of task-diverse and environment-diverse data collection, but it has its own distribution issues. Most actions in freeform episodes fall into a narrow set of grasp, pick, move, and place operations. This is not categorically an issue; general zero-shot pick and place is far from solved, and the vast majority of manipulation in the real world is mostly pick and place. However, it cuts off the long tail of much more dexterous manipulation, like fine-grained insertion tasks, which can end up underrepresented in a freeform dataset. To get the best of both worlds, we also overhauled the standard single-task loop to complement freeform collection.
Scripted collection
Supporting scripted collection required a dedicated tooling layer. We built Nomos1, a tooling platform that plays three important roles:1The ancient Greek daimon (not daemon) of statute and ordinance
- Generating diverse tasks. Nomos allows us to achieve remarkably high levels of task diversity. It tracks a growing catalog of physical objects alongside a dictionary of manipulation primitives, enabling us to generate millions of distinct tasks from combinations of objects and modifiers. We also designed it to never mint physically impossible tasks. To date, we have collected data for tasks involving nearly every possible pair of objects in our catalog.
- Categorizing manipulation. Nomos categorizes an object according to how a human hand would manipulate it, which enables accurate reasoning of coverage with respect to object types and their associated manipulation priors. This helps us separate coarse- and fine-grained manipulation, giving us a distributional knob to tune as tasks are served.
- Adapting collection to research needs. Nomos also allows us to change the distribution of the data we collect at a moment’s notice. If, for instance, we require greater representation of tool use, a different class of manipulation primitives altogether, or greater coverage of a particular object type, we can implement the request across our operation within minutes. This has been indispensable because our understanding of what data distributions are valuable has changed substantially as our model checkpoints have improved.
Task given by the dataset“Pick up the fake lemon, transfer it to the right gripper, and place it just to the right of the 10 of clubs.”
Failure data
We noticed early in our world model experiments that expert trajectories, especially when thoroughly cleaned of all possible failures and mis-grasps, are poor proxies for the real-world state distribution. Real robots need to gracefully recover when a mistake happens, but expert data is devoid of recovery paths from bad states to good states, so policies (both VLA and world model based) struggle. World models can naturally ingest and learn from failure data, since actions are provided by a separate planning layer like RP-1. Thus, we collect episodes that are full of missed grasps, premature drops, and general task failures and recoveries. Argus can segment out the failures and successes so that our models can learn how the world evolves during failure and how to recover from failure, but not to take actions that result in failure.
Activity in this excerpt“Recover dropped rings and return them to the peg rack.”
In-the-wild data
As useful as it is to have a centralized office space for data collection, it limits scalability of the operation. It also limits the environment distribution those table-mounted arms would see, which is reasonable but not sufficient long-term. We build our hardware to enable fully untethered, in-the-wild freeform data collection in any environment. Argus makes immediate labeling unnecessary, so operators can just perform arbitrary tasks with the grippers anywhere they want and produce usable data.
Activity in this excerpt“Color a character sheet with an orange-red pencil.”
Hardware
By running our own operation, we retain full control of the UMI gripper itself and can freely optimize it for our needs. To iterate on our hardware as fast as possible, we designed it to be fully modular. Our cameras are off-the-shelf and can be easily swapped if they fail, unlike many suppliers’ fully integrated UMI gripper solutions. This also allows us to update the grippers across our entire operation without replacing existing cameras.
Throughout the course of scaling operations, we deployed three iterations of grippers:
- Our first iteration of gripper hardware was designed quickly to get UMI grippers in hands ASAP and unblock iteration on downstream parts of the pipeline, such as firmware and data processing. It was rugged and functional, but too mechanically complex. It had six separate major components and 15 screws, and had unrefined finger pivots that used threaded screws as axles. The silicone gripper surface was added by a complicated casting and setting process that required significant effort and had a high failure rate. It was also uncomfortable to use. We produced six pairs before updating the design.
- For the second iteration, we did a complete overhaul of the structure. We reduced the main assembly to the housing and two gripper fingers, each of which used two press-fit bearings spaced apart to better support cantilever loads, and used metal dowels instead of screws for axles. We replaced the silicone with a 3D-printable TPU part. The resulting mechanism ran more smoothly, had substantially less play, and cut final assembly time to just 30 seconds. We produced 30 pairs before iterating further.
- For our most recent iteration, we further optimized for both mass production and capability in deployment. We redesigned the gripper to print finished parts with minimal assembly steps. For this gripper, we verified that it was capable of performing dexterous tasks in actuated teleoperation setups. This enabled us to achieve strong embodiment matching between UMI and deployment. Once the design was verified, we produced 60 pairs.
For our third iteration, we optimized production costs down to $11 per pair2. External production quotes hovered at $80–110 per pair, and required large purchase commitments and slow iteration speeds. By keeping production in house, we retain full embodiment control and also achieve otherwise unattainable unit economics.2Though this includes only materials and not labor, we also designed our grippers to snap together and assemble in a matter of seconds, so at scale the labor cost is minimal.
We iterated on the camera module completely separately from the gripper. We started by choosing an off-the-shelf camera with sufficient hardware capabilities. Though it had the right hardware, the stock software was designed for consumers and not large-scale robotics data collection. We require multi-camera synchronization, strong metadata attachment like task descriptions, the ability to display custom messages and buttons to operators, and a long tail of other software-level needs. To achieve this, we wrote a sophisticated firmware layer for our cameras specifically tailored for data collection.
To write our firmware, we decompiled and redesigned the native firmware binaries to support a system that allowed operators to record with n cameras simultaneously, perform a unique, tagged task for every episode, automatically discard suboptimal takes immediately, and collect rich task metadata. By running this software entirely on these integrated cameras, we completely avoided adding extra external sensors or compute.
By separating the camera stack from the end-effector and keeping the manufacturing process flexible, we can introduce new embodiments quickly, without rebuilding the collection system around them. As our robotic platforms evolve, the same infrastructure can support multiple end-effectors in parallel, giving us control over not just the distribution of tasks and environments in our data, but the distribution of embodiments used to interact with them.
Infrastructure
Our data operation creates a lot of data, and we needed an efficient way to manage it. This required setting up a significant amount of dedicated hardware infrastructure to properly ingest data.
We produce over 2 petabytes of UMI footage per month at current rates. Even storing 20 petabytes of data on standard hyperscale object storage would cost us roughly $5.2M/year, so it was clear that we needed a storage solution under our own management. Thus, in parallel to scaling our UMI operation, we built Hades3, a 20-petabyte (and growing) bare-metal storage cluster in a datacenter near our SF office, where we rent rack space and manage the storage ourselves. This arrangement cuts our data storage bill from over $5.2M/year to roughly $100,000 plus a fixed hardware investment that paid itself back in two months.3The underworld; stores the souls of our data
Hades solved our capacity problem, but we still needed a way to send data from our operation to our storage. Initially, we couldn’t ingest data fast enough due to internet bandwidth, so we needed a second storage cluster colocated with our operation. Styx4 is our second storage cluster with roughly a petabyte of local capacity and removes the operational bottleneck of local disk space.4The goddess of the river that flows into Hades
To further reduce bandwidth usage and decrease the latency between ingestion and our research team’s ability to use the data, we built Pallas5, a gaming PC with a couple of RTX 5080s. Although we evaluated more conventional compute options, we realized that a Minecraft-ready consumer GPU box gave us an absurd amount of throughput per dollar. Pallas decodes and reencodes footage at lower resolution at 1.9 times real-time throughput. Combined with private DIA, we drain up to 90 TB per day to rapidly support our training.5The husband of Styx
Data
We care deeply about three pillars of operational quality: cost, throughput, and distribution. Compared to other solutions like procurement from data vendors, our results give us the optimal tradeoff on these metrics.
Cost. A complete bimanual rig, including three cameras, two grippers, and mounting hardware, costs us approximately $650 per operator. Despite including an additional exocentric camera, our kit is less than half the cost of commercial rigs6. We have a clear path to reducing cost to as low as $100 per kit by further optimizing our choice of commoditized camera hardware to our exact needs and scaling up supply contracts. We also expect our future versions to be lighter and easier to use for extended periods. Coupled with advancements in throughput, our costs have dropped to just $10 per recorded hour.6Though not directly comparable in features, these rigs generally have extraneous features we don’t need, like binocular vision.
Kit cost, USD
Cost per recorded hour, USD
Throughput.In eight weeks, we collected over one million unique recorded tasks. We saw a dramatic increase in throughput thanks to rapid iterations of firmware, collection formats, our ingest pipeline, and operator experience. Today, we collect over 16,000 unique tasks per day at only 45 operators per shift. In particular, a large portion of this throughput is enabled by removing camera starts/stops and scene resets from our data collection, so very little time is lost to overhead. At our current pace, we're on track for 12 million unique tasks in 2027, and if we linearly extrapolate our scaling, as many as 25 million tasks.
Distribution. Thanks to freeform collection and Argus, the data distribution of our tasks and environments is extremely diverse. Our current data collection spans over 1 million unique tasks, themselves containing multiple subtasks and annotations each. We also collect a wide variety of world states through in-the-wild collection and annotated failure data through adversarial freeform. This broadens our coverage beyond the successful, task-scripted demonstrations common in UMI datasets.
Scaling
With the onset of in-the-wild freeform collection, our operation is immensely scalable. We can employ new operators with minimal operational overhead, cheaply and efficiently manufacture UMI kits, scale ingestion and processing with commodity hardware, and update desired data distributions with scripted tasks distributed directly to operators. We have a clear path to an app-based system of UMI data collection that ingests data from thousands or tens of thousands of unique operators around the world on our custom hardware.
We took a lot of nonconsensus bets to get our operation to this point, which we derived by both extrapolating trends and reasoning about our data needs from first principles. This enabled us to build this operation from scratch, designed for our needs and on a clear trajectory to generate tens of millions of hours of diverse data.
At Pantheon, we're working to create general purpose robotics foundation models, and we believe good models require good data. If you're training robotics models and you want to talk data, reach out to us at data@pantheon.inc.
Acknowledgements
Thanks to Humaid, Victor, Mo, Lilian, Daniel, Sixtus, XK, Sambhav, and Kunvar.
Works cited
- Chi, Cheng, et al. “Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots.” Robotics: Science and Systems, 2024.
- FastUMI-100K. Dataset card and released metadata. Hugging Face.
- HiFi-UMI-2K. Dataset card and released metadata. Hugging Face.
- Hy-UMI public release. Dataset card and released metadata. Hugging Face.
- Li, Eric. “Argus: An Open-Source Annotator for Robotics Data.” Pantheon Research, 2026.
- Pantheon. “Introducing Reinforced Planning (RP-1).” Pantheon Research, 2026.
- RealOmni. Dataset card and released metadata. Hugging Face.
- YUBI UMI Arena. Dataset card and released metadata. Hugging Face.













