We capture expert human demonstrations in physical environments, build RL environments from real workflows, and evaluate AI agents with domain experts.
Every dataset starts with a real person doing a real task in a real place. We structure what they did into training data models can learn from.
Egocentric and third-person capture of skilled workers performing manipulation tasks in homes, kitchens, warehouses and light assembly lines. Bimanual teleoperation episodes with synchronized stereo video, IMU and action logs. Dense action-level annotation optimized for VLA training.
Environments that mirror how work actually happens: e-commerce operations, customer support, logistics dispatch, manufacturing SOPs. Domain experts generate trajectories, preference data and rubrics inside the environment, so agentic models learn the task, not the benchmark.
Evaluation design, failure diagnosis and continuous monitoring by people who do the job the agent is replacing. We find where the agent fails, produce targeted training data for the top failure modes, and re-score as models and prompts change.
A capture network of real workplaces and a controlled studio, run as one pipeline with quality measured at every step.
We define the task family, environment, embodiment and success criteria with your research team.
Our AI recruiter screens skilled workers for the exact task. Experts, not generalist annotators.
Head-mounted and third-person rigs, or teleop in our studio. Segmentation, labels and QC follow the same day.
De-identified data lands in your US cloud bucket with a live pipeline dashboard. Next batch adapts to what your model got wrong.
We publish open benchmarks on real-world tasks so the field can measure what matters: whether a policy works outside the lab.
Open-source VLA policies evaluated on household tasks captured in real kitchens and laundry rooms, not simulation.
| Tasks | Folding, loading, sorting, wiping |
| Policies under test | π0, OpenVLA, GR00T, Octo |
| Metric | Success rate, 3 trials per task |
Bimanual assembly and packaging tasks from real production lines, scored by the line workers who trained on them.
Join as a partner lab →Computer-use agents on real e-commerce seller operations: listing edits, ad bids, returns and review handling.
Join as a partner lab →Three questions. We reply within one business day with a sample set and a quote.