Skip to content
AnalysisAI

The Robot Finished the Job. Who Was in Control?

Robots are moving into clothing returns. The people who handle their awkward moments reveal a harder question about the data used to train them.

NIST computer scientist Omar Aboul-Enein beside a silver and blue robotic arm and measurement apparatus in a laboratory.
NIST computer scientist Omar Aboul-Enein beside a mobile-robot test setup in an archival photograph published in 2022. The work measures how robots perform tasks as conditions change.R. Wilson / NIST · NIST public-domain work; worldwide reuse grant

At warehouses in Greven, Germany, and Świebodzin, Poland, robots have begun handling clothes sent back by Zalando customers. The two-armed machines identify, grasp and sort garments they have not been individually trained to recognize, according to an October 7 announcement by the retailer, the logistics company CEVA and the robotics company Sereact.

“Returns are the hardest manipulation problem in fashion logistics,” Ralf Gulde, Sereact’s chief executive and co-founder, said in the announcement.

People are still assigned work: supervising stations, dealing with exceptions and checking quality. The announcement does not disclose how often the robots need help, or their training records. That leaves a distinction unresolved: finishing the work and doing it independently are different achievements.

A warehouse has good reason to care about the first. A researcher trying to teach a machine has good reason to care about both. The difference becomes consequential when records of everyday work become examples for the next generation of robots. A person’s intervention can be an expense to reduce, a sensible precaution or a lesson worth saving. Sometimes it is all three.

An unusually explicit record of this handover appears in BodenAI’s public human-intervention dataset. It contains about 82 hours of household tasks performed by an R1Lite robot. In the publisher’s September 14 statistics, 3,068 of 3,347 recorded episodes included human intervention. A label attached to each video frame identifies whether a person or the robot’s control model was in charge.

That sounds like a great deal of help. But the same documentation reports that human control occupied 11.41 percent of the recorded time. Frequent intervention and continuous dependence are not the same thing. Nor is this a representative sample of all robot work: the collection specifically concerns intervention.

BodenAI also lists the limits. There are no episode-level success labels or structured reasons explaining why someone intervened. Some recordings end while the person still has control. The figures describe this robot’s model on these tasks, not the intervention rate of commercial robots in general.

A cautious operator could interrupt a maneuver that another would allow to continue. Someone under pressure to meet a production target might intervene rather than wait. Counting handovers is possible. Treating them all as identical failures would discard the human judgment behind them.

Imagine a packing station where a robot drops a soft bag across the edge of a container. A worker moves it onto a flat surface. The robot picks it up successfully on the next attempt. This is a hypothetical example, not an account from the Zalando sites. The final inventory record could be entirely accurate: the bag reached its destination. A video beginning after the worker’s adjustment could also show an entirely real successful pick.

Neither would explain how the machine got out of trouble.

The worker may have recognized that pulling harder would tear the bag. Or the adjustment may have been a shortcut to keep the line moving, even though the robot might eventually have managed. These possibilities imply different lessons. One teaches a physical limit; the other reflects a decision about how much time the business can afford to spend. A movement captured on camera does not necessarily reveal the judgment behind it.

There is a temptation to treat any human assistance as evidence that a robot has failed to live up to its billing. That can be too simple. A machine that handles routine work and occasionally calls for help could be useful long before it becomes capable of handling every case. Requiring complete independence might mean rejecting a system that already reduces a difficult job’s physical burden. The relevant comparison is with the work that people would otherwise have to do, including the new burden of supervising the machine.

Yet the opposite temptation is just as misleading. If assistance disappears into an overall completion rate, a customer cannot tell whether it is buying a mostly independent machine or a machine that keeps another person busy. The person has not disappeared from the cost of the work simply because the finished item is credited to a robot.

Researchers are also finding ways to teach recovery without waiting for a worker to rescue a real machine. In Workhorse, a paper submitted on October 6, Songbo Hu, Qiayuan Liao and their colleagues at the University of California, Berkeley trained a humanoid using human demonstrations. Five wearable trackers and a chest camera captured the demonstrator’s movements. One part of the robot’s system planned body movements; another carried them out.

The researchers altered their training examples to imitate mistakes those components could introduce when working together. In the project’s simulated box-sorting test, the full system succeeded in 77 percent of episodes without pushes. Removing that adjustment from the planner’s training reduced success to 20 percent. Videos separately show a real Unitree G1 recovering after a person pushes it or takes a box away. The numerical comparison comes from simulation, not a warehouse trial.

Deliberately producing mishaps on a working line would be a peculiar way to improve it. Research settings offer room for mistakes a customer would refuse to tolerate. But the trouble invented for training still has to resemble the trouble a machine will encounter.

On October 6, TwelveLabs released Pegasus 1.6, with support for video filmed from a person’s or remotely operated machine’s viewpoint. The company says the model can label actions, describe interactions and find unusual events in large video collections.

Jae Lee, its chief executive and co-founder, said in the company’s announcement that robotics teams could “train on real human experience instead of starting from scratch.” His examples included changing a grip and recovering after something slips.

That is an appealing prospect for anyone facing weeks of footage. Finding a possible failed grasp could save a specialist considerable viewing time. But a description of an arm changing direction cannot necessarily establish who issued the command. The camera might show the movement clearly while missing the remote operator entirely. A control log and a video answer different questions, even when the caption sounds confident.

Clocks still have to agree, and an action needs to be connected to the right task. A reviewer may have to determine whether the machine resumed or someone finished the job. These details separate a usable correction from an interesting clip.

Robotics has been wrestling with the gap between an impressive movement and reliable work for much longer than this week’s announcements. In a 2022 account, the National Institute of Standards and Technology described researcher Omar Aboul-Enein and colleagues testing arms mounted on mobile robots. Their measurements included positioning, task time and performance under different conditions. Moving to a new place was part of the test, not an inconvenience to remove before it began.

A more recent effort, RoboRecover, submitted on September 24, evaluates robots from situations reached partway through a task rather than only from preset starting positions. Its 2,000 simulated scenarios test what happens after execution has altered the scene. The authors report that strong performance from the original starting position does not determine how well a system recovers. Getting off to a good start and finding a way back are distinct abilities.

The distinction is sharper than a choice between real and artificial data. A real demonstration can begin from an unusually tidy setup. A simulated exercise can deliberately begin from an awkward one. Neither label, by itself, says how much the example will help a robot that has already made a mess.

The commercial question is less tidy than either score. A recovery that takes two minutes might be a technical success and an operational nuisance. Calling a person for five seconds might be cheaper. Conversely, a machine that repeatedly demands attention at unpredictable moments could be harder to staff than one whose work is slower but dependable. The total labor depends on when help is needed as well as how much.

That leaves a complicated role for the people watching the machines. Their adjustments can keep the business running today while supplying examples intended to reduce the need for those same adjustments tomorrow. The useful knowledge may be a modest physical act, such as loosening a fold, together with an explanation that never appears in the robot’s command history. Recording only the act risks losing the expertise.

A returned garment reaching the correct container is a useful result. It is also the end of a story. For a machine still learning the work, the consequential part may have happened several seconds earlier, when someone decided it was time to step in.

All Insights

More from SnowRock

Anthropic co-founder Dario Amodei gesturing during a discussion at TechCrunch Disrupt
Strategy

AI Is Getting Cheaper. Finished Work Is the Real Price.