The new capture uses English on-screen labels. It is a separate simulation run; the original three-trial results are reported below.
Public code and pretrained weights do not necessarily support the task you want to run. TeamZ tested microwave opening with Flex-π, a model that generates robot actions and future information, in simulation.
The initial environment did not reach its success criterion. With a different released model and environment whose training tasks include opening, one of three trials reached the official criterion. This is a record of checking task, robot and evaluation compatibility—not a physical-robot test, a G1 door-opening demonstration or a newly trained TeamZ model.
Watch the door and arm from two viewpoints
83.2 seconds at normal simulation-time speed, excluding inference waits. Angle labels come from the synchronized log. The final 0.088 seconds are outside the recorded frames; the log and terminal snapshot confirm the final angle.
This new capture is not added to the original three-trial results below. The criterion is approximately 54°, not fully opening the door.
The task and the checkpoint need to match
The initial LIBERO environment implements microwave opening, but this task was absent from the 40-task fine-tuning list for the checkpoint we used. A compound task involving closing a microwave is included; closing is not opening.
We started an exploratory test without making this correspondence a firm selection requirement. For reproducing a known task, it should be checked before running the model. Absence from the fine-tuning list alone does not establish that the model never encountered the task during pretraining.
What happened in the two environments
| Setting | LIBERO | RoboTwin |
|---|---|---|
| Robot | Panda arm | ALOHA-AgileX dual arm |
| Simulator | MuJoCo | SAPIEN |
| Checkpoint correspondence | Opening absent from fine-tuning list | Opening included in training tasks |
| Inference | Action-only / full joint | Full joint, also generating future information |
| Success criterion | Opening beyond approximately 74.5° | At least 60% of joint range, approximately 54° |
| Observed result | 0/3 in each mode | 1/3 |
The largest opening angle in LIBERO was about 51.6°, below its criterion. These results do not establish task mismatch as the sole cause of failure.
RoboTwin's three trials reached maximum angles of 54.0°, 42.7° and 6.5°. One passed the official test and two did not. Success does not mean fully opening the door or turning a knob to release a latch.
View the original test recording (82.6 s, normal simulation speed)
The model, robot, object, simulator and success criterion all differ between environments. The change from 0/3 to 1/3 is not evidence of a performance improvement caused by changing the model. These small samples also cannot establish general success rates or an advantage from generating future information.
Four checks before adopting a public model
1. Does the checkpoint support your task? Check the model card, training-task list and evaluation configuration, not just whether the simulator implements the task. Different checkpoints under the same project name may cover different tasks.
2. Do observations and actions match the robot? Check the robot, hand, cameras and joint or end-effector action representation. A fixed-arm model does not directly transfer to G1 with five-finger hands and walking. We did not test that transfer.
3. What does success mean? Here RoboTwin declares success at approximately 54°. Read what the evaluator measures rather than judging the video alone. Object mass and joint resistance are benchmark settings, not parameters calibrated against a real microwave.
4. Is video speed being confused with real-time performance? The recording follows simulation time and excludes model inference waits. Its playback speed is not demonstrated physical-robot throughput.
What TeamZ implemented
We prepared the public model and runtime, repaired dependency and inference issues, checked configurations and weights, recorded opening angles, retrieved videos and logs, and independently checked success conditions. All three RoboTwin videos were decoded to their ends and checked against the logs.
We verified an example of a released model running a supported simulation task. We did not establish reliable opening, physical-robot operation or applicability to G1.
TeamZ can join development teams to implement public-model environments, reproduction tests and data preparation. Tell us which model and robot you are considering and where implementation is getting stuck. Discuss implementation and reproduction testing.
Sources
Flex-π project · Official code · RoboTwin checkpoint · RoboTwin opening task
LET’S BUILD TOGETHER
Turn your robotics ideas into questions we can test.
Tell us about your task, equipment and current challenges. Through simulation and tests on real robots, we explore feasibility and technical challenges together.
Discuss joint development