Original experiment footage is shared with the Japanese edition. Some on-screen labels are in Japanese; methods, results and video descriptions are provided in English below.
After reducing G1’s gait jitter, we connected Visual SLAM to the walking configuration. G1 walks between shelves and cargo while camera images build a colored map. A separate, earlier short test also evaluated matching another run’s images to a saved map.
Follow-up: from Visual SLAM to Nav2 navigation. The results below describe the earlier mapping stage.
Walk between shelves and compare the map to the scene
Walking longer did not make the map easier to read
An earlier 88-second walk in front of shelves added points, but the route mostly saw floor and wall, making objects hard to recognize. We moved the start between shelves, followed them, turned at the far end and returned along the adjacent aisle.
We also changed the visualization from height-only coloring to left-camera image colors on points reconstructed from stereo disparity. Orange shelf beams, brown boxes and the far white wall can now be compared between video and map. The upper view’s bounds follow observed points. Ground-truth simulator shapes and positions are not used to construct the map.
This run reached all seven targets with 0 tracking losses across 1,757 images. The surface cloud contains 681,630 points; horizontal position RMSE is 6.29 cm and final stop-position error is 18.76 cm. These are localization errors against simulator truth, not measurements of shelf or box reconstruction accuracy.
70.30 seconds of simulation required 994.64 seconds of processing. This is not real-time validation. The route was predefined; automatic planning and obstacle avoidance were not implemented at this stage. Relocalization on this 70-second map was not tested. The following relocalization results come from a separate earlier run of about 20 seconds.
The first map and relocalization from another run
A trajectory alone did not show the spatial map
The first video showed image features and a motion trajectory, but not where shelves and walls were. We separated the visual-feature map used for localization from a surface map people can interpret. Stereo disparity supplies depth; cuVSLAM’s estimated pose aligns points. Floor and the right-hand wall are the main visible structures in the first test. Simulator ground truth is not fed into map generation.
Accumulate approximately 126,000 stereo points
Stereo images are 848 × 480 with a 50 mm baseline at 25 Hz. OpenCV StereoSGBM computes disparity and the surface map updates at 5 Hz. Inconsistent left/right matches are rejected and points from 0.4 to 10 m are accepted.
Points are grouped into 5 cm cells and retained after observations in at least three updates, yielding 125,935 saved points. The cell width is not a claim of 5 cm reconstruction accuracy. Distances between reconstructed surfaces and ground-truth geometry have not been evaluated.
| Metric | Result |
|---|---|
| Processed images / tracking losses | 518 / 0 |
| Errors at two target arrivals | 6.5 cm / 7.4 cm |
| Final stop-position error | 13.1 cm |
| Horizontal position RMSE | 6.24 cm |
| Saved surface points | 125,935 |
| cuVSLAM map graph | 14 nodes / 13 edges |
Position evaluation aligns only the initial camera pose, without whole-trajectory fitting or scale correction. 20.74 seconds of simulation took 138.85 seconds to process. Normal-speed video refers to simulation time, not real-time computation.
Locate another run in the saved map
A separate recording was tracked from its start in another process. The resulting visual positions supplied initial guesses for saved-map matching at 8, 10, 12, 14 and 16 seconds. All five locations matched, with horizontal RMSE of 5.09 cm. Shifting the guesses another 30 cm laterally also succeeded at all five locations, with RMSE of 5.08 cm. Ground truth was used only after matching for evaluation.
Matching uses cuVSLAM’s saved visual-feature map, not the displayed dense cloud. Conditions include the same starting region, processing of preceding images and a nearby pose guess. Five locations within one run are not five independent driving trials.
Verified capabilities and remaining work
We generated a map while walking, displayed position and reused a saved map with images from another run. Checks included decoding every video frame, reloading point-cloud files and recovering depth from known disparities.
Relocalization here is a replay test. At this stage, controlling a new run from a saved map, classifying traversability and avoiding obstacles were not implemented. Targets and aisles were known; this was not finished autonomy for arbitrary factory destinations. Physical verification remained future work.
Pose estimation, recognizable spatial structure and reuse of a saved map require separate checks. The subsequent navigation article develops traversability and map-based travel.
Environment and sources
Isaac Sim 5.1.0, Isaac Lab 2.3.0 and cuVSLAM v17, using the same matched Unitree model and public policy as the gait test. Measurements and videos are TeamZ’s September 6, 2026 test results.
NVIDIA PyCuVSLAM API · OpenCV StereoSGBM · Unitree model/policy revision used
LET’S BUILD TOGETHER
What are you working on?
Tell us about the software you need to implement or test. We can discuss the work, the development setup and how we could contribute.
Discuss your development needs