WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation

Samuel Zhen · Siwon Jo · Yanze Zhang · Wenhao Luo

Overview

Learned manipulation policies can complete complex tasks while still making unsafe motions. Many safety methods focus only on the end effector, so collisions involving the rest of the robot or an object it is carrying can go unprotected.

WBAG addresses this by modeling the full robot body and, after a grasp, adding the held object to the geometry that the safety controller protects.

Project overview. A three-minute introduction to WBAG and its motivation, framework, and results.

How It Works

WBAG maintains a safety model of the robot throughout the task. Before grasping, it protects the articulated robot body. After a grasp is detected, the geometry of the held object is added as well, so the protected system changes with what the robot is carrying.

This geometry is used to build differentiable control barrier function constraints. These constraints adjust the six-dimensional operational-space action produced by the vision-language-action policy only when needed to avoid unsafe motion.

Because the intervention happens after the policy produces an action, WBAG can be applied at inference time without retraining or modifying the policy.

WBAG pipeline: two RGB-D views and a frozen VLA policy feed perception-derived geometry, a whole-body and attached-geometry safety layer, and a safety-adjusted action.
Framework overview. RGB-D perception and precomputed robot geometry define whole-body safety constraints that expand to include the grasped object after attachment. A CBF-QP enforces these constraints by minimally modifying the pretrained VLA's operational-space action.

Results

We evaluate WBAG on SafeLIBERO, using 16 manipulation tasks drawn from four LIBERO suites, two obstacle configurations per task, and 50 randomized episodes per configuration, for 1,600 episodes total. All methods use the same frozen π0.5 policy and identical episode initializations.

We report three metrics: Scene Safety, whether the robot completes an episode without disturbing any evaluated non-task object; Safe Success, task completion without a safety violation; and Unsafe Success, task completion with a safety violation.

What we compare

  • Policy only runs the pretrained VLA without a safety filter.
  • AEGIS baseline protects an end-effector ellipsoid against a designated obstacle.
  • SAM3 EEF-MVEE changes the perception model while retaining the same end-effector safety representation.
  • SAM3 Scene EEF-MVEE expands obstacle coverage to the full eligible scene while retaining end-effector-only protection.
  • WBAG EEF + attached adds explicit attached-object geometry but still protects only the end effector.
  • WBAG w/o attached protects the full articulated robot but omits the grasped object.
  • WBAG combines scene-wide geometry, whole-body robot protection, and attached-object protection.
Metric Policy only AEGIS SAM3 EEF-MVEE SAM3 Scene EEF-MVEE WBAG EEF + attached WBAG w/o attached WBAG
Scene Safety 21.25% 70.87% 72.94% 82.75% 68.40% 77.50% 97.38%
Safe Success 19.00% 51.06% 55.56% 46.25% 46.96% 46.31% 59.38%
Unsafe Success 39.56% 11.25% 8.13% 4.31% 17.58% 14.69% 0.68%

These are the aggregate SafeLIBERO results across both obstacle levels.

Expanding scene coverage already improves safety: moving from SAM3 EEF-MVEE to SAM3 Scene EEF-MVEE raises Scene Safety from 72.94% to 82.75%, but Safe Success falls from 55.56% to 46.25%. Full WBAG reaches 97.38% Scene Safety while also increasing Safe Success to 59.38% and reducing Unsafe Success to 0.68%.

The ablations show that whole-body and attached-object protection contribute separately. Removing whole-body protection or omitting the attached object both reduce aggregate safety and Safe Success, supporting the need to model both the articulated robot and what it carries.

Acknowledgements

This work was supported in part by the U.S. National Science Foundation under
Grant 2530297.

Citation

@misc{zhen2026wbag,
  title  = {WBAG: A Whole-Body and Attached-Geometry Safety Framework for Vision-Language-Action Manipulation},
  author = {Samuel Zhen and Siwon Jo and Yanze Zhang and Wenhao Luo},
  year   = {2026},
  eprint = {2610.01083},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO}
}