TL;DR
SAFE-Pruner is a training-free, plug-and-play visual token pruning framework. It forecasts deep-layer token saliency from historical frames, achieving up to 1.89× speedup with less than 1.5% success-rate degradation.
Abstract
Real-time inference of vision–language–action (VLA) models is essential for robotic control. While visual token pruning has shown strong potential for accelerating inference, most existing methods mainly base pruning decisions on shallow-layer cues and risk discarding visual information required by deep layers. To address this issue, we propose SAFE-Pruner, a plug-and-play pruning framework that incorporates attention cues of future layers into pruning decisions. Specifically, we identify semantic attention consistency, the tendency that VLA models concentrate their attention probability mass on the same semantic entity across control timesteps. Based on this observation, we design a forward-looking strategy to forecast token saliency in deep layers, preventing the premature removal of critical tokens and enabling more stable acceleration. We further introduce a reference timestep refresh strategy that triggers updates upon attention shifts, improving forecasting accuracy and pruning reliability. Extensive experiments across diverse evaluation settings demonstrate up to 1.89× speedup with less than 1.5% degradation in success rate, while outperforming state-of-the-art methods by up to 1.9%.
Method
Experiments
Simulation Experiments
We evaluate SAFE-Pruner on LIBERO and SIMPLER across OpenVLA, OpenVLA-OFT, CogACT, and π₀.₅. The tables report task success together with model efficiency.
LIBERO
| Model | Method | Success Rate (%) ↑ | FLOPs (T) ↓ | Latency (ms) ↓ | ||||
|---|---|---|---|---|---|---|---|---|
| Spatial | Object | Goal | Long | Average | ||||
| OpenVLA | Vanilla | 84.6 | 86.6 | 78.2 | 53.2 | 75.7 | 1.862 | 48.24 |
| FastV | 81.8 | 70.6 | 71.8 | 44.6 | 67.2 | 0.811 | 34.63 | |
| SparseVLM | 80.6 | 72.4 | 68.2 | 46.6 | 67.0 | 0.876 | 34.32 | |
| DivPrune | 78.6 | 62.8 | 66.2 | 43.2 | 62.7 | 0.832 | 37.18 | |
| VLA-Cache | 78.2 | 71.2 | 70.6 | 45.8 | 66.5 | 0.941 | 35.21 | |
| VLA-Pruner | 82.0 | 84.4 | 77.8 | 52.6 | 74.2 | 0.793 | 36.47 | |
| SAFE-Pruner | 82.2 | 84.8 | 77.6 | 52.0 | 74.2 | 0.742 | 32.18 | |
| OpenVLA-OFT | Vanilla | 98.4 | 98.2 | 96.4 | 94.2 | 96.8 | 3.970 | 70.25 |
| FastV | 97.8 | 92.8 | 95.6 | 91.8 | 94.5 | 2.141 | 50.24 | |
| SparseVLM | 98.6 | 91.8 | 94.8 | 92.0 | 94.3 | 2.289 | 52.15 | |
| DivPrune | 97.2 | 90.0 | 93.4 | 89.6 | 92.6 | 2.126 | 53.87 | |
| VLA-Cache | 97.0 | 92.2 | 94.6 | 91.0 | 93.7 | 2.730 | 58.14 | |
| VLA-Pruner | 97.4 | 93.0 | 94.0 | 92.0 | 94.1 | 2.234 | 54.53 | |
| SAFE-Pruner | 98.0 | 98.0 | 96.2 | 93.4 | 96.4 | 1.722 | 37.22 | |
| π₀.₅ | Vanilla | 98.2 | 98.0 | 98.0 | 91.8 | 96.5 | 2.115 | 35.28 |
| FastV | 96.4 | 96.8 | 95.6 | 89.0 | 94.5 | 1.612 | 28.64 | |
| SparseVLM | 94.6 | 97.2 | 94.8 | 89.2 | 94.0 | 1.713 | 29.03 | |
| DivPrune | 93.8 | 96.6 | 92.4 | 88.0 | 92.7 | 1.576 | 30.55 | |
| VLA-Cache | 94.8 | 96.0 | 96.4 | 88.2 | 93.9 | 1.632 | 28.14 | |
| VLA-Pruner | 95.8 | 97.2 | 95.6 | 88.0 | 94.2 | 1.583 | 32.64 | |
| SAFE-Pruner | 97.0 | 98.8 | 96.0 | 90.2 | 95.5 | 1.482 | 24.33 | |
Results across OpenVLA, OpenVLA-OFT, and π₀.₅. SAFE-Pruner consistently delivers the strongest accuracy–efficiency trade-off across different VLA architectures.
SIMPLER · Visual Matching & Variant Aggregation
| Method | Visual Matching | Variant Aggregation | ||||
|---|---|---|---|---|---|---|
| Success (%) ↑ | FLOPs (%) ↓ | Speedup ↑ | Success (%) ↑ | FLOPs (%) ↓ | Speedup ↑ | |
| CogACT | 74.8 | 100.0 | 1.00× | 61.3 | 100.0 | 1.00× |
| FastV | 74.1 | 42.0 | 1.21× | 62.1 | 42.0 | 1.19× |
| VLA-Cache | 74.4 | 80.1 | 1.38× | 62.3 | 82.6 | 1.37× |
| SAFE-Pruner | 74.5 | 37.4 | 1.73× | 61.9 | 36.2 | 1.67× |
SAFE-Pruner achieves the strongest acceleration in both SIMPLER settings while maintaining success rates comparable to the unpruned CogACT baseline.
Real-World Experiments
Real-world evaluations use the Astribot S1 dual-arm platform and task-specific π₀.₅ policies.
| Method | Pick & Place | Throw Basketball | Pack Doll | Average ↑ | FLOPs (T) ↓ | Latency (ms) ↓ |
|---|---|---|---|---|---|---|
| π₀.₅ | 94% | 77% | 73% | 81.3% | 2.264 | 80.36 |
| FastV | 82% | 69% | 56% | 69.0% | 1.857 | 56.93 |
| SAFE-Pruner | 90% | 81% | 67% | 79.3% | 1.543 | 43.56 |
SAFE-Pruner provides a substantially better accuracy–efficiency trade-off than FastV, with 79.3% average success and 43.56 ms backbone latency.
Rollouts
Simulation Rollouts
Representative SAFE-Pruner rollouts from the four LIBERO evaluation suites.
Real-World Rollouts
Astribot S1 executes three representative manipulation tasks, demonstrating that SAFE-Pruner preserves action continuity across diverse interactions.
BibTeX
If you find SAFE-Pruner useful, please cite our work.
@article{ma2026safe,
title={SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation},
author={Ma, Shilin and Zhang, Chubin and Wang, Changyuan and Wang, Yuji and Wu, Yue and Wang, Zixuan and Tian, Jingqi and Zhu, Zheng and Tang, Yansong},
journal={arXiv preprint arXiv:2605.29662},
year={2026}
}


