機械人執行誤差下的 VLA 調適:自我補償與壓力測試
Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
研究提出 Self Compensating VLA,利用指令與實際執行動作的差異,在部署期間線上調適策略。方法以本體感覺回饋更新小型 LoRA adapters,不需任務獎勵或標籤,也不改動機械人控制器;研究另提出涵蓋七種情境的 RoboStress 模擬基準。
🤖 Robots don’t always move as commanded. We introduce Self Compensating VLA, which learns from this gap during deployment.
🧠 The policy adapts online using the difference between commanded and executed motion. We update small LoRA adapters using proprioceptive feedback, without task rewards or labels, while keeping the robot controller unchanged.
🧪 We also introduce RoboStress, a simulation benchmark with seven scenarios covering heavy payloads, thermal drift and mechanical wear. It models friction, backlash, compliance and gravity compensation errors at individual joints.
🚀 On two Piper arms, one new and one used for a year, our method improves average task success by over 30 percentage points with both π₀ and π₀.₅.
🎯 With objects unseen in the task demonstrations, π₀.₅ success increases from 16% to 64%.
📊 On RoboStress, our method outperforms the base policies, domain randomization and RobustVLA in all seven scenarios with both backbones.
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co