LATIDIA · Robótica
Más fácil decirlo que hacerlo: Desempaquetar la brecha de intención-comportamiento en los robots basados en LLM de Jailbreaking
arXiv:2412.16633v4 Tipo de anuncio: reemplazar Resumen: los robots basados en LLM utilizan modelos de lenguaje grande (LLM) como planificadores para traducir instrucciones de lenguaje natural en políticas como grasp(), move_to() y open_gripper(). J
WhatsApp ↗Telegram ↗
La noticia
arXiv:2412.16633v4 Announce Type: replace Abstract: LLM-based robots use Large Language Models (LLMs) as planners to translate natural language instructions into policies such as grasp(), move_to(), and open_gripper(). Jailbreak attacks on these robots extend the threat from generating malicious content to executing harmful behaviors. However, we find that existing jailbreak attempts against LLM-based robots that produce malicious-looking policies (intent jailbreaks) often fail to induce harmful physical actions by robots (behavior jailbreaks), due to robot-specific constraints, such as logical errors and hallucinated control APIs. In this paper, we demystify the intent-behavior gap and investigate its root causes to inform effective defenses. Our measurement study finds that current LLM jailbreak methods overlook robot-specific