LATIDIA · Ciberseguridad
HARDE: Optimización de los arneses de agentes para la detección de riesgos en tiempo de ejecución y el control de ejecución
arXiv: 2609.38291v1Tipo de anuncio: nuevo Resumen: Los agentes del modelo de lenguaje grande (LLM) son vulnerables a los riesgos de seguridad, como la inyección de instrucciones maliciosas o información engañosa, lo que motiva las defensas en tiempo de ejecución que evitan
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.38291v1 Announce Type: new Abstract: Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtime defenses that prevent unsafe action in execution across diverse risks while preserving benign-task utility. Existing system-level defenses either focus on risk detection rather than timely prevention or rely on predefined rules with limited flexibility across diverse risks. We propose a risk-aware harness that integrates LLM-based monitoring for flexible risk detection and structures monitor-guided execution around three core modules: trigger, monitor, and feedback, enabling targeted safety interventions while limiting disruption to benign task execution. To adapt the harness to different risks and deployment settings, we introduce