LATIDIA · Ciberseguridad
¿Cómo domar una hidra de varias cabezas? Dirección de seguridad multicategoría adaptable para modelos de lenguaje grandes
arXiv: 2609.34514v1Tipo de anuncio: nuevo Resumen: A medida que los modelos de lenguaje grandes (LLM) se generalizan cada vez más, evitar respuestas inseguras a mensajes dañinos es esencial para su implementación segura. Dirección de activación o
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.34514v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly widespread, preventing unsafe responses to harmful prompts is essential for their safe deployment. Activation steering offers an approach to improving LLM safety by modifying internal activations during inference without updating model parameters. However, a single prompt can involve multiple harm categories, and steering toward safety in one category may leave harmful content from another unaddressed. Despite advances in adaptive steering, existing methods do not explicitly coordinate steering direction and strength when multiple harm categories co-occur within a single prompt. To address this problem, we propose CAM-Steer, a Category-Adaptive Multi-category Safety Steering framework. Specifically, it estimates the risk associated