LATIDIA · Ciberseguridad
¡Los datos de alta calidad no significan seguros! Envenenamiento de LLM después de la selección de datos
arXiv: 2610.01367v1Tipo de anuncio: nuevo Resumen: Los modelos de lenguaje grande alineados con la seguridad siguen siendo vulnerables al ajuste fino en pequeños conjuntos de muestras dañinas o de aspecto benigno. Sin embargo, los estudios previos generalmente asumen que la poiso
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.01367v1 Announce Type: new Abstract: Safety-aligned Large Language Models remain vulnerable to fine-tuning on small sets of harmful or benign-looking samples. However, prior studies typically assume that poisoned samples directly enter downstream fine-tuning, overlooking quality-based selection in practical training pipelines. To fill this gap, we systematically evaluate both the filtering effects against poisoning and the downstream safety impact of retained data. The results reveal that selection removes many overtly harmful samples, yet some retained high-quality samples can still degrade model safety alignment possibly due to their harmful-like training-update patterns at the layer-wise gradient level. Together, these findings expose a practical vulnerability: safety-degrading influence can pass through quality-based selection via retained