LATIDIA · Investigación
¿Cómo se produce y es importante el maltrato de la IA del usuario en los sistemas conversacionales?
arXiv: 2609.13579v1Tipo de anuncio: nuevo Resumen: La investigación sobre seguridad a menudo se centra en los daños generados por los modelos, pero los usuarios también pueden dirigir hostilidad, coerción y presión adversaria a los modelos. Entender cómo y cuándo eso o
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.13579v1 Announce Type: new Abstract: Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5%