LATIDIA · Ciberseguridad
Evaluación y mejora de la robustez de los modelos de lenguaje grandes para introducir variaciones de secuencia
arXiv:2610.02432v1 Announce Type: new Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. Esta tesis desarrolla modelos,
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.02432v1 Announce Type: new Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-box attack that reaches an