UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

Ataques de decodificación controlada en Black-Box LLM

arXiv:2609.36956v1 Announce Type: new Abstract: Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Sin embargo, los enfoques existentes se basan en el acceso al modelo Weig

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con Ataques de decodificación controlada en Black-Box LLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.36956v1 Announce Type: new Abstract: Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled text. Reconstructing probabilities from sampled outputs offers a possible alternative, but finite sampling produces sparse and noisy estimates, while repeating this process at every generation step incurs substantial query costs. Our empirical observations suggest that large distributional changes along successful jailbreak trajectories are concentrated at a small subset of positions, motivating selective control. We introduce \method{}, a framework for jailbreaking through text-only continuation interfaces that permit repeated

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto