LATIDIA · Ciberseguridad
PrivDrift: Auditoría de fugas secretas de usuarios bajo la deriva del tema en conversaciones activas de LLM
arXiv: 2609.30094v1Tipo de anuncio: Cross Resumen: Los modelos de lenguaje grandes operan cada vez más como asistentes persistentes en la configuración orientada al usuario, de sesión compartida y aumentada por herramientas. Cuando los usuarios divulgan información confidencial
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.30094v1 Announce Type: cross Abstract: Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts even after the dialogue shifts to unrelated topics. We introduce \textbf{PrivDrift}, a benchmark for auditing whether user-disclosed secrets remain recoverable after conversational topic drift and persuasion-based probing. PrivDrift contains 1{,}000 controlled multi-turn dialogues with seeded secrets, content-dense drift turns, and standardized extraction probes. Across three LLMs with extended context windows, dialogue-level hybrid leakage remains substantial, ranging from 38.7\% to 54.6\%, and varies strongly by model, secret type, and persuasion intensity. Within the tested