UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

CredLeakBench: Evaluación de fugas de credenciales y recuperación en agentes de LLM

arXiv: 2610.08871v1Tipo de anuncio: nuevo Resumen: Los agentes del modelo de lenguaje se implementan cada vez más para automatizar las tareas digitales cotidianas, desde la administración de correos electrónicos y redes sociales hasta el manejo de cuentas bancarias y facturas, lo que permite a los usuarios

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con CredLeakBench: Evaluación de fugas de credenciales y recuperación en agentes de LLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2610.08871v1 Announce Type: new Abstract: Language model agents are increasingly deployed to automate everyday digital chores from managing emails and social media to handling banking and bills allowing users to step away from supervision. However, this capability also exposes sensitive information to phishing. Safe execution requires distinguishing malicious requests from genuine ones without simply refusing to act. Despite its practical importance, this problem remains underexplored and it is unclear whether current agents or existing defenses can achieve it. To study this problem, we first propose CredLeak-Bench, a comprehensive benchmark designed to evaluate how effectively and securely agents automate human workflows when confronted with phishing and identity verification. The benchmark covers

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto