UNA NUEVA PERSPECTIVA

LATIDIA

Preparando tu experiencia…

Tu lugar en este universo.

Con tu autorización. Las coordenadas se muestran sólo en esta página y no se guardan.

CONECTANDO FUENTES
← Actualidad

LATIDIA · Ciberseguridad

JEV como juez de seguridad de rastreo de agentes: una comparación empírica con jueces generativos de LLM

arXiv: 2609.34862v1Tipo de anuncio: nuevo Resumen: La evaluación de seguridad de los agentes que utilizan herramientas requiere juzgar las acciones en contexto, pero los jueces generativos agregan latencia, sobrecarga de explicación y fallas de validación de salida.

WhatsApp ↗Telegram ↗
Ilustración editorial relacionada con JEV como juez de seguridad de rastreo de agentes: una comparación empírica con jueces generativos de LLM
Ilustración conceptual de LATIDIA.

La noticia

arXiv:2609.34862v1 Announce Type: new Abstract: Security evaluation of tool-using agents requires judging actions in context, yet generative judges add latency, explanation overhead, and output-validation failures. We study whether JEV, a typed decision model, offers a useful alternative for retrospective trace classification. We evaluate JEV and four generative judges on four benchmark collections totaling 5,219 trajectories, using a common risk rubric and behavior-level labels. JEV attains a benchmark-averaged positive-class F1 of 77.8, compared with 74.1 for the strongest generative configuration, GLM-5.2, with valid-result coverage of 95.5\% and 94.4\%, respectively. Performance varies across datasets, with JEV leading on ATBench500 and MCPHunt and GLM leading on R-Judge and TraceSafe. Across the four benchmarks,

← Volver a los modelos

Cargando ficha del modelo…

LATIDIA / lectura con contexto