LATIDIA · Ciberseguridad
Modelos rápidos, evidencia lenta: una evaluación emparejada y autoauditada de los modelos de decisión del sistema 1 para arneses de agentes LLM
arXiv:2610.02267v1 Tipo de anuncio: cruz Resumen: Los arneses de agente toman muchas decisiones pequeñas y escritas por tarea: a qué modelo llamar, qué herramienta usar, si el texto recuperado es relevante, si una entrada lleva un inyecti
WhatsApp ↗Telegram ↗
La noticia
arXiv:2610.02267v1 Announce Type: cross Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points