LATIDIA · Ciberseguridad
Just Ask Jev: Aprendizaje de refuerzo para decisiones calibradas como un detector de disparos cero de fallas de alineación de IA
arXiv: 2609.29429v1Tipo de anuncio: Cross Resumen: Detectores de fallas de alineación pantalla implementó modelos de lenguaje y puntos de referencia de alineación de puntajes. La mayoría son jueces generativos que gastan un pase de decodificación en cada criterio,
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.29429v1 Announce Type: cross Abstract: Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44