LATIDIA · Ciberseguridad
HE-Guardrail: una barandilla homomórfica contra los ataques de jailbreak para la inferencia cifrada del modelo de lenguaje grande
arXiv: 2609.21484v1Tipo de anuncio: nuevo Resumen: El cifrado homomórfico (HE) se ha convertido en un enfoque prometedor para el aprendizaje automático que preserva la privacidad (PPML), que permite el cálculo directamente sobre los datos cifrados. En HE-base
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.21484v1 Announce Type: new Abstract: Homomorphic encryption (HE) has emerged as a promising approach to privacy-preserving machine learning (PPML), enabling computation directly over encrypted data. In HE-based PPML, a client submits an encrypted input to the server, which evaluates models such as large language models (LLMs) without access to the underlying plaintext. However, we identify a critical security vulnerability in this setting: HE-LLM inference is vulnerable to malicious clients that submit adversarial prompts, such as jailbreak attacks. The same confidentiality that protects benign clients also prevents the server from inspecting incoming prompts or generated responses, making adversarial attempts difficult to detect or block and potentially allowing successful attacks to remain