LATIDIA · Ciberseguridad
CounterSteer: Supresión de la inyección de aviso indirecta con dirección de activación
arXiv:2609.36570v1 Announce Type: new Abstract: Indirect prompt injection makes a LLM agent treat untrusted retrieved text as instructions. Presentamos CounterSteer, una defensa de tiempo de inferencia que suprime este comportamiento
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.36570v1 Announce Type: new Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages),