LATIDIA · Robótica
Búsqueda de política solicitada de preservación de la privacidad para el control robótico
arXiv: 2609.30554v1Tipo de anuncio: nuevo Resumen: los modelos de lenguaje grandes (LLM) han demostrado recientemente capacidades prometedoras como optimizadores de políticas en contexto para el aprendizaje de refuerzo (RL), lo que permite la unidad de búsqueda de políticas
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.30554v1 Announce Type: new Abstract: Large language models (LLMs) have recently demonstrated promising capabilities as in-context policy optimizers for Reinforcement Learning (RL), enabling policy search driven by both numerical reward signals and natural language reasoning. However, deploying such methods in practice requires transmitting raw policy parameters and rewards history to cloud-based LLM APIs, exposing proprietary control strategies to third-party service providers. To address this issue, this paper introduces Privacy-Preserving Prompted Policy Search (PP-ProPS), a framework that enables LLM-guided policy optimization while keeping policy and environmental parameters confidential. PP-ProPS encodes policy parameters and reward values using secret client-side transformations before they are included in each API request, ensuring that the LLM