LATIDIA · Investigación
Informe Técnico Pistis
arXiv:2609.28554v1 Tipo de anuncio: nuevo Resumen: Presentamos la familia de modelos Pistis, que comprende modelos de lenguaje grande multimodales de 27B y 9B parámetros construidos sobre Qwen3.6 y Qwen3.5, respectivamente, y desarrollados a través de un
WhatsApp ↗Telegram ↗
La noticia
arXiv:2609.28554v1 Announce Type: new Abstract: We introduce the Pistis model family, comprising 27B- and 9B-parameter multimodal large language models built on Qwen3.6 and Qwen3.5, respectively, and developed through a general and scalable post-training framework. The framework first establishes a strong foundation through large-scale multimodal supervised fine-tuning (SFT). Building on this SFT foundation, we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel post-training paradigm that tightly integrates on-policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives, rather than optimizing either in isolation or combining them in a static joint loss, IDRL enables more effective knowledge transfer, greater optimization stability, and more precise credit