Position: Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain

Published in ICML 2026 (Position Paper Track), 2026

This position paper argues that self-evolving LLM systems fail when the synthetic data they generate carries no learnable information gain across iterations. We propose three design principles for sustained self-improvement beyond the initial self-play plateau: asymmetric co-evolution across Proposer, Solver, and Verifier roles; capacity growth; and proactive information seeking.

Recommended citation: W Liu, S Qi, Y Du, Y He. (2026). "Position: Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain." Forty-third International Conference on Machine Learning (ICML 2026), Position Paper Track.
Download Paper