Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning Paper • 2510.25992 • Published 15 days ago • 41
Multi-Agent Evolve: LLM Self-Improve through Co-evolution Paper • 2510.23595 • Published 18 days ago • 10