Maestro: Learning to Collaborate via Conditional Listwise Policy Optimization for Multi-Agent LLMs
arxiv.org·11h
Flag this post

View PDF HTML (experimental)

Abstract:Multi-agent systems (MAS) built on Large Language Models (LLMs) are being used to approach complex problems and can surpass single model inference. However, their success hinges on navigating a fundamental cognitive tension: the need to balance broad, divergent exploration of the solution space with a principled, convergent synthesis to the optimal solution. Existing paradigms often struggle to manage this duality, leading to premature consensus, error propagation, and a critical credit assignment problem that fails to distinguish between genuine reasoning and superficially plausible arguments. To resolve this core challenge, we propose the Multi-Agent Exploration-Sy…

Similar Posts

Loading similar posts...