Controlling Collectives of AI Agents in Reasoning Space with Spatial Transformers
Frederic Vatnsdal, Roshan Gopal, Romina Garcia Camargo, Vijay Kumar, Alejandro Ribeiro
- Published
- Sep 23, 2026 — 15:08 UTC
Problem
Large Language Models (LLMs) exhibit significant performance degradation in multi-robot tasks as the size of the team increases. This paper addresses this gap in capability, particularly in the context of decentralized control of AI agents, which is critical for applications involving large collectives of robots. The work is presented as a preprint and has not undergone peer review.
Method
The authors propose COMPASS, a decentralized multi-robot architecture designed to enhance coordination among agents. A key feature of COMPASS is its use of a spatial transformer that aggregates multi-hop messages into a learned feedback token, facilitating effective communication among agents. The architecture is trained on scenarios involving up to 1024 robots, allowing it to handle large-scale coordination tasks. To mitigate biases in command interpretation, the input commands are structured with diversity, ensuring that the agents can adapt to a variety of instructions.
Results
COMPASS demonstrates a significant performance gain in producing cohesive flocking formations compared to a centralized frontier LLM policy and a language-only communication ablation. Notably, COMPASS exhibits the ability to generalize zero-shot to unseen instructions, effectively commanding flocks of agents up to 16 times its training scale. The available text does not report quantitative results.
Limitations
The authors acknowledge that the hand-engineered feedback mechanism, which incorporates raw state information in the language channel, can reduce the overall cohesion of the agent collective. This limitation suggests that further refinement of the feedback process may be necessary to enhance performance in more complex scenarios.
Why it matters
The implications of this work are significant for the field of multi-agent systems and robotics. By demonstrating that a decentralized approach can effectively control large collectives of AI agents, this research opens avenues for more scalable and robust applications in real-world scenarios, such as autonomous vehicle fleets, drone swarms, and collaborative robotic systems. The ability to generalize to unseen instructions also suggests potential for adaptability in dynamic environments.
By Callan Zhang · Sep 23, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
