Mixture of Agents (MoA). The key point is that MoA is not a new model architecture, but rather a method for improving the quality of AI responses. Several large language models process the same input independently of one another. Their responses are then evaluated and combined into a single, higher-quality response.
What is "Mixture of Agents"?
Most people today are looking for the best AI model. GPT-4, Claude, Gemini, or Llama. The reasoning behind this is simple: the most powerful model automatically provides the best answers.
But it is precisely this assumption that is increasingly being called into question. With Mixture of Agents (MoA) A new approach is emerging. The goal here is not to develop an even larger model. Instead, several existing models work together on the same task.
So the real issue isn't which model is the most intelligent, but rather how different models can work together effectively. Mixture of Agents, or MoA, but it is not a new model architecture; rather, it is a method.
Several language models are given the same question. Each one independently generates its own answer. These answers are then compared and combined to produce a better overall solution.
The idea comes from the paper „Mixture of Agents Enhances Large Language Model Capabilities“ by Together AI, which was featured as a Spotlight Paper at the ICLR 2025 was introduced.
What makes this special is that it requires no additional training. Instead of developing a new model, MoA makes smarter use of existing models.
The basic idea can be summarized simply:
A well-coordinated team can perform better than even the strongest individual player.
How MoA Works
The process consists of three steps. First, several models answer the same question independently of one another. This results in different approaches and perspectives. Next, other models read these answers. They identify strengths, weaknesses, or contradictions and use them to develop better versions.
Finally, a so-called Aggregator. It compiles all the answers into a single final solution. It is important to note that this is not a vote. The best answer doesn't win simply because it's the most common. Instead, the aggregator combines the strongest arguments into a new answer.

The aggregator is the real brain
Many people think of MoA as a majority decision.
But that would be far too simple.
A poor aggregator would simply average out multiple answers. The result would often be nothing more than an average.
A good aggregator, on the other hand, recognizes that,
- which answers are particularly convincing,
- which statements contradict each other,
- which arguments complement each other,
- and how this leads to a better overall solution.
The key difference, then, lies not in the number of models, but in the quality of their collaboration.
Why does this even work?
The most exciting finding from this research is surprisingly simple. Language models improve when they see the answers provided by other models—even if those answers aren’t perfect. Seeing other approaches helps them identify their own mistakes, avoid blind spots, and draw better conclusions.
Researchers call this effect Collaborativeness. People have long been familiar with this principle from teamwork. Good ideas often don't come about on their own, but through interaction with others.
Sakana Fugu – the pufferfish for MOA
In the original paper, a combination of open-source models achieved 65.1 percent on AlpacaEval 2.0 and thus surpassed GPT-4 Omni at 57.5 percent. However, the real message isn't in the number. Sakana AI goes one step further here. The company was recognized by, among others, Llion Jones, a co-author of Attention Is All You Need, founded.
Your System Fugu Replaces MoA's fixed architecture with an intelligent orchestrator.
He decides for each individual task,
- which models are used,
- what role they play,
- and how they work together.
Sometimes we need more creative thinkers, other times more thorough reviewers or specialized experts. So the team is put together anew each time.
That is exactly why the system is called Fugu, the Japanese pufferfish. It only inflates itself when necessary. In the same way, Fugu activates only the additional intelligence that is truly needed for the task at hand.
MoA offers several clear strengths. Multiple models provide different perspectives and make the results more robust. At the same time, it reduces dependence on a single provider. This became clear by 2026 at the latest, when export restrictions made certain models temporarily unavailable. Those who orchestrate different models spread this risk.
Of course, this approach also has its drawbacks. Using multiple models requires more computing power and more time. The aggregator can only generate the final answer once all the answers are available. This isn’t always suitable for real-time applications. In addition, researchers are currently debating an important question: Do we need models that are as diverse as possible, or is a particularly good aggregator sufficient?
There are many indications that, in the future, the quality of the orchestration will be the decisive factor.
What does this mean for businesses?
For years, the most important question was: Which model is the best? This question is slowly becoming less important.
The more strategically important question going forward is: How can I get different models to work together to produce better results? The competitive landscape is shifting, and it is no longer the largest model that wins, but rather the best-coordinated team. For companies, this means that orchestration is becoming a new core competency.
Key Points Summarized
- Mixture of Agents is a runtime method that allows multiple models to work together in layers without the need for retraining; it was introduced by Together AI at ICLR 2025.
- The architecture has two roles: proposers generate various initial responses, and an aggregator merges them into the final response.
- The key is collaborativeness: models improve when they use others' answers as a reference, even weaker ones.
- A purely open-source consortium outperformed GPT-4 Omni on AlpacaEval 2.0, proving that coordination can trump size.
- Sakana Fugu takes this principle a step further: a trained pufferfish conductor assembles a team for each task and recursively calls itself.
- The limitations are latency and an ongoing debate among researchers over whether diversity or aggregator quality is the deciding factor.
- The trend is shifting the bottleneck from model size to coordination logic. Orchestration is becoming a strategic priority.
And that’s why I remain convinced: The future no longer belongs to typing and clicking, nor to the largest single model, but to collaborative work with orchestrated systems in which people decide which forms of intelligence are brought together for which tasks. So if you’d like to talk, and if you’d like to collaborate with me, feel free to reach out. www.rogerbasler.ch
Source
Sakana AI. (2026). Sakana Fugu: A multi-agent orchestration system as a foundation model [Technical report]. arXiv. https://arxiv.org/html/2606.21228v1