More models, worse answers
A new 23-model study finds that growing the pool behind a multi-agent system usually drags it below its own best member. The reasons are older than LLMs: voting amplifies whatever the voters already are, and recognizing a right answer is harder than producing one.
There is an intuition I have carried into almost every AI system I have built: if one model is good, several different models checking each other must be better. Different labs, different training data, different blind spots. Put them on the same question, let them vote, and the mistakes should cancel. It is the same reason we ask for a second opinion. The literature seemed to agree. Sampling and voting scales with the number of agents [2], layered mixtures of open models beat GPT-4o on AlpacaEval [3], and models that debate each other reason better [4].
A paper posted this month tests that intuition harder than anything I have read, and mostly breaks it. "Mo' Models, Mo' Problems", from researchers at NVIDIA and the University of Copenhagen, builds multi-agent systems out of a pool of 23 open models and asks the question practitioners rarely ask out loud: which models should go in the pool, and how many? [1] The headline is blunt. In the authors' words, "more models nearly always decreases performance."
I want to walk through what they did, why the result is less surprising than it sounds once you line it up against older work on ensembles and two years of inference-scaling research, where the opposing evidence is right, and what I am changing in how I build.
The setup
The motivation is a selection problem. The paper counts close to three million models on Hugging Face as of July 2026, and nobody building a multi-agent system can evaluate more than a sliver of them. So the authors take 23 open models, from a 2B Qwen3.5 up to the 1.6-trillion-parameter DeepSeek-v4-Pro, spanning the Llama 3.1, Qwen3, Qwen3.5, Gemma 4, OLMo 3, and GPT-OSS families plus a handful of science specialists, and test eight ways of choosing a subset of size 3, 5, 10, 15, or 20 [1]:
- Size: the largest models, one per family.
- Family: every member comes from a single model family.
- LLM-chosen: GPT-5 with deep research is asked to pick the team.
- Accuracy: the top k models by single-shot accuracy.
- Correct-answer diversity: models whose sets of solved questions overlap the least.
- Error diversity: models whose wrong answers differ the most.
- Two blends that weight accuracy and each diversity measure equally.
Random subsets of the same size serve as the control. Everything is evaluated on three expert-level science benchmarks: Humanity's Last Exam (HLE), GPQA-Diamond, and the FrontierScience Olympiad set, with tools and retrieval switched off.
Each subset is then wired into three architectures. In routing, a trained router reads the question and sends it to exactly one model, so the choice happens before anything is generated. In plurality voting, every model answers five times, the final answers are embedded and clustered, and the largest cluster wins. In the judge setup, gpt-oss-120b reads the candidates and picks one, following the GenSelect recipe [5]. Every system is scored against the best single model inside its own pool. That is the right baseline, because the best member is what you would have shipped otherwise.
What they found
- 01Bigger pools lose. Across selection strategies and architectures, the gain over the best member is mostly negative, and it gets more negative as the pool grows from 3 to 20.
- 02Staying in the family loses least. The only groupings that improved on their best member with any regularity were single-family pools, Gemma 4 most often. The authors are careful to frame this as the smallest penalty rather than a big win.
- 03Clever selection did not beat luck. With an LLM judge, nearly all of the deliberately chosen subsets scored below random subsets of the same size. The best case for accuracy-based selection was roughly 2 points on FrontierScience and under 1 point on GPQA.
- 04Specialists did not specialize. Plain Llama-3.1-8B outscored the 8B science fine-tunes in the pool across every science domain, including the physics and chemistry questions those fine-tunes were built for.
The number I keep coming back to sits on the other side of the ledger. On HLE, the best single-model system scores 29.4% with one sample. Sample the same model five times and take a plurality vote, and it reaches 32.2%. Let the judge pick among five samples, and it reaches 36.5% [1]. That is about seven points with no second model anywhere in the system.
Why a committee can be worse than its best member
Voting amplifies whatever you already are
In 1785 the Marquis de Condorcet proved the theorem that every argument for ensembles quietly leans on: if each juror is independently right more than half the time, a majority vote approaches certainty as the jury grows [6]. The half that gets quoted less often is the mirror image. If each juror is right less than half the time, the majority converges on the wrong answer just as fast.
Now look at the benchmark. The best model in the pool is right 29.4% of the time on HLE, and most of the pool sits far below that. In a two-answer world this committee would be hopeless. Plurality voting works at all on open-form questions because there are many ways to be wrong: if correct answers agree with each other while wrong answers scatter, the correct cluster can win with a minority of the votes. That is the engine behind self-consistency [7], and behind the jump from 29.4 to 32.2. It also tells you exactly what breaks it, which is anything that adds wrong answers faster than right ones, or that makes wrong answers agree with each other.
The pool knows more than the committee can say
The paper also computes an oracle score: the system gets credit if any model in the subset was right. Oracle accuracy rises with pool size for every selection strategy, as it must. Realized accuracy goes the other way. The authors' explanation is that a diverse set of solutions pollutes the answer space and makes the correct one harder to isolate [1].
This gap has been measured before from the single-model side. In Large Language Monkeys, coverage (the share of problems solved by at least one sample) scales across four orders of magnitude of samples. On SWE-bench Lite it goes from 15.9% with one sample to 56% with 250, because unit tests can verify each attempt. Without an automatic verifier, majority voting and reward models plateau after a few hundred samples [8]. Generation keeps scaling, and selection stops.
Chen and colleagues explain the voting half. In compound systems that vote, accuracy can rise and then fall as calls are added, because extra votes help on easy queries and hurt on hard ones [9]. HLE is close to a benchmark made entirely of hard queries.
To see the mechanism in isolation I wrote a small simulation. Twenty simulated models range from 29% accuracy down to 4%, they answer 40,000 questions of varying difficulty, wrong answers are spread over thirteen options with one tempting distractor that attracts 35% of the errors, and a plurality vote decides, with ties going to the strongest model.
The oracle climbs from 29% to 56%. The vote over the mixed pool peaks at 33% with five models and then sinks to 24% with twenty, below the best single model. With a pool of similar-strength models the vote holds around 35%. I should be honest that my toy is kinder to small committees than reality was: it still shows a gain at three models, while most of the paper's real pools were already behind at that size. Real model errors are not independent once you account for difficulty, and clustering free-text answers or asking a judge to choose adds an error rate of its own. Treat this as a sketch of the mechanism and not as a replication.
Quality beats diversity, and we have known for a while
Mixture-of-Agents made the case for mixing: layers of different open models refining each other's drafts reached 65.1% on AlpacaEval 2.0, against 57.5% for GPT-4o [3]. Li and colleagues then asked whether the mixing was doing the work. Their Self-MoA variant aggregates repeated samples from only the single best model, and it beats the mixed version by 6.6% on AlpacaEval 2.0 and by 3.8% on average across MMLU, CRUX, and MATH [10]. Their diagnosis is a trade-off between quality and diversity in which quality dominates, because every weaker member lowers the average quality of what the aggregator has to read.
The classical ensemble literature reached the same place in 2003. Kuncheva and Whitaker compared ten ways of measuring diversity among classifiers and concluded that, outside special cases, the measures were of doubtful use for actually building better ensembles [11]. Twenty-three years later, the two diversity-based selection strategies in this paper had the weakest oracle potential, and the diversity signals did not translate into realized gains.
Those two correlations are the quiet center of the paper. The first says strong models mostly solve the same questions, so a second strong model buys you few new ones. The second sounds like good news for voting, since errors that scatter should cancel. But on a question where most of the pool is wrong, scattered errors mean the judge or the vote faces a ballot crowded with distinct, plausible candidates and perhaps one correct answer among them.
There is a tension here that I have not resolved. Kim and colleagues studied more than 350 models and found that on one leaderboard dataset, models agree 60% of the time when both are wrong, with larger and more accurate models showing highly correlated errors even across different architectures and providers [12]. Goel and colleagues found model mistakes becoming more similar as capability increases, and showed that LLM judges favor models similar to themselves [13]. A weak error correlation on HLE does not obviously fit with that. My guess is that the difference is in the measurement: a multiple-choice leaderboard offers a handful of wrong options, while HLE answers are free text with unlimited ways to miss. The paper does report that DeepSeek-v4 and Qwen3.5 make very similar errors to each other. And the judge-similarity result deserves a follow-up, given that the judge here is a GPT-OSS model and GPT-OSS models are also in the pool.
Where the other side is right
It would be easy to read all of this as a verdict that heterogeneous systems are a dead end. The strongest counterexample is from February. Yang and colleagues show that scaling homogeneous agents hits diminishing returns because their outputs are strongly correlated, that heterogeneous agents keep adding complementary evidence, and that 2 diverse agents can match or beat 16 homogeneous ones [14].
I think both papers are right, and they agree on the mechanism. What matters is the number of effective, independent channels of evidence, which is what Yang's K* metric estimates. They differ on whether an off-the-shelf pool gives you channels that are both independent and good. Yang's heterogeneity includes different prompts and tools as well as different models. The NVIDIA study fixes the prompt, disables tools, varies only the model, and runs on questions where most of the pool is far weaker than its best member. In that regime the quality gap swamps the diversity gain. The authors say as much: they do not claim heterogeneous systems are futile, only that reported failures may come down to a poor choice of candidates [1].
Task structure matters as much as the pool. Across 260 agent configurations, Kim and colleagues measured multi-agent effects ranging from +80.8% on decomposable financial reasoning to -70.0% on sequential planning [15]. And when Cemri and colleagues annotated more than 1,600 traces from seven multi-agent frameworks, their 14 failure modes fell into three groups: system design, inter-agent misalignment, and task verification [16]. Verification again.
Routing deserves its own note. RouteLLM shows that a learned router between a strong and a weak model can cut cost by more than 2x while holding response quality [17]. Notice what routing is good at there: cost, not beating the best model. A router has to predict which model will be right before seeing a single answer, and in this study only the diversity-of-correct-answers and family groupings showed any routed gain at all [1].
What the paper does not tell us
- It covers science reasoning only. Math and code, where cheap verifiers exist, are not tested, and the inference-scaling results suggest those are exactly the domains where bigger pools should pay [8].
- Tools and retrieval are off, and only before-generation and after-generation systems are tested. Debate and layered refinement, where models read each other mid-flight, are out of scope [3][4].
- There is one router implementation and one judge model. The automated judging was spot-checked on 100 HLE questions (93% agreement, Cohen's kappa of 0.63), and the authors note it may lean generous.
- Each family contributes only a few models, so the finding that families lose least rests on small samples.
- It is a preprint, and I read the HTML version. The per-strategy gains live in figures, not tables, which is why I quote so few of them here.
What I am changing
My own projects are not innocent here. NoteVision runs four different models. But each one owns a different job: layout, long-form reasoning, messy diagrams, normalization. No two models ever answer the same question, so nothing has to be reconciled. That is decomposition, and it is the regime where multi-agent gains are real [15]. A committee is a different thing: many models, one question, one answer to pick. This paper is about committees, and the following is what I take from it.
- 01Start with the best single model and sample it more than once. Measure vote-of-five and judge-of-five before adding anything. On HLE that alone was worth seven points [1].
- 02Score every pool against its own best member. A committee that beats its average member has proven nothing.
- 03If you add a model, add from the top. A member far below your best is a net source of wrong answers. Track both the oracle gain and the realized gain, and remember that only the second one ships.
- 04Spend on selection before you spend on generation. Unit tests, a checker, a rubric, a better judge. Where verification is cheap, more samples keep paying. Where it is not, they plateau [8].
- 05Use different models for different jobs, not for the same job.
- 06Re-run the comparison whenever a model changes. The ranking inside your pool is the input that everything else depends on.
The intuition I started with was not wrong so much as incomplete. Second opinions help when the second doctor is about as good as the first, and when someone competent reads both reports. Take away either condition and you have not built a panel of experts. You have built a louder room.
References
- 01Mo' Models, Mo' Problems: How to best select model pools when designing Multi-Agent Systems (Marjanović et al., 2026) · arXiv:2609.17306
- 02More Agents Is All You Need (Li et al., 2024) · TMLR · arXiv:2402.05120
- 03Mixture-of-Agents Enhances Large Language Model Capabilities (Wang et al., 2024) · arXiv:2406.04692
- 04Improving Factuality and Reasoning in Language Models through Multiagent Debate (Du et al., 2023) · arXiv:2305.14325
- 05GenSelect: A Generative Approach to Best-of-N (Toshniwal et al., 2025) · arXiv:2507.17797
- 06Condorcet's jury theorem · Wikipedia
- 07Self-Consistency Improves Chain of Thought Reasoning in Language Models (Wang et al., 2022) · ICLR 2023 · arXiv:2203.11171
- 08Large Language Monkeys: Scaling Inference Compute with Repeated Sampling (Brown et al., 2024) · arXiv:2407.21787
- 09Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems (Chen et al., 2024) · arXiv:2403.02419
- 10Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial? (Li et al., 2025) · arXiv:2502.00674
- 11Measures of Diversity in Classifier Ensembles and Their Relationship with the Ensemble Accuracy (Kuncheva and Whitaker, 2003) · Machine Learning 51(2)
- 12Correlated Errors in Large Language Models (Kim et al., 2025) · ICML 2025 · arXiv:2506.07962
- 13Great Models Think Alike and this Undermines AI Oversight (Goel et al., 2025) · arXiv:2502.04313
- 14Understanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity (Yang et al., 2026) · arXiv:2602.03794
- 15Towards a Science of Scaling Agent Systems (Kim et al., 2025) · arXiv:2512.08296
- 16Why Do Multi-Agent LLM Systems Fail? (Cemri et al., 2025) · arXiv:2503.13657
- 17RouteLLM: Learning to Route LLMs with Preference Data (Ong et al., 2024) · arXiv:2406.18665
Reach out
If something here resonated, I'd love to hear what you're building. Always open to a good conversation.