AI Engineering

Multi-agent consensus in financial analysis: beyond single-prompt LLMs

In financial analysis a single plausible fabrication is expensive, because it is the sentence that gets repeated in the meeting. Here is what the scaling research says about when multiple agents help, where they stop helping, and why the part that does the real work is not a model at all.

A vertical brass bar with a drift of blurred chart slips on one side and three clean cards on the other
Figure 1: What reaches the review is what survives the gate.

The monolithic wall in enterprise finance

Language models are fluent across text, code, and qualitative reasoning. Enterprise finance is a harder setting for them, because a balance sheet reconciles to the dollar and a working capital error is not a stylistic problem.

Consolidating an entire analysis into one prompt creates a single point of failure. General ledger detail, vendor terms, and bank reconciliations all go into one context, and everything that comes out rests on one model's reading of it. Cognizant's Babak Hodjat makes the architectural version of this argument directly: context size remains a limiting factor, and distributed, modular designs with built-in checks are more resilient than delegating every decision to one powerful model[2].

The deeper problem is not context length. It is that a lone prompt has no stage at which a claim must justify itself. If the model misreads a deferred revenue entry or applies the wrong coverage ratio, nothing stands between that mistake and the executive summary.

The failure mode is not that the model sounds uncertain. It is that it sounds entirely confident about something nothing supports.

What the scaling research actually found

The most useful evidence on when multi-agent systems help comes from Towards a Science of Scaling Agent Systems, a 2026 study across a wide set of tasks and architectures[1]. Three of its findings shaped what we built, and two of them are cautionary.

  1. The task shape decides the outcome. Relative performance against a single-agent baseline ranges from +80.8% on decomposable financial reasoning to −70.0% on sequential planning[1]. Splitting work across agents is not a general improvement. It helps on work that genuinely decomposes, and it actively hurts on work that does not.
  2. There is a ceiling, and it is low. Tasks where a single agent already exceeds roughly 45% accuracy see negative returns from additional agents, as coordination costs exceed the remaining improvement[1]. More agents is not a strategy.
  3. Coordination is superlinear. Turn count against agent count fits T = 2.72 × (n + 0.5)^1.724, R² = 0.974[1]. Every seat added past a small number spends budget on talking rather than on analysis.

The finding that matters most for design is about structure rather than count. The same study measures trace-level error amplification by architecture: independent ensembles amplify at 17.2, centralized coordination at 4.4[1]. A verification bottleneck is what contains error. Parallelism alone does not.

A clipped stack of sheets in focus, discarded drafts falling away behind it
Figure 2: A verification bottleneck is what contains error. Parallelism alone is not.

Five seats, read independently

Kalends runs recurring period reviews through four stages, and only one of them is a set of language models.

Detect is a deterministic exception engine. Structured accounting metrics and rule checks run against the financial records before any model reads anything. Threshold breaches, unusual movements, and relevant comparisons come out of software. The analysis does not get to choose what it looks at.

Analyze is five anonymous analytical seats reading that same exception set in parallel, without seeing each other's work. Each returns structured claim drafts rather than one unconstrained narrative. Five is a deliberate number, chosen against the scaling results above rather than in spite of them: enough independent readings for disagreement to carry signal, few enough that coordination does not eat the budget.

We do not claim that more seats produce better answers. The research says the opposite past a low threshold. What independent readings buy is disagreement, and disagreement is a signal a single confident narrative can never give you.

The verification gate is not a model

Verify is the stage that does the work, and it is software logic rather than another language model. Asking a model to check a model reproduces the original problem one level up.

The gate does three things:

  • It drops claims that nothing in the packet supports.
  • It checks every cited amount against the financial packet it came from.
  • It binds each surviving statement to the evidence it was derived from.

A claim that cannot be traced does not proceed. This is the centralized verification bottleneck the scaling study measured, applied to a financial packet: the point is not that five readers are wiser than one, but that a deterministic checkpoint keeps one reader's error from propagating[1].

Ratify is what survives, converged into one concise period review in which every published statement keeps its evidence. The word describes the system's verified consensus output. It is not a claim that a person has personally signed off on every line.

Why the claim type matters more than the agent count

The constraint we lean on hardest is not architectural. It is grammatical. An analytical seat may publish exactly three kinds of statement:

  • an observation based on a cited figure;
  • a threshold breach tied to a defined rule; or
  • a comparison between two or more cited movements.

That is the entire vocabulary. There is no causal claim type, so no model can publish a statement about why something happened merely because the explanation sounds right. That is a property of the output format rather than an instruction in a prompt that a model might drift away from.

It is worth being plain about what this does and does not buy. Dropping unsourced claims removes one route by which a plausible fabrication reaches a decision-maker. It does not make the surviving analysis correct, and it does not make the system incapable of error. We publish no accuracy or error-reduction figures of our own, because we have not measured any that would mean anything to you. The numbers in this article come from published research on agent systems in general, not from a benchmark of ours.

For the case against giving a model the authority to act on its own analysis, see what a digital CFO should actually do, which works through a CFO benchmark in which only 16% of agent runs survived eleven simulated years[3].

References & Academic Sources

  1. [1] Kim, Y., Gu, K., Park, C., Park, C., Schmidgall, S., Heydari, A. A., Yan, Y., Zhang, Z., Zhuang, Y., Liu, Y., Malhotra, M., Liang, P. P., Park, H. W., Yang, Y., Xu, X., Du, Y., Patel, S., Althoff, T., McDuff, D., & Liu, X. (2026). Towards a Science of Scaling Agent Systems. arXiv preprint arXiv:2512.08296. [arXiv:2512.08296]↩
  2. [2] Hodjat, B. (2025). Single Agent vs. Multi-Agent: Choosing the Right Architecture. Cognizant AI Lab, February 11, 2025. [Source Link]↩
  3. [3] Han, Y., Wang, Y., Qian, L., Li, H., Cao, Y., He, Y., Peng, X., Shen, N., Xu, Y., Chen, Y., Feng, D., Huang, J., Liu, X., Nie, J.-Y., & Ananiadou, S. (2026). Can LLM Agents Be CFOs? A Benchmark for Resource Allocation in Dynamic Enterprise Environments. arXiv preprint arXiv:2603.23638v1. [arXiv:2603.23638v1]↩

Frequently Asked Questions

Q1

Why does a single prompt struggle with a full financial packet?

A single prompt consolidates reading, calculation, and narration into one pass with no internal check. Everything it produces rests on one model's interpretation of a long context, and there is no stage in which a claim has to justify itself before it appears in the summary. A misread deferred revenue entry becomes a sentence in the executive summary with nothing standing between the two.

Q2

Do more agents produce a better answer?

No, and the research is explicit about it. The scaling study found that tasks where a single agent already exceeds roughly 45% accuracy see negative returns from adding agents, because coordination cost outruns the improvement. Turn complexity also grows faster than agent count, fitted at T = 2.72 x (n + 0.5) ^ 1.724. Adding seats past a small number spends budget on coordination rather than analysis.

Q3

What does Kalends actually run?

Four stages. A deterministic exception engine examines the financial records before any language model reads them. Five anonymous analytical seats then read the same exception set in parallel, without seeing each other's work, and produce structured claim drafts. A deterministic validation gate drops unsupported claims and binds the survivors to their evidence. What remains converges into one sourced period review.

Q4

What kinds of statements can the system publish?

Three, and only three: an observation based on a cited figure, a threshold breach tied to a defined rule, or a comparison between cited movements. There is no causal claim type, so a model cannot publish an explanation of why something happened on the strength of it sounding plausible. That is a property of the output format rather than an instruction in a prompt.

KK

Written by

Ken Koo

Founder

Decades of CFO experience creating and curating financial KPIs and multi-agent AI verification systems.

See a period review

Kalends produces a sourced monthly review of your financial records, built to be inspected and challenged before the meeting starts.

Request early access →