The 16% result, and what it does not mean
In March 2026 a research team introduced EnterpriseArena, a benchmark that puts a language model in the CFO's chair of a simulated consumer lending company and leaves it there for 132 monthly steps — eleven years[1]. Each month the agent can close the books, buy a costly market signal, or request equity or debt financing. Macroeconomic conditions shift underneath it, and the consequences of any decision arrive months after the decision is made.
Eleven models were tested. Only 16% of runs survived the full horizon, and five of the eleven never survived a single trial[1]. Scale did not rescue them: in the reported results a 9-billion-parameter model outperformed a 397-billion-parameter one[1].
The human expert baseline survived 100% of the time.
That gap is wide enough to read as a verdict on digital finance. It is not one. Look at what the humans actually did. They spent 94.3% of their actions closing the books and 3.4% on fundraising. They used 0.2 tools per month against the agents' 2.2. They kept more cash underneath them throughout — a low point of $14.3M against $9.6M for the agents[1]. The humans did not win on brilliance. They won on bookkeeping. They knew where the company stood every month, and they acted only when conditions justified acting.
So the paper is not evidence that software cannot help a finance function. It is evidence that giving one free-form autonomous model eleven years of unsupervised financial authority is the wrong design. Those are separate claims, and the second one is the useful one. We built Kalends on the second one.
What a digital CFO actually is
Define the role by the work rather than by the title.
A digital CFO worth having is a financial intelligence layer. It keeps the picture current. It surfaces what has moved and what has broken a rule. It remembers across accounts and periods, so recall does not depend on one person's memory of the last eighteen months. It compares independent readings of the same evidence rather than trusting a single narrative. And it produces decision support in which every statement can be traced back to the figure it came from.
What it does not do is commit the company to anything. It does not raise capital, change credit policy, approve spending, or move resources between departments. Those are acts of authority, and authority carries accountability that cannot be assigned to a model.
Two terms from the paper are worth translating, because they explain why the distinction is load-bearing.
Partial observability means the agent cannot see the whole company at once. It sees what it has recently bothered to look at. In the simulator, closing the books is the act of refreshing that view — which is why the humans did it with 94.3% of their actions[1].
ReAct is the standard loop the tested agents ran: reason about the situation in free-form text, take an action, observe the result, reason again. It is flexible, and the flexibility is the problem. Nothing in the loop requires the reasoning to be true, sourced, or current. Whatever the model finds plausible becomes the plan, and the plan becomes the action.
A governed digital CFO breaks that chain in two places. It constrains what the analysis is permitted to say, and it stops the analysis short of the action.
What a governed digital CFO is good for
The practical value is unglamorous, and it compounds.
- Consistent attention. A period review that happens every period beats a deep analysis that happens when someone has time. The benchmark's clearest signal is that regular closing is what kept the human experts solvent.
- Faster exception identification. Liquidity pressure, margin movement, customer concentration, and deviation from plan are visible in the records before they are visible in the business. A single customer drifting from 12% to 31% of receivables is a fact about the ledger months before it becomes a conversation about risk.
- Recall that does not depend on memory. A comparison across accounts and prior periods should not rest on whether the controller happens to remember what freight ran last October.
- A shared record before the discussion. Reconciliation is worth more before the strategy meeting than during it. Most arguments about the numbers are arguments about which numbers.
- More than one reading. Independent analyses of the same evidence disagree in useful ways. A disagreement between two readings is a flag. A single confident narrative is not.
- Findings you can challenge. A finding is only useful if a finance leader can open it, see the figure it rests on, and say no. Unsourced analysis cannot be argued with, only believed or ignored.
None of that requires the system to decide anything. All of it increases the review capacity of the person who does.

The five ways the agents failed
The paper's case studies are more instructive than its headline number, and they map onto failure modes any finance leader will recognize.
Stale information. One agent ran an initial exploration in month 0 and never revisited its analysis[1]. Everything it believed about the company was, by the end, eleven years out of date. Its view was not wrong when it was formed. It was wrong by the time it mattered.
Disengagement that looks like compliance. That same agent recorded a 99.1% pass rate on its monthly checks[1]. It passed almost every month and still ran the company out of cash. A pass rate measures whether an action was well formed. It says nothing about whether the actor was still paying attention. Do not read one as the other.
Analysis without action. Another agent used forecasting and market tools heavily and closed the books 0.0% of the time[1]. It generated a great deal of outlook and almost no position. The paper's conclusion is blunt: extensive tool usage without timely action proved no better than complete inaction.
Neglect of the internal record. The pattern underneath both cases is the same. External signal is more interesting than internal reconciliation, so the agents bought forecasts and skipped the close. The humans did the opposite, and the humans survived.
No buffer before the turn. Most failures clustered around months 40 to 60, when external conditions shifted from favorable to adverse[1]. The agents were not destroyed by the downturn. They were destroyed by having spent the good years without building a cushion. Only one model built a comparable buffer and survived the shift.
Four of those five are failures of attention rather than of intelligence. That is the part worth designing around.
Detect, analyze, verify, ratify
Kalends runs recurring period reviews through four stages. Each exists because free-form reasoning fails in a specific way.
Detect — a deterministic exception engine. Before any language model reads anything, structured accounting metrics and rule checks run against the financial records. Threshold breaches, unusual movements, and relevant comparisons come out of software, not out of a prompt. This is the answer to stale information and to the neglected close: the analysis cannot begin until the records have been examined, and the examination is not optional.
Analyze — five independent seats. Five anonymous analytical seats read the same exception set, in parallel, without seeing each other's work. Each produces structured claim drafts rather than one unconstrained narrative. More seats do not automatically produce a better answer, and we do not claim they do. The value is that five independent readings of the same evidence can disagree, and disagreement is information.
Verify — a deterministic validation gate. This stage is software logic, not another language model. It drops claims that nothing supports, checks every cited amount against the financial packet, and binds each surviving statement to the evidence it came from. A claim that cannot be traced does not proceed.
Ratify — a sourced period review. What survives converges into one concise review, and every published statement keeps the evidence it was derived from. “Ratified” describes the system's verified consensus output. It is not a claim that a person has personally signed off on every line.
The claim vocabulary is deliberately narrow. A seat may publish an observation based on a cited figure, a threshold breach tied to a defined rule, or a comparison between cited movements. That is the entire list.
That is a structural exclusion rather than a prompt instruction a model might drift away from. The most dangerous sentence in financial analysis is a plausible unsupported explanation, because it is the one that gets repeated in the meeting.
How the design answers the benchmark
The benchmark and Kalends do not do the same job. What follows is a comparison of design choices, not a scoreboard.
| EnterpriseArena | Kalends |
|---|---|
| Tests whether one autonomous agent can make consequential resource-allocation decisions over eleven years | Produces governed financial analysis and sourced decision support for people |
| Uses a single ReAct agent with memory and organizational tools | Uses deterministic detection, five parallel analytical seats, deterministic verification, and one consolidated review |
| Allows free-form reasoning to drive actions | Restricts publishable claims to observations, threshold breaches, and comparisons |
| The agent may act while relying on stale or incomplete internal information | The analysis begins with structured financial records and deterministic exceptions |
| A plausible unsupported explanation can influence the agent's plan | Unsupported statements are dropped, and surviving claims remain connected to evidence IDs |
| One model's output becomes the action | The final output is an inspectable review intended to inform human judgment |
What this does not prove
This is the part to read carefully.
Kalends has not been tested in EnterpriseArena. We have no survival rate, no benchmark score, and no measured comparison against any of the eleven models the paper evaluated.
Our architecture does not establish a higher survival rate or superior long-horizon resource allocation. It is not built for that task. It produces analysis for people; the benchmark tests an agent making binding commitments alone. Nothing about the first is evidence about the second.
Architectural safeguards close specific avenues and no more. A deterministic gate that drops unsourced claims removes one route by which a plausible fabrication reaches a decision-maker. It does not make the analysis correct, and it does not make the system incapable of error. Architecture is a design argument, not real-world performance evidence.
The paper studies a simulated consumer lender. Real companies carry more complexity than any simulator: people who behave unexpectedly, rare events with no history to learn from, and information that never reaches the general ledger at all.
And the responsibilities that matter stay where they are. Strategic decisions, fiduciary duty, and accountability to a board are human. The paper's own authors name the single-agent decision structure as a limitation, observing that a real CFO operates within a hierarchy involving a board of directors and department heads, and they warn that strong performance on the benchmark could be misinterpreted as evidence that agents are ready for real financial or operational deployment[1]. Both points are right.
Disciplined attention, not unsupervised control
The near-term value of a digital CFO is not autonomy. It is attention.
The human experts in that study did not win by out-thinking the models. They won by knowing where the company stood every month, and by acting only when the evidence supported acting. That is a discipline, and holding a discipline every period without drifting is something software is genuinely good at.
So the useful question is not whether an AI system can run a finance function. It is whether your finance function has current books, surfaced exceptions, recall across periods, more than one reading of the evidence, and findings a leader can challenge before the meeting starts. Build that, and the judgment stays where it belongs.
References & Academic Sources
- [1] Han, Y., Wang, Y., Qian, L., Li, H., Cao, Y., He, Y., Peng, X., Shen, N., Xu, Y., Chen, Y., Feng, D., Huang, J., Liu, X., Nie, J.-Y., & Ananiadou, S. (2026). Can LLM Agents Be CFOs? A Benchmark for Resource Allocation in Dynamic Enterprise Environments. arXiv preprint arXiv:2603.23638v1, March 24, 2026. [arXiv:2603.23638v1]↩
Every figure quoted above is taken from version 1, dated March 24, 2026. arXiv also serves a later revision dated May 16, 2026, retitled Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment, which reports 15.4% survival across 23 models and four agent frameworks. We cite v1 because it is the version this article was written against.
Frequently Asked Questions
What is a digital CFO?
A financial intelligence layer rather than an executive. A useful digital CFO keeps the financial picture current, surfaces material exceptions, holds recall across accounts and prior periods, compares independent readings of the same evidence, and produces decision support in which every statement can be traced back to the figure it came from. It does not commit the company to anything.
Can an AI CFO replace a human CFO?
No. Strategic judgment, fiduciary responsibility, and accountability to a board are human responsibilities and cannot be transferred to a model. The EnterpriseArena benchmark is a direct illustration: across eleven language models given eleven simulated years of autonomous financial authority, only 16% of runs survived the full horizon, while the human expert baseline survived every time. Kalends is decision support for a finance leader, not a substitute for one.
How is Kalends different from a single AI finance agent?
A single agent reasons in free-form text and then acts on that reasoning. Kalends separates the two. A deterministic exception engine examines the financial records before any language model reads them, five independent analytical seats draft structured claims in parallel, a deterministic validation gate drops unsupported claims and binds the rest to evidence, and the result is a review for people to inspect. The value comes from independent analysis plus deterministic verification, not from agent count.
Does Kalends make financial decisions automatically?
No. Kalends does not raise capital, change credit policy, approve spending, or allocate company resources. It produces a sourced period review intended to inform human judgment. The output is an analysis you can open, inspect, and disagree with.
