
How Many Agents Can One Finance Team Actually Supervise?
Supervision is a finite resource. A controller has a certain number of hours, a certain tolerance for context-switching, and a certain amount of attention that can be spent reading things carefully rather than quickly. Every agent placed into production draws on that pool. Deploy past it and oversight tends not to fail loudly — it degrades into approval without reading, which is the worst of the available outcomes.
The Question Is Being Asked Backwards
Most agent-deployment plans count agents. The unit that matters is not agents. It is exceptions, and the reviews they generate.
Take two hypothetical agents. The first handles two thousand transactions a month and escalates two per cent of them, producing forty items that need human judgement. The second handles four hundred transactions and escalates fifteen per cent, producing sixty. The second looks smaller on the measures a business case usually tracks — transaction volume, licence cost, process count — and consumes half again as much of the thing that is actually scarce.
(Those figures are illustrative, not drawn from a client engagement.)
This is why “how many agents” has no general answer, and why benchmarking against another organization’s agent count tells you very little. Two finance teams running the same number of agents can sit at quite different points relative to their ceiling. It is also why the sequencing question — which process to automate first — and the capacity question are really the same question approached from opposite ends.
What One Agent Actually Costs in Human Attention
The review of escalated exceptions is the visible cost, and it is rarely the only one. Four others are easy to leave out of a plan.
- Exception review. The residual the agent could not resolve, arriving at a rate that is difficult to predict accurately before go-live.
- Confidence monitoring. Somebody has to notice when performance drifts. An agent that was right most of the time in March and rather less often by July has not announced the change.
- The named supervisor role. In some shipped Dynamics 365 agents this is not a metaphor but a configured role. Microsoft’s Payables Agent in Business Central designates specific users as agent supervisors, and involves them when the agent is not confident how to register an invoice — with review depth depending on the agent’s configuration and its confidence in its own suggestions. It also holds certain actions pending a person: a vendor the agent creates is left blocked until someone unblocks it. Note this is a Business Central capability, and Finance and Operations customers have a different agent set, so do not carry the specifics across.
- Configuration drift. Thresholds, mappings and boundaries need revisiting as the business changes. This rarely gets scheduled.
- Audit evidence. Producing a defensible record of what was reviewed, by whom, and on what basis is work in its own right.
The design of that Payables Agent supervisor role is worth dwelling on, because it makes the shape of the cost explicit. Supervisory load is not a flat per-agent overhead. It is driven by how often the agent is uncertain — which means the same agent can be cheap or expensive to oversee depending entirely on the quality of what you feed it.
What Actually Determines the Ratio
Five variables move the number far more than agent count does.
- Exception rate. The dominant term, and mostly a consequence of data quality and process standardisation rather than of the agent itself. Reducing it is usually the highest-leverage move available, though not always — see reversibility below.
- Reversibility. An action that can be corrected afterwards permits asynchronous review: a person looks when they get to it. An irreversible action requires approval before it happens, which is synchronous, interruptive, and far more expensive per instance. This is a property of the action and of how the process around it is designed, which means it is more changeable than it first appears.
- Homogeneity. Five agents doing similar work in the same domain share a supervisor’s mental model. Five agents doing unrelated work in unrelated domains each demand their own. Supervisory load scales with variety, not only with volume.
- Observability. If aggregate performance is visible and trustworthy, review can happen at portfolio level. If it is not, the only available assurance is sampling individual transactions, which does not scale. This depends partly on what telemetry the platform exposes and partly on what you build around it.
- Blast radius. How far a mistake travels before someone catches it determines how much pre-emptive checking the process needs.
What is worth noticing is how few of these are fixed properties of the agent you bought. Most are consequences of your data, your process design and your instrumentation — which is why supervisory capacity is largely an engineering and governance outcome rather than a staffing one.
The Failure Mode Worth Designing Against

When review load exceeds review capacity, approvals do not usually stop. They get faster.
This is the outcome to design against, because it is hard to see in the metrics a programme typically reports. Throughput looks healthy. Approval rates look healthy. Exception queues are clearing. What may actually be happening is that someone is approving items they have not meaningfully examined, and the audit trail is recording an oversight step that did not really occur.
That is arguably worse than leaving the process unautomated. An agent running without supervision is a known risk that can be bounded and disclosed. An agent with documented but hollow supervision is an unrecognised risk wearing evidence of control, and it may not surface until something goes wrong and somebody reads the log carefully.
The tell is behavioural rather than a threshold: review time per exception falling while exception volume rises. If nobody measures how long a review actually takes, that signal is simply unavailable — which is an argument for instrumenting it before you need it.
Raising the Ceiling

The ceiling is not fixed, and none of what follows is an argument for deploying fewer agents. It is an argument for spending on capacity before spending on scope.
- Reduce the exception rate before adding the next agent. Often the highest-leverage move, and rarely sequenced first, because it looks like unglamorous data work rather than AI work.
- Buy reversibility where the process allows it. Designing so mistakes are catchable afterwards converts expensive synchronous approval into cheap asynchronous review. Where an action is genuinely irreversible and high-value, that gate should stay — the question is how many such gates the design actually requires.
- Cluster agents by domain under one supervisor, so a single mental model covers several agents rather than one each.
- Make the supervisor role explicit, named and resourced. Supervision assigned as a side-of-desk duty to whoever is least busy is supervision that gets skipped in a close week. Which specific decisions should stay with a person is a separate question, and one we have written about in where human oversight still belongs.
- Instrument review quality, not only review completion. Time-per-review and override rates indicate whether oversight is real. Approval counts do not.
- Define the criteria for withdrawing autonomy, not only for widening it. Most governance frameworks specify how an agent earns more scope. Fewer specify what triggers taking it back, and that asymmetry is where capacity quietly gets overcommitted.
Some agents also reduce supervisory load rather than adding to it. An agent that clears work a person was already reviewing by hand can be net negative on attention, and those are worth identifying early — though the honest caveat is that escalation rates are hard to predict precisely until the thing is running.
A More Useful Question
“How many agents can we supervise” invites a number that will not survive a close week. Three better questions:
- How many exceptions per month can this team review properly, at the current standard of properly?
- What is each candidate agent’s expected escalation volume, and what happens to the total when it is added?
- Which of these actions can be corrected afterwards, and which cannot?
Those are answerable with evidence, and the answers produce a deployment sequence rather than a target count.
Where This Leaves a Finance Leader
The instinct when agent capability arrives is to ask what else could be automated. The more productive question, at least for the next few deployments, is what would have to be true for the team to supervise twice as much — because that answer sets the pace of everything afterwards.
That is the work DAX Software Solutions tends to do before agent deployment rather than after it. Our AI readiness assessments look at ERP stability, data quality, integration and governance before agent work is scoped. Our operating-model work then settles the questions that actually determine the supervisory ratio: what an agent is permitted to do unattended, what it must put in front of a person, who is accountable for the outcome, and what evidence gets retained. Set those deliberately and the ceiling is high. Leave them implicit and it is wherever the busiest week puts it.
Agent capability in Microsoft Dynamics 365 is moving quickly. Specific agents, boundaries and processing limits change with each release, so verify current behaviour against Microsoft Learn rather than against any article, including this one. What is unlikely to change is the underlying constraint: the number of agents a finance team can run is the number it can genuinely oversee, and that number rises with better data and clearer accountability long before it rises with headcount.
If you are sizing a first or next agent deployment, the supervisory arithmetic is worth doing before the licence conversation, not after it. Get in touch with DAX Software Solutions if you would like help doing it honestly.