We Quote Tasks, Not Roles
The request arrives as a headcount question and the answer is never a headcount answer. A role is a bundle of tasks held together by a person who knows which one to do next, and that binding is the part nothing we build replaces.
The Question and the Better Question
The request was to replace a half-time back-office position with an agent. It is a legitimate business question and it is unanswerable as posed, because nobody, including the person doing it, has a written description of what that half-time actually consists of.
So we spent a day with them writing one. Eleven distinct tasks, with rough frequencies and durations. That list is the deliverable of the first phase now, and on three occasions it has ended the project there, which we consider a good outcome.
What the List Showed
Four tasks were clearly automatable: retrieving documents from a portal, extracting fields from delivery notes, reconciling two reports, and producing a weekly summary. Together they were about forty percent of the time.
Three were work that had drifted from other departments and mostly needed a process change rather than software. Four required judgement about exceptions, and each of those four was the reason the position exists, because the routine part is exactly what a person would automate first if they had the tools.
The Benchmark That Says This Plainly
Xu and colleagues published TheAgentCompany that December, a benchmark placing agents in a simulated company and measuring how they perform on consequential tasks resembling real work, with results showing a substantial share of tasks not completed.
We now quote numbers from work of this kind in proposals rather than demonstration videos. A benchmark that reports partial completion on realistic office tasks is a far better basis for a customer expectation than a recording of a task that went well.
| Task type | What we propose |
|---|---|
| Repetitive, verifiable, high volume | Automate. This is the four |
| Repetitive, no way to check the result | Automate with review, or not at all |
| Judgement about exceptions | Assist the person, do not replace |
| Coordination between people | Leave alone. This is the binding |
The Coordination Nobody Prices
A person doing eleven tasks decides the order, notices when one is unusual, and carries context between them. Split four of those tasks into automated steps and something has to reproduce that binding: someone still checks the queue, handles what the automation refused, and knows that a late delivery note means chasing a supplier.
In our experience that residue is between a fifth and a third of the time saved. We include it in proposals as a line item, because a customer who expects forty percent and receives twenty-eight percent has been misled by arithmetic that omitted the part that stayed human.
What We Told This Customer
That we could take about forty percent of the tasks and realistically deliver a net saving closer to a quarter, that the position should not be reduced but redirected toward the exception work, and that the four judgement tasks would get better rather than cheaper because the person would have time for them.
They went ahead on that basis. Eighteen months later the automated share is roughly what we projected, and the half-time position still exists, which we count as the proposal having been honest rather than as a failure to deliver.
Why We Refuse the Role Framing
Because it produces systems that fail invisibly. If a project is scoped as replacing a person, the automation is expected to handle everything that person handled, including the cases nobody wrote down, and it will handle them by producing plausible output rather than by stopping.
Scoped as tasks, each automated step has a defined input, a defined output and a defined behaviour when it cannot proceed. The unhandled cases arrive at a person by design instead of being absorbed silently, and the difference between those two arrangements is most of what separates a working system from an incident.
The One Case That Went the Other Way
A customer with a genuinely narrow role: one person, one repetitive task, high volume, fully verifiable. There the role framing and the task framing coincided, and we said so rather than inventing complexity to justify a longer engagement.
That is worth naming because our position can sound like a consultancy protecting scope. The distinction is empirical: count the tasks. If there is one, the question was fair. If there are eleven, the question was about a bundle that nobody had examined.
What We Do Not Claim
We do not claim agents cannot do office work. The four tasks we automated are office work, they run daily, and they are reliable. The claim is about scope: the unit that automates well is a task with a checkable result, not a job title.
We also do not claim our fifth-to-a-third coordination figure is precise. It comes from a handful of projects, we measure it after the fact rather than predicting it well, and it is the number we are least confident about in any proposal we write.
