01
Why this, why now
Type "jobs AI can't replace" into a search bar and you get a hundred confident lists: nursing, plumbing, therapy, the trades. They are not wrong, exactly. We do not find them the useful answer either, because whole jobs are the wrong unit and the interesting thing happens one level down, at the capability.
We published the first version of this dossier in July. Three weeks later a16z put out benchmark data that made one of its central comforts untenable, and Anthropic published a survey that made another of its arguments considerably stronger. Both are in this edition. The short version is that the machine got better at the doing and no better at the vouching, and the gap between those two things is where the human work went.
This dossier draws on the field experiment that coined the jagged frontier, on Anthropic's Economic Index tracking how people actually use these systems, on a meta-analysis of 106 studies of human-AI teams, on McKinsey's mapping of which skills automation touches, and now on a set of August measurements of what agents can and cannot finish inside a real company. The picture is less comforting than the listicles and more useful. It tells you what to build, what to stop defending, and where a human still decides.
02
Contents
1. Past the listicles
2. The jagged frontier
3. The frontier moved, and it moved fast
4. The mode most people actually use
5. When a human helps, and when a human hurts
6. The arithmetic of vouching
7. What the skills data shows
8. The durable core
9. Presence, trust and the things that need a body
10. Tacit knowledge, and the people who rate it highest
11. The pivot: from making the answer to judging it
12. Where we could be wrong
13. What this means for a Dutch business
14. Verification and sources
03
1. Past the listicles
The safe-jobs lists are comforting and mostly beside the point. They sort by profession, when the thing that gets automated is the task. A radiologist and a paralegal both do work that is partly deeply human and partly a rote lookup a model now does in seconds. The question we would ask is not which title is safe. It is which capability keeps earning as everything text-shaped gets cheap.
That reframing is the whole reason we wrote this. Once you stop defending job titles and start looking at capabilities, the evidence points somewhere unexpected: away from the prestigious cognitive work the twentieth century taught us to prize, and towards presence, tacit judgment, taste, and the plain fact of being accountable for a decision. Getting there means walking through the best studies available, starting with the one that showed the same tool making the same people both better and worse.
04
2. The jagged frontier
In 2023 a team from Harvard, Wharton, MIT and BCG ran what we still regard as the cleanest experiment on this. They gave 758 BCG consultants realistic consulting tasks, half with a top model, half without, and measured what happened. The work was peer-reviewed in Organization Science in 2025, which matters, because most of what circulates on this subject never is.
Inside what the model was good at, the effect was dramatic. Consultants with AI finished more tasks, worked faster, and produced work rated around 40 percent higher in quality. The lowest performers gained the most, so the gap between the best and the rest narrowed.
It does not end there, and the second half is the half that survived. The same researchers planted tasks sitting just outside the model's competence, close enough to look identical from the outside. On those, consultants using AI were about 19 percentage points more likely to reach the wrong answer than colleagues working without it. The model was fluently, confidently wrong, and the humans went along with it.
The authors named the shape the jagged frontier, and they report it as a general property rather than a quirk of one model: AI is brilliant at some tasks and quietly hopeless at others that look no harder, and the boundary is invisible unless you already know the work. Knowing where that line runs is itself a human skill. It is the one this whole dossier turns on, and it is the one that got more valuable this summer rather than less.
05
3. The frontier moved, and it moved fast
Here is the finding that should unsettle anyone who read the July edition and felt reassured.
On 10 August 2026 a16z published measurements of how well agents can actually operate a computer. On OSWorld-Verified, a benchmark of real tasks in real desktop applications, agents completed 85 percent of tasks, up from 42 percent a year earlier. Humans attempting the same tasks complete about 72 percent.
We read that pairing slowly, because the ordering is new. On this particular measure the machine is no longer approaching human performance. It has passed it, by a margin of roughly 13 points, on a benchmark specifically built out of ordinary office work.

Anyone whose confidence rested on "it still cannot really use the tools" needs a different argument. We leaned on that comfort in July more than we should have, and we would rather say so.
I went looking for that 72 percent, because a human baseline is the sort of number that gets quoted for years after nobody remembers how it was produced, and it does most of the work in the comparison. It comes from the same source as the agent figure, and I could not find it independently replicated. Human baselines on benchmarks like this are notoriously soft: they depend on whether the humans were experts or crowdworkers, whether they were paid by the task or by the hour, how long they were given, and whether they were allowed to look things up. So I would hold the 13-point margin loosely and the direction of travel firmly. An agent going from 42 to 85 in twelve months is not a measurement artefact, whatever the humans scored.
Two things stop this from settling the question, and both are more interesting than the benchmark.
The first is arithmetic, and section 6 does it properly. The second is what a16z found when it looked at why deployments fail anyway, which was not capability. a16z write that the model is no longer the main bottleneck; what decides the outcome is context, permissions, process knowledge, validation, escalation and error handling. The hard part, they argue, is not whether an agent can navigate an SAP screen. It is whether it understands how a particular company actually gets work done: the tribal knowledge, the internal terminology, the preferred formats, who to escalate to.
That is a benchmark-beating system failing on exactly the thing section 10 of this dossier calls tacit knowledge. The frontier moved. It moved past the part we were measuring, and stopped at the part we never wrote down.
06
4. The mode most people actually use
There is a lazy assumption that using AI means handing work to a machine and walking away. The usage data says otherwise, though the picture has moved since the July edition and deserves restating with its dates attached.
Anthropic's Economic Index measures how millions of real conversations unfold. In its January 2026 report, covering November 2025 data, the share of conversations classified as augmentation rose 5 percentage points to 52 percent while automation fell 4 points to 45. The March 2026 report found augmentation rising slightly again, in both the consumer product and API traffic. The June 2026 report does not publish a single headline split, so the honest statement is that augmentation has been gaining modestly for three consecutive reporting periods, and that we do not have a current number.
The distinction matters more than the exact percentage. Automation is one prompt and one answer: classify this email, extract these fields, accept the output. Augmentation is a loop, where the person asks, pushes back, corrects, iterates, and keeps the decision. Most knowledge work sits in the second mode, and the second mode is precisely where human judgment still moves the result.
There is a wrinkle underneath the aggregate that the July edition missed. API traffic is automation-dominant, and the share of transcripts tied to office and administrative support work rose 3 points to 13 percent by November 2025. So the consumer aggregate drifting towards augmentation and businesses quietly automating the back office are both happening, in the same dataset, at the same time. Which number you quote depends on which population you meant.
The June 2026 report adds texture that is easy to skip past and worth not skipping. Anthropic reports that chat and collaborative sessions run about 35 percent personal use on weekdays and spike to roughly 50 percent at weekends, which is a small finding with a large implication: a meaningful share of what looks like enterprise adoption is people using a work tool for their own lives, and any adoption figure that does not separate the two is measuring something other than what it claims. The same report notes that women make up 12 percent of its linked sample, use the coding product 7.3 percentage points less often, and show an automation share 7.3 points lower even after controlling for occupation. Anthropic does not offer a mechanism for that, and neither will we, but a gap that survives occupational controls is not a composition effect and it is not nothing.
The meta-analysis of human-AI teams adds a twist worth sitting with. Across 106 experiments, the teams gained most on tasks that involved creating something and lost ground on tasks that involved deciding something. The human edge is real, and it is uneven in a specific direction.
07
5. When a human helps, and when a human hurts
Here is the finding we would put in front of anyone who treats human-in-the-loop as a safety guarantee.
Averaged across those 106 studies, which Vaccaro, Almaatouq and Malone report in Nature Human Behaviour, human-AI combinations performed worse than the better of the human or the AI working alone. Putting a person next to the machine did not reliably help. Sometimes it hurt.
The pattern underneath is clean once you see it. Where the human was genuinely better at the task, the team gained. Where the AI was better and the human overrode or second-guessed it, the team lost. A reviewer who rubber-stamps whatever the model produces adds nothing but latency. A reviewer who cannot tell a good answer from a confident one adds risk.
Oversight, in other words, is a real skill with a real failure mode, and it does not arrive free with the job title. The people who stay valuable are the ones who know, task by task, whether they are the better judge. The uncomfortable version of the same sentence: a human in the loop can add negative value.
08
6. The arithmetic of vouching
Now the sum that the 85 percent headline hides, because we think it is the most useful calculation in this dossier and it takes one line.
A business process is a chain. Every step has to work for the process to finish. If an agent completes any given step 85 percent of the time, then a five-step process finishes about 44 percent of the time, and a ten-step process finishes about 20 percent of the time. The same agent, the same benchmark-beating reliability, and four in five processes fall over somewhere in the middle. a16z put it plainly: 85 percent still means 15 of 100 failed, and back-office work does not grade on a curve.
That is why the human did not disappear when the benchmark was passed. Somebody has to catch the fifteen.
And catching them is harder than it sounds, because the failures cluster where success cannot be checked immediately. a16z give the example of an insurance claim where you only learn the outcome days later, on a callback. A task that fails loudly is cheap. A task that fails silently, and is discovered three steps downstream by a customer, is the expensive kind, and it is exactly the kind an agent produces most confidently.
Which leads to the sentence from that piece that we would take as our thesis, and which arrived from an unrelated direction: the scarce resource is no longer writing the code, it is vouching for it.
There is a cost side too, and we will state it plainly because it sets the pressure everything else operates under. a16z estimate the running cost of a computer-use agent at roughly $6 to $8 an hour, in a range of $3 to $15, against about $10 an hour for offshore business-process outsourcing and $30 to $45 for US back-office labour. The economics are not marginal. They are decisive for the doing, and silent on the vouching, and that asymmetry is the whole shape of the next few years.
09
7. What the skills data shows
Zoom out from single tasks to whole skills and the same shape appears. The McKinsey Global Institute reports a Skill Change Index scoring how exposed some 6,800 skills are to automation over the next five years, published in November 2025. Where a skill sits on that curve is telling.
At the low-exposure end sit leadership, negotiation, coaching, management and communication: the skills that run on reading people and holding responsibility. At the high-exposure end sit invoicing, inventory management, data entry and routine coding: the skills that run on processing information to a defined output.
McKinsey's own framing we quote because it is more careful than most summaries of it. AI will not make most human skills obsolete, they argue, but it will change how they are used, and negotiation, problem solving and leadership may matter more as people work alongside agents and robots.
That last clause deserves attention, because it is a claim about direction rather than survival, and it is where the index gets interesting. A skill can be low-exposure and still lose value if the work around it disappears; a skill can be high-exposure and gain value if what remains of it becomes the bottleneck. Routine coding sits at the exposed end of McKinsey's curve, and the demand for people who can review code has, on the evidence of section 6, gone up rather than down. Exposure measures how much of a task a machine can perform. It does not measure what happens to the price of the part it cannot.
The skills that survive, on both readings, are the ones where the point was never the raw output.
10
8. The durable core
Pull the threads together and we are left with a short, specific list. Capabilities, not professions, which is why a plumber and a board chair can be durable for the same underlying reason.
They cluster into four.
Presence. Work where a body in a room, or the trust of another person, is the product. Care, therapy, the trades, the difficult conversation.
Tacit judgment. The know-how that lives in experienced practitioners and was never written down, so no model trained on it.
Taste and creation. The ability to originate, and to know what is good. This is where the meta-analysis found the human edge largest.
Accountability. Someone who will stand behind a decision when it goes wrong. No model can do this, not because of capability but because responsibility is a relationship between people, and a model cannot be a party to it.
Everything durable in the studies maps onto one of those four. The fourth has become more load-bearing since July, for the reason section 6 sets out: when the doing costs $7 an hour and the vouching does not, the person who can vouch is the person holding the valuable end.
11
9. Presence, trust and the things that need a body
The most automation-proof work in every ranking shares a feature that has nothing to do with intelligence. It needs a human body or a human bond to happen at all.
Anthropic reports the same pattern in reverse in its labour-market work. The most AI-exposed jobs are computer programmers, customer service representatives, data-entry clerks and financial analysts: higher-paid knowledge roles where the work is information in, information out. The least exposed are cooks, mechanics, lifeguards, bartenders and care workers, where hands and physical presence are the work.
This inverts the twentieth-century intuition that cognitive work was the safe, prestigious ground and manual work was precarious. For this wave of technology the opposite holds. A model can draft the financial analysis. It cannot rewire the house, calm a frightened patient, or be the person a customer trusts.
We should be precise about why, because "AI can't do physical work" is the wrong reason and it will date badly. Robotics is improving quickly, and a decent share of what a mechanic does will eventually be mechanically feasible. The durable part is narrower and stranger than dexterity. It is that for a whole class of work, the human being is not performing a service so much as constituting it. A frightened patient who is calmed by a machine that sounds calm has not been reassured in the sense that matters; a negotiation conducted on your behalf by something that cannot be held to its word is not the same transaction. Presence resists automation not because the movements are hard to copy but because the value was never located in the movements.
Which is also why we make this section narrower than the reassuring version. The trades are safer than financial analysis for now, and a good part of that margin is a robotics timeline rather than a permanent human advantage. The part that is permanent is the trust, and trust is a smaller share of most jobs than the people in them would like to believe.
12
10. Tacit knowledge, and the people who rate it highest
There is a deeper reason some expertise resists automation, and it is almost mechanical. These models learn from what has been written down. Tacit knowledge, by definition, was never captured on a page. It is the judgment a surgeon builds over ten thousand cases, the feel a mechanic has for an engine that sounds slightly wrong, the read a negotiator has on a room. It transfers through apprenticeship and repetition, so there is no corpus to absorb.
Until this summer that argument rested on reasoning rather than measurement. Anthropic's June 2026 Economic Index report, drawing on survey responses collected between mid-May and early June alongside usage logs, supplies a measurement, and it is a neat one. Workers with 15 or more years of experience estimate that AI can handle roughly 10 percentage points fewer of their tasks than workers in their first year estimate about theirs.
Two readings compete and both are worth holding. The generous one is that experienced practitioners are correct: they have accumulated judgment that genuinely does not automate, and they can see it from the inside. The unkind one is that senior people overestimate their own irreplaceability, which would be an unremarkable finding about human beings. Anthropic notes the first reading, and the a16z deployment data leans the same way, since the thing that actually broke those deployments was company-specific process knowledge rather than reasoning ability. Two independent sources pointing at the same unwritten stuff is not proof. It is better evidence than the argument had in July.
One AI researcher put the labour version of this sharply on the Dwarkesh podcast in December 2025: human labour is valuable precisely because it is not cheap to train. The skills that took a person a decade to acquire, and cannot be handed over in a prompt, stay scarce. The paradox is that the more AI commoditises explicit, written-down knowledge, the more the unwritten kind is worth.
13
11. The pivot: from making the answer to judging it
If there is one shift to take from all of this, it is a move in what a valuable person does.
For most of the knowledge economy, value meant producing the answer: the analysis, the draft, the code, the plan. AI is very good at producing answers and getting cheaper at it by the month. The scarce skill is one step up, at judging answers: knowing which is right, which is confidently wrong, and which question was worth asking.
What is new since July is that we are no longer inferring this from a set of studies. It is what practitioners report when their deployments fail, in a piece written by people with money on the outcome. When a16z write that the scarce resource is vouching rather than writing, they arrived at this dossier's conclusion from the deployment side, which is a better test of an idea than agreement from the same direction would be.
This is the See, Understand, Adopt discipline pointed at a person's own role. See where the frontier actually runs for your work, honestly, task by task, and update it, because section 3 is what happens when you do not. Understand where your judgment beats the machine and where the instinct to override is ego. Then adopt a working style where you orchestrate and check rather than manually produce, with the human contribution concentrated exactly where the machine is weak.
The people this goes well for are not fighting the tool. They have climbed to the part of the work it cannot reach.
14
12. Where we could be wrong
A dossier that only reassured would be dishonest, so we put four counterweights against our own case, one of which cuts against its own companion volume.
The frontier moves, and section 3 is the proof. Tasks safely outside it in 2023 are inside it now. A benchmark that agents lost by 30 points a year ago they now win by 13. Durability is relative and temporary. Any capability listed as safe in section 8 should be re-tested annually, and this dossier will be wrong about something within a year.
The labour-market panic may be ahead of the evidence. The Budget Lab at Yale compared AI-exposed and unexposed occupations and found that there was no clear indication of labour-market effects attributable to AI, with no detectable rise in aggregate unemployment among highly exposed workers since ChatGPT launched. I could not open the full paper, so I am relying on its published summary and would hold the detail loosely. Several other reviews reach similarly null or modest aggregate conclusions. That is a real check on the displacement narrative, including on parts of our own entry-level dossier, and it belongs here rather than in a footnote.
The people closest to the tools are the least worried, which is awkward for everyone. Anthropic reports that survey respondents delegating the most work to Claude expressed the highest optimism about their pay, job security and prospects, and that self-rated job-loss probability came in around 10 percent, slightly below actual US annual separation rates of roughly 13 percent. There are at least two readings. Heavy users may see genuine complementarity that non-users cannot. Or heavy users may be selected for enthusiasm and poorly placed to judge their own exposure. The survey cannot separate these, and neither can we.
The bottom rung is still the weak point. Concern about young workers is concentrated in the same data, and entry-level knowledge roles are exactly where tacit judgment used to be built. If the rungs disappear, the durable senior judgment of the future has nowhere to grow from. That argument survives the Budget Lab check, because a null aggregate effect is compatible with a sharp compositional one.
Our position holds the hopeful and the hard together. A specific set of capabilities is genuinely durable, and that is real ground. The line defining it moves, unevenly and often faster than is comfortable, and standing on durable ground means moving with it.
15
13. What this means for a Dutch business
We do not end on a list of safe jobs to hire for. We end on a way to design work.
Put people where the four durable capabilities live, on the judgment, the relationships, the tacit calls and the accountability, and let the machine carry the produce-the-answer load underneath. The mistake is the other way around: keeping people busy on output that AI now floods, and bolting a token human review onto the end. That is precisely the negative-value human-in-the-loop the research warns about, and it is the most common shape we see in practice.
Four moves we would make, and the fourth is new this quarter.
We would redesign roles around judgment and presence rather than throughput. The unit of redesign is the task, not the job title, which is the same reason section 1 exists.
Invest in tacit, apprenticed skills that cannot be prompted, and protect the entry-level rungs where they are learned, even when automating them looks cheaper this quarter. This is the point where a short-term cost decision does long-term structural damage, and it is nearly always taken by someone who will have moved on before the damage is visible.
Train people in the genuinely new skill of oversight: how to tell a good answer from a confident one, task by task. This is teachable, it is currently taught almost nowhere, and section 5 says the alternative is a reviewer who adds risk.
And we would count the vouching, not the doing, in the business case. If an automation removes ten hours of production and adds six hours of checking, that is the real number, and it is the number that decides whether the project pays. Ask specifically what happens when a step fails silently and who finds out.
Our move: design work so people spend their time on presence, taste, tacit judgment and accountability, the four things AI cannot flood, and let the machine carry the output beneath them. Then check, every year, that the four are still four.
16
14. Verification and sources
This dossier draws on live web research and a personal archive of more than 15,000 sources. The notes below flag confidence and the material caveats.
| Claim | Confidence | Note |
|---|---|---|
| Jagged frontier: 758 BCG consultants; ~40% higher quality inside the frontier, ~19pp more errors outside | High | Dell'Acqua et al. (Harvard, BCG, MIT, Wharton), 2023; peer-reviewed in Organization Science, 2025. |
| OSWorld-Verified: agents at 85% task completion, up from 42% a year earlier, against ~72% for humans on the same tasks | Medium-high | a16z, 10 August 2026. A benchmark, not a workplace: OSWorld tasks are bounded and individually scored, which is exactly the property the compounding argument in section 6 exploits. The human baseline comes from the same source and we have not seen it independently replicated. |
| A process of n steps at 85% per-step reliability finishes 0.85ⁿ of the time: ~44% at five steps, ~20% at ten | High as arithmetic, illustrative as a model | Our calculation. It assumes steps are independent and equally reliable, which real processes are not; retries, checkpoints and easier steps all push the number up. Treat it as a shape, not a forecast. |
| Agent running cost ~$6-8/hour (range $3-15) against ~$10/hour offshore BPO and $30-45/hour US back-office labour | Medium | a16z, 10 August 2026. Vendor-adjacent source with an interest in the comparison; the direction is well supported, the precise levels less so. |
| Deployment bottleneck is context, permissions, process knowledge, validation, escalation and error handling rather than model capability | Medium-high | a16z, 10 August 2026, from deployment experience rather than controlled study. |
| Augmentation ~52% vs automation ~45% on the consumer product, November 2025 data | High for its period, not current | Anthropic Economic Index, January 2026 report. The March 2026 report found augmentation rising slightly again; the June 2026 report publishes no single headline split. There is no current published figure, and the July edition of this dossier presented the November number as though it were one. |
| API traffic is automation-dominant; office and administrative support tasks rose 3pp to 13% of transcripts by November 2025 | High | Anthropic Economic Index, January 2026 report. |
| Workers with 15+ years of experience estimate AI can handle ~10pp fewer of their tasks than first-year workers | Medium-high | Anthropic Economic Index, June 2026 report, survey run mid-May to early June 2026. This is self-assessment, not measurement of what actually automates, and it is equally consistent with accurate tacit expertise and with senior overconfidence. Reported here with both readings. |
| Heavy delegators report the highest optimism on pay, job security and prospects; self-rated job-loss probability ~10% against US separation rates ~13% | Medium | Anthropic Economic Index, June 2026 report. Self-selected sample of Claude users; not a general workforce survey. |
| Human-AI teams on average worse than the best of either; gains on creation, losses on decisions | High | Vaccaro, Almaatouq & Malone, meta-analysis of 106 experiments, Nature Human Behaviour. |
| Skill Change Index: leadership, negotiation, coaching at low exposure; invoicing, data entry at high | High | McKinsey Global Institute, "Agents, robots, and us", November 2025, covering ~6,800 skills. |
| Most AI-exposed jobs (programmers, customer service, data entry) versus least (cooks, mechanics, care) | High | Anthropic labour-market impact report, 2026. |
| No clear aggregate labour-market effect of AI detectable to date; no measurable rise in unemployment among highly exposed occupations since late 2022 | Medium | The Budget Lab at Yale, plus several concurring reviews. We could not open the full paper and are relying on its published summary, so the detail should be held loosely. Included because it cuts against this dossier's own direction and against parts of our entry-level dossier. |
| Young workers squeezed first; entry-level hiring in high-exposure roles down | Medium | Anthropic and concurrent labour research, 2025-26. In tension with the Budget Lab finding above; a null aggregate effect is compatible with a compositional one, which is how we read it. See the entry-level dossier. |
| "Human labour is valuable precisely because it is not cheap to train" | Framing, not a finding | Paraphrased from AI-researcher commentary, Dwarkesh podcast, December 2025. |
| AI-resistance rankings by field (therapy, trades, care highest) | Low-medium | Aggregated from 2026 career analyses; illustrative of the pattern rather than one authoritative source. Retained with that label. |
The framed figure in this dossier is the McKinsey Global Institute's own chart, credited in its caption. Charts labelled "BFF" are our own, drawn from the sources named beneath them.
Dossier as pdf
Download this dossier
The full dossier as a pdf, with every figure and the source list. Fill in your details and the download starts right away.