AI productivity is the headline nearly every organization wants, and HR is often the function asked to make it happen: rolling out AI tools across departments, retraining managers, and increasingly, running AI through its own processes, from screening resumes to drafting performance reviews. The numbers on AI’s productivity impact are real in specific ways. But a run of recent research, spanning agentic coding, controlled experiments with knowledge workers, and studies of how people actually learn while using AI, suggests that raw productivity gains understate the full cost of these tools, a cost that lands squarely in HR’s lap either way. Some of what looks like time saved is really time moved, from doing the work to judging whether the work was done correctly. Call it the discernment tax: the cognitive cost of deciding when to trust AI output, when to override it, and how to keep the expertise sharp enough to tell the difference.
The Productivity Case Is Real
Let’s start with what’s working. An analysis of roughly 400,000 coding sessions collected between October 2025 and April 2026 found that the tasks people hand to AI agents kept growing in value, with the average session worth about a quarter more by the end of the window than at the start. The share of sessions spent simply fixing broken output fell by nearly half, as usage shifted toward more complex, end-to-end work. A separate field experiment with knowledge workers at a management consulting firm found even larger gains on tasks well suited to AI: workers using the tool completed 12.2% more tasks, worked roughly 25% faster, and produced work rated meaningfully higher in quality than those working without it. On paper, this is exactly the AI productivity story most organizations expect to hear.
But Only Within a Boundary, and the Boundary Isn’t Visible
The same consulting experiment also complicates the story. Workers operating outside AI’s comfort zone, on a task deliberately designed to require judgment the tool didn’t have, were 19 percentage points less likely to reach the correct answer than workers without AI at all. What makes this genuinely tricky is that the AI-assisted answers, right or wrong, were rated as more coherent and persuasive than the human-only ones. AI doesn’t just get some tasks wrong; it gets them wrong in a way that reads as confident and well-argued, which makes the error harder to catch. Researchers call this a jagged technological frontier: AI capability doesn’t map cleanly onto how difficult a task looks to a human, so two tasks that feel equally hard can sit on opposite sides of where AI actually helps or hurts. Knowing which side of that line you’re on, task by task, is itself a skill, and it’s one the tools don’t teach.
Productivity Today Can Cost Skill Tomorrow
There’s a second, slower version of the same tax. A randomized trial with junior software engineers learning a new skill, some with AI assistance and some without, found that AI barely moved completion time but significantly changed how much people actually understood afterward. On a quiz covering material they had used minutes earlier, the AI group scored 17% lower than the group that worked unassisted, roughly the gap between a grade B and D. The steepest drop showed up in the questions that tested the ability to catch and diagnose errors, the exact skill needed to catch AI’s mistakes later on. Crucially, the outcome depended heavily on how people used the tool. Participants who leaned on AI purely to produce answers scored worst. Those who used it to ask follow-up and conceptual questions, treating it more like a tutor than a typist, scored as well as or better than people working without AI at all. The tool itself didn’t determine the outcome; how it was used did.
The New Division of Labor Puts Judgment at the Center
A look at how people actually use agentic AI tools day to day shows where this is heading. Across hundreds of thousands of real sessions, people retain about 70% of the planning decisions (what to do, what counts as done) while AI takes on roughly 80% of the execution decisions (how to do it). The people who get the most out of this arrangement aren’t necessarily the most technically fluent. They’re the ones with the deepest understanding of the problem itself: domain expertise, not tool proficiency, is what predicts whether a session succeeds. Sessions led by domain experts reach verified success more than twice as often as sessions led by novices, and when something goes wrong, novices abandon the task at three to four times the rate of everyone else. As AI absorbs more of the execution, the human role concentrates into exactly the part that’s hardest to automate: knowing enough to steer the work and to catch it when it goes sideways.
Naming the Tax
Put these findings together and a pattern emerges. AI genuinely speeds up and improves output on tasks within its capability. But it also shifts a real cost onto the human side of the ledger: the effort of figuring out which tasks are safe to hand over, the effort of catching errors that now arrive dressed up as confident, polished answers, and the slow erosion of the expertise that makes good judgment possible in the first place, if that expertise isn’t deliberately maintained. None of this shows up in a speed metric. All of it shows up in the quality of decisions made months later. That is the discernment tax, and it tends to be paid disproportionately by whoever in the organization is expected to supervise AI-assisted work without necessarily having had the chance to build the judgment to do it well.
The Same Pattern Is Already Inside HR
None of this is abstract for HR teams. AI already touches core HR workflows: screening resumes and ranking candidates, drafting job descriptions and policy language, summarizing engagement survey results, and increasingly, drafting first-pass performance reviews. Some of these tasks sit comfortably inside the frontier. Structured, rules-based screening against clear criteria is the kind of task AI tends to handle well, and the productivity gains there are genuine. Others sit outside it. Judging cultural fit, reading the nuance in an employee relations case, or weighing the human context behind a performance dip are exactly the kind of judgment calls where a confident, polished AI answer can be persuasive and yet quietly wrong. The same discernment tax facing engineering and consulting teams applies to HR’s own tools, and the same principle applies to managing it: know which HR tasks are inside the frontier and which aren’t, and build in a human check specifically where the stakes and the ambiguity are both high.
What This Means for HR and People Leaders
The practical implication isn’t to slow AI adoption down, in HR or anywhere else. It’s to stop measuring productivity purely as output and start building in the conditions that keep discernment strong, both across the organizations HR supports and inside HR’s own practice.
- Protect time to build judgment. Especially for people early in a role or a domain, unassisted practice builds the debugging and verification instincts they’ll later need to supervise AI, whether that’s an engineer, a manager, or an HR generalist.
- Train the interaction, not just the tool. Asking AI to explain its reasoning produces very different learning and retention outcomes than asking it to simply produce the answer. This applies to how recruiters, managers, and HR business partners are trained on AI tools, not just how engineers are.
- Hire and develop for domain expertise. The research is consistent that understanding the problem, not comfort with the software, predicts who gets good outcomes from AI. That holds for HR’s own hiring bar as much as for any other function.
- Build review points around higher-stakes decisions. The cost of a confidently wrong answer is highest precisely where oversight is thinnest, which in HR often means final hiring calls, disciplinary action, accommodations, and compensation decisions.
- Turn the same scrutiny on HR’s own AI tools. The tools HR rolls out to the business are worth auditing with the same jagged-frontier question HR is asking everyone else to ask: which of our own AI-assisted decisions are inside the frontier, and which need a human in the loop by design?
This tension, between AI’s real productivity gains and the judgment required to use them safely, is exactly the kind of question being asked with “Redefining Productivity with AI: Beyond Output Volume” – one of the Table-Top Peer Discussions at this year’s Horizon Summit. AI can make you faster today. Whether it makes you sharper tomorrow depends on how deliberately you build the judgment to use it well.