We often treat effort as evidence of worth. But “effort” hides two very different things — and AI is keeping the counterfeit while discarding the real one.
A painting the artist bled over for a decade seems worth more than one dashed off in an afternoon, even when we can’t tell them apart on the wall. A lawyer who returns your contract in six minutes has somehow earned less than one who sends it back six hours later, thick with tracked changes. When we can’t judge quality directly (and we often can’t), we reach for the next best thing: how much it seems to have cost.
The instinct has a name. Psychologists call it the effort heuristic: we infer quality from apparent labor. (Its original demonstrations have had a mixed record in later replications, so hold it as a tendency, not a law.) Its stranger cousin is the labor illusion, described by Ryan Buell and Michael Norton in 2011: travelers preferred a flight-search site that made them wait while it narrated what it was doing over an identical site that answered instantly. The slower one felt more thorough. Watching a system sweat, we assume it tried harder.
That is effort as performance: a signal read off the surface of the thing. And the signal fails constantly, because visible strain and real quality come apart all the time.
Two kinds of effort
There is another kind of effort, and it is nearly the opposite of a signal: it doesn’t advertise value, it produces it. Learning scientists have names for this. Robert Bjork’s “desirable difficulties” are the struggles that slow you down while you learn but leave the knowledge deeper and more durable. Manu Kapur’s “productive failure” finds that students made to wrestle with a problem before being shown the method understand it better than those handed the solution first. The strain is doing the work.
Not all effort is this kind. Plenty of struggle is just flailing, and grinding at midnight can teach you nothing. That’s the flaw in the heuristic: it can’t tell productive struggle from mere performance, or from wasted motion, and reads all three as worth. Which is the seam AI has learned to work. Left to its defaults, it strips out the productive kind and counterfeits the performed one.
Smarter, but none the wiser
In two studies from Aalto University, hundreds of people solved logical-reasoning problems with an AI assistant. Their scores rose. Their self-assessments rose more — they overestimated by about four points; in the first of the two studies, that was roughly double the inflation measured in an earlier cohort working without AI. The usual Dunning-Kruger pattern, where the least skilled misjudge themselves most, didn’t just shrink; in the AI condition it disappeared. Everyone was overconfident, evenly.
The line that traveled was that the most “AI-literate” participants were the worst calibrated — and it deserves the asterisk the headlines dropped. Literacy there was self-reported; the people who rated their AI knowledge highest were the most overconfident about the task. That might mean fluency breeds false confidence, or simply that rating yourself an expert is already the thing the study was measuring. Tellingly, self-rated literacy predicted more confidence but not higher scores. Either way, AI inflates how good we think the work is without making us any better at telling. You over-trust the product because you never built the skill that would let you doubt it.
The opposite error
There is a second possibility, harder to measure, and it looks like the reverse of the first.
I know this one from the inside. Some weeks I have Dunning-Kruger days and impostor days back to back: create something with AI and feel sharp, create something with AI and feel like I subcontracted the part of me that used to be good at this. The experience is common enough to have a shape: you ship your best work and feel less ownership of it. Not “I’m better than I am,” but “I didn’t really do this.” It reads as the opposite of overconfidence, which is why the two are usually drawn as ends of a single line. But they aren’t one line. Overconfidence is a gap between your confidence and your competence; the fraud feeling is a gap between your success and your sense of owning it. AI can widen both at once, because it removes the experience that would have closed either one.
That is what productive struggle actually leaves behind: a receipt. Not “I worked hard enough” — “I can explain why this works, reproduce it without the tool, adapt it when the situation shifts, and notice when it’s wrong.” That is evidence you can point to the next time the doubt shows up. Grinding at midnight writes no such receipt. Neither does a fluent draft you couldn’t defend without the tool that wrote it.
There’s a folk model of learning (not what Dunning and Kruger actually measured, but true to how mastery feels) where you crest a hill of early confidence, drop into a “valley of despair” as you see how much you don’t know, and only then climb toward real skill. The valley is where the receipt gets written. AI can pave it over: you ship the competent work without the descent, and so never accumulate the evidence that the competence is yours.
This is why the reassurance usually offered here gets slippery. Many approaches to impostor feelings work by testing the thought “I didn’t really do this” against the evidence and finding it false. AI can muddy that test, because authorship now really is shared, in degrees you can’t cleanly settle. The thought isn’t simply a distortion anymore; it isn’t simply true either. It just stops being answerable, which is worse.
The obvious objection is that you did work — you prompted, steered, edited. You did. But the fear underneath the impostor feeling was never about pride; it’s about exposure, and “I worked hard this time” says nothing about next time. What quiets that fear is stable evidence of ability, and effort-as-input isn’t that. So AI offers a strange bargain: a cause for your success that is external, so it can’t feed your sense of ability, but reliable, because the tool will be there next time too. You may never be exposed. You also never bank the evidence that would make exposure stop mattering.
The machines learn to sweat
The newest systems don’t just answer; they display the signs of deliberation: pauses, tool calls, intermediate steps, and increasingly elaborate narrations of how an answer was reached. We read effort as worth; now it can be shown to us on cue.
The early evidence is pointed. At CHI 2026, researchers at NYU randomized 240 people to interact with the same underlying model under one of three response delays (two, nine, or twenty seconds) while they worked through everyday writing and advice tasks. The longer waits were rated more thoughtful than the two-second one; usefulness peaked around nine. The model hadn’t done more work; the interface had simply made people wait. Among those who noticed the delay, about a third read it as the AI “thinking.” And the researchers used a plain interface with no reasoning display at all, so this is the effect of bare time alone. They name the larger risk directly: temporal cues that simulate cognitive effort and lead users to over-attribute reasoning and competence to the system, what recent research has begun calling performative deliberation.
A visible reasoning trace is that bare pause with the work spelled out. Two things make it a bad quality signal. First, the same careful-looking reasoning can wrap a correct answer and a confidently wrong one, so its thoroughness cannot tell you which you got. Second, the trace is not guaranteed to represent what actually produced the answer; research on these reasoning chains finds they can be unfaithful — and in Anthropic’s tests, the unfaithful chains were substantially longer than the faithful ones. It’s the open kitchen where the chef you’re watching chop isn’t the one who cooked your dish.
The struggle worth keeping
None of this makes AI the villain. I should know: I co-founded a company that reimagines courses and degree programs for a world where everyone has access to AI, which makes me one of the people who chooses these defaults. The incentive is exactly what you’d guess: friction looks like a bug until you remember it was the product. AI can pave over the valley — or it can build handrails down into it. Ask it for a hint instead of an answer, to test your recall, to argue the other side, to find the flaw in your draft, and it manufactures desirable difficulty rather than removing it. The useful friction has to be put back on purpose. We try to build the handrail versions: a tutor that asks before it tells, that cold-calls instead of lecturing. It is easier to praise struggle in an essay than to defend it in a roadmap meeting.
We had one word for two efforts. The performed kind was always only a proxy, which is why it was always easy to counterfeit. The productive kind was the real article: it built the competence, and it wrote the receipt.
AI can remove the friction through which we learn to judge the work, while manufacturing the friction through which we judge the machine. It gives us performed effort in abundance and, unless we resist its defaults, quietly takes the productive kind away. The outputs really are better, which is what makes the trade easy to miss. What’s left is the discipline the heuristic let us skip all along: judging the work itself, with no sweat to go by.
