The Judgment Collapse

· 17 min read · AI workforce
#1 Beyond Ideas

My son Leo was supposed to be in bed twenty minutes ago. Instead he was running laps between his bedroom, the bathroom and the living room in his underwear, bare feet slapping the tiles, voice pitched at the frequency six-year-olds reserve for maximum parental exhaustion. I caught him on the fourth pass. Teeth took negotiation and showcasing what happens if you do not take care of your teeth (I recently got an extraction). Pajamas took bribery. By the time he was under the covers, eyes fighting gravity with the stubbornness of someone who believes sleep is a personal insult, I was on the couch with my phone. Too tired to be useful. Not tired enough to sleep.

I know… I shouldn’t have been scrolling. I already know how it will end - with me sleeping too little hours. But I was, in that mindless way you do when the evening has cost you more energy than it should and you need five minutes of someone else’s thoughts. An article from Tom’s Hardware Italia caught me. The title: ""Avere idee” è sempre stato facile e sopravvalutato, e l’AI lo dimostra” (Having ideas has always been easy and overrated, and AI proves it.) Written by Valerio Porcu, one of their senior editors.

I’ve sat through enough workshops, innovation sprints, and brainstorming sessions to know the feeling Porcu was writing from: the suspicion that all that ideation theater was a waste of time, and that AI just made it obvious.

His argument: AI has driven the cost of generating ideas to near zero. Only execution matters now. The sooner organizations accept this, the better. He anchored the whole piece on a quote from Terence Tao, the Fields Medalist at UCLA, who told the Dwarkesh Podcast that “AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero.” He then drew a clean line: AI generates, humans judge and execute. He identified organizational change as necessary but didn’t specify what kind. And his most interesting claim, that AI compresses the timeline for distinguishing builders from talkers, went completely unsubstantiated.

Good quote. Clean argument. The comment section split between people who found it obvious and people who found it reductive. I got intrigued by the articled, yet I found it a bit “incomplete”. So I did what I do when a piece cites a quote and a famous name: I pulled the source and started reading (well, watching in this case).

The Tao Problem

Porcu’s entire opening rests on Terence Tao. The citation comes from the Dwarkesh Podcast episode “Kepler, Newton, and the true nature of mathematical discovery,” published March 20, 2026. It’s a circa 90 minutes conversation about how Kepler discovered the laws of planetary motion and what that history tells us about how AI will change science. Suggestion: carve out some time to listen to it!

Porcu extracted one sentence and applied it to business.

The actual conversation was about something else entirely. Here is what Porcu left out.

Tao was talking about science, not startups. The discussion was about scientific discovery. Tao listed a dozen components of the scientific process, from problem identification through hypothesis formation to peer communication, and noted: “The ones we celebrate are these eureka genius moments of idea generation.” He said hypothesis generation may no longer be the bottleneck, but his reasoning was domain-specific. He was describing the shift from hypothesis-first to data-first science. He was not making a claim about corporate strategy.

The automobile analogy is the core of Tao’s argument. Porcu ignores it entirely. Tao compared AI’s impact on mathematics to the automobile’s impact on cities. Cars were faster than anything before them. But they clogged roads built for people, horses, and carriages. New roads made fast travel possible but created urban sprawl and environmental damage. Only thoughtful urban planning could have united both worlds. Tao’s point: the existing infrastructure of mathematics (journals, conferences, mentoring, peer review, citation systems) is like those old narrow roads. Built for human-speed work. Human proofs may be slow, but they generate valuable side effects: researchers develop expertise, map mathematical terrain, and document instructive dead ends along the way. AI-assisted proofs can be fast and correct, but they lose exactly these side effects. At the UCLA/IPAM event co-hosted with OpenAI the same month, Tao argued that the real payoff from AI will come from redesigning workflows around it, not from inserting AI into old ones.

AI lacks cumulative reasoning. Tao distinguishes between what he calls “artificial cleverness” and actual intelligence. AI cannot build up from partial progress. When two human mathematicians collaborate, they start from zero, test ideas, modify them, map what works and what doesn’t in a living way. AI does not do this. On the AI-solved Erdos problems: “It either solves something or it doesn’t. It cannot plant a flag halfway up a cliff and build from there the way a human mathematician would.”

Verification is social, not just technical. Tao said that assessing ideas “depends on the future. It depends also on the culture and society, which ones get adopted, which ones don’t… It may never be something that you can just reinforcement learn.” This is the opposite of a simple “evaluation is the new bottleneck” claim. Evaluation may be fundamentally resistant to automation, because it requires historical context and cultural awareness that no scoring function captures; it requires the kind of forward-looking judgment that learns from the past without being trapped by it.

The pen-and-paper claim is precise. When Dwarkesh Patel asked Tao what year he’d be twice as productive from AI, Tao refused the framing: “Productivity, I think, is not quite a one-dimensional quantity.” His papers now include more code, more plots, more numerics, because AI made those components cheap. But the core of his mathematical work, the hardest step, still happens on pen and paper. AI hasn’t sped up the actual work so much as opened up new possibilities alongside it.

The validated prediction. In 2023, Tao predicted that by 2026 AI would be “a trustworthy co-author if used correctly.” In March 2026, he confirmed: “I’m pretty pleased” with how that aged. On his blog, he documented using AI tools to confirm numerical plausibility of an approach and to recognize the right technique for a sub-problem, then switching to pen and paper for the core insight. Co-author, not replacement.

Compared to what was cited in the article that kickstarted my endless scrolling and research, Tao’s actual position is far more interesting. I think he’s arguing for institutional redesign to handle verification at scale, and he’s explicit that the social and cumulative dimensions of evaluation may never be automatable. So the real argument might be more like: the entire infrastructure through which we develop, evaluate, and transmit judgment is about to be disrupted, and nobody is redesigning it fast enough.

That I think it is a much, much harder problem than “generating ideas is cheap with AI”.

What I found when I kept opening tabs in my browser and reading is that the boundary Porcu treats as settled, the line between AI generation and human judgment, is dissolving from both directions at once. From one side, AI is colonizing parts of the judgment layer. From the other, it is arming fakers with tools to produce execution artifacts that look real but are hollow underneath. The result is something messier than the meritocracy of executors Porcu imagines.

When AI Learns to Judge

Start with the judgment side. Porcu draws a clean line: AI generates, humans select, humans execute. The line assumes judgment sits safely on the human side of the fence.

A Harvard Business School study by Rembrand Koning and colleagues tested exactly this assumption. They ran a field experiment with 640 small business owners in Kenya. Half got access to a GPT-4-powered business advisor via WhatsApp. The other half got written guides from the International Labor Organization.

The headline result: no statistically significant average difference in performance between groups. AI advice, on average, made no measurable difference.

But underneath that average sat a different story. High-performing entrepreneurs who got AI access saw profits rise by 10 to 15 percent. Low performers who got the same access saw results drop by roughly 8 percent. The gap between winners and losers didn’t narrow. It widened.

The difference was not in the questions they asked or the answers they got, because both groups received similar guidance from the same model. Instead the gap came from who could pick which advice to follow. Successful entrepreneurs pursued business-specific recommendations (generators for neighborhoods with blackouts, cold sodas at car wash queues). Struggling ones defaulted to generic suggestions (cut prices, run ads) that couldn’t reach the deeper problems in their businesses.

The judgment gap was about the so called operator, not the tool.

So much for AI as the great equalizer. If prior experience determines how much value you extract, the tool amplifies existing advantages rather than flattening them. That is a distribution problem, not a technology one, and it’s a thread I’ll pick up in one of the following articles.

But there is a second layer I think Porcu doesn’t touch, and it is that AI isn’t just generating options for humans to evaluate but it is starting to design the evaluation environment itself.

Researchers at MIT Sloan and TCS have documented what they call “intelligent choice architectures”: systems where AI doesn’t just answer questions but restructures how choices get presented to decision-makers. Companies are deploying them now. The AI injects unconventional options into brainstorming sessions. It reframes risky acquisitions as “growth opportunities” to counteract executive risk aversion and vice versa. It learns which presentation formats nudge leadership toward bolder decisions. So, instead of replacing judgment, the AI reshapes the environment in which judgment happens. I think this is a harder thing to notice while it’s occurring.

McKinsey’s own framework for agentic AI, published in June 2025, makes a quiet concession that goes further. They classify organizational decisions by risk and complexity, and their analysis suggests that a large portion of what companies call “judgment” is actually low-risk pattern matching. In simple words: it is repeatable and predictable. The kind of thing a well-trained model handles without breaking a sweat. Decisions that looked like (senior) management or executive wisdom often turn out to be pattern matches once you strip away the rank and the conference room. It was always automatable; we just called it wisdom because a person was doing it.

So the question would be how much of what we called judgment was bureaucracy we dignified with the word wisdom.

When Fakers Learn to Build

Now flip to the other direction. While AI colonizes the judgment layer from above, it is simultaneously handing fakers a better toolkit from below.

In January 2026, the SF Standard reported on a Y Combinator startup (Winter 2025 batch) that had attracted significant buzz for its AI-powered smart glasses, branded as a “soul computer.” The founder promised personalized AI recommendations delivered through the lens display. Book endings for novelists. Food choices at grocery stores. The price: $799. The demo video: gorgeous. Well, it was CGI.

The glasses were “design mockups.” The company had raised $4.42 million, which based on some research I did seems to be not enough to manufacture a consumer electronics product at scale. Matthew Dowd, founder of VR company Wild, publicly accused them of “fraudulently representing their product.” The founder responded with conviction and sincerity, which is exactly what founders of vaporware companies do.

Another case. Humane AI raised over $230 million before its AI Pin arrived to reviews so devastating the company was eventually sold to HP for $116 million. The ratio between what was raised and what was recovered tells you everything about the gap between demo and delivery.

But here is what has changed: the cost of producing convincing execution artifacts. Product School’s 2026 guide to AI prototyping documents how tools like Bolt and Lovable now let teams build interactive, database-connected prototypes in hours. Bolt hit $40 million in annual recurring revenue in five months. Lovable reached $200 million ARR and a $6.6 billion valuation by the end of 2025. You do not need a technical background anymore but just an idea and a couple of days (preferably without your kids around).

That changes the calculus for anyone evaluating startups, products, or talent. A polished demo used to mean someone had built something real. Now it might mean someone spent a Saturday with Bolt. The surface tells you less than it used to.

Porcu argues that AI compresses the timeline for distinguishing builders from talkers. I wish he was right, because what I read points somewhere less comforting: both real and fake execution are getting cheaper to produce, and the signals we used to rely on to separate them are degrading. Yanko Design had to publish a guide before CES 2026 to help people distinguish genuine AI products from marketing wrappers. When you need a field guide to spot the fakes at the world’s biggest consumer electronics show, the “execution exposes pretenders” narrative needs revising.

The Signal That Broke

Both threads converge on a problem that I have not found in Porcu’s article. Signal collapse.

A 2026 analysis from CodeConductor puts it plainly: “Speed used to signal competence. Now it simply signals access to the same tools everyone else has. The real differentiation is moving elsewhere: judgment, focus, validation.” When everyone can spin up an MVP in a weekend, the question stops being “how quickly can this be built?” and becomes “why was this worth building at all?” That second question is a judgment question. The very thing Porcu’s article treats as settled.

And the tools themselves may be eroding the capacity to answer it.

MIT Technology Review reported that engineers who rely heavily on AI coding tools are losing the instincts they once had. One developer found that after working without AI tools on a side project, “things that used to be instinct became manual, sometimes even cumbersome.” A study by the nonprofit METR put numbers on the phenomenon: experienced developers believed AI made them 20 percent faster. Objective testing showed they were actually 19 percent slower. The developers were confident they were more productive but they were measurably less so.

This is where Tao’s automobile analogy stops being a metaphor for mathematics and becomes a map for everything else. He warned that building faster roads for AI-speed work means losing the side effects of human-speed work. In mathematics, those side effects are expertise, terrain-mapping, instructive dead ends. In software engineering, they are instinct and pattern recognition, the ability to smell when something is wrong before you can explain why. In business, they are the judgment you accumulate through years of making small decisions in ambiguous situations and watching what happens next. The junior analyst who manually builds a financial model develops instincts about which assumptions break first. The associate who drafts and redrafts a client memo learns what persuades and what falls flat. The junior consultant who work and rework for days on the same four or five slides learns how to master PowerPoint and craft nice looking slides. The junior auditor who spends nights and weekends on Excel files, evidences, controls learns the methodologies and standards. Automate those tasks and you get speed, but you also lose the training ground. I once pushed a junior colleague to speed up delivery by leveraging some good crafted prompts I had prepared. After a couple of days, she came back to me saying she was not learning anything, just copy and pasting outputs from an LLM to a spreadsheet for my review. And that was not ok.

So… The tools that make execution cheap may simultaneously erode the judgment required to evaluate what they produce. Call it what it is: a design flaw in how we are adopting them.

Tao said verification “may never be something that you can just reinforcement learn.” He was talking about mathematical proof. But the same logic applies to market judgment, product taste, and the thousand quiet assessments that separate a working company from a decorated corpse. These are social and cumulative skills. They resist automation not because the technology falls short, but because they depend on context that doesn’t sit in any training set.

The organizations that built entire industries around the premise that idea generation is the hard part (a subject I’ll turn to in another article I am writing) had it wrong. But so do the people who think the answer is simply “execute better.” Execution is being hollowed out as a signal. The judgment required to direct it is distributed in ways that have nothing to do with technology. And the people most likely to develop that judgment are, for reasons that are structural rather than meritocratic, a narrower group than the “just execute” camp wants to acknowledge.

What Started as Source-Checking

Our house was quiet and silent by the time I finished watching the Tao interview and checking some parts of the transcript. Leo had surrendered to sleep. The article which kickstarted hours of browsing and scrolling was still open on my phone (thank you Valerio!). It wasn’t wrong in its foundations. Ideas have always been cheap. AI does make that cheaper. And yes, organizations that confuse brainstorming with production are wasting everyone’s time.

But I think that the clean narrative, where AI generates and humans judge and the best executors rise to the top, is (unfortunately for us) more comforting than it is accurate. It looks like that the boundary between generation and judgment is dissolving. The signals we relied on to identify real execution are degrading. The tools we are building to make work faster may be quietly eroding the judgment we need to direct that work. And the race to automate tasks in the workplace is having impacts on the younger generations. And the one source Porcu anchored his argument on was saying something far more interesting than he let on: that the real challenge is redesigning the institutions through which we develop and verify judgment at a speed that doesn’t leave everyone behind.

I started that evening with some scrolling to “turn off” my brain, which turned into a source-checking exercise, which then turned into something I needed to further research and to write about. Not because Valerio Porcu was wrong. Because he stopped too early. He identified the right problem (ideas are overvalued) and pointed at the right replacement (execution, judgment, verification) but didn’t ask who gets access to the conditions that develop those things. He didn’t notice that the source he built his argument on was warning about exactly this gap. And he didn’t follow his own logic to the place where it gets uncomfortable: that “execution wins” only works as a narrative if the playing field for developing execution skills is level. It isn’t. It’s getting less level. And the tools that were supposed to flatten it may be tilting it further.

The hardest question Porcu’s article never asks is also the one that unfortunately matters most for Leonardo’s and his brother Ettore generation: if judgment is what counts now, who gets to develop it?

G.

All views expressed here are my own and do not represent the opinions or positions of my employer or any organization I am affiliated with.

AIL: 0 1 2 3 4 5

Giulio wrote the core content and analysis. claude-opus-4.6 / Anthropic (primary contributor) and other AI models supported with research, sounding board, refinement, and structural editing.