Skip to content

Leadership1 publisher2 min readPublished

Itamar Gilad draws an AI danger zone where output outruns the reviewer's judgement

An essay by the former Google product manager argues that LLMs lift what a person can produce faster than what they can evaluate. Every curve in it is labelled speculative, and it reports no measurement of the gap.

The Board Room · Leadership desk

Photograph accompanying Itamar Gilad draws an AI danger zone where output outruns the reviewer's judgement
Photo: itamargilad.com

What happened

  • Itamar Gilad, a former Google product manager, argues that a manager who cannot perform every task in her org at the same level should still be able to judge and critique the work.
  • He argues that AI raises what a person can produce above what they can assess, so people now generate work they are poorly equipped to judge.
  • He sorts AI-performed tasks into a Safer Zone, where the user can see output is sub-par and iterate, and a Danger Zone, where the risk includes creating security vulnerabilities in the product.
  • He describes roles overlapping already: PMs and designers contributing production code and internal tools, and developers using chatbots for market segmentation, prioritization and spec generation.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision The gate on delegating a task to AI becomes whether someone in the approval chain can grade that specific output, which is an org-design answer settled before any tool choice.
  • cost Review time from the few people who can grade a domain becomes the scarce input, and it is now spent on output generated by people in other functions.
  • exposure The approver carries a defect that was invisible at the moment of sign-off, and discovery moves downstream to whoever maintains the code or lives with the decision.
  • constraint Because the curves are speculative, the framework can be used to sort which tasks need an outside grader, but it cannot tell a board how often unreviewed AI work is failing.

Supervision works because judgement is the wider of the two capabilities. Gilad makes the point with a list of things he could not do himself: "As a product manager at Google, I could write decent product requirement documents, produce sub-par marketing copy and designs, barely knew how to add to the very complex Gmail codebase, and had no clue how to create a balance sheet for my business unit," he wrote [2]. He could still review a designer's work and give useful feedback on it [3]. "We're generally able to judge more than we can do," he wrote [3].

The reversal he describes is domain-specific. Gilad does not claim the tools make anyone an expert: "I don't think AI gets you to expert level on anything, but you can perform far more tasks at medium/OK level (which is sometimes all you need), and yet more at a below-average level," he wrote [6]. He adds that output quality depends heavily on the context the user supplies and on the user's ability to guide the model [14]. A user outside their own area of competence has no reliable way to tell whether they guided it well. His own example of the failure is blunt: "A PM may generate bad production code and a developer may choose bad ideas, and neither can tell the difference" [15].

Gilad anticipates the objection a manager will raise in the room. "You may argue that smart, responsible people can recognize their limitations and seek help when they step outside their comfort zone," he wrote [10]. His answer is that decades of psychological research suggest otherwise, and that cognitive biases make people overconfident, most of all in areas where they do not know what they do not know [11]. The passage names one of them, the better-than-average effect, under which people tend to overestimate their own qualities and abilities compared with others [12].

Even in the tasks a person can grade, Gilad writes that "the risk is not zero because sometimes we fail to check" [8]. That sets the practical question for anyone signing off on AI-assisted work this quarter. It is not whether the model is good enough. It is whether anyone in the approval chain holds the competence to grade the specific output, and whether the process forces that person to look at it.

The framework sorts work; it does not price it. Gilad labels all the graphs in the article speculative [13], and the piece reports no measurement of how wide the doing-versus-judging gap runs in any function [16].

What to watch

  • Any measured rate of AI-assisted output failing review outside the producer's own domain would turn this framework into a finding.
  • Approval policies that name a domain grader for cross-function AI work, rather than leaving sign-off with the producer.
  • A published post-incident account tracing a shipped defect to code generated by someone outside the engineering function.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories