Skip to content

Science1 publisher3 min readPublished

Expecting an AI grader pushed job seekers to embellish in a University of Georgia experiment

Applicants who believed software would score their one-way videos stretched the truth more often, the rating agent scored them no lower for it, and telling them what it measured brought the exaggeration back down.

The Scientist · Science desk

Illustration accompanying Expecting an AI grader pushed job seekers to embellish in a University of Georgia experiment

What happened

  • A University of Georgia study in Information Systems Research found that job seekers who knew an AI would evaluate their recorded interview answers reported exaggerating considerably more than those expecting a human.
  • An analysis of the participants' videos and answers found the same pattern the applicants had reported about their own behavior, rather than contradicting it.
  • Candidates who were told an AI would review them and what it looked for reported and displayed the same authenticity as candidates who expected a human reviewer.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • exposure Candidates who describe their record accurately are being ranked by a scorer that, in this study, could not separate them from candidates who inflated the same qualifications, so the accurate ones absorb the cost.
  • decision Employers who withheld process details as anti-gaming insurance now have a narrower option to trial: publish which features and traits are scored, without publishing the model.
  • contradiction Lakhiwal points to existing research showing that applicants who know the evaluation criteria will game them, while his own group found criteria disclosure restored authentic behavior, so what gets disclosed decides which way the effect runs.
  • constraint Any validation done on videos recorded for human reviewers describes behavior a tool will no longer see once candidates know a machine is grading them.

Participants reported more exaggeration when they believed an AI agent rather than a human would evaluate them [3], and an analysis of the videos and the answers themselves found the same pattern [4]. Two channels agreeing rules out the dullest explanation on offer, that people merely describe themselves differently when a machine is in the room rather than behaving differently in front of it [15].

Why the behavior moves is, in Lakhiwal's telling, a matter of missing information rather than character. A candidate applying to a company they badly want has little practical room to decline the format and no picture of how it works [14], which he says leaves applicants disoriented [19]. So they guess at the criteria and then over-serve the guess; Lakhiwal describes participants "throwing the kitchen sink at the situation to try to give the 'evaluator' what it was looking for" [13]. Asked whether embellishing qualifications was ethical, they told the researchers it was necessary to perform well in an unpredictable environment [11].

The scoring half is where the two kinds of evaluator come apart. The AI agent the team used, which the researchers describe as industry-favored, gave candidates who stretched the truth the same scores as candidates who earnestly described having the same qualifications [5], while human evaluators looking at the same videos generally penalized the embellishers and reserved their highest ratings for those who seemed authentic [6]. Hold qualifications constant and the rating moves with the identity of the judge [16].

The intervention that worked was narrow. One group was told an AI would review the video, and then told what it was looking for: facial expressions, verbal sentiment and specific keywords, scored against teamwork, job-related abilities, work style and personality [9]. That group reported and displayed the same authenticity as the group expecting a human reader [10]. The model itself stayed undisclosed. Lakhiwal's framing is that he does not need to know which model is analyzing his video [20], which is a much smaller ask than opening the system.

How big any of this runs is left open. The published account gives "hundreds" of online job seekers and a "considerable" increase in embellishment, with no effect size, no exact sample count, and no numeric gap between the AI and human ratings [17]. One agent was tested [5], so insensitivity to embellishment is a property of that scorer, not a law about automated scoring. And the outcome measured is a rating rather than a hire, so the study cannot say what the distorted signal costs an employer in the people it eventually seats.

A recruiter who wants to know whether this holds in their own funnel has a cheap local check: rescore a sample of AI-graded videos with trained human raters and see whether the two orderings agree. In this study, the two orderings diverged [16].

What to watch

  • The full Information Systems Research paper, which should carry the effect size, the exact sample, and the AI-versus-human rating gap the announcement leaves qualitative.
  • Whether other vendors' scoring agents show the same indifference to embellishment, since the study tested one agent.
  • Whether any employer publishes its scoring criteria to applicants and reports what happens to the resulting score distribution.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories