Skip to content

Leadership1 publisher3 min readPublished

The bias evidence for agent-to-agent hiring comes from the vendor selling the agents

Matt Wilson's AI agents represent close to half a million professionals to around 5,000 companies, and the bias experiment he cites as the reason to trust audited screening sits in his own company's white paper.

The Board Room · Leadership desk

Photograph accompanying The bias evidence for agent-to-agent hiring comes from the vendor selling the agents
Photo: recruitingfuture.com

What happened

  • His company's agents represent close to half a million professionals and work with around 5,000 companies, building matches through in-depth conversations instead of keyword searches.
  • An experiment described in Jack and Jill's own white paper found that both humans and out-of-the-box large language models consistently exhibit bias when ranking job applications.
  • The episode notes say human judgment remains essential to hiring decisions and that the agents exist to maximise the time both sides spend in the right conversations.

Compiled by The Board RoomSomething wrong?How this is made

Why it matters

  • decision A buyer weighing another screening layer this quarter has to decide whether its current stack can produce a bias measurement at all, because the vendor's case turns on measurability.
  • constraint A hiring process that has never been measured cannot supply a baseline. A committee has nothing to hold a vendor's audit up against.
  • exposure The finding covers off-the-shelf models as well as people, so employers who bolted a general-purpose model onto resume sifting are carrying the same risk they were told specialist tooling creates.
  • contradiction Independent auditing is offered as the route to trust while the only cited evidence is the seller's own paper, and that gap decides how much weight the bias result can bear.

Divide the two network figures and you get roughly 100 represented professionals for every company on the platform [4][1]. A pool that shape does not generate a thousand applications per opening, so the promise of fewer conversations is partly a description of the inventory. The promise of better ones is the contested half, and what supports it in the episode is a forecast that broad adoption of agents on both sides will produce lower volumes of conversations but much higher quality ones [9].

Matt Alder, who hosts the show, framed the problem in terms of signal. "AI has made it easier than ever to apply for a job and easier than ever to contact a candidate. And the result is a level of noise that both sides are starting to tune out of," Alder said [7]. Wilson's version of the same point is narrower: most AI at the top of the funnel supercharges existing processes, producing more applications and more outbound that both sides increasingly tune out [2].

The case for the audited alternative rests on measurability. An experiment described in Jack and Jill's white paper found that both humans and out-of-the-box large language models consistently exhibit bias when ranking applications, and that a well-designed, audited AI system can be measured and improved in ways human-based processes cannot [5][6]. The episode notes put transparency and independent auditing at the centre of how trust in AI matching gets built [8]. The published material does not name an auditor or describe the experiment's sample.

The company selling audited matching is also the company reporting that the unaudited alternatives are biased. That conflict does not settle the question, because the weak point sits at the buyer's end. If a hiring team cannot produce a bias measurement of its own screening decisions from last year, the comparison in front of a procurement committee is one measured claim against one unmeasured process. The measured claim wins by default, whether or not it deserves to.

Fairness criteria and an audit trail are, in my view, cheaper to write into a specification before a tool starts ranking candidates than to reconstruct after two quarters of scoring. The argument Wilson makes about noise points the same way: adding automation to a process you have not characterised gives you more output from the same design [3]. Employers running an off-the-shelf model over resumes are inside the bias finding as stated [5].

On the division of labour, the claim is modest. The episode notes say human judgment remains essential to hiring decisions, and that the purpose of the agents is to maximise the time both sides spend in the right conversations [10]. Episode 827 of Recruiting Future went out on recruitingfuture.com in September 2026 [12].

What to watch

  • An external replication of the bias result by someone other than Jack and Jill would move it from vendor claim to evidence a procurement team can cite.
  • Whether any applicant tracking or screening vendor publishes a bias measurement of its own past decisions that a buyer could compare against.
  • Whether the 5,000-company, half-million-professional network grows enough to test the forecast of fewer but better conversations.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories