Back to blog
hiring

every vendor says their AI is unbiased. Check anyway.

Bias in hiring is older than AI. But AI changes the shape of the problem. It can remove human variance, or it can scale it. The difference comes down to what the platform was built to evaluate, what the team has tested, and what they will tell you. Here is the practical view.

Tl;dr

Bias is not a checkbox. A serious vendor will tell you What they evaluate, what they ignore, and how they audit. You should test for disparities on your end too. The goal is not a perfect system. The goal is a system that is honest about what it sees, transparent about how it scores, and improvable when something looks off. Here is what good looks like.


Where bias comes from in the first place

In human screening, the largest single source of bias is something most teams do not name. It is Interviewer variance. The same candidate gets a different evaluation from a different person, on a different day, asked different questions. That variance is not random. It correlates with attributes the interviewer was never supposed to factor in.

Resume signals add another layer. Names, schools, employment gaps, formatting choices, all of it shapes who gets a callback before anyone has spoken to anyone. These are well documented effects, not edge cases.

AI sits on top of this. It can either remove these signals from the evaluation, or it can learn them from training data and amplify them at scale. Which one happens depends entirely on how the system was built.


What unbiased actually means

"Unbiased" is not "the model treats everyone the same" in some abstract sense. That phrase is too easy to say. A useful definition has three parts.

First, the evaluation framework Does not see protected attributes or strong proxies for them. With aperture, the AI does not see names, accents, gaps, or formatting. It sees how the candidate reasons about a structured set of behavioral prompts.

Second, scoring is Identical across candidates. The same six dimensions, the same rubric, the same conditions. Interviewer variance is removed because there is no interviewer variance to remove.

Third, outcomes are Tracked for disparities by group, Openly. A system that nobody is checking is a system you cannot trust, no matter how good the intent.


What to ask the vendor

Most vendors will tell you they are unbiased. Very few will tell you what they tested, what they measured, or what they would do if a disparity showed up. These are the questions that actually surface that.

Questions worth asking

What does the model see during evaluation, and what does it explicitly ignore
What disparity tests have you run, against what reference groups, and how often
Can you show the dimensions a score is built from, and the evidence behind each one
What would you do if a customer surfaced a disparity in their pipeline
Can we audit our own outcomes against your scoring
How do you handle accents, non-native speakers, and non-traditional career paths

What good practice looks like

A serious platform looks the same for every candidate. The same structured behavioral interview, the same questions adapted to the same set of dimensions, the same scoring rubric, the same conditions. No resume signals enter the evaluation. No interviewer variance. No shortcuts.

Scoring should be tied to specific evidence from the conversation. Not "the model felt this candidate was strong" but "this answer demonstrated this dimension at this level for this reason." every score traces back to behavioral evidence. No black boxes.

This is how aperture is built. λ-CORE scores six dimensions, cognitive reasoning, domain knowledge, communication, behavioral indicators, collaboration, and adaptability. Every output is a score, a confidence interval, and a pool rank percentile, all traceable back to what the candidate actually said. That traceability is the precondition for any honest conversation about bias.


What to check on your end

Trust but verify. The vendor's word is a starting point, not the answer. You have data the vendor does not. Use it.

On your side of the funnel

Track outcomes by group at every stage, application to interview to offer to hire
Compare pass-through rates across groups for drift, both relative to each other and to your prior baseline
Look at scores at the boundaries, the top of the shortlist and the cut line
Do not rely on the vendor's word alone, audit your own pipeline regularly
Flag patterns early, when the sample is small enough to investigate properly

What to do if you find a disparity

Do not panic. Do not silently kill the tool either. A disparity in your funnel is information. It is the first step in fixing something, not the last word on whether the system works.

Investigate the source. Is the disparity larger or smaller than your prior baseline. Does it appear at the AI step, or upstream in sourcing, or downstream in human review. Is it concentrated in a specific role, a specific stage, a specific period.

Then work with the vendor to fix it. A real partner will respond with data, with a hypothesis, and with a plan. A vendor that goes quiet, or insists nothing could possibly be wrong, has told you something important. The goal is not a perfect system on day one. The goal is a system that gets more honest over time. That is the standard worth holding any tool to, AI or human.

For the research behind why disparities hide in pooled averages in the first place, and how aperture is built to avoid the mechanism entirely, see Does AI hiring cause racial bias, and how does aperture avoid it.

Published march 1, 2026 by aperture team 9 Min read