Is AI biased? How it happens, and how to check
AI bias sounds like a political argument. It is mostly an arithmetic one — and you can find it yourself in about a minute, once you know where to look.
Here is the report card for a system that sorts job applications. It was checked against 125 decisions a human had already made, and it agreed with 112 of them.
90% correct — looks strong, right?
The headline number is fine. One group is being got wrong seven times out of ten, and the average absorbs it quietly — partly because that group is the smallest one tested.
Then the cause. Of the examples the system learned from, applicants resembling past hires were about 65%. Applicants with a career gap were about 1%.
Training an AI on historical decisions teaches it to repeat history, including the unfair parts. Nobody has to intend anything for this to happen, and nobody in the process has to hold a single opinion about career gaps. The arithmetic does it on its own.
The version you have already met
If a hiring system feels remote, here is the same shape in something you have probably raised your voice at. A voice assistant, scored on how often it understands a spoken command:
90% correct
Same 90% headline, same 30% underneath it. That accent was roughly 1% of the training audio. The system is not hostile to it. It has barely heard it.
This is also where the fix gets misunderstood. You do not solve it by gathering thousands more recordings in the accent it already handles — that adds nothing that was missing. You certainly do not solve it by asking people to speak differently, which is hiding the problem rather than fixing it. You solve it by adding the examples that were never there.
Two mistakes that are easy to make
Both matter more than they look, and the second is the half most explanations skip.
“The lowest number is always the problem”
Sometimes a system is genuinely fair. Take one scored across four conditions: 93%, 90%, 89%, 90%. There is a lowest bar, and it means nothing — the conditions differ slightly and the whole spread is four points. Anyone who learns to hunt for a victim in every dataset has learned suspicion, not measurement, and suspicion is not a skill. Fair results exist, and being willing to say so is part of reading these honestly.
“Any gap is unfairness”
Zero out of two looks alarming and means almost nothing. A small gap, or a group with only two or three results behind it, is not yet evidence of anything. Declining to reach a verdict on a tiny sample is the grown-up half of this, and small differences and small samples both deserve a shrug.
The question worth carrying
Any system that judges, sorts, recommends or scores can be interrogated with one question, and it does most of the work:
It works on a hiring screen, a lending decision, a photo app that struggles with some faces, a voice assistant that mishears you, a feed that never shows you certain things. It is a better question than “is this AI biased?”, because it is answerable, and because it points straight at the fix.
Read next
Was this useful?
Would you want your child to know this?
FutureMinds has a level where a report card looks great until you break it apart — and the child has to find who is being let down. Take the 2-minute AI Smart Challenge yourself first — five real questions from the app, no signup.
Take the Challenge →
Join the conversation
Comments are read by a person before they appear — usually within a day or two. Nothing is published automatically.