Bias and Fairness
An AI system can be accurate on average and still fail one group of people far more often than another. Bias can enter at every step from the choice of target to the feedback loop, and deleting a sensitive column does not remove it. Fairness is a choice leaders make for each decision, test by group and keep monitoring.
After this chapter you can
- Explain why high overall accuracy can hide large differences between groups.
- Identify where bias enters, from the choice of what to predict to the feedback loop.
- Explain why removing sensitive attributes does not remove bias, and why testing needs them.
- Recognize that fairness criteria can conflict and must be chosen for each decision.
- Name the EU and US legal anchors for AI used in hiring and match testing depth to impact.
In 2018 two researchers, Joy Buolamwini and Timnit Gebru, tested three commercial systems that classify a person’s gender from a photograph of their face. Each system looked accurate when its results were reported as one number. Then the researchers split the results by skin type and gender. For lighter-skinned men, no system erred more than 0.8 percent of the time. For darker-skinned women, the worst system erred 34.7 percent of the time, roughly one face in three1.
Nobody designed those systems to fail one group. The researchers found that widely used face datasets were overwhelmingly made up of lighter-skinned people, and a headline accuracy figure is often not broken down far enough to show what that does. That combination, an uneven failure that no one intended and an average that hides it, is the core of AI bias, and it is why an executive approving an AI system needs to ask more than whether it is accurate.
Accurate is not the same as fair
The previous chapter, Accuracy, Hallucination and Reliability, asked how often a system is right and on what kind of work. This chapter asks a different question: when the system is wrong, who is it wrong about?
Bias, in this sense, is a systematic pattern that produces consistently different outcomes for particular people or groups. It does not require intent, and often there is none. Fairness is the judgment about whether those differences are appropriate and justified for this particular decision. A film recommendation that suits one audience better than another is a minor matter. A screening tool that rejects qualified applicants from one group more often is a serious one, with legal duties attached. The stakes of the decision set the standard.
Two consequences follow. First, fairness is not a property of a model alone. It belongs to the whole decision: what the system was asked to predict, the data it learned from, where the cut-off was set and what people do with the output. Second, an accuracy figure, however high, cannot by itself answer the fairness question. Accuracy is one number for everyone. Fairness needs an answer for each group the decision affects.
Where bias enters
The instinct is to look for bias inside the model. The US National Institute of Standards and Technology (NIST) takes a wider view. Its bias standard describes three sources: systemic bias carried in institutions and historical records, statistical bias in data and methods, and human bias in the way people build and use systems. It stresses that bias can enter at any stage of an AI system’s life, not only during training2. The NIST AI Risk Management Framework lists “fair, with harmful bias managed” as one of the marks of a trustworthy system for the same reason3.
The chain is worth walking slowly. The problem comes first: a system asked to optimize the speed of hiring will pursue speed, not quality of hire, and do so very efficiently. Data and labels come next. Historical data records historical decisions, so a label such as “good employee” or “likely to default” may encode the habits of the managers and lenders who made those calls. The threshold, where the cut-off score sits, decides who is wrongly flagged and who is wrongly missed. People decide when to trust the output and when to override it. And feedback closes the loop: the people a system selects today generate the success data it learns from tomorrow, so a gap can widen on its own.
The first link is the one leaders own, and it is easy to get wrong because a convenient stand-in looks like the real target. A widely cited 2019 study of a US health-care algorithm found that predicting future costs instead of health needs under-selected Black patients for extra care, because less money is spent on them at the same level of need; the researchers estimated that the correct target would have raised their share from 17.7 percent to 46.5 percent4. Nothing in the model was malicious, and race was not an input. The bias lay in the choice of target, which is a leadership decision, usually made in a project brief long before any data scientist is involved.
Deleting the column does not delete the signal
A common executive response to bias is to remove the sensitive attribute: take gender, ethnicity and age out of the data, and the system cannot discriminate. It can. Other variables carry the same information. A proxy is a variable that correlates with a characteristic you care about, and real data is full of them. Postal code tracks income and often ethnicity. The college someone attended can track gender or background. Career gaps track caregiving. The words people choose, the clubs they list and even their first names carry signals a model can learn.
Tools built on language models have not escaped this. In a 2024 laboratory test of three text-embedding models, names alone moved resume rankings, favoring white-associated names in most test conditions5. NIST’s profile for generative AI lists harmful bias among the risks these systems carry for the same reason: they learn from vast amounts of text that records how people have written about one another6.
Proxies can also be chosen directly. As The AI Risk Landscape described, the Dutch tax authority’s childcare-benefits risk model used nationality as one of its indicators; no clever inference was needed. Whether a proxy is chosen or learned, the response is the same, and it is counterintuitive. To find out whether outcomes differ by group, you usually need the sensitive attribute, held separately and securely, so you can test. Deleting it often makes the problem harder to see without making it smaller. How such data should be held is a privacy question, taken up in the next chapter.
The average hides who carries the errors
The face-classifier result is the pattern to look for. A headline figure can be strong while one group carries most of the errors, and the only way to see it is to split the results. A system that looks 92 percent accurate overall has not shown it is fair until someone has asked how it performs for each group it affects.
The next question is which kind of error. A false positive wrongly flags someone; a false negative wrongly misses someone. In fraud detection, a false positive blocks a legitimate customer’s card. In hiring, a false negative means a qualified candidate never gets an interview and never learns why. Moving the threshold trades one kind of error for the other, and it can shift which group bears them. Before approving any system that makes decisions about people, ask for three things: performance overall, performance by group, and who experiences which kind of error.
Fair by which definition?
Once results are split by group, the next question is uncomfortable: fair by which definition? Several reasonable ones exist. Equal selection asks whether groups receive a positive outcome at similar rates. Equal opportunity asks whether, among people who truly qualify, each group is selected at a similar rate. Equal error rates asks whether wrong flags and wrong misses fall on each group alike. Equal reliability asks whether a positive prediction means the same thing whichever group it is about.
These can point in opposite directions, as a well-known 2016 dispute over a court risk score showed: an analysis found that wrong high-risk flags fell on Black defendants at nearly twice the rate of white defendants, while the tool’s maker replied that its scores predicted reoffending about as well for both groups7.
Both sides were measuring something real. Shortly afterwards, Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan proved why the argument could not be settled by more analysis: when groups have different underlying rates of the outcome, a score cannot be equally reliable for every group and also spread its errors equally, except in trivial cases8. Someone has to choose.
That turns a technical question into a leadership one. The wrong question is “what is the fairness metric?” The right one is “which fairness objective fits this decision, given the cost of each kind of error, and who has signed off on it?” Better data often improves accuracy and fairness together, so fixing the data comes first. When a real trade-off remains, it should be explicit, owned by a named leader and written down, not buried in a vendor’s default setting.
Hiring AI already carries legal duties
For most uses of AI, fairness is a matter of trust, reputation and good business. For hiring, credit and access to essential services, it is also a matter of law. Hiring is the clearest case, and three anchors matter most to an organization operating in Europe or the United States.
A law on paper is not the same as a law that bites, and New York City’s experience is instructive for any leader who assumes compliance equals safety. A year after enforcement began, researchers checked 391 employers and found that only 18 had posted a bias-audit report and 13 a candidate notice; the law lets each employer decide whether its own tool is in scope10. In December 2025 the New York State Comptroller audited the city agency responsible for enforcement and judged it ineffective. The agency had surveyed the websites and bias audits of 32 companies and found a single issue; the Comptroller’s auditors, reviewing the same companies, found at least 17 instances of potential non-compliance. Nine of the auditors’ 12 test calls to the city’s 311 line about these tools never reached the agency at all11.
The lesson for executives is not that the rules can be ignored. The Comptroller’s findings put pressure on the city to enforce more firmly, and the EU regime is far broader. It is that a bias audit is the floor, not the ceiling. A system that passes a narrow legal test can still fail the people it screens, and “we were never asked to test by group” is no longer a defensible answer.
Test in proportion to impact
None of this means every AI system needs a fairness review board. The depth of testing should match what is at stake.
For a low-impact system, such as recommending articles, basic checks before launch are enough. For a medium-impact system, such as tailoring prices or offers, ask for performance by group before launch and monitoring after. For a high-impact system, such as hiring, credit or eligibility for a service, ask for a formal fairness assessment, a person who reviews and can overturn the output, and a way for the people affected to question a result and reach a human.
The sequence before launch is simple: define the groups and the outcomes that matter, measure, compare, investigate the differences, mitigate and test again. Sometimes the best mitigation is not technical at all; it is a redesign of the workflow so that the system recommends, a person reviews and a person decides. After launch, keep watching. Populations, behavior, data and policy all change, so a system that was fair at launch can drift. How these checks fit into approval processes and registers belongs to Module 07; the point here is the habit of asking.
Story: a recruiting model that learned the past
In October 2018 Reuters reported on an experiment inside Amazon12. In 2014 a team had begun building a tool to automate the review of CVs, rating candidates from one to five stars, much as shoppers rate products. The model was trained on ten years of CVs submitted to the company, and most of those came from men, which reflected the make-up of the technology industry.
By 2015 the company realized the system was not rating candidates for software and other technical roles in a gender-neutral way. It penalized CVs that included the word “women’s”, as in “women’s chess club captain”, and it downgraded graduates of two all-women’s colleges. It also favored verbs such as “executed” and “captured”, which appeared more often on men’s CVs. Nobody told it to do any of this. It had learned what successful applicants had looked like in the past.
The engineers edited the model to be neutral to those particular terms. But, as Reuters put it, that was no guarantee the system would not devise other ways of sorting candidates that could prove discriminatory. The team was disbanded by early 2017. Reuters reported that recruiters had looked at the tool’s recommendations but never relied on them alone; the company said the tool was never used by its recruiters to evaluate candidates.
Every link of the chain is in this story. The label was past hiring decisions. The data under-represented one group. Deleting the obvious signal left the proxies in place. And the decisive safeguard was organizational: people kept the decision, and the company was willing to stop. An organization buying a screening tool today may have neither the in-house team that spotted the problem nor the habit of looking for it.
What this means for leaders
Bias is not a technical defect to be delegated to data scientists. It starts with the choice of what to predict, it lives in data that records your organization’s own history, and the decision about which kind of fairness matters is a value judgment that belongs to the business. Four habits follow. Look at results by group before you look at the average. Ask what the labels really measure. Keep the sensitive data you need to test, under proper control. And treat fairness as something to monitor, not a box ticked at launch.
Check yourself
- Removing gender and ethnicity from the data makes a model fair.
- A system with high overall accuracy can still fail one group far more often than another.
- Bias in AI comes mainly from the model’s algorithm.
- Common fairness criteria can conflict, so someone has to choose which one fits the decision.
- Passing a legally required bias audit shows a hiring tool is fair.
- Once a system has been tested for fairness before launch, the job is done.
Reflection: find the hidden group
What comes next
Testing for bias often means holding sensitive data about people, and every AI system processes information that someone wants kept private. The next chapter, Privacy and Confidential Data, looks at where that information goes when it enters an AI system, who can see it, and what leaders should require before confidential data goes in.
Laws referenced
Not legal advice. Laws change; verify before relying on this, and consult counsel for decisions.
EU AI Act · EU
Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744
Risk-based rules. Prohibited practices include social scoring, untargeted scraping of facial images, and emotion recognition in workplaces and schools (with narrow exceptions). High-risk systems (Annex III: biometrics, safety components of critical infrastructure such as energy, water and traffic, employment and worker management, credit, education, essential services, law enforcement, migration, justice) need risk management, data governance, documentation, logging, human oversight, human oversight that keeps people able to understand the system, notice automation bias (over-reliance on its output), override it or stop it (Art. 14(4)), appropriate accuracy, robustness and cybersecurity (Art. 15), automatic logging of events (Art. 12), a provider quality-management system (Art. 17) and conformity assessment. An Annex III system is not high-risk if it poses no significant risk of harm, for example a narrow procedural or preparatory task that does not replace human assessment; systems that profile people are always high-risk, and a provider relying on this exception must document it and register (Art. 6(3)). Deployers of high-risk AI must use it as instructed, assign competent human oversight, monitor its operation, keep logs for at least six months and report serious incidents (Art. 26); employers must inform workers' representatives (Art. 26(7)). Public bodies, private providers of public services, and deployers of credit-scoring or life and health insurance pricing systems must carry out a fundamental-rights impact assessment before first use (Art. 27). Providers must run post-market monitoring (Art. 72). A deployer that puts its name on a high-risk system, substantially modifies it, or changes its purpose so that it becomes high-risk takes on the provider's obligations (Art. 25(1)). A substantial modification (Art. 3(23)) of a high-risk system needs a new conformity assessment, unless the change was pre-determined and documented at the first assessment, as with planned continuous learning (Art. 43(4)). Providers of general-purpose AI models (from 2 Aug 2025) must keep technical documentation, have a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publish a sufficiently detailed summary of training content (Art. 53). Research, testing and development before a system is placed on the market or put into service is outside the Act, except testing in real-world conditions (Art. 2(8)). Since the 2026 Omnibus, the Art. 4 AI-literacy duty is an obligation of effort (take measures to support literacy), not of result. Fines reach EUR 35 million or 7% of global turnover for prohibited practices.
- 2024-08-01 — Entered into force
- 2025-02-02 — Prohibited practices (Art. 5) and the AI-literacy duty (Art. 4) apply
- 2026-07-27 — Omnibus softens Art. 4: providers and deployers must take measures to support AI literacy; no specific level must be guaranteed
- 2025-08-02 — General-purpose AI model obligations apply; governance and penalties regime in place
- 2026-08-02 — Transparency duties (Art. 50) apply: disclose AI interaction, label synthetic and deepfake content (marking for generative systems already on the market: 2 Dec 2026)
- 2027-12-02 — High-risk obligations for Annex III systems (e.g. hiring, credit, education, essential services) - moved from 2 Aug 2026 by the 2026 Omnibus
- 2028-08-02 — High-risk obligations for AI in products regulated under Annex I
Last verified 2026-10-06 · official text
General Data Protection Regulation · EU
Regulation (EU) 2016/679
Personal data is any information relating to an identified or identifiable person, directly or indirectly, including by an identifier such as an online ID (Art. 4(1)). Lawful basis and purpose limitation (Arts. 5-6); processing special-category data, including biometric data used to identify a person, health data and data revealing ethnicity, is prohibited unless a specific exception applies (Art. 9); data protection by design and by default (Art. 25); processors such as AI vendors may act only under a written contract with required terms and sufficient guarantees (Art. 28); transparency to data subjects (Arts. 13-14); right not to be subject to a decision based solely on automated processing with legal or similarly significant effects (Art. 22); breach notification to the supervisory authority within 72 hours (Art. 33) and to individuals without undue delay when the risk is high (Art. 34); data protection impact assessment for high-risk processing (Art. 35). Fines up to EUR 20 million or 4% of global turnover.
- 2018-05-25 — Applies
Last verified 2026-10-08 · official text
NYC Local Law 144 (automated employment decision tools) · US - New York City
NYC Local Law 144 of 2021; DCWP rules
An automated tool that substantially assists hiring or promotion decisions needs an independent bias audit within the past year, a published summary of results, and notice to candidates at least ten business days before use. Penalties USD 500 to 1,500 per violation.
- 2023-07-05 — Enforcement began
Last verified 2026-10-06 · official text
References
- Joy Buolamwini and Timnit Gebru. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of Machine Learning Research 81:77-91 (Conference on Fairness, Accountability and Transparency). 2018.
- Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt and Patrick Hall. Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, NIST SP 1270. NIST. 2022.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. NIST. 2023.
- Ziad Obermeyer, Brian Powers, Christine Vogeli and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464):447-453. 2019.
- Kyra Wilson and Aylin Caliskan. Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES) 7. 2024.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
- Jeff Larson, Surya Mattu, Lauren Kirchner and Julia Angwin. How We Analyzed the COMPAS Recidivism Algorithm. ProPublica. 2016.
- Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan. Inherent Trade-Offs in the Fair Determination of Risk Scores. arXiv (Innovations in Theoretical Computer Science 2017). 2016.
- European Commission, AI Act Service Desk. AI Act, Annex III: High-risk AI systems referred to in Article 6(2). European Commission. 2024.
- Lucas Wright, Roxana Mika Muenster and colleagues. Null Compliance: NYC Local Law 144 and the Challenges of Algorithm Accountability. ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). 2024.
- Office of the New York State Comptroller. Enforcement of Local Law 144 - Automated Employment Decision Tools (audit of the NYC Department of Consumer and Worker Protection). Office of the New York State Comptroller. 2025.
- Jeffrey Dastin. Amazon scraps secret AI recruiting tool that showed bias against women. Reuters. 2018.
Further reading
- Reva Schwartz, Apostol Vassilev, Kristen Greene, Lori Perine, Andrew Burt and Patrick Hall. Towards a Standard for Identifying and Managing Bias in Artificial Intelligence, NIST SP 1270. NIST. 2022.
- Ziad Obermeyer, Brian Powers, Christine Vogeli and Sendhil Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366(6464):447-453. 2019.
- Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan. Inherent Trade-Offs in the Fair Determination of Risk Scores. arXiv (Innovations in Theoretical Computer Science 2017). 2016.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1. NIST. 2024.
Sources last verified 2026-10-08.