Proxy Discrimination Risks in Automated Resume Screening Systems
Screening algorithms can discriminate through data fields that correlate with protected traits.

A candidate submits a resume with no name field visible to the hiring algorithm, no box checked for race or gender, no date of birth entered anywhere in the application. He still gets filtered out before a human ever opens the file, and the reason has nothing to do with the system failing to work as designed. You don't need protected-class data to discriminate when you screen resumes automatically. It only needs fields that correlate with protected classes closely enough to reproduce the same sorting, and most resumes are full of them.
A lot of hiring software, and the people who buy it, assumes that stripping out race, gender, and age fields solves discrimination. That assumption misreads how statistical correlation operates inside a trained model. A system doesn't need to see "race" to sort by race if it has access to a dozen other fields that move together with race in the underlying population. This holds true for employers acting in complete good faith, with no intent to exclude anyone. The mechanism runs on math, not motive, and that is what makes it hard to catch and harder to defend against after the fact.
Scale is what turns this from a flaw into a structural problem. A human reviewer with a bias toward certain schools or against certain neighborhoods only makes a limited number of bad calls across a career. But an automated screen applies the same flawed logic to every application that crosses the platform, the same way, thousands of times over, before anyone notices the pattern. The rest of this piece works through how that happens: the statistical mechanics of proxy variables, the role of historical training data, a newer wrinkle involving AI-written resumes, the documented harms that have resulted, and the legal and regulatory response now taking shape around all of it.
How proxy variables link neutral fields to protected characteristics
A proxy variable doesn't have to be designed to discriminate. It only has to correlate with a protected characteristic strongly enough that filtering on the proxy produces the same demographic skew as filtering on the protected trait directly. That is the entire mechanism, and it requires no intent on anyone's part, just enough statistical overlap between an ordinary data field and a legally protected one.
ZIP code is the clearest example because residential segregation in this country has never been randomly distributed. Where someone lives tracks closely with race and with household income, because housing policy has shaped both for a long time. A system that scores applicants down for living in certain postal codes is, in effect, running a geographic filter for race, whether or not anyone involved in building the system intended that outcome. This is not a hypothetical concern. A 2025 lawsuit against Sirius XM alleges that an AI screening tool used ZIP codes, schools, and employment history as proxies for race. The allegations have not been ruled on, but the case shows how plaintiffs are now framing these claims in court, treating the proxy relationship itself as the discriminatory act.
School name works the same way, and it arguably carries more hidden dimensions. Where someone went to school correlates with race, with socioeconomic background, and with access to disability services during their education. A filter that favors graduates of certain universities over others doesn't make just one judgment. It's making three at once, along race, class, and disability status simultaneously, under the appearance of evaluating academic pedigree.
Employment gaps carry their own layered correlation. A gap in someone's work history might reflect caregiving responsibilities, which fall disproportionately on women, or medical leave, which implicates disability status, or a stretch of unemployment concentrated in lower-income communities. None of these appear as "gender" or "disability" in the data. They show up as a blank space on a timeline, and a system trained to penalize blank spaces penalizes all three populations without ever asking about any of them.
Graduation year is a smaller but sharper example. It reveals age almost directly, so a date that looks like simple biographical detail becomes a filtering signal, and it produces age discrimination without the system ever touching a field protected under the Age Discrimination in Employment Act.
Names carry the same risk through linguistic and cultural association. A 2024 University of Washington study by Wilson and Caliskan found that a language-model retrieval system strongly favored white-associated names, and it selected Black-associated names only 8.6% of the time across a substantial sample of CVs. And these effects compound when proxies appear together. ZIP code and employment history combined create a stronger correlation with protected characteristics than either produces alone. An audit that checks features one at a time, in isolation, can miss the effect that only appears when several ordinary fields interact.
How biased training data becomes algorithmic hiring rules
Proxy variables are one entry point for bias, with another lying upstream in what the system was trained to reproduce. When an AI model learns from a company's past hiring decisions, it absorbs whatever patterns those decisions carried, and that includes patterns shaped by discrimination that happened long before the software existed. The system then applies those patterns at scale, treating them as though they were neutral predictors of who performs well on the job.
Amazon's scrapped recruiting tool from 2018 remains the clearest illustration. Trained on a decade of resumes submitted mostly by men, the algorithm learned to penalize the word "women's" and to downgrade graduates of women's colleges. Nobody programmed it to do either thing. The training data had defined what a "good candidate" looked like in gendered terms, and the model simply learned that definition and reproduced it.
The more pressing concern today is how routine this practice has become across the industry. Eightfold, one vendor in this space, describes training its models on more than tens of millions of historical candidate-position pairs with known outcomes. That scale is the point: if the historical outcomes used to train a system carried bias, the resulting model doesn't just repeat that bias, it applies it with more consistency and at a volume no human reviewer could match. Research examining the validity of LLM-based resume screening has started asking a pointed question: are these systems scoring candidates on substantive, job-relevant signals, or on superficial features that happened to correlate with who got hired in the past? The answer matters for fairness, and it matters just as much for whether these tools predict job performance.
A newer bias vector: AI self-preferencing among candidates who used AI to write their resumes
A third mechanism has emerged more recently, and it follows directly from the logic of the first two. A screening model's judgment is shaped not only by its training data but by the particular linguistic patterns of the underlying language model it runs on, which creates a new channel for bias having nothing to do with ZIP codes or school names.
Candidates who use the same underlying model as the screener to write their resumes are substantially more likely to get shortlisted because their writing sits closer to the model's own output patterns rather than because their qualifications are any stronger. An analysis by Xu, Li, and Jiang found that candidates using the same large language model as the evaluator were 23 to 60% more likely to be shortlisted than equally qualified applicants who submitted human-written resumes. Access to premium AI writing tools isn't distributed evenly across candidates, so this produces a new proxy: AI-tool access, which tracks with socioeconomic status, educational background, and digital literacy, each of which tracks in turn with protected characteristics.
What makes this mechanism especially hard to catch is that it can pass straight through a conventional bias audit. An audit built to check for demographic disparities in final outcomes won't flag a screener that favors AI-polished writing style, because the disparity it produces tracks access to technology, not race or gender directly, even if the downstream effect lands on the same populations. The deeper lesson here is that removing old proxy fields doesn't close the door on proxy discrimination. New proxies form as the technology changes, so any audit built around yesterday's risk factors will miss tomorrow's.
Documented outcomes of proxy discrimination across race, gender, age, and disability
The mechanisms above are not theoretical. They have produced measurable, documented harm across race, gender, age, and disability, and the direction of that harm is not consistent from one system to the next, which is itself a significant finding.
On race, the FAIRE study found that a language-model retrieval system favored white-associated names the large majority of the time and selected Black-associated names only a small fraction of the time, across a substantial sample of matched CVs. Because the underlying resumes were matched for qualifications, the gap can't be explained by any difference in candidate quality.
On age and gender together, the clearest case remains the EEOC's settlement with iTutorGroup. An automated system there auto-rejected female applicants over 55 and male applicants over 60, a pattern that led to a $365,000 settlement and stands as the most concrete documented instance of age-by-gender proxy filtering turning into legal liability.
Race and gender don't always move in the same direction, which matters for how employers should think about fairness testing. Research published through VoxDev found that AI hiring tools favored female applicants over Black male applicants who had identical qualifications. Bias direction depends entirely on which groups are being compared, and that means intersectional analysis isn't optional if an employer wants to know what its system is actually doing.
Disability deserves more attention than it usually gets in these conversations. Employment gaps tied to medical leave or to treatment and recovery are routinely treated as negative signals by screening systems, disadvantaging applicants with disabilities even though those same gaps are legally protected activity under disability-rights law.
Taken together, these findings carry a warning for any employer tempted to treat a clean bias test on one dimension as proof of overall fairness. A system can show no disadvantage against one group and still run bias against a different group, or against an intersectional combination that a simpler test never examined. A study titled "Fairness Is Not Enough" ran an extensive battery of comparisons and identified what its authors called an "Illusion of Neutrality": some models post statistically flat racial bias scores because they match resumes on superficial keywords. Apparent neutrality, in those cases, masks a system that isn't doing its job.
The landmark cases that are setting the legal boundaries of vendor and employer liability
The documented harms above have moved from research findings into active litigation, and the central legal question courts are now working through is who bears liability when AI screening discriminates, and the early answers are extending that exposure beyond employers to the vendors who build the software.
The most consequential case currently active is Mobley v. Workday, filed in 2023 and still ongoing. The lead plaintiff, Derek Mobley, is Black, over 40, and lives with anxiety and depression. He alleges he was rejected across many applications at companies using Workday's platform, and he often received automated rejections within an hour of applying or in the middle of the night. His claim is that the system sorted through proxies, including employment gaps, in ways that disadvantaged him on the basis of race, age, and disability at once.
The rulings in the case, as of August 2026, have steadily widened its scope. In July 2025, a court ordered Workday to identify which employers had enabled its HiredScore AI screening features. In early 2026, the court confirmed that applicants aged 40 and older can challenge AI screening tools under disparate-impact doctrine, a standard that doesn't require proof of intentional discrimination, only proof of disproportionate effect. On June 22, 2026, Judge Rita F. Lin rejected Workday's argument that it was "merely a tool provider" and allowed claims under a state fair-employment statute and a federal disability-rights law, built on a proxy-discrimination theory, to proceed. None of this amounts to a final finding that Workday's systems discriminated. Workday disputes the allegations, and the case continues. What matters for the industry is the reasoning behind the vendor-liability ruling: a tool that makes hiring cuts on an employer's behalf is treated as acting as the employer's agent. If that reasoning holds up and spreads to other courts, every company that builds or buys AI recruiting software takes on a category of legal exposure that didn't exist before.
A second case, already settled, shows what that exposure looks like in dollar terms. The EEOC's 2023 action against iTutorGroup, built on the same age-and-gender auto-rejection pattern described above, produced a $365,000 settlement, the first EEOC enforcement action to result in payment specifically for AI-driven age discrimination. It established that an automated rejection is just as actionable under the ADEA as a decision made by a human.
One development inside the Mobley litigation cuts against the push for transparency and deserves to be understood clearly. A court declined to turn over Workday's bias-testing data to the plaintiffs after Workday represented that its own counsel had curated that data and that the testing had been conducted as legal advice rather than for ordinary business purposes, which shielded the material under attorney-client privilege. The incentive this creates is straightforward and troubling: route internal bias testing through corporate counsel, and the resulting data becomes harder to reach in litigation. Guidance already circulating among compliance practitioners recommends exactly that approach. The legal system's own evidentiary rules may end up discouraging the open, transparent bias auditing that would actually catch problems like these.
The regulatory landscape: federal retreat, state advance, and the EU's hard requirements
The regulatory response to all of this has not moved in one direction. Federal enforcement of disparate-impact liability for AI hiring has narrowed as a matter of executive policy, reducing the pressure that used to come from federal agencies pursuing these claims as a matter of course. State law has moved the opposite way, and the EU AI Act imposes hard requirements of its own on systems used for employment decisions. A patchwork of overlapping state rules and binding international requirements, taken together, is more demanding on employers and vendors than any single federal standard would have been on its own. Companies operating across state lines, or doing business in the EU, now have to satisfy the strictest rule that applies to them in any jurisdiction where they hire. The practical compliance bar has risen even as federal enforcement has pulled back.
Sources
- AI Bias in Hiring: Algorithmic Recruiting and Your Rights
- FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
- Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
- AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights
- Measuring Validity in LLM-based Resume Screening
- Unlawful Proxy Discrimination: A Framework for Challenging Inherently Discriminatory Algorithms
- AI-Assisted Hiring in 2026: Managing Discrimination Risk
- Gender, Race, and Intersectional Bias in Resume Screening via Language Model Retrieval


