
For job hunters in Phoenix, the AI gatekeeper may be learning the wrong lesson almost instantly. New research suggests some language models can watch a handful of hiring outcomes, invent a stereotype, and then funnel entire groups toward or away from certain jobs—even when every candidate is equally capable.
Researchers from Princeton University and the University of Chicago tested language models in a multiround virtual hiring game involving fictional demographic groups and dozens of jobs. As outlined in the study, every group had the same underlying chance of success, but early random results pushed the models into increasingly narrow assignments.
In the comparison reported by Arizona’s Family, the models scored about 1.83 on the study’s segregation scale, compared with 0.84 for human participants. The higher the score, the more strongly the system sorted groups into particular types of work.
The mechanism is less futuristic than it sounds: The models generalized from sparse, noisy feedback. One failed outcome involving a fictional group and a particular job could cause the system to avoid making similar assignments later, even though the failure was random and the candidates were equally qualified.
The paper also found that newer and more capable reasoning models sometimes produced stronger stratification, not less. In other words, better at drawing conclusions did not necessarily mean better at knowing when a conclusion was completely premature.
Phoenix Employers Are Already Using AI Interviews
The warning lands in Phoenix as AI moves deeper into the employment process. Arizona’s Family reported that Coinbase uses AI voice conversations with applicants, while Zapier uses AI avatars and says the technology lets it interview up to five times as many candidates.
Travis Laird, a Phoenix-area jobs expert with Robert Half, told the station that avatar interviews can give applicants a chance to show personality and context, but he also urged employers to treat the software like a new team member. That means coaching it, reviewing its work, and keeping a person involved rather than allowing an automated score to become the final word.
The Researchers Found One Fix That Actually Worked
The researchers tested more than 10 interventions, including telling models to be fair and changing how they reasoned through the task. According to the paper, the intervention that consistently reduced stratification added a diversity incentive to the model’s objective.
That finding matters because it shifts the focus from whether a model can explain a decision to what the system is actually rewarded for doing. A model can produce a polished explanation for a biased pattern and still be optimizing the wrong goal.
The Legal Risk Still Belongs To Employers
The experiment does not prove that a Phoenix employer has illegally discriminated against applicants, and the researchers caution that real-world hiring feedback is slower and more complicated than the simulation. But the U.S. Equal Employment Opportunity Commission says employment tests and selection procedures can violate federal law if they disproportionately exclude protected groups and are not job-related or legally justified.
Arizona’s Attorney General’s Office likewise lists race, color, national origin, sex, religion, age, disability and genetic information among protected categories in employment. The state guidance makes clear that hiring discrimination does not become acceptable simply because a software vendor—or a chatbot—made the recommendation.
For applicants, the practical takeaway is straightforward: An AI interview or résumé screen may be part of the process, but it is not proof that the process is neutral. For employers, the study is a reminder that preventing bias requires more than adding the word fair to a prompt; it requires meaningful oversight, testing and accountability.









