The cleanest way to test hiring bias is also the simplest. Researchers write resumes that are identical in every way that should matter, change only the name at the top, and mail them to real job postings. Then they count who gets called back. Because the qualifications are held constant by design, any difference in the response rate has one remaining explanation. This method has been run repeatedly since the early 2000s, across cities, industries, and decades, and the results have been consistent enough that the debate has moved from whether the gap exists to how large it is and where it lives.

The study most people cite came out in 2004. Two economists sent about 5,000 resumes to help wanted ads in Boston and Chicago, assigning names that prior surveys showed were strongly associated with white or Black applicants. Resumes with white sounding names received roughly 50 percent more callbacks. The researchers translated that gap into a rough equivalent of eight additional years of work experience, which is a striking way to state it. They also found that improving a resume helped white names considerably more than it helped Black names, meaning the return on building a better record was not the same for everyone applying.

A much larger study landed almost twenty years later and changed what we know in an important way. Researchers sent roughly 83,000 fictitious applications to more than 11,000 openings at 108 of the largest employers in the country. On average, applications with distinctively Black names received about nine percent fewer contacts. That is a smaller gap than the 2004 figure, which matters, but the headline average buried the more useful finding. The gap was not spread evenly across employers. It was concentrated in a relatively small share of companies, with a minority of firms producing most of the total measured difference.

That concentration is the part worth sitting with. It means the problem is not a diffuse fog hanging over every hiring desk in America. It is a set of identifiable places where something in the process is going wrong, sitting alongside a majority of employers whose callback rates showed little or no gap at all. The researchers also found the differences were larger in customer facing roles and clustered in particular industries, including auto sales and some retail segments. When a problem concentrates like that, it becomes fixable in a way that a vague cultural explanation never is.

Applicants figured out the pattern long before the studies confirmed it. A 2016 study looked at what researchers called resume whitening, where minority candidates soften or remove signals of their background by using a nickname, dropping an affinity group from the activities section, or listing an award without the identifying part of its name. Whitened resumes drew callbacks at more than twice the rate of unwhitened ones in that experiment. The same study found something worse. Job ads that advertised a commitment to diversity did not reduce the advantage of whitening, and applicants who trusted those ads and left their resumes intact still lost callbacks.

The cost is not just the interview that never happened. A callback gap at the first screen compounds through the whole arc of a career, because the job you do not get is also the experience you do not build, the network you do not join, and the salary base every future offer gets anchored to. Someone sending 40 applications instead of 30 for the same number of interviews is spending real hours that nobody reimburses. Multiply a nine percent contact gap across a labor market and it stops being a rounding error. It becomes a durable difference in earnings that never traces back to any single decision anyone remembers making.

The fix people reach for first is name blind screening, and it helps at exactly one stage. Several national governments and large employers have run it, and it moves the first gate while doing nothing about interviews, referrals, or offers, where names are visible again. The stronger move is measurement. Track your own callback rate by stage, compare it against your applicant pool, and look at whether a handful of managers or one requisition type accounts for most of the spread. The concentration finding suggests most organizations are not looking at a company wide culture problem. They are looking at a few specific places where the numbers do not hold up, and those can actually be found.