In 2003 two economists ran an experiment that people still argue about. Marianne Bertrand and Sendhil Mullainathan built roughly 5,000 fake resumes and mailed them to about 1,300 real job ads in Boston and Chicago. The resumes were built in matched pairs with the same schools, the same job history and the same skills. The only thing that changed was the name at the top. Half carried names common among white applicants, such as Emily and Greg. Half carried names common among Black applicants, such as Lakisha and Jamal.

The white sounding names drew about 50 percent more callbacks. Put in plain terms, one group needed roughly ten resumes to get a call and the other needed about fifteen. The study was published in the American Economic Review in 2004 and it changed the shape of the argument, because it removed the usual explanations. Nobody could point to school quality, work gaps or interview manner, since the paperwork was identical by design. The only variable in the room was the name. The paper was the same. The name was not. That was the whole test. It has been run many times since then.

The finding that stung hardest was about credentials. The researchers also varied resume quality, adding real experience and better skills to some applications. Stronger resumes with white sounding names got a meaningful bump in callbacks. Stronger resumes with Black sounding names got almost none. In other words, the return on doing everything right was not the same for both groups. The common advice to simply be twice as good ran into a wall in the data.

That was more than twenty years ago, and the obvious question is whether it still holds. A 2017 review in the Proceedings of the National Academy of Sciences pooled every field experiment of this kind run in the United States since 1989. It found that discrimination against Black applicants at the callback stage had not measurably declined over almost three decades. A larger study published in 2021 sent tens of thousands of applications to Fortune 500 companies and found the same average gap, with wide variation between firms. Some large employers showed no gap at all. Others showed a very large one, which is the part that matters.

That variation is the useful finding, because it means the outcome is not fixed. Firms with structured hiring showed smaller gaps. Structured hiring means the job description is written before applications open, the same questions are asked in the same order, and answers are scored against a rubric rather than a feeling. Firms that let individual managers scan a pile and follow instinct showed larger gaps. The difference is not about intent. It is about how much room the process leaves for a snap judgment made in six seconds.

Blind screening helps at the first stage and does not solve the whole problem. The classic case is orchestra auditions, where screens raised the share of women advancing. Stripping names, schools and addresses off a resume before the first read produces a similar effect. The limitation is that the name reappears at the phone screen and the interview, so a blind first pass moves the bottleneck rather than removing it. It is still worth doing, because getting into the room is the step where most candidates are lost. The gap is not just a first read. It can follow a name down the line. That is why one fix on its own is not enough.

For job seekers this is a hard subject, and the honest answer is uncomfortable. Some applicants shorten or change how their name appears on a resume, and research suggests it can raise callbacks. That is a personal decision and no one owes anyone else that adjustment. The more durable moves are the ones that route around the pile entirely. Referrals, direct contact with a hiring manager, and work that can be seen before you are screened all reduce how much the first six seconds decide. None of that is fair. It is just what the data shows. The load should not sit on the person applying.

For anyone who hires, the checklist is short and unglamorous. Write the scorecard before you open the role. Strip identifying details from the first read. Ask every candidate the same questions and write down scores before you discuss anyone. Track callback rates by group and look at the number rather than your impression of yourself. None of that requires a program, a budget or a statement. It requires a process that does not depend on a good mood on a busy afternoon. That is not a big lift. It is just a habit that has to hold. Write it down and stick to it.

Sources: American Economic Review, Proceedings of the National Academy of Sciences.