It has long been known that Artificial Intelligence (AI) can be biased. It can absorb the subconscious belief systems of its programmers. AI can also reflect the worldview embedded in the content used to train it.
We have even known for more than a decade that AI can be sexist or racist. In 2015, Amazon discovered that its new AI-based recruiting engine discriminated against women. The team had designed programs to review job applicants’ resumes with the aim of mechanising the search for top talent.
AI and systemic racism
The computer models learned to vet applicants by observing patterns in resumes submitted over the previous 10 years. As a result, the system found mostly male applicants. This was inevitable because women were under-represented in the tech industry.
Then, the system taught itself to prefer male candidates, since men appeared more likely to be hired. By the beginning of 2017, the team behind the system disbanded. The case left the world with a clear example of the dangers of automating hiring.
In 2019, a study reviewed 3.2 million mortgage applications and 10 million refinance applications from major US home loan providers. The researchers found evidence of racial discrimination in both face-to-face lending and AI-based lending.
The National Bureau of Economic Research reported that Black and Latino applicants had higher rejection rates of 61%, compared to 48% for everyone else. They also paid up to 7.9 basis points more in interest. That translated into an annual “race premium” of over $756 million.
A pattern of racial bias in treatment recommendations
Many expected such bias to have disappeared by the mid-2020s. This was a time when society was supposedly hyper aware of the dangers. However, a new study led by Cedars-Sinai Health Sciences University has now found racial bias in treatment recommendations from leading AI platforms for psychiatric patients. This study provides one of the clearest recent examples of AI racism in healthcare.
“The findings highlight the need for oversight to prevent powerful AI applications from perpetuating inequality in healthcare,” said the institution. Cedars-Sinai aims to advance groundbreaking research and educate future leaders in medicine, biomedical sciences and allied health sciences.
Investigators studied four Large Language Models (LLMs). These are AI algorithms trained on enormous amounts of data. In medicine, LLMs are attracting attention for their ability to evaluate and recommend diagnoses and treatments quickly, according to the University.
The study found that the LLMs, when presented with hypothetical clinical cases, often proposed different treatments for psychiatric patients when African American identity was stated or implied. This contrasted with cases where race was not indicated.
AI racism in mental health
This result highlighted the seriousness of AI racism in healthcare, especially in mental health. Diagnoses were relatively consistent, yet treatment recommendations varied. The findings, published in the peer-reviewed journal NPJ Digital Medicine, were startling.
Most of the LLMs showed some form of bias when dealing with African American patients. At times, they made dramatically different recommendations for the same psychiatric illness and otherwise identical patient. This bias was most evident in cases of schizophrenia and anxiety.
According to Dr Elias Aboujaoude, Director of the Program in Internet, Health and Society at Cedars-Sinai and corresponding author of the study, the disparities were clear.
The study uncovered a wide range of examples:
Two LLMs omitted medication recommendations for an attention-deficit/hyperactivity disorder case when race was explicitly stated. Yet, they suggested medication when racial details were missing.
Another LLM suggested guardianship for depression cases with explicit racial characteristics.
One LLM focused heavily on reducing alcohol use in anxiety cases only for patients explicitly identified as African American or with a common African American name.
Racial bias in training content
Aboujaoude suggested that the LLMs reflected the bias found in their extensive training content. He explained that future research should develop strategies to detect and quantify bias in AI platforms and training data. He also recommended creating LLM architectures that resist demographic bias and establishing standardised protocols for clinical bias testing.
His colleague, Dr David Underhill, chair of the Department of Biomedical Sciences at Cedars-Sinai, proposed a way forward:
The findings of this important study serve as a call to action for stakeholders across the healthcare ecosystem. We must ensure that LLM technologies enhance health equity rather than reproduce or worsen existing inequities. Until that goal is reached, such systems should be deployed with caution and with consideration for how even subtle racial characteristics may affect their judgment.
The Cedars-Sinai study underscores the urgent need to confront AI racism in healthcare. This step is necessary not only to protect patients but also to shape fair and equitable AI systems for the future.