The push to remove race as an input variable from clinical algorithms, following well-documented harms in kidney function estimation, lung function testing, and vaginal birth after cesarean calculators, has moved from policy debate into direct empirical testing. The 2025 evidence is more complicated than the "remove race, fix bias" framing that dominated the initial policy response.

What did the breast cancer risk model study actually test?

A 2025 study in npj Digital Medicine directly evaluated the Breast Cancer Surveillance Consortium's 6-year cumulative advanced breast cancer risk model, a tool used clinically to decide screening frequency and eligibility for supplemental imaging like MRI. The researchers built a race-naive version of the model, removing race and ethnicity as inputs, and then measured performance across racial and ethnic subgroups compared to the original, race-inclusive model. Source: Effect of race and ethnicity on advanced breast cancer risk prediction model performance, npj Digital Medicine 2025.

The key finding: removing race and ethnicity from the model did not uniformly improve calibration or discrimination across subgroups, and in some subgroups, performance changed in ways that could affect who qualifies for supplemental imaging. This is an important empirical correction to the assumption, common in early policy commentary, that race-naive modeling is automatically the fairer approach. The relationship between race as a variable, the biological and social factors it is proxying for, and model performance across groups is more complex than a simple inclusion-or-exclusion decision.

What is the broader conceptual argument here?

A widely cited 2024 to 2025 analysis in Science Advances, "Use of race in clinical algorithms," lays out the underlying tension precisely: race is frequently used in clinical algorithms as an unexamined proxy for biological difference, when what it is often actually capturing is the downstream effect of structural inequities, differential access to care, environmental exposure, or measurement artifacts in the original data the algorithm was trained on. Source: Use of race in clinical algorithms, Science Advances.

A related 2025 review, "Clinical Algorithms and the Legacy of Race-Based Correction," traces how several race-based corrections, once considered standard clinical practice, were built on thin or outdated evidence and persisted for decades before being reconsidered, a cautionary case study for how long a flawed default can survive inside clinical guidelines once it is embedded in calculators and order sets. Source: Clinical Algorithms and the Legacy of Race-Based Correction, PMC.

What does the algorithmic detection research show?

A 2025 study in the Journal of Medical Internet Research took a different, more technical approach: rather than debating whether to include or exclude race, the researchers developed a method to detect and characterize implicit and explicit racial biases in health care datasets using subgroup learnability analysis, essentially testing whether a model's errors are systematically worse for certain subgroups even when race is not an explicit input. Source: Detecting racial biases in health care datasets, JMIR 2025.

This is a meaningful methodological advance because it addresses the core failure mode of naive fixes: a model can be "race-blind" on its input variables and still be racially biased in its output errors, if other correlated variables (zip code, insurance type, prior utilization patterns) carry the same signal race used to carry. Removing the labeled variable does not remove the underlying correlation structure in the training data.

What does the fairness literature recommend instead?

A 2025 debate paper in BMC Medicine, "Navigating fairness aspects of clinical prediction models," argues that most clinical algorithms in use today have not undergone thorough fairness evaluation at all, regardless of whether they include race as a variable, and proposes a structured framework for evaluating calibration, discrimination, and error rates across subgroups as a standard part of model validation, not an optional add-on. Source: Navigating fairness aspects of clinical prediction models, BMC Medicine 2025.

Approach testedWhat the 2025 evidence shows
Remove race as an input variableDoes not uniformly improve subgroup calibration; effects vary by model and condition
Subgroup learnability analysisCan detect bias that persists even after race is removed, via correlated proxy variables
Structured fairness evaluation frameworksRarely applied as a standard part of model validation today
Historical race-based correction factorsSeveral persisted for decades on thin original evidence before reconsideration

What this does not prove

This research does not prove that including race in clinical models is preferable to excluding it, the breast cancer risk model study found problems with removal, not evidence that inclusion was correct. It does not provide a universal fix applicable across all clinical algorithms, since the right approach appears to be condition- and model-specific. It does not mean every race-based clinical calculator still in use is unjustified, some corrections rest on stronger physiological evidence than others, and it does not resolve the practical question of what health systems should do with legacy calculators while better-validated alternatives are developed.

The takeaway

Removing race from a clinical algorithm is not a fairness fix by default, it is a modeling change that requires the same subgroup validation any other change would require, and the 2025 evidence shows that validation is still rarely done. The more durable direction the research points to is structured subgroup performance auditing for every clinical model, race-inclusive or not, rather than a single categorical rule about the variable itself.