Sure, it’ll be fine: Whose body counts as a body?
There is a device called a pulse oximeter. Everybody knows it. It is one of the most important medical devices in history, as it provides a quick, painless, and non-invasive way to measure blood oxygen saturation. Low saturation can be life-threatening; the reading lets us understand who needs immediate intervention. However, for a lot of people, this reading is wrong.
In 2020, during the covid pandemic, researchers at the University of Michigan examined around fifty thousand paired measurements and found Black patients were three times as likely as White patients to have dangerously low oxygen that their oximeter did not catch. The device said they were fine. This was due to the device and how it was calibrated. The red light that passes through skin gets absorbed not only by oxygen in blood, but by melanin, a natural pigment that alters the shade of our skin. Most pulse oximeters were designed and calibrated years ago without high melanin levels in mind, which in turn led to efficacy issues for people with darker skin. The machine was doing exactly what it was designed to do, for the people it was built for.
The frustrating thing about it is that this was not a new finding. In 1990, a publication identified that readings are biased based on skin colour. That’s 30 years of knowing there are efficacy limitations with this device. It took a pandemic centred around respiration and oxygen-saturation levels for this problem to be highlighted again. For those 30 years, nothing happened, even though the research was there and a vast majority knew there were health inequalities when it came to pulse oximeters.
For most of us, we believe the medical devices that we interact with just work. Why would we disagree with them? They’re like a black box that tells the doctor whether I’m okay or not. The doctor would know if there were an issue. We believe the output must be the hard truth. And I’ve built these things myself. As a biomedical engineer who’s made multiple physiological recorders, I can tell you: a machine with perfect, reliable output is near impossible. At least not on the first few attempts.
But this is where the black box idea crumbles. A medical device is not a window into everyone’s condition; it is a product: validated on a narrow slice of the population, assumed for everyone else. This leads to several questions about the people chosen for the validation. How different is my body from theirs? Why were they the ones chosen for validation? Will my doctor know if my readings are incorrect? When a medical device is validated on a select few, it works exactly as intended on those select few. No alarms telling the doctor the reading is unreliable for everyone else. It fails silently. Silently on the excluded population, without anyone realising.
This is not unique to pulse oximeters. It is the framework behind the claim “generalisable,” and it is applied everywhere.
Before the pulse oximeter finding, there was a Nature paper. Its authors demonstrated that AI could classify skin cancer as accurately as board-certified dermatologists. The results looked good, but then the same issue persisted. The AI model was trained on skin-cancer images from majority-white populations: the US, Europe, and Australia. The model performed considerably worse when it came to darker-skinned individuals with skin cancer. The exact same issue, with a common pattern.
The same issue persists with EEGs, the device I work with in my PhD and startup. EEGs are predominantly used for neuroscience research and epilepsy diagnosis. For the EEG to function, its electrodes require a connection with the scalp to record brain activity; EEGs have been tested and validated on populations with thin, straight hair. For individuals with coarse, curly hair or protective hairstyles, the hair itself is the barrier. The electrode can’t reach the scalp. The signal comes back poor, the data gets flagged, and they drop out. This predominantly affects Black populations. Again, same pattern. If Black individuals are included but their data is labelled “poor quality,” they are then excluded later down the line, sometimes by someone who doesn’t realise who they’re excluding. One systematic-review found only five out of eighty-one neuroscience studies included Black participants. Five. That’s most of neuroscience assuming how the brain functions for everyone from a slice of the population.
I’m not going to list every example, as the pattern has persisted for decades and reaches well beyond medical devices. This pattern has a name. In a 2021 Lancet Digital Health viewpoint, Ibrahim and colleagues called it health data poverty: the inability of individuals, groups, or populations to benefit from a discovery or innovation because there isn’t enough data that adequately represents them. The pulse oximeter, the skin-cancer classifier, the EEG. They’re three instances of the same documented problem.
So rather than adding examples, I want to look at why health data poverty keeps happening, and what would actually fix it.
I believe the language we use perpetuates this health data poverty. Inclusion, fairness, the right thing to do. These are all moral issues. If I don’t have the money or the time for it, these moral reasonings will go in the bin. And that is an honest response from someone who has rushed their PhD and not looked at their research with this lens.
The moral reasoning is not enough as it is subjective and is seen as optional, which is why I believe scientific validity should be the priority. When most instruments do not function on a frequently excluded segment of the population, it can be seen as unfair. That shouldn’t be the first point. It is wrong, unreliable, and biased. This is oftentimes the quiet part. And what is worse is that we sell these devices with the labels “correct, reliable, robust,” which in itself is scientifically incorrect.
Reframing the problem
Reframing this from a moral standpoint to a scientific one is essential. Adequate data on every group is the precondition for science being true, not kind.
A quick note on words, because I am about to use them and they are not interchangeable. Race is a social construct. It categorises people on appearance, and these categories were never based on science. Ancestry is biological heritage: your genetics, tracing your family line. Phenotype is the observable trait itself which is derived from genetics and environmental factors, such as skin tone or hair texture.
This matters because the failures in this piece are not all the same failure. The pulse oximeter and EEG both had a problem with phenotype. Carbamazepine, a drug that we will discuss later, has a problem with ancestry. The gap in who gets to see a neurologist is a problem with race, in the social sense. Yet race is usually the only one of the four we record.
Research based on homogenous data is not complete research. Extrapolation is not a scientifically sound method. Yet from conversations I have had with senior members of the field, extrapolation apparently becomes acceptable the moment ancestry or race is involved.
For most readers, this is obvious. You can’t make a claim on data that was not measured. This is why I wanted to write this piece: to break it down for the people who strongly disagree with my stance. The argument I hear runs like this: “That health inequality has nothing to do with my research. All bodies are generalisable, so my work isn’t affecting it.” There’s no data showing differences between groups, so they conclude no real difference exists. Absence of evidence becomes evidence of absence.
The scientific response should be simple: we have no data on this, so let’s get some and answer the question. Instead, the conversation usually stops at an assumption, that the sample in front of us is the ‘norm,’ and that whatever we find generalises to everyone we didn’t measure. It’s an easy assumption to make if you’ve never had reason to question it. I made it myself for most of my PhD. But it’s still an assumption, and it’s the one doing the damage.
This is why the moral argument does not work. Saying it is the right thing to do makes recruiting diverse populations a suggestion. At the end of the day, a suggestion is optional, and it does not say anything about the validity of the science.
There is a second assumption behind the previous. “Anything to do with instrument outcomes is the implementer’s job; I just do the research”. However, what happened with the pulse oximeter, the skin-cancer diagnosis, and the EEG epilepsy diagnosis. The validation was run on a narrow sample of the population. That makes the findings unreliable for everyone outside of it. The implementation then inherited the flaw, because someone chose the narrow slice to validate on and called it generalisable to everyone.
Narrow validation cuts the other way too. A result we have seen in one population gives us unknown scope. How do other populations change our results? How does this impact our understanding of the disorder? To achieve actual generalisability, we need further understanding of disorders and how they behave differently, which in turn improves outcomes for everyone.
The structure: how the bias gets built in
Why did the “absence of evidence means evidence of absence” standpoint become the “norm”?
It would be nice if there were some big fat cats at the top of MedTech and academia building these barriers on purpose. Then we’d just hit them with the facts, they’d change their minds, and health data poverty would be corrected for excluded populations. That’s not how it works. The truth is complex and boring: these barriers are systematic and no individual’s fault. What’s worse is that each segment of the exclusion pipeline has its own individual justification.
Recruitment comes first in time, and I will come to it. I am starting with data collection because it is the part that gets skipped, and it is what I am most familiar with. We have to use an instrument to record data. Who is excluded from this recording, and what data meets the threshold of “good quality”? If there are accessibility barriers with the hardware, it doesn’t matter if you want to join the study; if the device is not compatible, you are excluded from participating, or if your signal will not meet the threshold for entry, your data is excluded after the fact.
In the oximeter example, a Black participant is not excluded from anything. The data is recorded, and it just carries the error we covered earlier. With EEG, if your hair is too thick or curly, the electrodes can’t make reliable contact with your scalp. The data is recorded, but the impedance, the electrical resistance between the electrode and the scalp, is too high. Signal quality drops and the data is excluded. Oftentimes the analyst is excluding data due to this catch-all “poor-signal quality” label, with no clue the exclusion tracks specific phenotypes. This is acquisition bias: systematic exclusion that happens at the moment of measurement, anonymously, and most of the time with no user input.
Look outside of data collection and the exclusion pattern is repeated. Epilepsy studies are commonly recruited through a neurologist. Black adults in the US are about 30% less likely than white adults to see an outpatient neurologist. Then it narrows again at every step: referral to monitoring, recruitment to a trial, and if the research is genomic, comparison against reference cohorts that are roughly 95% European ancestry. At every stage, there is a valid reason: data quality was bad, access to participants, available reference data. The cumulative effect of this pipeline is homogenous datasets that are propped up as “generalisable” when there is no evidence to back this claim for everyone.
Many technologies try to solve the exclusion pattern. Dry electrodes can fix the hair barriers, larger training data for AI bias, federated data will fix representation. Each one is real progress on one hole, but a single band-aid on a bucket full of holes will not stop the leaks. It is an important start, but not a solution that targets health inequalities.
From measurement to treatment
Excluded populations feel the brunt of the technology not at the measuring stage, but at the treatment stage.
Epilepsy is predominantly diagnosed with EEG and is one of the most common serious neurological condition in the world. For treatment, antiseizure drugs such as carbamazepine are prescribed, and have been for decades, as it is a cheap and effective tool to minimise symptoms.
However, carbamazepine is by no means a miracle drug. Quite the opposite. In patients carrying an allele (HLA-B*15:02), an immune-system variant, carbamazepine can trigger a skin reaction leading to the skin blistering and detaching from the body. In its milder form, Stevens-Johnson syndrome, it kills up to 10% of the people who get it. In its severe form, toxic epidermal necrolysis, up to half. Why did they let this drug through approval if it could lead to deaths? It was because this allele was relatively common in South and East Asian populations and almost non-existent in populations of European descent (under 0.1%) and African-American descent (around 0.1–1%).
How was the drug validated and tested? Most of the populations that used the drug did not have the HLA-B*15:02 allele, making the drug look effective with minimal side effects. A silent risk passed through. Silent for the reasons previously stated: you can’t find a difference in a group you didn’t measure. In this case, the homogenous dataset led to deaths that could have been prevented.
In this situation, there was a fix once people looked into it. In 2007, the FDA issued warnings for carbamazepine and recommended genetic screening before prescribing it to patients of Asian ancestry. A large prospective study in Taiwan then showed that screening for the allele and prescribing a different drug to carriers essentially prevented the reactions. When Singapore made the screening standard of care, skin-detachment reactions caused by the drug fell by 92%. All because they corrected the implementation.
This makes you think: we caught this case because it was dramatic: skin peeling off, and deaths with no other explanation. Is there a quieter version? A current treatment that causes side effects in an excluded population, but not large enough to notice a relationship? We would never know for a lot of them, as we wouldn’t have data in under-sampled groups, due to the health data poverty we are currently facing. We are still extrapolating, just like carbamazepine before 2007.
In 2019, a team built a genetic risk score for epilepsy using European-ancestry data. When they tested it against a large Japanese-ancestry population, it didn’t predict reliably. A polygenic score is built on statistical proxies, and those proxies depend on the ancestry of the sample they were derived from, not because epilepsy is a different disorder in Japan. The knowledge didn’t transfer.
The important part is that we only know this because someone checked. Most models built on one population are never tested against another; they’re simply assumed to be automatically generalisable. We need far more of this kind of checking, because the alternative, as carbamazepine showed, can be lethal.
Recruiting more people, properly
Wouldn’t the solution just be to “recruit more people from different ancestry groups”? This would fix the data issues and make datasets diverse. It’s the right step, but if applied poorly it may not fix anything, for the three reasons below.
Before the reasons, two words I need to keep apart, because we use them as if they mean the same thing. Diverse means the group(s) are present. Representative means they are present in the same proportions as some reference population, usually the country. Neither one means there are enough people in any single group to draw a conclusion about them, and that last part is the thing that actually matters.
The first is not a flaw in the fix, it is that we are not applying it. As we have become more aware of the importance of representation, enrolment has got worse. Across US antiseizure-medication trials, Black enrolment fell from about 20% in 2007–2013 to about 8% in 2014–2019. That’s a 60% drop, during the exact period when diversity in research was being emphasised.
Secondly, a nationally representative sample matches the country’s demographic profile, which sounds rigorous, and still doesn’t explain the health inequalities for that group. If a group is 13% of the country, they’re 13% of your sample, and that’s almost never enough people to tell you whether a treatment behaves differently between groups. They’re present in the recruitment figures and absent from the analysis.
What you actually need is sufficient numbers in each group to support a conclusion about that population. And for a specific disorder, the target should be the disease’s profile, not the country’s. Black Americans are roughly twice as likely to be hospitalised for status epilepticus (prolonged, life-threatening seizures) as White Americans. If they carry more of the disease burden, matching their share of the general population under-samples the people who actually have it.
The third is the point we started off with and impacts all of the previous points. We can’t recruit our way through an acquisition bias. If the instrument you use cannot read data above a quality threshold, it’s the equivalent of not recording the data in the first place. If all of the devices still fail on the same phenotypes, as stated earlier, you still have the bucket full of holes: you are just patching the other side, the recruitment side when there is a glaring hole in the bottom. We need to fix the instrument, and the study design, and the reference cohort (a holeless bucket) because if you don’t, you have just changed at what point exclusion occurs.
Recruiting more people is the correct direction; I just wanted to highlight that it can’t be the only direction.
Why be so negative
This has been a very negative perspective piece, especially for my first one. Sadly, I think that is appropriate for the topic at hand, arising from the lack of scientific validity we have come to normalise. So instead of rambling further, I want to pose actionable guidelines that can tackle a significant chunk of health data poverty.
For every first-line drug, instrument or treatment, the field should be able to state (not assume) whether its efficacy and safety hold across all major ancestry groups of the people who have the condition. This should be the standard, and it is a lower bar than it sounds. Note what I am not asking for. Health inequalities are often labelled “multifactorial”: biology, care quality, access, and everything else. Untangling those is the work of an entire discipline, and causal inference there is legitimately hard. I am not asking anyone to pinpoint a cause. Establishing that a difference exists is enough to change what we do next, and we cannot even do that when the groups were never adequately measured.
Now this is the part where I feel more positive. The discipline that studies this already exists. Social epidemiology and structural health research have been untangling these effects for over a century. The problem is that neuroscience, device design, and clinical trials largely don’t interact with this field. The demographic profile of a disease should be understood before we design and validate an instrument for it, not discovered afterwards in a disparities paper. That’s a solvable coordination failure, and it’s a more interesting research programme than ‘being inclusive’, it’s publishable, it’s fundable, and it happens to require the data we’re not collecting.
At the end of this, we still have the silent issue of instrument accessibility. This is the part my team and I work on. Since I have been working at Synaptive, I have been thinking about how we can build an implementation model that incentivises robust medical devices, rather than relying on goodwill. I don’t want to turn the ending of this piece into a sales pitch, so to be quick and brief: we are building EEG electrodes designed to get a reliable signal across every hair type, so that a participant’s data is not excluded due to “poor quality.” We only work on one piece of this massive problem. We are aware that every instrument created carries silent decisions on whose body counts as a body, with most of these decisions made before we were alive. We want to fix that.
I promise I’m done
We started with pulse oximeters and their bias toward skin tone. From what I have shown, it looks like all devices not trained on all users carry similar biases. This is the framework medical devices are made from: finding the most accessible populations to test and validate the product on, and believing this represents everyone.
This still applies to new tools. Yet we are promised that the new device is better, more reliable, works for everyone. When you look behind the curtain, the same framework is used while the same inequalities persist. Not because the tools do not improve in some metric or form, but because the framework is making the same assumptions: “Do we know it will work for people of different backgrounds?” “Sure, it’ll be fine.”
