Product1 distinct publisher3 min readPublished
A computer engineer writing in Fast Company says dermatology models latch onto background skin tone instead of the lesion, which turns training-set composition into a spec question for anyone shipping skin imaging to patients.
The Product Desk · Product desk

Compiled by The Product DeskSomething wrong?How this is made
A pattern matcher has no idea which pixels you care about. Trained on photographs of light-skinned patients, it learns which visual features travel with which disease, and the tone of the skin around a lesion is one of the loudest features in the frame [3]. That is why the GPT-4 result indicts the feature the model is using rather than its average accuracy: the mole itself was untouched and only the background moved [6]. The border check that dermatology actually teaches lost to the surrounding pigment [7].
The gap between what is pitched and what gets done sits in the deployment story. The case for these tools is reach, with lifesaving screening in remote areas and underresourced communities where dermatologists are scarce [13]. Performance runs the other way, strongest on the tone that dominates the shared university and hospital image libraries these models are built from [11][12]. Melanoma on pigmented skin is already found later and survived less often [8].
For whoever owns the rollout, the useful grid has two axes: whether a clinician sees the output before the patient acts on it, and which direction the error runs. Consumer app, false positive, and a harmless dark spot buys panic and an appointment [9]. Consumer app, false negative, and a melanoma reads as reassurance with no second reader anywhere in the loop, which is the case when someone asks a general chatbot instead of a clinic [9][10]. Clinician tool, false positive, and someone gets a biopsy they did not need. Clinician tool, false negative, and the miss is the same, except a human can be warned in advance that background tone is a known failure mode. The clinician-assist cell is the only one of the four with a brake, but app stores are stocking the consumer version instead [1].
What this account does not give a buyer is a size for the gap. There are no accuracy percentages, no per-tone sample sizes, and no stated share of darker-skin images in the training libraries [14]. It does hand over a test that costs an afternoon. Take the images your tool currently gets right, darken the surrounding skin without altering the lesion, and count how many answers move [4][6]. If they move, you have measured your own exposure without waiting for a vendor benchmark or a peer-reviewed table.
The recommendation comes with its price attached. Hold autonomous consumer triage until you can report performance by skin tone, and ship the clinician-assist version with the background-tone failure named in the interface rather than buried in release notes. The cost of that call lands on the patients the tool was pitched to help first, who continue waiting because no dermatologist is available [13]. Training-set composition is a spec line. A product shipped without it delivers its worst performance to its stated beneficiaries [2][12].
Ranked by verification strength, evidence, and original report placement.
The author describes an AI model as a pattern-matching engine that can be thrown off by the background color of a person's skin: instead of learning to look at the lesion, it picks up the color of the surrounding skin as a clue, so its predictions degrade to guesses based on skin color.
The author says that because skin color was the more prominent feature, the AI became focused on the dark pigment and ignored standard medical rules used to identify cancer, such as checking whether the mole's borders are irregular.
AI skin-checking tools ship in two forms: smartphone apps anyone can download to scan their own skin at home, and software programs designed for clinicians to use in a doctor's office.
The author and colleagues trained an AI model on photographs of known skin conditions in light-skinned patients, then digitally manipulated the images to resemble darker skin tones; the clinical condition in the photo had not changed, but the model's ability to recognize it deteriorated sharply.
Atopic dermatitis causes discoloration that appears pink on light skin but gray or violet on darker skin, and the author's research and that of other groups shows models would reliably classify the pink marks but might not identify the darker colors as signs of the condition.
In a 2024 study, the authors presented OpenAI's GPT-4 with an image of a completely benign mole and digitally darkened the skin around it while keeping the mole exactly the same; GPT-4 classified the spot as malignant melanoma.
Distinct publishers with included, body-backed reporting in this cluster.
1 article · September 6, 2026
Follow any of these and your For You feed starts watching them — no settings page required.
science
Text watermarks land on 2 December. The detection they imply does not.1 distinct publisher
product
Meta's Mac app is a data connector wearing a chatbot's clothes6 distinct publishers
science
Claude's watermark is a compliance artefact, not a cheating detector1 distinct publisher
leadership
OpenAI asks Judge Stein to measure copying by output rather than by corpus3 distinct publishers
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
One essay, the author's own experiments
Everything rests on a single Fast Company piece written by the computer engineer who ran the studies it describes, including the 2024 GPT-4 test. The experiments are recounted, not cited: no journal, no sample size, no accuracy delta, no protocol a reader could check. The mechanism is coherent and matches what dermatology image libraries are known to contain, which keeps this above thin, but no part of it was verified outside the author's own account.
No deployment or usage figures
The piece opens on 'a slew' of new tools and then names none of them. No download counts, no clinic installs, no regulatory clearances turn up anywhere in it, and the only shipped system put to a test is GPT-4. Home use of skin scanners may be substantial; this reporting offers no way to establish that.
Wording is strong; the magnitude is not measured
'Massive blind spot' and 'measurably worse care' are carrying weight that numbers would otherwise carry. The most quotable finding, a benign mole reclassified as melanoma once the surrounding skin was darkened, is one image in one test, with no stated frequency, tone range or condition coverage. The direction of the problem is credible and consistent with how these datasets were assembled; the certainty of the prose runs ahead of what is shown.
A researcher writing about his own research line
The author studies these tools in clinical settings, and the fix the essay points toward, generating synthetic images of conditions on darker skin, is work his own group published a positive result on. That is a legitimate thing for a specialist to write and it is also an argument for the field he works in. Fast Company runs the byline without any statement of funding or commercial ties, and gives no tool maker a chance to answer.
Direction seems credible; the size is unknown
Two things pull against each other. The failure mode described, a classifier keying on background pigment rather than lesion morphology, is the sort of shortcut behaviour imaging models are known to fall into, so the qualitative finding is easy to accept. Every number a reader would want to act on is missing, and no second party has checked the GPT-4 result. Fairly confident the gap is real; not confident about its size in any product a patient can install.