Large language models are moving fast into health care. They are answering questions, helping with triage, supporting appointments, and, in some settings, shaping how patients are understood before a clinician even speaks to them.
That makes their hidden biases more than a technical flaw. It makes them a clinical concern.
A new study published in Nature Health has raised a sharp warning. Popular models may look relatively fair when assessed with standard stigma questionnaires, yet still produce biased, stigmatising responses when placed inside realistic health-related scenarios.
In other words, an AI system can appear clean on paper, then behave very differently once context is introduced.
The findings matter because stigma in health care is not a minor issue. It can delay diagnosis, reduce trust, distort treatment decisions, and discourage people from seeking help at all.
When the condition involves HIV, hepatitis B, schizophrenia, or other mental health disorders, the consequences can be especially serious. A chatbot that subtly signals danger, blame, incompetence, or social distance may reinforce the very barriers public health systems are trying to remove.
Researchers tested six major language models in two distinct ways. First, the models were given standard stigma surveys similar to those used with human respondents. These questionnaires are common fairness checks. They ask direct questions and measure explicit prejudice. On these tests, the models performed surprisingly well. In fact, they scored as less stigmatic than the average human benchmark, which included more than 56,000 people.
That result might sound reassuring. It is not the full story.
The second part of the study used contextual judgement tasks. Researchers created 51 everyday story scenarios and asked the models to predict how each situation would unfold. The key detail was simple. Every version of a story was identical except for one health-related label. A person might be described as healthy in one case, then as having schizophrenia, HIV, or hepatitis B in another. Across English and Chinese prompts, the models generated 61,200 decisions that were then examined for bias.
This design gets closer to real life. People do not usually experience stigma through a questionnaire. They experience it through assumptions, reactions, and subtle judgements in conversation. That is where the models began to show their weakness.
The bias was not random. It followed patterns linked to the condition involved. HIV and mental health conditions were often associated with danger, social distance, or caution. Some responses implied that others should keep away. Others suggested distrust or fear. By contrast, physical conditions such as back pain and high blood pressure were more likely to draw pity, or assumptions that the person was less competent.
That difference is important. It shows that the models were not simply using neutral language to describe illness. They were attaching social meaning to diagnosis. That social meaning is exactly what stigma is.
Language also played a role. Chinese prompts produced more stigma-congruent responses than English prompts, particularly in mental health scenarios. This points to a major challenge for global health AI.
A model that seems relatively stable in one language may behave differently in another. That is a practical problem for hospitals, helplines, public health tools, and patient-facing systems operating across multilingual settings.
There was, however, one notable exception. When models were prompted to reason step by step before answering, the bias decreased markedly. This suggests that the way a model is instructed can influence whether it falls back on stigma-linked shortcuts. That is a useful finding, although it should not be mistaken for a cure. Better prompting may reduce risk. It does not remove the need for testing, oversight, and careful deployment.
The study’s central message is straightforward. Traditional fairness evaluations are not enough. A model may answer direct stigma questions in a socially desirable way, then reveal prejudice once placed in a realistic narrative. That gap between declared fairness and contextual behaviour is where much of the danger lies.
The researchers responded by outlining nine mitigation strategies. Among them were individualisation, which encourages the model to focus on the person rather than reducing them to a diagnosis, and relevance filtering, which tells the system to ignore health status when it is not pertinent to the task. These are practical steps. They can be built into prompts used by clinicians, administrators, and developers.
The study also suggests a larger operational change. Health-care institutions could adopt prompting toolkits to reduce stigmatic outputs across languages and settings. That would help standardise safer use. More fundamentally, developers could carry out prerelease bias audits using contextual judgement tasks before launching medical AI products. That seems like a sensible minimum requirement. If a model will be used in health care, it should be tested the way it will actually be used.
This is where the paper feels especially relevant. Many current AI audits still rely on benchmark-style tests that reward polished answers. But health communication is not a classroom exercise. It is messy, relational, and loaded with social meaning. A model that can sound fair without being fair in context is not ready for clinical use.
The implications stretch beyond individual chatbots. As generative AI tools become part of patient screening, mental health support, appointment handling, and public health messaging, their tone can shape behaviour. A reassuring answer can increase engagement. A cold or suspicious one can drive people away. For conditions already burdened by shame, misinformation, or discrimination, even small distortions matter.
The study also fits into a broader conversation about how AI inherits human bias. Models learn from enormous datasets produced by societies that already contain prejudice. If those data reflect stigma around particular illnesses, the model may internalise those patterns, even when its surface-level responses seem polished and neutral. That is one reason health-related AI deserves special scrutiny. The stakes are unusually high.
Importantly, the research does not say that all medical AI is harmful. Nor does it suggest that every response from every model will be biased. It does show that fairness cannot be assumed from appearance alone. It must be measured where it counts, under conditions that resemble real use.
That distinction may seem subtle. It is not. A chatbot that performs well in a direct survey and poorly in a realistic case can still mislead users, clinicians, and developers. In a health system, misleading confidence is risky. The most dangerous tools are often the ones that seem safe until they are tested properly.
There is also a policy message here. Regulators and developers should treat contextual stigma testing as part of responsible evaluation, not an optional add-on. Health systems considering AI deployment should ask a simple question: does the model remain fair when the situation becomes specific, personal, and emotionally charged? If the answer is no, it should not be used as if it were ready.
The study’s findings are timely because adoption is accelerating. AI tools are already being used to draft patient messages, support administrative workflows, and advise on symptoms. Their reach is growing faster than many safeguards. That gap needs attention now, not after a problematic system has already been scaled.
In that sense, the new work is less about one paper and more about a change in standards. Fairness in health AI should not be judged by self-contained answers alone. It should be judged by context, language, condition, and consequence. If the model behaves differently when a diagnosis is mentioned, that difference must be treated as a serious design flaw.
For health care, the lesson is plain. Bias can hide behind polished language. It can also hide behind good scores. The question is not whether an AI sounds unbiased in a questionnaire. The question is whether it stays unbiased when the clinical story becomes real.
That is the test these researchers have now brought into sharper focus. It is a more demanding test. It is also the right one.























