An AI That Invents a Citation and an AI That Outs a Person Are Not the Same Risk
An AI that invents a citation and one that outs a person are not the same risk. For nonprofits and healthcare, the difference is measured in safety.
The governance question is not whether your AI can fail. It is what failure costs, and who pays.
There are two kinds of AI failure, and we have been governing them as if they were one.
The first kind is by now familiar. In 2023 a federal judge fined two lawyers who had filed a brief citing cases that ChatGPT had invented. The cases looked real. The citations were formatted correctly. They simply did not exist. The pattern has since become common enough that a researcher now maintains a public database of court filings built on fabricated, machine-generated citations, and it has logged more than sixteen hundred of them. Courts have responded with fines, mandatory training, referrals to state bars, and the occasional stricken filing.
This is a real failure with real consequences. A lawyer's standing takes a hit. A firm spends a week auditing its filings. Someone's name appears in an article like this one. And yet, look closely at how the harm travels from the model to the world, and you find the path crowded with people. The lawyer who signed the brief was supposed to read it. Opposing counsel reads it too. So does a clerk, and then a judge. Several trained professionals stand between the model's mistake and any damage it can do, and when the damage lands, it lands on the person who failed to check. The organization apologizes, refiles, and survives. We already have a name for this kind of risk. We call it quality control.
Now hold that failure next to a different one. Not the AI that fabricates a citation, but the AI that decides, correctly or incorrectly, that a person is gay, or transgender, and surfaces that conclusion to someone who was never meant to have it. There is no clerk in that path. There is no second reader. The harm does not move through a courtroom full of professionals trained to catch it. It goes straight to the person, and it does not refile.
What prompts this now is a new report from GLAAD, Build for Everyone, which documents these harms across the AI lifecycle in more detail than I will. Much of the evidence below is drawn from it, and it is worth reading in full. I am writing for the organizations I work with most closely, which are nonprofits and healthcare providers, because these are the two failures they are most likely to meet and least equipped to absorb. A mission-driven organization holds exactly the data that makes the second failure dangerous. It serves exactly the people who cannot afford it. And it is, as a rule, adopting AI tools faster than it is governing them, because the tools are cheap and the staff is thin. Two areas deserve their attention before anything else.
The first is inference. We praise modern AI for synthesis, for its talent at assembling a coherent picture out of scattered, unremarkable signals. That is the capability working as designed. It does not pause when the picture it assembles is a person's sexual orientation or gender identity. A model can infer those traits from things the person never disclosed: who they talk to, where they go, what they search, the cadence of their language, the content they linger on. In 2021, after Spotify described technology that would predict a listener's emotion, gender, and age from the sound of their voice, the digital rights group Access Now pointed out the obvious problem, which is that a device built to listen that closely is a device that is always listening. The same year, an investigation by The Markup found that Google had been letting advertisers withhold job ads from people whose gender it could not classify as male or female, which is to say from many nonbinary people. The inference engine does not need anyone to opt in. It builds the profile whether or not the person knows the profile exists.
Here is the part a privacy policy tends to miss. The inference does not have to be correct to cause harm, and being correct does not make it safe. A wrong inference outs someone who is not what the model decided they were, and they inherit a label they did not choose and cannot easily refute. A right inference outs someone who is exactly what the model decided, and who had made a deliberate, sometimes life-preserving choice not to say so. Both land on the person. Neither is embarrassment.
Put that in the setting where these readers actually operate. A health system runs an AI tool over its patient records to flag people for outreach. A youth-services nonprofit adopts a chatbot that quietly retains every conversation. A clinic shares a screen, a device, or a billing summary with a family member. In more than sixty countries, a same-sex relationship is a crime, and a profile assembled by a model is evidence. In a growing number of jurisdictions closer to home, the same inference can cost a person their healthcare, their legal recognition, or their physical safety. A 2025 Stanford study found that leading AI companies train on the conversations people have with their chatbots by default, and that some retain those conversations indefinitely. The data an organization thought it was holding briefly, for a narrow purpose, does not necessarily stay brief or narrow. This is not a privacy inconvenience. It is exposure, and exposure is the thing these organizations exist to prevent.
Inference is also no longer where these systems stop. The tools are moving from answering questions to taking actions, from describing the world to making decisions inside it. An automated system that screens housing applications, routes job candidates, or prioritizes patients does not merely hold an inference about a person. It acts on one. When a model that has quietly decided something about a person's identity is also the model deciding whether that person advances, the assumption and the consequence collapse into a single step, with no one in the room to notice the assumption that drove it.
The second area is what comes out of a model that learned from the worst of what people have written. Let me state the fair version of the counterargument first, because it is mostly true. A large language model learns the patterns in an enormous body of human text, and human text carries human contempt. The model is not malicious. It is a mirror. But a mirror that speaks back with fluent authority, in the steady voice of a helpful assistant, is not a neutral surface, and the reflection it returns is not harmless. A 2024 UNESCO study found that one widely used model generated negative content about gay people in roughly seventy percent of trials, and an earlier model did so in roughly sixty percent. The output was not subtle. It described gay people as criminals, as freaks, as lives not worth living.
There is a structural reason this is hard to fix and easy to ship. Hate is not static. New slurs, new dog whistles, and new conspiracy theories arrive constantly, and a model trained on a snapshot of the past does not recognize the language of the present. The bias does not hold still long enough to be filtered once and forgotten. It has to be tracked the way we track any other live threat, which is a responsibility most organizations deploying a downstream tool do not know they have inherited.
The most instructive recent example is also the most relevant to a healthcare audience. In April 2025, after one major company announced that its AI would present multiple sides of contested issues without passing judgment, GLAAD's researchers found the model recommending conversion "therapy" to a person asking about unwanted same-sex attraction. Every major medical and mental health organization in the United States has condemned that practice. The United Nations has compared it to torture. The model offered it as one option among several, in the calm register of a clinician. As GLAAD put it, treating anti-LGBTQ pseudoscience as one side of a legitimate debate does not inform anyone; it lends a settled falsehood the credibility of an open question. The failure here is not that the model said something offensive. It is that the model said something dangerous in the tone of something trustworthy.
Healthcare deserves its own line on this, because the harm is so concrete. Researchers at the Oxford Internet Institute have warned that models which collapse gender into biological sex can produce advice that is simply wrong for the person in front of them, offering care to a transgender woman as though her body matched an assumption it does not, or misjudging the needs of a cisgender woman whose physiology does not match the model's stereotype of one. An organization that points such a tool at drafting patient resources, triaging a message, or answering a question after hours has not bought a convenience. It has installed a source of clinical advice it never vetted, aimed at the people least able to absorb advice that is wrong.
This is the distinction I most want a board, an executive director, or a chief medical officer to hold onto, because it changes how the risk should be governed. The hallucinated citation and the biased recommendation are failures of the same underlying technology, but they are not failures of the same magnitude, and they should not share a governance plan. The citation has layers of human review between the error and the harm, and the harm, when it occurs, is professional and recoverable. The outing and the biased recommendation have no such layers. The harm reaches the person directly, and it is not professional. It is a closeted teenager whose parent now knows. It is a patient who followed guidance built for a body that was not theirs. We are good at pricing the first kind of failure. We can name the fine, the billable hours, the cost of the audit. The second kind cannot be priced, because you cannot file a correction on it, and you cannot give a person their safety back once it is gone. These are not two points on one scale. They are two different scales, and governing them with one policy means under-protecting the one that can actually hurt someone.
Which brings us to the part that is genuinely a governance problem and not a technology problem. Most organizations meet this moment by writing a policy. The policy says the organization uses AI responsibly, reviews its tools, and respects privacy. That document is not a control. A policy is a statement of intent. A control is something a specific, named person can be measured against and can fail, with an owner who answers for it when they do. "We protect sensitive data" is paperwork. "We do not collect or infer sexual orientation or gender identity, the system is built so it cannot surface that inference, and this engineer owns that boundary" is a control. "We use AI carefully in clinical contexts" is paperwork. "No AI-generated guidance reaches a patient without a named clinician approving it, and we test this tool against known failure cases every quarter" is a control.
The temptation, especially for a resource-constrained organization, is to reach for compliance and call it governance. For these two risks, that reach comes up empty. The United States has no comprehensive federal data-protection law, which means that on the questions that matter here, what a model may infer about a person and how long it may keep what it learns, compliance asks almost nothing of you. Compliance is a floor, not a ceiling, and on this particular floor there is very little standing between your tool and the person it could expose. The governance has to come from the organization itself, or it does not come at all.
None of this is an argument against using these tools. Used inside the right boundaries, they can help an under-resourced organization do more for the people it serves, which is the whole reason to adopt them. Protecting the people in your data is not the opposite of enabling the mission. It is the precondition for it. So the ask is narrow. Before you deploy an AI tool, sort the ways it can fail by who pays when it does. Separate the failures that cost your organization face from the failures that cost a person their safety, and refuse to govern them with the same sentence. The first list deserves a review step. The second deserves a control with a name on it, because we are not, in the end, protecting data. We are protecting people, and in the work these organizations do, the people in the data are often the ones with the least room to absorb our mistakes.
Source: GLAAD, Build for Everyone: A Framework for LGBTQ Representation and Safety in AI (2026).