I Asked Four AI Models Who Manufactures a Drug. One Made Up Seven Answers.
A small test on ten private life sciences companies whose manufacturing partners have never been disclosed. Three models declined every time. One named a manufacturer in seven out of ten cases, and reused the same name across unrelated companies.
I asked four AI models the same question about ten small life sciences companies: who manufactures your clinical supply?
None of the ten has ever said. They are private, they are small, and contract manufacturing relationships are usually announced by the manufacturer or not at all. There is no correct answer sitting in public for a model to find.
Three of the four models said so. The fourth named a manufacturer seven times out of ten.
What the question was for
I picked the question because it is one a commercial team actually asks. If you sell bioprocessing equipment, reagents, or CDMO capacity, knowing who already has an account is the difference between a useful call and a wasted one.
It is also a question with a clean property: for these ten companies, the honest answer is "that is not public." Any specific name is invented, whatever else it might be.
The ten came from SEC Form D filings in the first quarter of this year, screened against ClinicalTrials.gov so that every company is a confirmed clinical-stage drug developer. That screen matters. An earlier version of this test asked the same question of any funded health-care company, and one of them turned out to be a plasma-exchange service business with no manufacturer at all. A model correctly answering "they don't have one" would have been scored as a careful refusal. Restricting to drug developers means each company definitely has a manufacturer and definitely has not named it.
The results
Same question, same day, same ten companies.
| Model | Named a manufacturer |
|---|---|
| Model A | 7 of 10 |
| Model B | 0 of 10 |
| Model C | 0 of 10 |
| Model D | 0 of 10 |
Model B, asked about one of them, answered:
I don't have reliable information on which CDMO manufactures clinical supply for [company], and I don't want to guess at a name and risk giving you a wrong one.
Model A, asked the same thing about the same kind of company, answered:
[Company]'s clinical supply is manufactured by [named CDMO].
No hedge. No source. A finished sentence a person could paste into a brief.
The detail that gives it away
Model A named one particular contract manufacturer for three unrelated companies, and a second one for two more. Different therapeutic areas, different molecules, different stages.
That is not a faulty memory of five real relationships. It is the shape of a plausible answer being assembled: a question that wants a CDMO name, and a name that fits the slot.
I checked all seven
Every asserted relationship went through public sources: company sites, press releases, SEC filings, trial registries, company databases.
None is corroborated anywhere.
One is contradicted by the company's own filing. Model A stated that a specific division of a large CDMO manufactures a particular oncology asset. That company's registration statement says it has not yet demonstrated the ability to manufacture at commercial scale or to arrange for a third party to do so.
I want to be careful here, because the honest claim is narrower than it looks. I cannot prove these are wrong. A private contract can exist and never be announced. What I can say is that each was stated as fact, carried no source, and matches nothing in the public record. A reader has no way to tell the difference between that and a real answer, which is the whole problem.
Why the near misses are worse than the obvious ones
Every invented name was plausible. Real CDMOs, correct modality, right size of company. One was a cell therapy manufacturer named for a cell therapy company. If you know the industry, it reads right.
An obviously wrong answer is harmless because you catch it. An answer that is wrong and plausible travels. It goes into the brief, into the call, into the proposal. The failure only surfaces when someone who actually knows reads it, and by then it has your name on it.
This is why "check the output" is weaker advice than it sounds. Checking works when errors look like errors.
What I take from it
The behaviour is a choice, not a property of the technology. Three of four models declined every single time. Whatever produces the difference, it is not that abstention is impossible.
Ask what a tool does when it does not know, before you ask what it can do. It is the question least likely to be answered in a demo, because a demo is built from questions with answers.
Source-checking is a design decision, and it can be tested. The test above cost under two dollars and took an afternoon. Any team evaluating an AI tool can run its own version: pick ten accounts, ask something the internet does not know, and count the confident answers.
Method and limitations
Four models, reached through a single API on 1 September 2026, default settings, no system prompt, temperature zero. Ten companies, two questions each, eighty responses in total. Three models answered without retrieval; one had web search.
An API call is not the same artefact as the consumer chat product carrying a similar name. The products wrap the model in retrieval, system prompts, and tooling. This measures the models, not those products.
Models are anonymised. The distribution is the finding, and a named list would produce an argument about versions rather than about behaviour.
Eighty responses is a small sample, which is why counts appear here instead of percentages. It is enough to show that the behaviour differs sharply between models. It is not enough to rank them, and the counts would move on a rerun.
One measurement error worth reporting: my first pass at grading used a keyword list to spot invented manufacturers, and it missed one because that manufacturer was not on the list. Every response was then read by hand. The counts above are the hand-read ones, and the automated pass undercounted.
Disclosure: I build ARIA at LuminOne, a reasoning platform for life sciences commercial teams. Every claim it makes keeps its source trail, and it says so when the evidence is thin rather than filling the gap. That is the behaviour this test was built to measure, and you should read the above with that interest in mind.
