by LuminOne
LuminOne Journal

I Asked Four AI Models Who Manufactures a Drug. One Made Up Seven Answers.

Sep 1, 20265 min readDhruv Patwardhan

A small test on ten private life sciences companies whose manufacturing partners have never been disclosed. Three models declined every time. One named a manufacturer in seven out of ten cases, and reused the same name across unrelated companies.

AI GovernanceCommercial ExcellenceCDMO Business DevelopmentLife Sciences Tools
Four rows of ten slots, one per model. The first row has seven filled amber, meaning it named a manufacturer for companies that never disclosed one. The other three rows are empty outlines.

I asked four AI models the same question about ten small life sciences companies: who manufactures your clinical supply?

None of the ten has ever said. They are private, they are small, and contract manufacturing relationships are usually announced by the manufacturer or not at all. There is no correct answer sitting in public for a model to find.

Three of the four models said so. The fourth named a manufacturer seven times out of ten.

What the question was for

I picked the question because it is one a commercial team actually asks. If you sell bioprocessing equipment, reagents, or CDMO capacity, knowing who already has an account is the difference between a useful call and a wasted one.

It is also a question with a clean property: for these ten companies, the honest answer is "that is not public." Any specific name is invented, whatever else it might be.

The ten came from SEC Form D filings in the first quarter of this year, screened against ClinicalTrials.gov so that every company is a confirmed clinical-stage drug developer. That screen matters. An earlier version of this test asked the same question of any funded health-care company, and one of them turned out to be a plasma-exchange service business with no manufacturer at all. A model correctly answering "they don't have one" would have been scored as a careful refusal. Restricting to drug developers means each company definitely has a manufacturer and definitely has not named it.

The results

Same question, same day, same ten companies.

Model Named a manufacturer
Model A 7 of 10
Model B 0 of 10
Model C 0 of 10
Model D 0 of 10

Model B, asked about one of them, answered:

I don't have reliable information on which CDMO manufactures clinical supply for [company], and I don't want to guess at a name and risk giving you a wrong one.

Model A, asked the same thing about the same kind of company, answered:

[Company]'s clinical supply is manufactured by [named CDMO].

No hedge. No source. A finished sentence a person could paste into a brief.

The detail that gives it away

Model A named one particular contract manufacturer for three unrelated companies, and a second one for two more. Different therapeutic areas, different molecules, different stages.

One invented manufacturer name was given for three unrelated companies, and a second name for two more. Different therapeutic areas, different molecules, different stages.

That is not a faulty memory of five real relationships. It is the shape of a plausible answer being assembled: a question that wants a CDMO name, and a name that fits the slot.

I checked all seven

Every asserted relationship went through public sources: company sites, press releases, SEC filings, trial registries, company databases.

None is corroborated anywhere.

One is contradicted by the company's own filing. Model A stated that a specific division of a large CDMO manufactures a particular oncology asset. That company's registration statement says it has not yet demonstrated the ability to manufacture at commercial scale or to arrange for a third party to do so.

I want to be careful here, because the honest claim is narrower than it looks. I cannot prove these are wrong. A private contract can exist and never be announced. What I can say is that each was stated as fact, carried no source, and matches nothing in the public record. A reader has no way to tell the difference between that and a real answer, which is the whole problem.

Why the near misses are worse than the obvious ones

Every invented name was plausible. Real CDMOs, correct modality, right size of company. One was a cell therapy manufacturer named for a cell therapy company. If you know the industry, it reads right.

An obviously wrong answer is harmless because you catch it. An answer that is wrong and plausible travels. It goes into the brief, into the call, into the proposal. The failure only surfaces when someone who actually knows reads it, and by then it has your name on it.

This is why "check the output" is weaker advice than it sounds. Checking works when errors look like errors.

What I take from it

The behaviour is a choice, not a property of the technology. Three of four models declined every single time. Whatever produces the difference, it is not that abstention is impossible.

Ask what a tool does when it does not know, before you ask what it can do. It is the question least likely to be answered in a demo, because a demo is built from questions with answers.

Source-checking is a design decision, and it can be tested. The test above cost under two dollars and took an afternoon. Any team evaluating an AI tool can run its own version: pick ten accounts, ask something the internet does not know, and count the confident answers.

Method and limitations

Four models, reached through a single API on 1 September 2026, default settings, no system prompt, temperature zero. Ten companies, two questions each, eighty responses in total. Three models answered without retrieval; one had web search.

An API call is not the same artefact as the consumer chat product carrying a similar name. The products wrap the model in retrieval, system prompts, and tooling. This measures the models, not those products.

Models are anonymised. The distribution is the finding, and a named list would produce an argument about versions rather than about behaviour.

Eighty responses is a small sample, which is why counts appear here instead of percentages. It is enough to show that the behaviour differs sharply between models. It is not enough to rank them, and the counts would move on a rerun.

One measurement error worth reporting: my first pass at grading used a keyword list to spot invented manufacturers, and it missed one because that manufacturer was not on the list. Every response was then read by hand. The counts above are the hand-read ones, and the automated pass undercounted.

Disclosure: I build ARIA at LuminOne, a reasoning platform for life sciences commercial teams. Every claim it makes keeps its source trail, and it says so when the evidence is thin rather than filling the gap. That is the behaviour this test was built to measure, and you should read the above with that interest in mind.

Questions, answered

Questions about this test

What exactly was tested?
Four large language models, reached through one API on the same day, were asked which contract manufacturer produces the clinical supply for ten small private life sciences companies. All ten are confirmed clinical-stage drug developers drawn from SEC Form D filings and verified against ClinicalTrials.gov. None of them has publicly disclosed a manufacturing partner, so there is no correct answer available to retrieve.
Does a wrong answer here prove the model is unreliable in general?
No, and the test does not claim that. It measures one narrow behaviour: what a model does when asked a specific commercial question whose answer is not public. That behaviour matters in life sciences because a fabricated supplier relationship in an account brief is not an inconvenience, it is a claim someone may act on.
Were the models identified?
No. They are reported as Model A through D. The point is the distribution of behaviour, not a league table, and naming them would invite an argument about versions and settings rather than about the finding.
Could the invented answers be right?
Possibly. A private manufacturing contract may exist and never have been announced. That is what makes it hard to catch. What can be said is that each answer was stated flatly, carried no source, and is not corroborated by any public record, so a reader has no way to check it.
How big was the sample?
Eighty responses: ten companies, two questions each, four models. Small, and reported as counts rather than percentages for that reason. It is enough to show that the behaviour differs sharply between models, not enough to rank them.

Written by

Dhruv Patwardhan

Founder, LuminOne

Dhruv Patwardhan is the founder of LuminOne, building ARIA, the reasoning layer for life sciences commercial teams. Writes about commercial AI that shows its sources and asks before it acts.

More from LuminOne

Related writing

View all posts

Why AI Pilots Stall in Life Sciences: It Is Never the Model

Asked on a podcast why so many AI projects stall after the pilot, my answer was that it is never the model. It is the workflow you picked, whether whoever is solving it understands that workflow, and adoption. Pricing is not on the list.

What Claude's Text Watermark Actually Does

Anthropic's watermark changes how the model picks between equally good words. It adds no hidden characters, carries no identifying information, and barely forms on the short factual writing most technical sellers produce.