Analysis

AI and scientific research: why small models win

A laboratory is asked to “do AI”, and it looks first at the largest models. It is often the smallest that wins, and it fits in-house.

A person leaning into a microscope eyepiece, their face lit in red
The right model is the one that fits your task

In the scientific organisations we work with, the conversation about AI almost always starts the same way: which is the best model, and what does the subscription cost. The question sounds reasonable. Yet in nine cases out of ten it leads to sending research data to a foreign supplier for tasks a far more modest model would handle just as well, on a machine sitting in the corridor. That is the paradox of AI in scientific research: power is bought, relevance is chosen.

The 2026 edition of Stanford’s AI Index brings, on AI and scientific research, a set of results that all point the same way (Stanford HAI, AI Index 2026): size is not the right criterion, specialisation wins, and performance is very uneven across fields. For a laboratory, that is excellent news — and it changes the conversation.

AI in scientific research beats chemists, and fails at simple tasks

The report offers two images to be held together. On ChemBench, a set of more than 2,700 chemistry questions, the best models surpass the average of human experts (AI Index 2026, p. 233). The same passage notes that they stumble on elementary tasks, which no chemist would.

As soon as more than answering is required, the scores collapse. On ReplicationBench, which consists of reproducing the results of an astrophysics paper, models stay below 20% (AI Index 2026, p. 237). On UnivEARTH, one hundred and forty Earth-observation questions, agents reach 33%, and the code they produce fails in nearly six cases out of ten (AI Index 2026, p. 248).

There is no contradiction. Answering a closed question in a field well covered by the literature is one task; carrying out long reasoning, with real data and tools, is another. For a laboratory, the conclusion is operational: AI is judged task by task, never as a whole. It is the same discipline described in the article that opens this series.

There is a simple way to benefit from this straight away. Take thirty questions from your field whose answers you know, mix in five cases where the right answer is “we do not know”, and run them through the model being offered to you. In half a day you will learn what no sales demonstration will tell you: where it is good, where it invents, and whether it can admit that it does not know.

Small models beat the giants

This is the most striking result of the edition. MSAPairformer, with 111 million parameters, surpasses the best previous methods on variant effect prediction, for a fraction of the compute budget (AI Index 2026, p. 261). GPN-Star, with 200 million parameters, does better than a 40-billion-parameter model — about two hundred times larger — on several tasks of the same kind (AI Index 2026, p. 265).

The wording matters: “on several tasks”, not “everywhere”. A small specialised model does not replace a large generalist one across the board; it beats it where it was designed to work. That is exactly a laboratory’s situation: it does not need a model capable of everything, but one capable of its task.

Cost follows size. A model of a few hundred million parameters fits in the memory of a well-equipped desktop machine, runs without a subscription and, above all, can be frozen: you keep the version used, replay it identically, and can explain a result two years later. In a regulated setting, that is no mere convenience.

One technical point matters, because it avoids a pointless project. Moving to a compact model does not mean training it. In the vast majority of cases, it means running a published model as it is and giving it access to the right documents — yours. Specialised training belongs to a research laboratory, with the skills and compute that go with it. Confusing the two makes teams give up when they could have had the essentials in three weeks.

Specialisation beats generality

The report covers a protein design contest targeting the Nipah virus. Of 1,026 proposals tested experimentally, 99 actually bind to the target — and none neutralises it (AI Index 2026, p. 265). The specialised method entered in the contest obtains the best results, with a caveat we want to report: it had submitted only nine proposals.

That double finding is the most honest in the chapter. Yes, dedicated approaches take the lead. No, assisted design has not yet produced the intended effect on this target. A laboratory reading it will find grounds to launch a project, and grounds to refuse a promise: between “binds” and “neutralises” lies the whole difference between a screening result and a drug.

The report adds an observation that runs against the rest of the sector: AI models for science come mostly from academic and public institutions, and many are the fruit of international collaborations (AI Index 2026, p. 233). Where generalist AI is more than 90% produced by industry, science therefore keeps some autonomy — and published models, which you can run yourself.

In health, real gains and thin evidence

Clinical note-taking tools show substantial gains: up to 83% less documentation time, according to US health systems reporting their own results (AI Index 2026, p. 277). The figure is American and self-reported; it is worth citing, not transposing as is to a French practice.

The review that follows is harsher. Of more than 500 clinical AI studies reviewed, only 5% use real clinical data, and nearly half rely on exam-style questions (AI Index 2026, p. 278). In other words, the literature mostly assesses the models’ ability to pass tests, rarely their effect on patients. That is the kind of gap a scientific committee spots in thirty seconds, and that a sales pitch happily leaves out.

What this review suggests for negotiation is immediate. Faced with a tool sold as validated, three questions are enough: what data was the evaluation run on, how many cases, and has the result been reproduced by an independent team? If the answers come down to exam questions and a demonstration, the tool may be useful, but it is not proven. The distinction is familiar to any scientific committee; oddly, it gets lost when the subject becomes software.

What this opens up for sensitive data

Put the three findings end to end: a compact model is often enough, it runs on ordinary hardware, and scientific models are published more often than elsewhere. The consequence is direct for anyone handling preclinical data, regulatory files or unpublished results: it becomes possible to do serious AI without anything leaving the perimeter.

The uses best suited to this shift are not the most spectacular. Searching ten years of internal studies, classifying incoming publications, extracting values from a test report, preparing a summary an expert will validate: these are bounded, checkable tasks whose value lies in the fact that they bear on your data, not on the world’s. The legal framework and the layers to control are detailed in our article on digital sovereignty; the hardware and operations side, in running an LLM inside the business.

The race for size is a vendors’ debate. For a laboratory, the useful question is narrower and more demanding: does this model, on this task, with my data, do better than what we do today, and can I prove it? The results gathered this year show the answer can be yes far more often than people think — and without handing anything to anyone.

Common questions

What is a small language model?

A reduced-size language model — a few hundred million to a few billion parameters — designed for a family of tasks rather than for everything. It fits in memory on ordinary hardware, costs little to run, and can be frozen on a precise version.

Is AI overtaking researchers?

On some closed tests, yes: the best models exceed the expert average on a chemistry exam of more than 2,700 questions. On replicating an astrophysics paper, they stay below 20%. Performance depends entirely on the task.

Can AI be used on confidential research data?

Yes, provided the model runs on infrastructure you control and its version is frozen. Published compact models make that option realistic: the data does not leave the perimeter, and a result can be replayed identically in an audit.

Source: Stanford HAI, Artificial Intelligence Index Report 2026, published on 29 June 2026 under a CC BY-ND 4.0 licence. Page numbers refer to the PDF file. The reported gains on clinical note-taking come from US health systems and are self-reported. This article is an analysis by Maeliom; it is neither a translation nor an adaptation of the report, and is neither affiliated with nor endorsed by Stanford University.


Next article

AI and jobs: who will train tomorrow’s seniors?

Read

A transformation to support?