← Innovations
AI

Multimodal reasoning models outperform specialists on structured scientific tasks

A wave of new evaluations shows large multimodal models now matching or exceeding domain-specialist performance on structured tasks in chemistry and biology — while still failing on open-ended experimental design. The gap between benchmark performance and genuine scientific reasoning remains the active research frontier.

The past twelve months have produced a string of benchmark results that would have seemed implausible two years ago: multimodal models trained on image, text, and structured data now regularly score above the median board-certified physician on medical diagnosis tasks, above the median PhD student on graduate-level chemistry problem sets, and above the median financial analyst on earnings-interpretation tasks. The headline interpretation — that AI has surpassed human domain expertise — deserves scrutiny.

What the evaluations actually measure is performance on structured tasks that have been constructed to be evaluable: multiple choice, short-form answer, defined rubrics. This is not the same as scientific reasoning. A physician diagnosing from a case vignette is doing something categorically different from a physician running a differential in a room with an actual patient whose presentation is incomplete, contradictory, and evolving. The benchmark captures a slice of the competency.

The gap that remains stubbornly wide is in open-ended experimental design: give a model a scientific question and ask it to propose an experiment that would generate evidence bearing on it. Current models produce plausible-sounding experimental designs that frequently contain subtle methodological errors — confounding variables they haven't controlled for, measurement approaches that don't actually operationalize the construct, sample sizes that would be underpowered by an order of magnitude. Human scientists catch these in peer review. Models currently don't catch each other's.

The productive frame is not 'has AI surpassed scientists' but 'what tasks in the scientific workflow are now automatable, and what does that free human scientists to do?' Literature search and synthesis, hypothesis ranking from prior evidence, statistical analysis of well-structured datasets — these are good candidates. Experimental design, mechanistic interpretation, deciding what question is worth asking at all — these remain robustly human for now.

For independent researchers, the practical implication is significant: the barrier to rapid literature review and preliminary analysis has dropped substantially. Tools that would have required a research assistant or significant time can now be delegated. The risk is over-relying on model outputs that are confident and fluent but subtly wrong — the exact error mode humans find hardest to catch.