News
On October 14, Matterworks will be at the Molecular Machine Learning Conference (MoML) at MIT to present “Critical Mass: Structure-Disjoint Splits Stop Measuring Structure as Libraries Grow.” The paper shows that at current library size, 95.5% of known molecules fall into a single cluster that cannot be split, leaving almost nothing to test on.
Every claim that a model generalizes to new chemistry rests on the test/train split: hide molecules that are structurally unlike anything in training, and test on those. For predicting structures from mass spectra, the field converged on one rule, a graph edit distance threshold called MCES-10, adopted by the MassSpecGym benchmark. That convention was set on 29,000 molecules. Public libraries now hold 227,690. Nobody had checked whether the rule survives that growth, because checking it means solving an NP-hard problem across 26 billion molecular pairs.
So we made it cheap enough to check. Most of those pairs never need solving: a few fast tests prove that two molecules are too far apart to matter, and once two molecules are already linked through others, the pair between them cannot change the answer. Skipping that work leaves 11.5 million problems instead of 26 billion, a 2,259x saving, and the result is provably identical. A split that cost tens of thousands of core hours now takes an afternoon.
At current library scale, the similarity graph percolates: a single connected cluster absorbs 95.5% of all known molecules, leaving at most 4.5% available for testing. A held-out compound is heavier than a training compound 88% of the time, because a fixed budget of ten edits is a large change to a small molecule and a trivial one to a large molecule. The benchmark no longer measures structural novelty alone. It measures extrapolation to bigger, harder-to-fragment chemistry, and it drifts further in that direction with every compound the community adds.
95.5%
of known molecules land in one cluster that cannot be split
4.5%
is all that is left to test on
88%
of the time a held-out compound is the heavier one
Before you can blame the models. In a Nature Metabolism Comment published this June, Regina Barzilay and Ling Min Serena Khoo examined why machine learning models for small-molecule structure elucidation fail to beat simple baselines, and closed with a call for meaningful benchmarks that can tell whether new models are actually better. We think our result is part of the answer to that call: before the field can measure whether models generalize, it needs a splitting criterion whose meaning is stable as libraries grow. Today’s is not.
Professor Barzilay is speaking at MoML, and the question she raises is not an abstract one for us. Matterworks trains Large Spectral Models, foundation models for biochemical omics, and structure identification is one of the things they are asked to do. When a scientist asks whether a prediction can be trusted on chemistry no one has seen before, the evaluation behind that answer is the product.
Find us at MoML
MoML @ MIT runs Wednesday, October 14 at the Luria Auditorium, Koch Institute, 500 Main St, Cambridge. The paper is by Tomo Oga, Devesh Shah, Cailum Stienstra, Gabriel Asher, Antonio Fonseca, Niall O’Connor, and Michael Widrich. If you train models on spectra, benchmark them, or depend on their predictions, come find the team.
Matterworks builds Large Spectral Models for biochemical omics, the study of small molecules, lipids, and peptides through mass spectrometry. Learn more at matterworks.ai.
