News

A drug's approval year predicts its toxicity better than its chemical structure

A drug's approval year predicts its toxicity better than its chemical structure

On October 14, Matterworks will be at the Molecular Machine Learning Conference (MoML) at MIT to present “Toxicity Prediction Tools Do Not Generalize to Novel Chemistry.” The paper tests today’s toxicity-prediction tools on a panel of approved drugs, with careful controls for data leakage, so that we can measure how well they perform on genuinely novel chemistry. Applied as released, none did much better than a coin flip.

Computational tools that predict whether a drug will be toxic from its chemical structure are widely used in drug discovery, and they report strong accuracy on standard benchmarks. However, these benchmarks typically split data at random or by scaffold, so test molecules often resemble the ones a tool was trained on. In practice, though, when these tools are used on new molecules with chemistry they have never encountered, and their performance in that context is much less clear.

To investigate that performance, we built a panel of 569 approved drugs, focusing on recent approvals, and kept only those that are structurally dissimilar to the data used to train the tools. This is much closer to the real-world scenario, where these tools are asked to assess the toxicity of new drugs still in development. Each drug in the panel is labeled by regulatory outcome: toxic if it carries a black-box warning or was withdrawn from the market. We tested 35 prediction approaches, from classical models to modern molecular encoders and published toxicity tools, and checked each tool’s training data for overlap with our panel to guard against leakage.

The published tools did not reliably predict which drugs in the panel were toxic. Their scores ranged from 0.40 to 0.585, close to or below what a coin flip would produce. This includes tools built specifically to predict drug-induced liver injury, a toxicity of particular interest to industry. These tools perform well on their own training data, but none cleared chance on our panel. For reference, a predictor that knows only the year a drug was first approved, and nothing about its structure, scored 0.785. More complex models did not help either: trained directly on our drug panel, modern molecular encoders did no better than a simple fingerprint that just records which chemical substructures a molecule contains.

These results point to the input, not the model. A molecule’s structure describes what it is, but whether it harms a patient also depends on how much of it reaches which tissues, what the body turns it into, and the patient’s biological state. Every tool we tested was given only a molecule’s structure, with no measurement of how a drug is metabolized or how a biological system responds to it, both of which are essential context for estimating a drug’s toxicity.

Matterworks trains Large Spectral Models, foundation models for biochemical omics. This paper highlights exactly why that work is important. If the binding constraint for toxicity prediction is the input rather than the model, the appropriate remedy is to measure something else. A candidate is LC-MS/MS profiling of the biological sample, which captures the parent drug, its metabolites, and the perturbed endogenous metabolome together, giving an exposure- and metabolism-aware readout in place of a static structure. Our work makes testing that hypothesis possible.

Find us at MoML

MoML @ MIT runs Wednesday, October 14 at the Luria Auditorium, Koch Institute, 500 Main St, Cambridge. The paper is by Circe Hsu, Martin Weiss, Cailum M. K. Stienstra, Amy Caudy, and Antonio H. de O. Fonseca. If you build toxicity models or decide which compounds to advance on their output, come find the team.

Matterworks builds Large Spectral Models for biochemical omics, the study of small molecules, lipids, and peptides through mass spectrometry. Learn more at matterworks.ai.