Introducing de novo Chemotyping with Pyxis
We taught Pyxis™ to determine the structure and concentration of untargeted small molecules, lipids, and peptides at the scale, speed, and economics comparable to sequencing. This includes annotating the biochemical dark matter to discover new biochemistry not represented in traditional LC-MS spectral libraries. Pyxis analysis has been rigorously benchmarked against human PhD expert-labeled structural IDs and concentrations for more than 100,000 validation tests not seen in any training. These capabilities were tested and improved through early-access scientists from across the pharma, biotech, contract development & research, and academic research hospital settings.
89.2%
Metabolite recall on NIST SRM 1950 plasma.
364 of 408 metabolites recovered. Embedding-based retrieval through the LSM cuts the high-confidence false-positive rate from 8.2% to 0.9%.
0.98
Concentration slope with no analyte-matched standard.
Median error of about 22% across six matrices and five instrument classes, without a calibration curve per analyte.
99%
Exchange bounds inside the physiologically valid range.
In genome-scale flux balance analysis. Rank-based scaling from relative data puts 62% of bounds above published per-tissue uptake, which is physically impossible.
Faster answers for early decisions
Whole biochemomes in the same amount of time as a whole genome or transcriptome.
>10X faster
Versus conventional untargeted metabolomics.
Onsite comparison by top-ten global biotech, October 2025.
MANUAL
weeks
PYXIS
one read
Any
scale
From a handful to a full population cohort.
Pyxis makes reading every sample affordable, so you can work at any scale, from a handful to a full population cohort. Read them all end to end and more of the biology comes into view, for better decisions about health and disease.
How it works: LSMs make biochemical omics machine readable
Most of the biochemome is dark matter
Targeted methods in LC-MS analytical chemistry filter the data to a pre-specified set of molecules. Even untargeted methods are blind to any molecules that cannot be matched to a reference library. The roughly 90% of instrument signals that fail to match go unannotated. This is the biochemical dark matter.
So, we taught machines to read the raw signal
10 billion
Expert-curated spectra.
Expert-curated, and the largest training set of its kind. The current-generation LSM learns from it what a fragmentation pattern means, which is how Pyxis names molecules it has never been shown.
The foundation model for biochemistry
Pyxis employs the LSM to encode raw instrument signals into an information-dense machine representation.
Fine-tuned expert models
Pyxis then decodes the signal using a suite of fine-tuned expert models for structure determination, concentration determination, property prediction, and so on.
Measurement to meaning, and collaboration
Pyxis employs biological context to challenge and improve annotation and concentration hypotheses, and then works interactively with scientists to provide high-impact interpretation.
De novo chemotyping reads what no library contains
Genotyping reads a genome without needing to have seen that genome before. Chemotyping does the same job for the biochemome: it establishes what every molecule in a sample is and how much of it is present. Doing it de novo means Pyxis does not need a reference spectrum for a molecule in order to name it, and that is what lifts the ceiling off the layer.
Conventional untargeted work assigns a structure to roughly one in ten features, because identification runs by matching against a catalogue of spectra someone acquired earlier. Everything absent from the catalogue stays anonymous. De novo chemotyping infers the molecule from the fragmentation pattern itself, which extends coverage from the compounds that have reference spectra toward the whole of known chemical space.
The conventional way
Spectral similarity searching
Computes a similarity score between the analyte spectrum and a fixed library of known spectra. This method fails due to incomplete library coverage, low abundance species, mixtures and impurities, and many technical limitations.
Coverage stops at the catalogue.
How Pyxis works
Generative spectrum to structure
We trained the first commercially available generative structure annotation models to overcome spectral library matching. Pyxis refines every annotation hypothesis using biological and biochemical context.
Reads the pattern, names what no library contains.
01
Read the spectrum, know the molecule
Pyxis reads the raw spectrum directly, and works out what each molecule is from the pattern, whether or not it has ever been catalogued.
• One read across molecule classes.
• Structure generated from the spectrum, never capped by a catalogue.
• Nothing discarded before Pyxis has used it.
02
Know the molecule, know how much
Pyxis models past ionization and matrix effects to put a biochemical concentration on every feature, one you can compare across molecules and across samples.
• A concentration estimate, not just a relative peak area.
• Comparable across molecule classes and across samples.
03
Know biochemistry, map biology
Structure and concentration still leave a list. Pyxis maps each molecule to its pathway, then reads the pathway signature as a phenotype and a mechanism.
• Each molecule in context of its pathway.
• Phenotype and mechanism, biological readout.
NGS unleashed the genomic revolution. Imagine what Pyxis de novo chemotyping will do for the biochemome.
Engage with Pyxis on a new experiment or discover the novel biology hidden in data you already have.