Expertise Brief

Pyxis Learns Lipid Structure: De Novo ID for Lipidomics

Lipidomics is producing some of the most interesting biology in the last decade.

Lipids are identifying early predictors of Alzheimer’s [5], refining cardiovascular risk stratification [4], and mapping how individual lipid species drive cancer biology [6]. The same pattern runs through oncology, neurology, and immunology [7, 8, 9, 10].

Yet, in practice, reference library coverage sets the ceiling on which species can be identified, and, for a large portion of the lipid space, experimental reference spectra simply do not exist. As a result, the scope of discoverable biology remains constrained.

Most lipids can’t be identified

Lipid identification has a lot of room to grow. LIPID MAPS catalogs over 60,000 lipid species [1], but public spectral libraries contain experimental MS2 for only a small fraction of them. Even within existing libraries, the tools disagree. A 2024 cross-platform comparison found that MS-DIAL and Lipostar, run on the same dataset with default settings, agreed on only 14% of lipid identifications. Including MS2 data raised agreement to 36% [3].

Part of why this has persisted is that reference data is genuinely hard to generate. Building libraries requires synthesizing or purchasing lipid standards, which is expensive and slow. And the combinatorial space is large enough that comprehensive coverage through standards alone isn’t realistic.

Many tools sidestep the library problem with rules-based identification, matching spectra against predefined fragmentation patterns. But these rules rely on the presence or absence of diagnostic ions without accounting for their relative intensities, making it difficult to distinguish true matches from coincidental fragment overlaps.

The first fully de novo lipid ID model

Today, we’re expanding our Large Spectral Model into the first commercially available de novo lipid identification model, capable of identifying lipids no reference library has ever contained.

The model learns the relationship between lipid structure and fragmentation behavior directly from spectral data. Given a new MS2 spectrum, it generates the lipid structure most likely to have produced it without consulting a library. It’s the same model behind our metabolite ID work, extended to a new chemical domain.

Expanding a well-characterized lipidome by 40%

We tested Pyxis on the raw data from AdipoAtlas, a comprehensive lipidomics study of human white adipose tissue across lean and obese individuals [11]. The original authors identified over 500 unique lipid species, using three independent lipidomics software tools with manual curation, making AdipoAtlas one of the most thoroughly annotated public lipidomics datasets available.

Pyxis covered every lipid class in the original study and identified 222 lipid species not in the original report, a 40% expansion of the characterized lipidome from the same raw files.

222

lipid species identified beyond the original report

40%

expansion of the characterized lipidome from the same raw files

60,000

lipid species catalogued in LIPID MAPS, with experimental MS2 for only a small fraction

New lipids are only worth finding if they carry signal. We ran differential abundance analysis on the 222 Pyxis-only species, comparing obese and lean individuals in both subcutaneous (SAT) and visceral (VAT) adipose tissue. Many showed clear differences between groups. Fold changes above 2x, p-values below 0.05.

Volcano plots of obese versus lean log2 fold change against negative log10 p-value for the 222 Pyxis-only species in SAT and VAT.

Figure 1. Differential abundance analysis comparing obese and lean individuals in both subcutaneous (SAT) and visceral (VAT) adipose tissue. Human adipose, AdipoAtlas ST001738 [11].

Recovering a known biomarker

One of these IDs is PC 32:3, a known biomarker highly relevant to the biology of obesity that was not reported in the original AdipoAtlas study. PC 32:3 is a lower-abundance phosphatidylcholine, easy to miss against dominant PC species; but it’s been a fixture of commercial targeted lipidomics panels for years, with a steady accumulation of evidence tying it to insulin resistance and metabolic disease.

Prior work has linked PC 32:3 to insulin resistance in cultured adipocytes [12], ranked its ratio to related PC species among the most predictive metabolite ratios for insulin resistance across cohorts [13], and shown that higher LPCAT3 activity consumes PC 32:3 and related species to produce the polyunsaturated phospholipids that characterize obese, dysfunctional membranes [14].

In the AdipoAtlas data, PC 32:3 shows clear differential abundance between obese and lean groups across both SAT and VAT, consistent with the prior literature and surfaced here by a model that didn’t need to have seen it before.

Split violin distributions of normalized PC 32:3 abundance in SAT and VAT for lean and obese groups.

Figure 2. Per-sample abundance distributions for PC 32:3 in SAT and VAT, split by obese (orange) and lean (blue) groups. PC 32:3 was identified only by Pyxis and not reported in the original AdipoAtlas study, yet shows clear differential abundance between metabolic phenotypes, consistent with prior evidence linking this lipid to insulin resistance and obesity. Human adipose, AdipoAtlas ST001738 [11].

Getting started

The C18 AdipoAtlas data is loaded in Pyxis as a demo. Sign up today at app.matterworks.ai/sign-up to explore the data behind this post.

If you want to run your own data, simply sign up and upload. We’ve also put together a biological interpretation guide (below) with some tips on using the results you see in Pyxis.

Pyxis Lipidomics: Biological Interpretation Guide

1. Introduction

Pyxis lipidomics provides MS/MS-based lipid identification using our Large Spectral Model for AI-powered de novo prediction. This guide helps you interpret Pyxis lipid results for biological analysis, including what the outputs mean and where known limitations require caution.

2. Understanding Lipid Classification Levels (L1-L3)

Pyxis reports lipid identifications at three hierarchical levels, derived from LIPID MAPS:

  • L1: Lipid class (e.g., PC, PE, TAG, SM)

  • L2: Sum composition (e.g., PC 36:4), head group + total carbons and double bonds

  • L3: Molecular species with bond type (e.g., PE P-36:4)

“Uncategorized” lipids: Some identifications appear under an “Uncategorized” section in the Lipids tab. They are very likely lipids or lipid-like molecules, but could not be automatically placed into the LIPID MAPS classification hierarchy. This can happen because LIPID MAPS is not fully comprehensive. Some genuine lipids are absent from its database.

Tip: If you need to map Pyxis output to external databases, note that there is a many-to-one relationship: multiple analytes can map to a single L1/L2/L3 combination.

3. Scores

Each lipid ID is assigned a score: High, Medium, or Low.

  • The score is based on the maximum structural similarity score across all samples for a given identification

  • The same thresholds apply to both retrieval (library-matched) and de novo (AI-predicted) identifications; however, for the most rigorous comparisons, compare like to like: retrieval-to-retrieval and de novo-to-de novo.

  • A high score reflects structural match quality, but does not resolve all ambiguities (see Isomers section below)

4. Isomer Limitations

4a. Structural isomers (acyl chain composition). Isomers with the same sum composition but different individual acyl chains are collapsed to a single sum-composition name. For example:

  • PE(10:0_18:0) and PE(12:0_16:0) are both reported as PE 28:0

Pyxis does not resolve sn-position or individual chain-length composition.

4b. Ether lipids: Plasmenyl (P-) vs. Plasmanyl (O-). This is a key known limitation. Pyxis reports ether lipid bond type at L3 (e.g., PE P-36:4 vs. PE O-36:5), but cannot reliably distinguish between:

  • Plasmenyl (P-), vinyl ether linkage (plasmalogen)

  • Plasmanyl (O-), alkyl ether linkage

These species often share the same precursor mass and head group. The sn-1 neutral loss fragment that would distinguish them is frequently too weak to contribute meaningfully to the score.

Internal benchmarking shows that P- vs. O- assignment is essentially a coin flip, even at high confidence scores. Therefore, we recommend grouping them together as “ether” lipids rather than interpreting the P-/O- distinction as meaningful in any downstream analysis.

A note on chromatography: Plasmenyl (P-) and plasmanyl (O-) species can often be separated chromatographically due to fundamental differences in hydrophobicity. In principle, RT differences could help disambiguate them.

5. Pooled MS/MS samples: Current Limitations

A common study design collects MS/MS on pooled samples (for identification) and MS1 on individual samples (for quantitation and group comparisons).

Current limitation: Pyxis provides identifications only for MS/MS files. There is currently no mechanism to:

  • Extract peak areas from MS1-only data

  • Backfill lipid IDs from pools onto individual samples

This means differential abundance analysis (e.g., generating volcano plots comparing disease vs. healthy) requires MS/MS data on samples from each condition, not just on a pooled QC sample.

Workaround for pool-only designs: If your MS/MS data is limited to pools, you can use the precursor m/z (M/Z Min, M/Z Max columns, or the M/Z column when available) and retention time (First Observed RT, Last Observed RT columns) from the Pyxis CSV export to set up targeted extraction of those lipids in external software against your MS1-only sample files. This enables you to obtain peak areas for differential analysis across your individual samples, using Pyxis IDs as your target list.

We plan to address this limitation in an upcoming release, to remove this workaround, and provide peak areas natively in Pyxis.

6. Adduct Interpretation

  • Each identification includes a predicted adduct type (e.g., [M+H]+, [M+Na]+, [M-H2O+H]+)

  • Adduct information is not always displayed on XICs in the UI. To verify an ID, check that the reported adduct is consistent with the displayed precursor m/z for that compound class

7. Internal Calibrant Ions (Lock Mass)

If your instrument uses scan-to-scan internal calibration (e.g., fluoranthene at m/z 202.07 on Thermo Exploris), these calibrant ions appear in the raw spectra. This can lead to:

  • Unexpected peaks in mirror plots that don’t appear in the vendor viewer

  • Potential false matches near calibrant m/z values

Practical advice: Be aware of your instrument’s internal calibrant(s). If you see unexpected IDs at or near known calibrant m/z values, they may be artifacts.

8. Multi-Method and Multi-Column Data

If your study uses multiple LC methods (e.g., HILIC + RP, or C18 + C30):

  • Process each method as a separate sample set. Combining data from different chromatographic methods causes the same analyte to appear at very different retention times, making results difficult to interpret.

  • Cross-column discrepancies (e.g., a lipid appears significant on one column but not the other) can be a useful signal for flagging incorrect or isomeric IDs.

Equally, multi-polarity/ionization mode should provide an additional means of validating an ID. In this case, we expect a given ID to show the same RT in both modes. Furthermore, one should make the most out of having spectra from both polarities, as often these provide complementary, not redundant, information.

Identifying lipids de novo is one area of expertise in a repertoire that continues to grow as Pyxis learns from new spectra. Pyxis is the Matterworks co-scientist for interpreting omic data and predicting phenotypic biology.

Data: AdipoAtlas raw files for human white adipose tissue, C18 reversed-phase chemistry, Metabolomics Workbench ST001738 [11]. Identifications generated de novo from MS/MS spectra without reference library matching. Differential abundance compares obese and lean individuals in subcutaneous and visceral adipose tissue, at fold changes above 2x and p-values below 0.05. Content reproduced from the PyxisLabs Insights post of April 20, 2026, by Sam Burian, with the biological interpretation guide of the same date. Questions about this analysis: info@matterworks.ai

References

  1. LIPID MAPS: update to databases and tools for the lipidomics community. Nucleic Acids Research 52, D1677 (2024). academic.oup.com/nar/article/52/D1/D1677

  2. Kind, T. et al. LipidBlast in silico tandem mass spectrometry database for lipid identification. Nature Methods 10, 755-758 (2013). doi:10.1038/nmeth.2551

  3. von Gerichten, J. et al. Challenges in lipidomics biomarker identification: avoiding the pitfalls and improving reproducibility. Metabolites 14, 461 (2024). doi:10.3390/metabo14080461

  4. Hilvo, M. et al. Development and validation of a ceramide- and phospholipid-based cardiovascular risk estimation score for coronary artery disease patients. European Heart Journal 41, 371-380 (2020). doi:10.1093/eurheartj/ehz387

  5. Mapstone, M. et al. Plasma phospholipids identify antecedent memory impairment in older adults. Nature Medicine 20, 415-418 (2014). doi:10.1038/nm.3466

  6. Ogretmen, B. Sphingolipid metabolism in cancer signalling and therapy. Nature Reviews Cancer 18, 33-50 (2018). doi:10.1038/nrc.2017.96

  7. Wang, N. et al. Lipid metabolism drives dietary effects on T cell ferroptosis and immunity. Nature (2026). nature.com/articles/s41586-026-10193-4

  8. Liu, L. et al. NCBP2 drives colorectal cancer growth and metastasis through LIPG-mediated lipid droplet accumulation. Communications Biology (2026). nature.com/articles/s42003-026-09903-5

  9. Damiza-Detmer, A. et al. Lipid alterations and endothelial dysfunction are associated with multiple sclerosis pathophysiology. Scientific Reports (2026). nature.com/articles/s41598-026-44767-z

  10. Lu, J. et al. Lipidomic profiling identifies key pathways and a 5-lipid panel with high diagnostic efficacy for ischemic stroke. Scientific Reports (2026). nature.com/articles/s41598-026-42918-w

  11. Lange, M., Angelidou, G., Ni, Z., Criscuolo, A., Schiller, J., Blüher, M. & Fedorova, M. AdipoAtlas: a reference lipidome for human white adipose tissue. Cell Reports Medicine 2, 100407 (2021). doi:10.1016/j.xcrm.2021.100407. Data: Metabolomics Workbench ST001738.

  12. Böhm, A. et al. Metabolic signatures of cultured human adipocytes from metabolically healthy versus unhealthy obese individuals. PLOS One 9, e93148 (2014). doi:10.1371/journal.pone.0093148

  13. Molnos, S. et al. Metabolite ratios as potential biomarkers for type 2 diabetes: a DIRECT study. Diabetologia (2017). pmc.ncbi.nlm.nih.gov/articles/PMC6448944

  14. He, M. et al. Inhibiting phosphatidylcholine remodeling in adipose tissue increases insulin sensitivity. Diabetes 72, 1547 (2023). diabetesjournals.org/diabetes/article/72/11/1547