Expertise Brief
New Expertise Brief
De novo identification of oxylipins
The Pyxis co-scientist identifies oxylipins de novo from untargeted acquisitions and tests each annotation against independent evidence to reach a scientifically defensible result.
Summary
Pyxis™ is the Matterworks co-scientist for interpreting omic data and predicting phenotypic biology. Pyxis’ expertise spans several domains, drawing on a set of expert models at our disposal. One of those domains is biochemical omics, the study of small molecules, lipids, and peptides by mass spectrometry, and this brief shows one aspect of that work: the de novo identification of oxylipins. Pyxis brings PhD-level analytical chemistry expertise to it and works the way careful scientists do: by making an annotation hypothesis and then testing against independent evidence before accepting it. Against a curated oxylipin dataset spanning human serum and microalgae, Pyxis recovered 77% of the reported oxylipins in serum and 74% in algae, and resolved oxo-fatty-acid positional isomers where the fragmentation allowed. Pyxis then tested those annotations: they organized cleanly by lipid class, they followed the retention-time behavior expected from their structures, and they reproduced the omega-3 and omega-6 oxidation patterns of each matrix. The result is a scientifically defensible analysis, reached the way a trained expert would reach it.
Challenging and refining Pyxis’ conclusions
Pyxis interprets omic data and connects it to phenotypic biology, working across several domains, one of which is biochemical omics, where our analysis is powered by the Matterworks Large Spectral Model (LSM), a foundation model trained on more than ten billion raw MS/MS spectra across more than 1.9 million biological contexts, spanning small molecules, lipids, peptides, and proteins [5]. From that training Pyxis learns structural representations directly from fragmentation patterns, annotating compounds that lie outside existing libraries and estimating concentrations without per-analyte calibration standards [5,6]. Within biochemical omics Pyxis’ repertoire already covers structure identification, absolute quantitation, batch-effect correction, and phenotype prediction [5,6]. The work described here is one area of expertise within that repertoire.
Pyxis works the way careful analysts work. An annotation is a hypothesis, and a hypothesis is worth only as much as the evidence that survives an attempt to break it. Pyxis states the confidence each spectrum supports, then challenges the call, looking for evidence that could corroborate or refute it: physical measurements Pyxis did not use to make the annotation, and biological patterns that should hold only if the annotation is real. As partners, we challenge the same calls, and Pyxis defends or revises each one against the evidence we raise. Rigor here is an exercise in poly-intelligence, human scientists and an AI co-scientist reaching conclusions together that neither would reach alone. The sections that follow show that method applied to a chemistry that has been difficult for conventional tools, the oxidative modification of fatty acids.
Why oxylipins are hard to identify
Oxylipins are oxygenated derivatives of polyunsaturated fatty acids (PUFAs). In animals they arise through the cyclooxygenase, lipoxygenase, and cytochrome P450 pathways, together with non-enzymatic peroxidation, producing prostaglandins, thromboxanes, leukotrienes, hydroxy and hydroperoxy fatty acids, epoxides, and related mediators of inflammation and physiology [1,2]. The same chemistry drives rancidity in fish oils and other PUFA products, where hydroxyl, hydroperoxy, keto, and epoxy groups change flavor, nutritional value, and shelf life.
Two properties make these molecules hard to annotate. First, many oxidized species are absent from spectral libraries, and non-enzymatic oxidation products in particular are under-represented in public repositories [1]. Second, oxidation generates positional isomers that share a molecular formula, and therefore a precursor mass, so accurate mass alone leaves them unresolved and structural detail depends on the fragmentation spectrum [3,4]. A method that ranks spectra against a fixed library will miss species outside that library and will report only the level of detail the library holds.
Reading oxidation from the spectrum
For an oxidized lipid, Pyxis assigns the class, chain length, degree of unsaturation, and number of oxygen additions directly from the fragmentation spectrum, and resolves the position of a modification when the fragmentation carries that information. The oxidative modifications Pyxis recognizes are the ones that dominate PUFA chemistry: hydroxyl, hydroperoxy, keto, and epoxy additions (Figure 1).

Figure 1. Parent omega-6 and omega-3 PUFAs (top) and the common oxidative modifications Pyxis identifies (bottom): hydroxyl, hydroperoxy, keto, and epoxy additions, with the added oxygen highlighted. Parent structures are rendered from canonical SMILES; the modification row shows representative motifs on a fatty-acyl chain.
Benchmark against a curated oxylipin dataset
To measure oxylipin identification against an outside standard, Pyxis was evaluated on a public dataset built to catalog oxylipin diversity by tandem mass spectrometry (the NEO-MSMS dataset), covering human serum and microalgae [1]. The authors reported 125 unique oxylipins across the two matrices (82 in serum, 75 in algae, with 32 shared). Pyxis processed the raw LC-MS/MS data, and the resulting annotations were compared against the reported set by InChIKey after the reported common names were resolved to structures. Of the 125 reported oxylipins, 105 resolved to an InChIKey and formed the comparison set. Twenty compounds could not be resolved to a structure in PubChem or LIPID MAPS, a measure of how many reported oxylipins remain outside standard structure databases.
Matrix
Reported (with InChIKey)
Identified by Pyxis
Recovery
Serum
69
53
77%
Algae
66
49
74%
Among the serum matches, 30 were at high structural similarity (at or above 0.91) and 17 at medium similarity (0.75 to 0.91); in algae, 35 were at high and 12 at medium similarity.
Confidence Assignment Strategy
A careful chemist reports the level a spectrum supports and no more, and Pyxis holds to the same discipline. Most annotations in this dataset are confident at the sum-composition level: a defined class, chain length, double-bond count, and oxygen count, for example FA 20:5;O1 for a mono-oxygenated eicosapentaenoic acid. At this level a match indicates a species with the same composition as the reported oxylipin, which is consistent with the general limits of MS/MS annotation for oxidized lipids [3,4]. Where the fragmentation distinguishes isomers, Pyxis commits to the specific structure. In this dataset the oxo-fatty acids met that bar: 7-oxo-DHA, 13-oxo-DHA, and 17-oxo-DHA were identified at the isomer level, consistent with the distinct fragmentation of keto groups at different positions along the chain. Pyxis reports these as isomer-level calls and the rest at sum-composition level, so a reader knows exactly what each annotation claims.
For monitoring oxidation in omega-3 and omega-6 products, sum-composition-level detection of oxygenation state is enough to flag the onset and extent of oxidation and to track it across production and storage. Isomer-level calls add mechanistic detail where the spectra support them.
A first cross-check: class organization
Pyxis then looked beyond the reported oxylipins to the full set of lipids in the samples. After filtering to lipid-class analytes and applying a blank subtraction (five times the maximum blank peak area), 421 unique compounds remained. Of these, 293 (70%) carried oxygen modifications, 81 (19%) matched at high structural similarity (at or above 0.91), and 54 (13%) were both oxidized and high-similarity (Figure 2). The samples are rich in oxidized lipids, and Pyxis annotates them at scale.

Figure 2. Of 421 blank-filtered Pyxis detections, the share that carry oxygen modifications, the share that match at high structural similarity, and the share that are both.
A larger set invites the question of whether the annotations hold together. Pyxis’ first check is internal: do the assignments organize the way real lipid chemistry organizes? To answer it Pyxis used a standard first-pass tool from lipidomics, a plot of Kendrick mass defect (KMD) against retention time [3,4]. KMD arranges lipids by structural similarity, so each class should fall in its own region and a misassignment should stand out (Figure 3). The annotations span eight lipid classes: fatty acids (FA), free fatty acids (FFA), sterols (ST), bile acids (BA), acylcarnitines (CAR), fatty-acyl esters of hydroxy fatty acids (FAHFA), wax esters (WE), and steryl esters (SFE). Each fell where its chemistry places it. The fatty acids form a dense central band. Sterols sit at high KMD and later retention, reflecting the ring system, with bile acids nearby, consistent with their shared steroid backbone. Short-chain, highly oxygenated free fatty acids form a separate cluster at low retention time and low KMD. The minor classes appear at the periphery. The set organizes as a lipidomics analyst would expect, which is the first sign the assignments are sound.

Figure 3. Kendrick mass defect (KMD, H base) versus retention time for the blank-filtered detections, colored by lipid class and sized by double-bond count. Each class occupies its own region, consistent with its chemistry.
An independent test: retention time
Internal coherence is encouraging, but a rigorous analyst looks for a measurement that could actually refute the result. Retention time is that measurement. It is recorded during the chromatographic run and plays no part in the spectrum-based annotation, so it is free to disagree. In reverse-phase chromatography, retention reflects hydrophobicity: longer chains elute later, and added oxygenation raises polarity and shifts a species toward earlier elution. If Pyxis had assigned the wrong chain length, degree of unsaturation, or number of oxygens, the retention times would not follow these trends.
They do. Plotting retention time against chain length for the free fatty acids, faceted by double-bond count and colored by the number of oxygens, shows both relationships together (Figure 4). Within each facet, retention time rises with chain length along a smooth path, and for a given chain length the more oxygenated species elute earlier. The same behavior held for the esterified fatty acids and the bile acids. Because the retention data could have contradicted the annotations and instead corroborated them, it raises confidence in the calls rather than merely restating them. Fewer than one percent of the detections sat away from the expected trend, and Pyxis flagged those for closer inspection rather than letting them pass.

Figure 4. Retention time versus chain length for the free fatty acids (n=124), faceted by double-bond count and colored by number of oxygens. Retention rises with chain length, and added oxygenation shifts a species toward earlier elution.
The oxidation of omega-3 and omega-6 fatty acids
A third line of evidence is biological. If the annotations describe real molecules, the oxidized lipids in each sample should reflect the biology that produced them, and a signature that appears only in the matrix where the enzymes exist is hard to explain by chance matching. The omega-3 and omega-6 families are the substrates most relevant to oxidation in physiology and in PUFA products, so Pyxis examined them closely. Across the high-similarity detections, oxidation states ranged from unmodified fatty acids to species carrying four or more oxygens, and the distribution differed between the two matrices (Figure 5).

Figure 5. Distribution of oxygen substituents among high-similarity detections (at or above 0.91) in serum (n=60) and algae (n=63).
The omega families showed several patterns that match known biology (Figure 6). Serum carried more omega-6 oxylipins than algae (27 versus 17), consistent with the mammalian eicosanoid pathways that act on arachidonic acid (20:4) and linoleic acid (18:2) [2] and with the high omega-6 content of a typical Western diet. Algae carried a modestly higher count of omega-3 oxylipins (28 versus 25), consistent with the role of microalgae as primary producers of EPA and DHA.
The per-fatty-acid breakdown placed docosahexaenoic acid (DHA, 22:6) at the top of oxidative diversity in both matrices, with 12 species each, in line with its six double bonds and its susceptibility to oxidation. Eicosapentaenoic acid (EPA, 20:5) and stearidonic acid (SDA, 18:4) followed. On the omega-6 side, serum was led by linoleic acid (18:2) with 10 species and arachidonic acid (20:4) with 9.
The clearest biological test came from the matrix comparison. Eicosadienoic acid (20:2n-6) and adrenic acid (22:4n-6), both mammalian elongation products of linoleic acid, were present in serum (4 species each) and almost absent in algae (1 species each), which is exactly the pattern expected if Pyxis is reading real biology rather than matching indiscriminately, since algae lack the mammalian elongation machinery. The oxidation-state profiles also differed: serum omega-3 oxylipins were dominated by mono- and di-oxygenated species, with about two-thirds carrying one or two oxygens, while omega-6 oxylipins in algae shifted toward more heavily oxygenated forms, three oxygens being the most common. Each of these patterns is one Pyxis could have gotten wrong and did not.

Figure 6. Omega-3 and omega-6 oxylipins identified by Pyxis in serum and microalgae. Top: distribution of oxygen substituents by omega family and matrix. Bottom: unique compounds per parent fatty acid, colored by number of oxygens.
Scope and considerations
Part of a defensible analysis is naming where it could be wrong. Pyxis’ annotations reflect the composition of the underlying reference space, which is weighted toward mammalian lipids, so algal-specific species may be under-represented and recovery in non-mammalian matrices can be affected. Isomer-level assignment is reported where the fragmentation resolves it and at the sum-composition level otherwise. Pyxis states these bounds so that the results are read within them, as a co-scientist would.
Expanding analytical expertise
Pyxis identifies the common oxidative modifications of fatty acids, including hydroxyl, hydroperoxy, keto, and epoxy additions, from LC-MS/MS data. Because the identification draws on learned spectral representations, it reaches species that fixed libraries omit, and it works without per-analyte calibration standards. Pyxis does not stop at the annotation, but states the confidence it deserves, checks that the whole set organizes coherently, tests it against a measurement that could refute it, and confirms that the biology reads true. This is how a well-trained analytical chemist reaches a conclusion worth standing behind, and it is how Pyxis reaches conclusions.
For manufacturers of omega-3 and omega-6 products, this supports monitoring of oxidation across production and storage. For research, it supports profiling of oxylipin signatures associated with inflammation and disease [1,2]. Identifying oxylipins is one area of expertise within Pyxis’ biochemical omics work, which is itself one domain of a wider role: interpreting omic data and connecting it to phenotypic biology. The set of expert models at our disposal continues to grow as Pyxis learns from new data.
To work with Pyxis on oxidized lipid profiling, contact info@matterworks.ai
References
[1] Elloumi A, Mas-Normand L, Bride J, et al. From MS/MS library implementation to molecular networks: exploring oxylipin diversity with NEO-MSMS. Scientific Data. 2024;11:193. doi:10.1038/s41597-024-03034-4
[2] Eccles JA, Baldwin WS. Detoxification cytochrome P450s (CYPs) in families 1 to 3 produce functional oxylipins from polyunsaturated fatty acids. Cells. 2022;12(1):82. doi:10.3390/cells12010082
[3] Lerno LA, German JB, Lebrilla CB. Method for the identification of lipid classes based on referenced Kendrick mass analysis. Analytical Chemistry. 2010;82(10):4236-4245. doi:10.1021/ac100556g
[4] Richardson LT, Neumann EK, Caprioli RM, Spraggins JM. Referenced Kendrick mass defect annotation and class-based filtering of imaging MS lipidomics experiments. Analytical Chemistry. 2022;94:5504-5513. doi:10.1021/acs.analchem.1c03715
[5] Asher G, Delmar MC, Campbell JM, Geremia J, Kassis T. LSM1-MS2: a foundation model for MS/MS, encompassing chemical property predictions, search, and de novo generation. chemRxiv. 2024. doi:10.26434/chemrxiv-2024-k06gb-v3
[6] Ferro LS, Wong AYL, Howland J, et al. A scalable approach to absolute quantitation in metabolomics. bioRxiv. 2024. doi:10.1101/2024.09.09.609906
Pyxis and Large Spectral Models (LSMs) are technologies of Matterworks. · matterworks.ai
Keep reading
Pyxis identifies oxylipins de novo and tests each annotation against independent evidence.
Discover the meaning in your measurements
Engage with Pyxis on a new experiment or discover the novel biology hidden in data you already have.