INSIGHT SERIES
Two published reference points on the same raw data
In 2020, a Stanford-led team built one of the most detailed pictures yet of human pregnancy metabolism: 784 weekly blood draws from 30 women, profiled by untargeted LC-MS, used to construct a metabolic clock that predicts gestational age (GA) directly from a blood sample [1] . This hand-selected panel of 5 metabolic features, cross-validated and fit with elastic-net regression, tracked GA nearly as accurately as first-trimester ultrasound, the clinical gold standard [1] . Two years later, a team including several of the original authors reprocessed this same raw dataset to benchmark a new deep-learning method, deepPseudoMSI, against a standard machine-learning baseline [2] . Taken together, these results give two published reference points against which to test Pyxis.
The same raw files, processed through an untuned pipeline
We reprocessed the same raw files through the predictive-benchmarking pipeline of Pyxis™, the Matterworks co-scientist for omic data. It is the same pipeline we run, untuned, across dozens of heterogeneous public studies. Running the files through it took less than an hour and was essentially zero-touch. We matched the original paper’s own cohort structure, training on the discovery cohort (21 subjects) and testing on the test set (9 subjects, analyzed in a separate year) [1] . On that held-out cohort, the resulting model reached an R² of 0.936.
R² = 0.936 on the held-out Test Set 1 cohort, compared with 0.91 for the original metabolic clock and 0.79 for the published deep-learning method.
on the held-out Test Set 1 cohort, compared with 0.91 for the original metabolic clock and 0.79 for the published deep-learning method.
Comparison with the published deep-learning result
The 2022 reanalysis [2] evaluated two methods on this identical dataset, using subject-aware 5-fold cross-validation in which every subject is held out exactly once and results are averaged across folds. A Random Forest baseline built on traditional XCMS-derived features reached an R² of 0.76, and the proposed deepPseudoMSI method, a convolutional neural network trained on images derived from the raw LC-MS data, reached an R² of 0.79. On the same raw data, the best predictive model from Pyxis reaches an R² of 0.936.
Table 1. Comparison of model results on the GA prediction task. Pyxis and Liang et al. use the identical train/test cohort split; the figures from Shen et al. come from a separate, random 5-fold evaluation.
Approach Model Split R² Liang et al. 2020 Elastic-net clock (5 metabolites) Discovery train / Test Set 1 test 0.91 Shen et al. 2022 Random Forest (traditional features) Random 5-fold CV 0.76 Shen et al. 2022 deepPseudoMSI (CNN) Random 5-fold CV 0.79 PyxisLabs Best predictive model from Pyxis Discovery train / Test Set 1 test (same as Liang et al.) 0.936
The comparison is asymmetric in one respect. The result reported here comes from a single split that crosses a real acquisition-year boundary, while the figures from Shen et al. are averaged over a random 5-fold evaluation that pools all years together. That asymmetry makes the split used here the harder test.
Figure 1. Predicted against actual gestational age in the held-out cohort. Out-of-sample predictions from the Pyxis predictive model track the 1:1 line across the full gestational range, with R² = 0.936 and a mean absolute error of 1.68 weeks.
On the same raw data and the same predictive task, the model built by Pyxis performs better than both the published baseline and the published deep-learning method developed to improve on it.
Higher accuracy than the original metabolic clock, obtained without per-study curation
The original clock scored an R² of 0.91 on the Test Set 1 cohort, with a mean absolute error of 2.11 weeks and an RMSE of 2.76 weeks [1] . Liang et al. also ran a second, harder validation, Test Set 2, on 8 additional subjects recruited three years later; that cohort is absent from the raw data deposited for this study, so it falls outside this comparison. On the same held-out cohort, the predictive model from Pyxis reaches an R² of 0.936 with a mean absolute error of 1.68 weeks, about 20% lower than the original clock.
mean absolute error on the held-out cohort, about 20% below the 2.11 weeks reported for the original metabolic clock, and reached without the curation that clock depended on.
Accuracy is one part of the comparison. The analytical work required to reach it is the other. The original clock depended on QC-based instrument drift correction, batch correction, and manual curation of metabolic features, carried out for this study specifically.
Figure 2. Batch correction and curation steps behind the original metabolic clock. [1] Gold-outlined steps are manual, per-study curation work.
The predictive models built by Pyxis require none of those steps. Evaluated on the exact same cohort, the best model reaches a 20% lower mean absolute error while requiring less time and less specialist effort.
What the comparison establishes
The scope here is narrow by design: one dataset, one endpoint, and one held-out cohort defined by the original authors. Within that scope, a general-purpose interpretation of the raw files, produced with no dataset-specific engineering, supports a predictive model that exceeds two published models built for this task, one of them backed by extensive manual curation. For laboratories holding archived LC-MS data, the practical consequence is reach: files that have already been acquired can be reinterpreted and modeled in under an hour.
For the New Insight Series, we asked Pyxis to revisit published studies where advanced LC-MS analysis was performed and, in each case, to look for results that the original analysis may have left on the table. The pregnancy metabolome of Liang et al. is one such study.
Data: Metabolomics Workbench ST001430 (784 plasma samples, 30 subjects, untargeted LC-MS). Original study: Liang et al., Cell 181, 1680–1692 (2020). Deep-learning benchmark: Shen et al., Briefings in Bioinformatics 23, bbac331 (2022). Figure 1 is from the Pyxis reanalysis of ST001430. Figure 2 summarizes the workflow described in Liang et al.
References
[1] Liang, L., Rasmussen, M-L. H., Piening, B., et al. Metabolic dynamics and prediction of gestational age and time to delivery in pregnant women. Cell 181, 1680–1692 (2020). doi:10.1016/j.cell.2020.05.002
[2] Shen, X., Shao, W., Wang, C., et al. Deep learning-based pseudo-mass spectrometry imaging analysis for precision medicine. Briefings in Bioinformatics 23(5), bbac331 (2022). doi:10.1093/bib/bbac331

