Insight Series

The ovarian cancer signal was in the unannotated lipidome

A published serum lipidome study of 426 women with gynecologic disease (Metabolomics Workbench ST002521) reached its conclusions from 3.4 percent of the molecular features the instrument detected. Working from the same raw files, the same patients, and the authors’ own train and test assignment, Pyxis analyzed the complete detected feature set and improved discrimination of ovarian cancer from benign gynecologic disease at every comparison tested, with the largest gains in early-stage disease.

Clinical context

Reported serum classifiers for ovarian cancer are commonly trained against healthy donors. That comparison does not correspond to the clinical decision. A woman presenting with an adnexal mass requires determination of whether that mass is malignant, which governs referral for subspecialty surgical management. CA-125 and the Risk of Ovarian Malignancy Algorithm (ROMA) are the established serum tools for this assessment. In a multicenter study of 649 women with adnexal masses requiring surgery, CA-125 alone achieved a specificity of 0.787 and sensitivity of 0.747 in premenopausal women, while ROMA achieved a specificity of 0.926 and sensitivity of 0.707 [2]. Sensitivity near 0.71 to 0.75 is the principal limitation of both.

The cohort benchmarked here contains no healthy comparator. All subjects carry gynecologic pathology. This is a more difficult discrimination than cancer against health and a closer analogue of the clinical question. The original authors quantified the difference directly: the lipid panel that achieved an AUC of 0.94 against healthy controls achieved 0.82 against benign gynecologic disease [1].

Cohort

Study ST002521 comprises serum from 426 subjects, acquired on a Thermo Orbitrap ID-X Tribrid with reversed-phase separation in both ionization polarities [1]. All subjects were sampled during surgery. The study’s healthy control samples were collected under different conditions and were excluded by the original authors as confounded, which is why the malignant-versus-benign comparison reported here is internally consistent with respect to collection.

Clinical assignment

Subjects

Class

Ovarian cancer

273

Malignant

Cervical cancer

60

Malignant

Benign ovarian tumor

59

Benign

Benign uterine tumor

31

Benign

BRCA positive

3

Predisposed, undiagnosed

Table 1. Cohort composition of ST002521. The published analysis used an age-matched subset of 325 subjects, comprising 208 ovarian cancer subjects and 117 non-ovarian-cancer subjects, with benign tumors and cervical cancer pooled as the comparator.

Annotation coverage in the published analysis

The published workflow extracted 24,297 features in positive ionization mode and 5,485 in negative mode, giving approximately 29,800 molecular features after de-isotoping and de-adducting. Matching against a curated in-house spectral library, with manual review, identified 994 lipid species across 22 subclasses, together with 22 gangliosides assigned by molecular formula and exact mass [1].

All downstream analysis in the published study was performed on those identified lipids. The remaining 96.6 percent of detected features could not be matched to a library entry and was excluded. This unannotated fraction is biochemical dark matter: molecular signal that was measured, paid for, and retained, but was inaccessible to an analysis restricted to compounds that can be named.

3.4%

of detected molecular features were annotated against a curated spectral library and carried into the published analysis.

96.6%

were detected but unannotated and therefore excluded. Pyxis analyzed the complete detected feature set.

Benchmark on the published train and test split

The authors deposited their per-sample train and test assignment, permitting exact reproduction of the benchmark. The comparison below uses the same 227 subjects for training and the same 98 subjects held out, with the same task definition: ovarian cancer against a comparator pooling benign tumors and cervical cancer. The held-out set comprises 64 ovarian cancer subjects and 34 non-ovarian-cancer subjects, the same subjects on which the published performance figures were reported. The single difference is the proportion of the acquisition carried into analysis.

Metric on the shared 98-subject test set

Published [1]

Pyxis

AUC

0.85

0.92

Sensitivity

0.75

0.95

Specificity

0.82

0.71

Molecular features informing the result

1,016

Complete detected feature set

Table 2. Performance on identical subjects. Published figures are the authors’ best reported test-set result. The Pyxis configuration was selected on a validation split partitioned from the training set; the test set was not used for selection. A random-label control on the same split returned an AUC of 0.54.

At the default decision threshold the improvement is concentrated in sensitivity, which increases from 0.75 to 0.95, while specificity is lower at 0.71. Because the analysis returns a per-subject probability, the operating point is a choice rather than a fixed property. Selecting a threshold on the validation data to target higher specificity yields a test-set sensitivity of 0.78 and specificity of 0.91, exceeding the published result on both axes at once. Sensitivity and specificity can be traded against each other according to the clinical cost of each error type.

Same-organ comparison

A result on the pooled comparator admits an alternative explanation. The comparator group contains cervical cancer and benign uterine disease in addition to benign ovarian disease, so a classifier could achieve separation by learning anatomical site rather than malignancy. This was tested directly.

Restricting the comparison to ovarian cancer against benign ovarian tumor removes the anatomical cue: both groups arise in the same organ, and malignancy is the remaining difference. Discrimination was retained, with an AUC of 0.89 and sensitivity of 1.000, against a random-label control of 0.47 on the same subjects. Specificity in this comparison was 0.60 on 15 benign subjects. The discriminating signal is therefore attributable to malignancy rather than to tissue of origin.

Early-stage discrimination

Cancer stage is absent from the deposited metadata but was recovered from the authors’ public code repository and mapped to the deposited samples, giving stage assignments for 208 ovarian cancer subjects. Early-stage detection is the setting in which a serum test carries the greatest clinical value, and it is where the published lipid panel performed least well.

Two benchmarks were run with models trained specifically for early-stage discrimination, matching the design of the corresponding published analyses. In both, configuration was selected on a validation split and the test set was untouched until the configuration was fixed.

Comparison

Metric

Published [1]

Pyxis

Early-stage OC vs benign disease

AUC

0.86

0.93

Sensitivity

0.81

0.81

Specificity

0.69

0.82

Early-stage OC vs full comparator

AUC

0.75

0.92

Sensitivity

Not reported

0.81

Specificity

0.45

0.88

Table 3. Early-stage performance on held-out subjects. Held-out sets contain 31 early-stage ovarian cancer subjects with 17 benign comparators, and 31 with 34 comparators, respectively. Random-label control returned an AUC of 0.49.

The gain is largest where the published panel was weakest. Against the full comparator, specificity rises from 0.45 to 0.88 at comparable sensitivity. The information that distinguishes early-stage ovarian cancer from benign gynecologic disease was present in the acquisition and was not reachable through the annotated fraction.

Benign subjects with elevated malignant probability

Per-subject probabilities allow examination of individual disagreements between prediction and clinical assignment. Among held-out benign subjects, these disagreements are not distributed evenly across the benign group.

Of nine benign ovarian tumor subjects in the held-out set, three carry malignant probabilities above 0.5, two of them at 0.98 and 0.97. The median probability among confirmed ovarian cancer subjects is 0.97. Of eight benign uterine tumor subjects, none exceeds 0.5, and their median probability is 0.05. The same pattern appears in the independent same-organ comparison, in which six of fifteen benign ovarian tumor subjects exceed 0.5.

The organ specificity of this pattern is the relevant observation. Classification error distributed at random would appear across the benign group; instead it is confined to benign tumors of the ovary, the organ in which the malignancy under discrimination arises. The subgroups are small, and the specific subjects involved differ between the two comparisons, so the pattern rather than any individual subject is what replicates.

The scope of this observation should be stated precisely. It establishes that a subset of women diagnosed with benign ovarian disease have serum lipidomes more similar to the malignant group than to the remainder of the benign group. It does not establish that those women had undetected malignancy. ST002521 records clinical assignment at the time of sampling and contains no outcome or follow-up data, so early disease cannot be distinguished from benign biology with a similar lipidome profile. That distinction requires prospective follow-up.

Per-subject probabilities across clinical assignments

Because the analysis returns a probability for every subject rather than a label alone, the held-out set can be examined subject by subject. Figure 1 plots all 98 held-out subjects from the head-to-head benchmark against their clinical assignment, with ovarian cancer subjects separated by recovered stage.

Benchmark figure for discriminating ovarian and cervical cancer from benign gynecologic disease in the serum lipidome

Figure 1. Predicted probability of malignancy for each held-out subject. Each point is one subject from the 98-subject held-out set; horizontal bars are group medians. Ovarian cancer subjects are separated by stage assignment recovered from the original authors’ code repository. The dashed line marks the default decision threshold. Benign subgroups contain 9 and 8 subjects respectively, so the pattern rather than any individual subject is the observation.

Three features of the distribution are worth noting. Early-stage and advanced ovarian cancer sit at comparable heights, with medians of 0.958 and 0.972, which is the stage result of the previous section shown directly. The two benign groups behave differently from one another: benign uterine tumors cluster low, with a median of 0.053 and no subject above the threshold, while benign ovarian tumors are split, with three of nine above the threshold and two of those inside the range occupied by confirmed malignancy.

The third feature is the source of the specificity loss at the default threshold. Cervical cancer subjects account for seven of the ten comparator subjects placed above the threshold, which follows from the published design pooling cervical cancer into the non-ovarian-cancer comparator. A classifier separating malignant from benign disease will place malignancies of another organ on the malignant side, and under this task definition each one is scored as an error.

Differences from the published analysis design

The head-to-head benchmark reproduces the published design in order to isolate a single variable. Two features of that design were inherited rather than selected. The comparator pools benign tumors with cervical cancer, so the task measures separation of ovarian cancer from all other subjects rather than separation of malignant from benign disease. The cohort was also age-matched to 325 subjects, leaving 101 deposited samples unanalyzed, including all BRCA-positive subjects. Analyses addressing each are in progress.

Scope of this report

This is a benchmarking study of predictive signal in an existing dataset. It is not a diagnostic claim, and nothing reported here describes a test ready for clinical use. Any candidate arising from this analysis would require independent validation in separate cohorts, prospective evaluation, and the full regulatory path that applies to diagnostics. The question addressed is narrower and prior to all of that: how much of the predictive information present in a biomarker discovery dataset survives conventional analysis, and how much is discarded.

Limitations

The comparator throughout is gynecologic pathology rather than health, so these results apply to characterization of a known adnexal mass and not to screening in an asymptomatic population. Held-out subgroups are small, with 34 non-ovarian-cancer subjects in the main comparison and fewer within each benign category, which places substantial uncertainty on the specificity estimates. All samples were acquired on a single instrument within a single study, so instrument and site effects cannot be separated from biological effects. Results derive from single fixed train and test partitions rather than repeated resampling. Comparisons against CA-125 and ROMA performance are drawn from a separate cohort and are indicative rather than direct. Any application of these findings would require validation in independent cohorts.

Summary

Pyxis is the Matterworks co-scientist for omic data, developed to interpret molecular measurements and relate them to phenotypic biology. The improvement reported here does not derive from a different cohort, a larger cohort, or a different instrument. It derives from analyzing the same acquisition in full.

The implication extends past this one study. Biomarker discovery datasets are generated at considerable cost, and the analyses built on them typically reach only the fraction of the measurement that can be matched to a spectral library. The predictive information in the remainder does not disappear. It goes unused, and in a clinical domain where a missed malignancy carries the consequence it does here, that is a cost borne by patients.

Data: Metabolomics Workbench ST002521 (human serum, Orbitrap ID-X Tribrid, reversed phase, MS1). Benchmark figures use the original authors’ published per-sample train and test assignment and recovered stage assignments, both obtained from their public code repository. Configuration was selected on validation splits partitioned from the training data; test sets were not used for selection. Published comparison figures are as reported in the source study.

For questions about this analysis, contact info@matterworks.ai.

References

  1. Sah S, et al. Serum Lipidome Profiling Reveals a Distinct Signature of Ovarian Cancer in Korean Women. Cancer Epidemiology, Biomarkers & Prevention. 2024;33(5):681. doi:10.1158/1055-9965.EPI-23-1293

  2. Choi HJ, Lee YY, Sohn I, Kim YM, Kim JW, Kang S, Kim BG. Comparison of CA 125 alone and risk of ovarian malignancy algorithm (ROMA) in patients with adnexal mass: a multicenter study. Current Problems in Cancer. 2020;44(1):100508. doi:10.1016/j.currproblcancer.2019.100508