SomaScan™ 11K Assay is now Illumina SomaScan Discovery. Learn more.

Opening the black box: Building and evaluating machine learning models for proteomics

Machine learning can be a powerful way to extract multivariate signals from high-dimensional proteomics data, but only when models are built and evaluated with rigor. In this on-demand webinar, explore a practical framework for developing predictive and explanatory models, with an emphasis on study design choices that help prevent data leakage, reduce overfitting, and improve reproducibility.

Learn how to structure training and test splits, use resampling approaches such as cross-validation, and select features in ways that account for correlation, batch structure, and common confounders. The session also covers how to choose and interpret performance metrics for regression and classification, including AUC, RMSE, sensitivity, and specificity, as well as how to sanity-check models so results hold up on new samples, not just the dataset used for training.

The webinar also provides guidance on clearly communicating model results, with practical approaches to interpretation and reporting for proteomics research. For SomaScan™ Assay customers, the session highlights how these best practices translate to common SomaSignal™ Tests use cases and outputs.

Watch this on-demand webinar to better understand when machine learning is the right next step and how to apply it responsibly to generate insights you can trust.

By the end of this session, you will be able to:

  • Define the goal of a multivariable model in proteomics (prediction vs. explanation) and select an appropriate modeling strategy.
  • Design robust validation plans (train/test splits and cross-validation) that minimize data leakage and overfitting.
  • Apply feature selection approaches that account for high dimensionality, correlation, batch structure, and common technical confounders.
  • Choose and compare commonly used algorithms for regression and classification, and evaluate performance with fit-for-purpose metrics.
  • Interpret and communicate model results transparently so findings can be reproduced and trusted on new data.

Erin Hales PhD

Erin Hales, PhD

Illumina

Erin Hales received her PhD in animal biology from the University of California at Davis studying the genomics and metabolomics of an equine neurodegenerative disease. She applied her learning of statistics and animal biology to the Golden Retriever Lifetime study as a post-doc at Morris Animal Foundation. Driven by her desire to push science forward she joined the applied modeling team at SomaLogic before transferring to her current bioinformatics support scientist role.

Yehonatan-Elon PhD

Yehonatan Elon, PhD

OncoHost

Dr. Yehonatan Elon has over 15 years of experience in the data and algorithms industry, with a proven track record in end-to-end biomarker development. Dr. Elon holds a PhD in physics from the Weizmann Institute of Science, was director of new technologies at BrightSource Energy, and VP of Research at MeteoLogic and Feedvisor.

Opening the black box: Building and evaluating machine learning models for proteomics

A presentation by Erin Hales, PhD and Yehonatan Elon, PhD

Share with colleagues

More webinars

WebinarBeyond the trees: See the bigger biological picture with pathway enrichment analysis

Proteomic data holds enormous potential—but the biology is rarely contained in a single protein. When you’re handed a list of differentially abundant proteins, the next challenge is turning that list into a clear, testable story about mechanisms and processes.

Learn more

WebinarFinding the signal: Identifying reproducible biomarkers in high-plex proteomics

In biomarker discovery, the challenge is rarely a lack of data. Rather, it is knowing how to separate meaningful biological signal from technical distraction. This webinar focuses on how to use univariate analysis as a practical and powerful entry point for biomarker discovery in high-plex proteomics studies.

Learn more

WebinarMore than the sum of its parts: How harmonized proteomic data reveals meaning across disparate clinical cohorts

Proteomics data generated across sites, time points, and workflows can be difficult to compare using standard normalization alone. This webinar shows how harmonization aligns datasets into a shared biological framework and reveals signals across studies and cohorts, highlighted through the GNPC’s analysis of 40,000 patient samples.

Learn more

Explore webinars in our interactive viewer