Our paper on integrating data from 2+ multiplexed assays of variant effect (MAVE) for the same gene is available now: doi.org/10.1186/s130... We also have a Shiny webtool where you can plug in data and compare a few different integration approaches. A big thank you to all co-authors... 1/3
Combining multiplexed functional data to improve variant classification - Genome Medicine
Background With the surge in the number of variants of uncertain significance (VUS) reported in ClinVar in recent years, there is an imperative to resolve VUS at scale. Multiplexed assays of variant effect (MAVEs), which allow the functional consequence of 100s to 1000s of genetic variants to be measured in a single experiment, are emerging as a powerful source of evidence which can be used in clinical variant classification. Increasingly, multiple published MAVEs are available for the same gene, sometimes measuring different aspects of variant impact. When multiple functional roles of a gene need to be considered, combining data from multiple MAVEs may provide a more comprehensive measure of the consequence of a genetic variant, which could impact variant classifications. Methods We curated published datasets from five MAVEs for the gene TP53, two MAVEs for LDLR and two MAVEs for PTEN. Statistical methods (principal component analysis), unsupervised learning (k-means clustering), and supervised learning (Naïve Bayes and random forest classifiers) were used to integrate multiple MAVE datasets. The utility of MAVE integration methods were assessed using standard metrics (sensitivity, specificity, etc) as well as evidence strength in a putative variant classification framework. Results Here, we provide guidance for combining such multiplexed functional data, incorporating a stepwise process from data curation and collection to model generation and validation. We also present a web applet that allows users to test various methods for combining score sets from multiple assays, calculate integrated functional scores for all variants, and assess whether combining data enables the application of stronger evidence for pathogenicity or benignity. In general, supervised learning methods such as random forest led to improved variant classification as compared to any individual MAVE dataset. Conclusions By following the steps outlined herein with appropriate guardrails, researchers can maximize the value of MAVEs, strengthen the functional evidence for clinical variant classification, and potentially uncover novel mechanisms of pathogenicity for clinically relevant genes.
doi.org