Cookies on this website

We use cookies to ensure that we give you the best experience on our website. If you click 'Accept all cookies' we'll assume that you are happy to receive all cookies and you won't see this message again. If you click 'Reject all non-essential cookies' only necessary cookies providing core functionality such as security, network management, and accessibility will be enabled. Click 'Find out more' for information on how to change your cookie settings.

BackgroundWith the surge in the number of variants of uncertain significance (VUS) reported in ClinVar in recent years, there is an imperative to resolve VUS at scale. Multiplexed assays of variant effect (MAVEs), which allow the functional consequence of 100s to 1000s of genetic variants to be measured in a single experiment, are emerging as a powerful source of evidence which can be used in clinical variant classification. Increasingly, multiple published MAVEs are available for the same gene, sometimes measuring different aspects of variant impact. When multiple functional roles of a gene need to be considered, combining data from multiple MAVEs may provide a more comprehensive measure of the consequence of a genetic variant, which could impact variant classifications.MethodsWe curated published datasets from five MAVEs for the gene TP53, two MAVEs for LDLR and two MAVEs for PTEN. Statistical methods (principal component analysis), unsupervised learning (k-means clustering), and supervised learning (Naïve Bayes and random forest classifiers) were used to integrate multiple MAVE datasets. The utility of MAVE integration methods were assessed using standard metrics (sensitivity, specificity, etc) as well as evidence strength in a putative variant classification framework.ResultsHere, we provide guidance for combining such multiplexed functional data, incorporating a stepwise process from data curation and collection to model generation and validation. We also present a web applet that allows users to test various methods for combining score sets from multiple assays, calculate integrated functional scores for all variants, and assess whether combining data enables the application of stronger evidence for pathogenicity or benignity. In general, supervised learning methods such as random forest led to improved variant classification as compared to any individual MAVE dataset.ConclusionsBy following the steps outlined herein with appropriate guardrails, researchers can maximize the value of MAVEs, strengthen the functional evidence for clinical variant classification, and potentially uncover novel mechanisms of pathogenicity for clinically relevant genes.

More information Original publication

DOI

10.1186/s13073-026-01715-w

Type

Journal article

Publication Date

2026-07-01T00:00:00+00:00

Addresses

K, e, n, , a, n, d, , R, u, t, h, , D, a, v, e, e, , D, e, p, a, r, t, m, e, n, t, , o, f, , N, e, u, r, o, l, o, g, y, ,, , N, o, r, t, h, w, e, s, t, e, r, n, , F, e, i, n, b, e, r, g, , S, c, h, o, o, l, , o, f, , M, e, d, i, c, i, n, e, ,, , C, h, i, c, a, g, o, ,, , I, L, ,, , U, S, A, .