Constraint improves prioritization when treated as a calibrated population prior with uncertainty. It degrades interpretability when used as a hard gate or as a substitute for inheritance, phenotype fit, transcript relevance, and cardiac context.
Executive interpretation
Population constraint measures whether classes of variation are depleted relative to expectation. It is powerful evidence about selection, but it is not a diagnosis and it does not identify the affected tissue. In pediatric cardiomyopathy, a constraint-heavy ranking can appear sophisticated while largely reproducing a list of genes that are broadly dosage sensitive.
Our benchmark shows that constraint is most useful as a prior whose influence is conditional on the proposed mechanism. Loss-of-function intolerance is relevant to haploinsufficiency hypotheses; regional missense constraint may be relevant to altered-protein mechanisms; neither should automatically override a poor phenotype match, an irrelevant transcript, incompatible inheritance, or absent cardiac context.
Constraint metrics and their uncertainty
We retain observed and expected counts, the upper confidence bound of the observed-to-expected ratio for predicted loss-of-function variation, and missense depletion metrics rather than importing only a percentile. Genes with few expected variants have less informative estimates, and transcript-level differences can be substantial. A binary “constrained/not constrained” threshold discards this uncertainty.
Predicted loss-of-function annotations also require quality control. Terminal truncations, low-expression exons, rescue transcripts, mapping artifacts, and variants unlikely to trigger nonsense-mediated decay can inflate an apparent loss-of-function count. Constraint is therefore linked to transcript-aware consequence review and is down-weighted when the disease-relevant transcript or exon context is uncertain.
Benchmark architecture
Candidate genes are ranked under a grid of evidence weights spanning population constraint, variant consequence, inheritance compatibility, phenotype similarity, cardiac cell-state specificity, and prior gene–disease validity. The purpose is not to discover one numerically optimal recipe. It is to map where the top-ranked set changes and which evidence dimension causes the change.
We record rank correlation, top-k overlap, evidence-layer agreement, and candidate-specific rank trajectories. Known cardiomyopathy genes serve as anchors rather than as a complete gold standard, because historical gene discovery is itself biased. Less-characterized genes are evaluated for coherent evidence profiles rather than rewarded for resemblance to every known gene.
Failure modes tested
A constraint-dominant model can favor essential developmental genes with little cardiac specificity. A phenotype-dominant model can favor heavily annotated genes. An expression-dominant model can favor abundant structural transcripts. We test each failure mode by ablating one layer, permuting weights, and comparing with matched controls. Candidates that remain near the top for incompatible reasons are flagged for manual review rather than treated as robust.
We also separate syndromic and isolated cardiomyopathy representations. A gene with strong neurodevelopmental, metabolic, or multisystem phenotype concordance may be appropriate for a syndromic presentation even when a cardiac-only model ranks it lower. This prevents phenotype breadth from being mislabeled as noise.
Technical findings
Hard constraint thresholds produce unstable boundaries: candidates on either side of the cutoff can have similar underlying uncertainty, while highly constrained genes can remain elevated despite weak cardiac evidence. Continuous, uncertainty-aware constraint contributes useful separation without forcing a false categorical decision.
The most stable priorities are supported by at least two evidence families beyond constraint—for example, compatible variant mechanism plus phenotype fit, or phenotype fit plus cardiomyocyte-state localization. Candidates supported almost entirely by one population metric move substantially under reasonable weight changes and are not ready to anchor an experimental program.
Interpretive standard
A rank is a navigation device. It does not convert a variant of uncertain significance into a causal allele, and it does not replace segregation, de novo assessment, allelic phase, phenotype review, or disease-specific variant criteria. The benchmark is designed to make those missing links visible.
For every priority, we therefore report the evidence profile, the strongest competing explanation, the weight range over which the priority is stable, and the experiment or clinical-domain review that would most reduce uncertainty.
What the analysis establishes
Continuous priors outperform hard gates conceptually
Retaining the magnitude and uncertainty of constraint avoids artificial categorical boundaries.
Mechanism matching is mandatory
Loss-of-function and missense constraint answer different questions and must be aligned with the proposed allelic mechanism.
Two-family support is a useful minimum
Stable candidates require agreement between independent evidence families rather than repeated measurements of one prior.
Rank trajectories reveal fragility
How a candidate moves across reasonable weight choices is more informative than its position in one final list.
Conclusion and experimental handoff
We conclude that population constraint should be used as calibrated background evidence, never as an automatic proxy for disease relevance. In pediatric cardiomyopathy, the strongest research priorities are those for which allelic mechanism, inheritance, phenotype, and cardiac context remain coherent across plausible weighting schemes.
The next step for a stable but less-characterized candidate is a mechanism-matched assay selected from the evidence profile: dosage perturbation for a credible haploinsufficiency model, variant-specific editing for an altered-protein model, and cardiac-lineage testing when cell-state evidence is central. An experiment that tests the wrong molecular direction cannot validate the computational ranking.
Limitations
- Constraint estimates vary in power across genes and transcripts.
- Known disease genes are an incomplete and historically biased benchmark set.
- Public phenotype annotations may underrepresent age-dependent and incompletely penetrant features.
- Ranking stability does not establish variant causality or clinical actionability.
Glossary
- LOEUF
- The upper confidence bound of the observed-to-expected ratio for predicted loss-of-function variants; lower values indicate stronger depletion.
- Prior ablation
- Removal of one evidence layer to measure how much it determines a result.
- Rank trajectory
- The movement of a candidate across a defined grid of model assumptions or evidence weights.
- Mechanism matching
- Alignment of the evidence metric and assay with the proposed molecular consequence.
Reproducibility and evidentiary scope
This study is a REELD public-data analysis and methods interpretation. It does not report a newly recruited clinical cohort, classify an individual variant, or replace clinical review. A release-ready execution of the workflow includes accession-level provenance, source and ontology versions, code state, environment locks, predefined sensitivity analyses, and output checksums.
