Primary analytical observation

Specific terms provide discrimination, broad ancestors protect recall, and neither view is sufficient alone. The most reliable ranking is accompanied by a term-influence map showing which observations, ontology relationships, and propagation choices determine candidate order.

Executive interpretation

Phenotype-driven prioritization is often presented as though the phenotype list were a neutral input. It is not. The choice between a broad term and a specific descendant changes information content, candidate overlap, and the influence of missing or uncertain observations. Ontology version and propagation rules are therefore part of the model.

We conclude that broad and specific representations should be reported together. Specific terms sharpen the differential when they are well supported; ancestor expansion protects against vocabulary mismatch and incomplete annotation. A ranking is trustworthy only when the analyst can identify which terms carry it and how it changes under plausible alternate representations.

Phenotype representation

Each observation is retained with term identifier, label, ontology release, onset, frequency, modifier, negation, and certainty when available. Exact terms are never replaced silently by ancestors. Instead, the analysis constructs parallel representations: exact-only, bounded ancestor expansion, full ancestor expansion, and information-content weighted expansion.

Absent phenotypes are modeled separately from unobserved phenotypes. A documented absence can reduce support for a disease model when the feature is expected and ascertainment is reliable. Missing information is not negative evidence. Conflating the two produces false precision.

Similarity models

We compare most-informative-common-ancestor similarity with normalized measures that account for the information content of both terms. Pairwise term similarities are aggregated under best-match and directional schemes because a symmetric score can hide whether a candidate disease explains the observed phenotype or merely shares a few broad ancestors.

Annotation frequency and ontology topology are release dependent. Information content is therefore computed from a named corpus, and both the ontology graph and annotation corpus are version locked. Recomputing one without the other changes the model even if the phenotype list is unchanged.

Stability analysis

We remove each phenotype term in turn, collapse it to successive ancestors, and perturb uncertain observations. Candidate ranks are compared with Spearman correlation, top-k overlap, reciprocal-rank change, and candidate-specific influence. These metrics expose both global stability and local swaps among the highest priorities.

We also calculate redundancy between terms. Two observations may appear independent while sharing most of their ancestors and annotations. Without redundancy control, repeated description of one organ system can overwhelm a rarer but more discriminating feature from another system.

Technical findings

Specific terms generally increase separation among closely related disease models, but they also create fragility when based on uncertain interpretation or sparse annotations. Broad terms recover semantically adjacent candidates and reduce vocabulary mismatch, but they can flatten meaningful distinctions. The optimal representation is therefore conditional on observation quality and the decision being supported.

The most concerning rankings are those dominated by one broad, highly connected term or by a cluster of redundant terms. Stable rankings distribute influence across multiple phenotype branches and retain a recognizable candidate core under bounded ancestor changes. Instability is not automatically a failure; it identifies which clinical observation or ontology relationship requires review.

Reporting standard

A phenotype-driven analysis should publish the exact term set, ontology release, annotation corpus, propagation rule, similarity function, aggregation function, treatment of negation, and term-influence analysis. A screenshot of a phenotype list is not a reconstructable model.

We recommend presenting a dual ranking: a specificity-forward view for discrimination and a bounded-ancestor view for recall. Agreement defines a stable core; disagreement defines the review queue.

What the analysis establishes

Specificity and recall are complementary

Exact terms sharpen discrimination while ancestor expansion reduces vocabulary mismatch.

Term influence should be visible

Candidate order can be dominated by one observation, and that dependence must be reported.

Missing is not absent

Unrecorded phenotypes should not be used as negative evidence.

Ontology release is a model version

Graph and annotation changes can alter results even when the observed terms do not change.

Conclusion and practical handoff

We conclude that there is no responsible single phenotype ranking without a representation audit. The appropriate output is a stable candidate core, a sensitivity envelope, and a term-influence map that shows why candidates move.

For downstream review, analysts should prioritize the observations whose clarification would most change the ranking. This converts ontology sensitivity into a practical phenotyping plan rather than treating it as computational noise.

Limitations

  • Disease annotations are incomplete and uneven across rare conditions.
  • Information content depends on the selected annotation corpus and release.
  • Clinical observation quality cannot be recovered computationally from an imprecise source record.
  • Semantic similarity ranks explanatory proximity; it does not establish a molecular mechanism.

Glossary

Resnik similarity
Similarity defined by the information content of the most informative common ancestor shared by two ontology terms.
Best-match average
An aggregation that pairs each term with its most similar counterpart before averaging.
Term ablation
Removal of one term at a time to quantify its effect on the result.
Sensitivity envelope
The range of plausible rankings produced by predefined reasonable modeling choices.

Reproducibility and evidentiary scope

This study is a REELD public-data analysis and methods interpretation. It does not report a newly recruited clinical cohort, classify an individual variant, or replace clinical review. A release-ready execution of the workflow includes accession-level provenance, source and ontology versions, code state, environment locks, predefined sensitivity analyses, and output checksums.