A phenotype ontology is part of the model
Terms, ancestors, and information content shape the answer. They belong in the methods, not in a footnote.
Our position
A phenotype ontology is not a neutral thesaurus placed between observation and algorithm. It is a versioned graph that encodes decisions about categories, relationships, granularity, and knowledge. Once a phenotype list enters semantic similarity, gene prioritization, cohort matching, or disease classification, the ontology is part of the computational model.
That means ontology release, annotation corpus, term selection, propagation, negation, onset, frequency, and similarity function belong in the methods. A ranking without this information cannot be reconstructed and should not be presented as though it followed directly from the patient or disease description.
Observation and representation must remain distinct
The observation is what was assessed and found. The representation is the ontology term chosen to encode it. Those are related but not identical. Two observers can map the same clinical description to different term depths, and two ontology releases can place the same term in a different neighborhood.
A responsible record retains source language, selected term, mapping confidence, modifiers, and alternatives when ambiguity matters. Silent replacement of a specific observation with a broad ancestor may improve recall while erasing the distinction required for mechanism. Silent over-specification can create precision the source never supported.
Ancestors improve recall and create dependence
Ontology propagation allows a specific observation to support comparisons at broader conceptual levels. This is useful when two records use different granularity. It also creates correlated features: a single observed term may contribute itself and many ancestors. Treating every propagated node as independent evidence multiplies one observation.
Propagation should therefore be bounded and explicit. Information-content weighting can reduce the influence of common ancestors, but information content is itself corpus dependent. There is no universal specificity value detached from the annotations used to estimate it.
Missing, absent, uncertain, and unassessed are not synonyms
A documented absent phenotype can be informative when the feature is expected, age appropriate, and adequately assessed. A missing field carries no such meaning. An uncertain observation should not be forced into either category. Algorithms that collapse these states produce confident but biologically misleading differentials.
The same applies to onset and frequency. A feature expected only later in life cannot be used as strong negative evidence in a young individual. A rare feature can be highly discriminating when present but weak evidence against a disease when absent. The phenotype model must preserve this asymmetry.
Similarity scores hide direction
A symmetric similarity score can conceal whether a disease model explains the observed features or whether the two records merely share broad ancestors. Directional coverage—how much of the observed phenotype a candidate explains and how much of the candidate phenotype was assessed—can reveal this difference.
Composite scores also hide term influence. Two candidates can receive similar scores for completely different reasons. Review requires a term-by-candidate explanation showing the most informative matches, unmatched observations, documented contradictions, and sensitivity to term removal.
Release drift is scientific drift
Ontologies add terms, revise definitions, and change parent-child relationships. Annotation files add diseases, update frequencies, and alter evidence. A result can therefore change without a new observation or a code change. If the release is not locked, the model has changed silently.
We recommend periodic release-delta analysis for active projects. When rankings change, the responsible graph or annotation update should be identified. Knowledge-base evolution is not noise; it is new model input and should be treated with the same provenance discipline as a new genomic reference.
A minimum reporting standard
Publish the exact term set, source descriptions, mapping confidence, ontology release, annotation corpus, propagation depth, similarity metric, aggregation rule, treatment of missingness and negation, and leave-one-term-out analysis. If a ranking changes materially under a reasonable alternate representation, publish both views.
The correct endpoint is not the appearance of stability. It is knowledge of the stability envelope and the clinical observations most likely to reduce it.
Positions we apply in review
- Phenotype encoding is a modeling step and must be auditable.
- Exact observations and propagated ancestors should remain distinguishable.
- Missing information must not be treated as an absent phenotype.
- Information content is corpus and release dependent.
- Every phenotype ranking should include term-influence and release-sensitivity analysis.
Our conclusion
Our editorial conclusion is that phenotype ontology choices are substantive scientific assumptions. A ranking without its representation model is not portable evidence. REELD will publish the phenotype graph, not merely the final list, and will use sensitivity to identify where better observation is more valuable than more computation.