Cosine Similarity Conflates Clinically Distinct Cancer Variants: A Case for Typed-Graph Retrieval in Precision Oncology Decision Support

Cosine Similarity Conflates Clinically Distinct Cancer Variants: A Case for Typed-Graph Retrieval in Precision Oncology Decision Support

Abstract

Abstract
Cancer variant interpretation increasingly relies on retrieval from biomedical knowledge bases, with cosine similarity over neural text embeddings now the dominant retrieval substrate. Whether these embeddings preserve the entity-level distinctions that variant interpretation requires that BRAF V600E and V600K are distinct alleles, that EGFR L858R is a sensitizing mutation while T790M is a resistance mutation has not been systematically measured. We hypothesize that cosine-similarity retrieval over biomedical embeddings conflates clinically distinct cancer variants at high rates, while a typed-graph approach in which each variant is a discrete node preserves variant identity by construction. We constructed a benchmark of 9 cancer variant pairs known to have differential FDA-approved therapy indications or distinct molecular biology, curated from theCIViC clinical evidence database and primary clinical literature. Pairs included BRAF V600E vs V600K (melanoma), EGFR L858R vs T790M (NSCLC, the canonical sensitivity-vs-resistance pair), EGFR exon 19 deletion vs L858R, KRAS G12C vs G12D (only G12C has FDA-approved targeted therapy), KRAS G12C vs G12V, ERBB2 amplification vs activating point mutation, two PIK3CA hotspot pairs, and NTRK1 fusion vs point mutation. For each pair we computed cosine similarity across three open-source embedding models (PubMedBERT, MedCPT, BGE-large-en-v1.5) over three text formats (short, medium, long). A 6-pair positive control verified embeddings recognize equivalent variant-form pairs (e.g., "EGFR L858R" vs "EGFRp.L858R") as similar. Across the medium format (gene + variant + tumor type), **100% of clinically distinct variant pairs had cosine similarity [≥] 0.95 under both biomedicalencoders** (PubMedBERT, MedCPT). The general-purpose encoder (BGE-large-en-v1.5) conflated 33% in medium format but rose to 100% with added clinical context. At the more stringent {tau} = 0.99 (averaged across formats), PubMedBERT conflated 56%of pairs and MedCPT conflated 22%. The biomedically pre-trained encoders performed worse, not better, than the general-purpose encoder. A typed-graph retrieval baseline (each variant a discrete node) achieves zero conflation by construction. The conflation behavior is a property of the embedding architecture, not a coverage gap fixable by domain fine-tuning. We argue that bioinformatics applications that route on variant identity clinical decision support, variant-trial matching, pharmacogenomic recommendation should use typed-graph retrieval, not vector retrieval, as the routing substrate.
View original →