Abstract
Abstract
Allostery is increasingly understood as the propagation of dynamic information through a protein, yet the computational descriptors of that flow are computed one structure at a time and carry no transferable, sequence-level prior. Here we build an alphabet of allostery: a dictionary of local contact words, short sequence windows anchored by three to four-residue spatial cliques, each carrying a distribution of net Gaussian network model transfer entropy scores pooled over a non-redundant set of Protein Data Bank structures. The alphabet consists of 131,611,766 unique words drawn from 212,860,934 clique observations. Projecting any protein's sequence and structure onto this dictionary yields a per-residue allosteric track, net TE, source, sink, and switch channels, with no system-specific fitting. Validating against the Allosteric Database, annotated allosteric-site residues behave as transfer entropy sinks (information receivers): the sink channel discriminates sites from the rest of the protein with pooled ROC-AUC = 0.543 over 646,629 residues (permutation z = 19.1), an effect small in magnitude but overwhelmingly significant and robust to word-frequency leakage. The directional channels are mechanistically informative in a two-state experiment on nine canonical allosteric proteins: source residues predict the largest apo[->]holo conformational rewiring (meta-analytic Spearman {rho} = +0.106, positive in 7/9 proteins) while sink residues mark the most conformationally stable positions ({rho} = -0.105, 8/9). Sinks thus mark where allosteric signal is received and sources mark where it drives motion. Finally, we mine the most context-variable words into a compact, hydrophobic-enriched switch vocabulary that we propose as a design dictionary for engineering allosteric mechanisms.