Abstract
Abstract
Motivation: The ongoing revolution in genome sequencing is delivering an unprecedented number of genome assemblies to global repositories, resulting in an overwhelming amount of data imported to UniProt in the form of proteomes. To manage this growth sustainably, there is a need for a systematic workflow to select the best proteomes. Results: We propose a novel pipeline for cellular organisms to select the best Reference Proteomes, i.e. those that best represent the protein space of a species. The pipeline uses a clustering algorithm based on MMseqs2 to select the minimum number of Reference Proteomes whilst maximising the representation of the protein space for each species. Additionally, we aligned our viral Reference Proteomes with the exemplar genome set defined by the International Committee on Taxonomy of Viruses. Because this method ensures that all species are represented with at least one Reference Proteome, the UniProt Knowledgebase increased the number of Reference Proteomes of 36% and covering 34% more species in the Tree of Life. The UniProt Knowledgebase will mainly retain proteins from Reference Proteomes and therefore this method reduces the overall number of proteins by 43%, leading to a more concise yet representative knowledgebase.
### Competing Interest Statement
The authors have declared no competing interest.
National Human Genome Research Institute, https://ror.org/00baak391
Office of the Director, https://ror.org/00fj8a872
National Institute of Allergy and Infectious Diseases
National Institute on Aging, https://ror.org/049v75w11
National Institute of General Medical Sciences
National Institute of Diabetes and Digestive and Kidney Diseases
National Eye Institute
National Cancer Institute
National Heart, Lung, and Blood Institute
National Institutes of Health, U24HG007822
European Molecular Biology Laboratory
State Secretariat for Education, Research and Innovation