KEGG OC (KEGG Ortholog Cluster) is a novel database of ortholog clusters (OCs) based on the whole genome comparison. The OCs were constructed by applying a novel clustering method to all possible protein coding genes in all complete genomes, based on their amino acid sequence similarities. KEGG OC has the following original features in terms of coverage, efficiency, and usability. First, it consists of all fully sequenced genomes of a wide range of organisms from three domains (eukaryotes, bacteria, and archaea). Second, it is computationally efficient to calculate OCs, which makes it possible to regularly update the contents. Third, it is compatible with the KEGG database, which provides an easy way to link the OCs with KEGG PATHWAY, BRITE functional hierarchies, KEGG MODULE, KEGG MEDICUS, and many more.

OC search

ID or terms query

ex. hsa:362, K04517, M00170, pyruvate dehydrogenase org=bsu


Amino acid sequence query

ex. hsa:5162

Statistics

Number of organisms:4,717
Number of eukaryotes:352
Number of bacteria:4,117
Number of archaea:248
Number of sequences:20,132,178
Number of clusters:1,833,756

Cluster typeSingletonMultipleTotal
domain common0808808
eukaryotes, bacteria03,9703,970
eukaryotes, archaea01,6901,690
bacteria, archaea04,8304,830
eukaryotes440,495231,829672,324
bacteria630,673430,4601,061,133
archaea49,70039,30189,001
Total1,120,868712,8881,833,756

ver. 2017-06-05