Editor’s summary
Only a small fraction of known proteomic space has been studied by experimental protein structure methods. The AlphaFold Protein Structure Database (AFDB) thus represents a truly massive expansion in the landscape of protein structural models. Such a large database represents a major challenge for categorization and analysis. Lau et al. developed an analytical pipeline based on deep learning for sectioning predicted protein structure models into discrete domains and classifying them. The resulting Encyclopedia of Domains expands and connects existing domain families, and there are also a large number of domains with folds that have not yet been observed in experimental structures. —Michael A. Funk
Structured Abstract
INTRODUCTION
The AlphaFold Protein Structure Database (AFDB) represents a major achievement in computational biology, containing high-quality predictions for more than 214 million protein structures. These structures are composed of domains, independently folding units that are critical to understanding protein function and evolution. The vast scale of the AFDB makes it difficult to accurately identify and classify these domains, thus limiting our ability to fully exploit its potential for biological discovery.
RATIONALE
To address this challenge, we developed The Encyclopedia of Domains (TED), a comprehensive resource for the systematic identification and classification of protein domains detectable within the AFDB. Traditional sequence-based methods, such as those used by Pfam and Gene3D, are unable to detect distant evolutionary relationships or entirely novel domains, but TED integrates multiple deep learning and advanced structural comparison techniques to uncover a much broader range of domain architectures, providing a more complete map of protein structure diversity.
RESULTS
TED identifies nearly 365 million domains across more than more than 1 million taxa, substantially expanding the scope of domain detection in the AFDB. This includes more than 100 million domains that were previously undetected by sequence-based methods, showcasing TED’s ability to leverage structure to increase domain coverage. We found that 77% of nonredundant domains were similar to existing CATH domains, validating the effectiveness of our approach.
Our study further highlights previously unobserved interactions between domains in the AFDB, uncovering more than 10,000 new structural interactions between superfamilies. We assessed the diversity of domain-packing geometries in interactions common to TED and CATH, and found that they were for the most part similar, with only around 5.4% of interactions showing a greater variety in TED. Overall, TED expands the set of known domain superfamily interactions to about three times the number seen in experimental structures, providing a means to greatly enrich our understanding of how multidomain proteins have evolved.
Moreover, TED has revealed thousands of putative new folds, substantially expanding the known fold space. This includes uncovering previously uncharacterized architectures, providing valuable new targets for structural and functional analysis. TED’s analysis of domain distribution across the Tree of Life has uncovered patterns of fold exclusivity and shared architectures among different lineages, suggesting important evolutionary trends. For example, certain folds were found to be highly conserved across all superkingdoms, indicating their essential roles in cellular functions, whereas others were lineage specific, pointing to unique evolutionary adaptations.
TED facilitates a deeper understanding of the structural diversity within the AFDB by identifying highly symmetric and repetitive domain architectures, which are often associated with specialized biological functions. Previously unknown domains, including unique beta-propeller and alpha-helical structures, represent new opportunities for studying protein evolution and function. Additionally, TED has demonstrated the ability to detect new variations in domain packing, further enriching the structural data available in the AFDB.
CONCLUSION
TED is a resource that not only fills critical gaps left by previous sequence-based methods, but also opens up new avenues for research in structural and evolutionary biology. It will be continually updated to keep in step with future releases of the AFDB. The comprehensive domain annotations provided by TED will facilitate a wide range of functional analyses, from drug discovery to the study of protein evolution. By revealing previously hidden relationships within the vast protein structure space, TED sets the stage for future breakthroughs in understanding the complex interplay among protein structure, function, and evolution.

TED workflow.
TED consists of 365 million domains identified from the AFDB by taking a consensus from three state-of-the-art domain segmentation methods. Nonredundant domains are assigned with CATH labels, and unassigned domains are subjected to a novel fold identification pipeline. Overall, we found several thousand novel folds, as well as novel high-symmetry fold clusters.
Abstract
The AlphaFold Protein Structure Database (AFDB) contains more than 214 million predicted protein structures composed of domains, which are independently folding units found in multiple structural and functional contexts. Identifying domains can enable many functional and evolutionary analyses but has remained challenging because of the sheer scale of the data. Using deep learning methods, we have detected and classified every domain in the AFDB, producing The Encyclopedia of Domains. We detected nearly 365 million domains, over 100 million more than can be found by sequence methods, covering more than 1 million taxa. Reassuringly, 77% of the nonredundant domains are similar to known superfamilies, greatly expanding representation of their domain space. We uncovered more than 10,000 new structural interactions between superfamilies and thousands of new folds across the fold space continuum.
Register and access this article for free
As a service to the community, this article is available for free.
Access the full article
View all access options to continue reading this article.
Supplementary Materials
The PDF file includes:
Other Supplementary Material for this manuscript includes the following:
MDAR Reproducibility Checklist
- Download
- 166.35 KB
References and Notes
1
M. Varadi, S. Anyango, M. Deshpande, S. Nair, C. Natassia, G. Yordanova, D. Yuan, O. Stroe, G. Wood, A. Laydon, A. Žídek, T. Green, K. Tunyasuvunakool, S. Petersen, J. Jumper, E. Clancy, R. Green, A. Vora, M. Lutfi, M. Figurnov, A. Cowie, N. Hobbs, P. Kohli, G. Kleywegt, E. Birney, D. Hassabis, S. Velankar, AlphaFold Protein Structure Database: Massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res. 50, D439–D444 (2022).
2
M. Varadi, D. Bertoni, P. Magana, U. Paramval, I. Pidruchna, M. Radhakrishnan, M. Tsenkov, S. Nair, M. Mirdita, J. Yeo, O. Kovalevskiy, K. Tunyasuvunakool, A. Laydon, A. Žídek, H. Tomlinson, D. Hariharan, J. Abrahamson, T. Green, J. Jumper, E. Birney, M. Steinegger, D. Hassabis, S. Velankar, AlphaFold Protein Structure Database in 2024: Providing structure coverage for over 214 million protein sequences. Nucleic Acids Res. 52, D368–D375 (2024).
3
N. Borkakoti, J. M. Thornton, AlphaFold2 protein structure prediction: Implications for drug discovery. Curr. Opin. Struct. Biol. 78, 102526 (2023).
4
J. Durairaj, A. M. Waterhouse, T. Mets, T. Brodiazhenko, M. Abdullah, G. Studer, G. Tauriello, M. Akdel, A. Andreeva, A. Bateman, T. Tenson, V. Hauryliuk, T. Schwede, J. Pereira, Uncovering new families and folds in the natural protein universe. Nature 622, 646–653 (2023).
5
I. Barrio-Hernandez, J. Yeo, J. Jänes, M. Mirdita, C. L. M. Gilchrist, T. Wein, M. Varadi, S. Velankar, P. Beltrao, M. Steinegger, Clustering predicted structures at the scale of the known protein universe. Nature 622, 637–645 (2023).
6
N. Bordin, I. Sillitoe, V. Nallapareddy, C. Rauer, S. D. Lam, V. P. Waman, N. Sen, M. Heinzinger, M. Littmann, S. Kim, S. Velankar, M. Steinegger, B. Rost, C. Orengo, AlphaFold2 reveals commonalities and novelties in protein structure space for 21 model organisms. Commun. Biol. 6, 160 (2023).
7
R. D. Schaeffer, J. Zhang, K. E. Medvedev, L. N. Kinch, Q. Cong, N. V. Grishin, ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM2. PLOS Comput. Biol. 20, e1011586 (2024).
8
I. Sillitoe, N. Bordin, N. Dawson, V. P. Waman, P. Ashford, H. M. Scholes, C. S. M. Pang, L. Woodridge, C. Rauer, N. Sen, M. Abbasian, S. Le Cornu, S. D. Lam, K. Berka, I. H. Varekova, R. Svobodova, J. Lees, C. A. Orengo, CATH: Increased structural coverage of functional space. Nucleic Acids Res. 49, D266–D273 (2021).
9
A. L. Cuff, I. Sillitoe, T. Lewis, O. C. Redfern, R. Garratt, J. Thornton, C. A. Orengo, The CATH classification revisited—Architectures reviewed and new ways to characterize structural divergence in superfamilies. Nucleic Acids Res. 37 (Database), D310–D314 (2009).
10
H. Cheng, R. D. Schaeffer, Y. Liao, L. N. Kinch, J. Pei, S. Shi, B.-H. Kim, N. V. Grishin, ECOD: An evolutionary classification of protein domains. PLOS Comput. Biol. 10, e1003926 (2014).
11
A. Bateman, E. Birney, L. Cerruti, R. Durbin, L. Etwiller, S. R. Eddy, S. Griffiths-Jones, K. L. Howe, M. Marshall, E. L. L. Sonnhammer, The Pfam protein families database. Nucleic Acids Res. 30, 276–280 (2002).
12
A. Bateman, L. Coin, R. Durbin, R. D. Finn, V. Hollich, S. Griffiths-Jones, A. Khanna, M. Marshall, S. Moxon, E. L. L. Sonnhammer, D. J. Studholme, C. Yeats, S. R. Eddy, The Pfam protein families database. Nucleic Acids Res. 32, D138–D141 (2004).
13
J. Lees, C. Yeats, J. Perkins, I. Sillitoe, R. Rentzsch, B. H. Dessailly, C. Orengo, Gene3D: A domain-based resource for comparative genomics, functional annotation and protein network analysis. Nucleic Acids Res. 40, D465–D471 (2012).
14
C. A. Orengo, A. D. Michie, S. Jones, D. T. Jones, M. B. Swindells, J. M. Thornton, CATH—A hierarchic classification of protein domain structures. Structure 5, 1093–1109 (1997).
15
T. E. Lewis, I. Sillitoe, N. Dawson, S. D. Lam, T. Clarke, D. Lee, C. Orengo, J. Lees, Gene3D: Extensive prediction of globular domains in proteins. Nucleic Acids Res. 46, D1282 (2018).
16
A. G. Murzin, S. E. Brenner, T. Hubbard, C. Chothia, SCOP: A structural classification of proteins database for the investigation of sequences and structures. J. Mol. Biol. 247, 536–540 (1995).
17
N. K. Fox, S. E. Brenner, J.-M. Chandonia, SCOPe: Structural Classification of Proteins—extended, integrating SCOP and ASTRAL data and classification of new structures. Nucleic Acids Res. 42, D304–D309 (2014).
18
C. Hadley, D. T. Jones, A systematic comparison of protein structure classifications: SCOP, CATH and FSSP. Structure 7, 1099–1112 (1999).
19
R. Day, D. A. C. Beck, R. S. Armen, V. Daggett, A consensus view of fold space: Combining SCOP, CATH, and the Dali Domain Dictionary. Protein Sci. 12, 2150–2160 (2003).
20
A. M. Lau, S. M. Kandathil, D. T. Jones, Merizo: A rapid and accurate protein domain segmentation method using invariant point attention. Nat. Commun. 14, 8445 (2023).
21
J. Wells, A. Hawkins-Hooker, N. Bordin, B. Paige, C. Orengo, Chainsaw: protein domain segmentation with fully convolutional neural networks, bioRxiv (2023) p. 2023.07.19.549732.
22
K. Zhu, H. Su, Z. Peng, J. Yang, A unified approach to protein domain parsing with inter-residue distance matrix. Bioinformatics 39, btad070 (2023).
23
M. van Kempen, S. S. Kim, C. Tumescheit, M. Mirdita, J. Lee, C. L. M. Gilchrist, J. Söding, M. Steinegger, Fast and accurate protein structure search with Foldseek. Nat. Biotechnol. 42, 243–246 (2024).
24
M. Steinegger, J. Söding, MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat. Biotechnol. 35, 1026–1028 (2017).
26
D. Du, Z. Wang, N. R. James, J. E. Voss, E. Klimont, T. Ohene-Agyei, H. Venter, W. Chiu, B. F. Luisi, Structure of the AcrAB-TolC multidrug efflux pump. Nature 509, 512–515 (2014).
27
C.-H. Tai, R. Paul, K. C. Dukka, J. D. Shilling, B. Lee, SymD webserver: A platform for detecting internally symmetric protein structures. Nucleic Acids Res. 42, W296–W300 (2014).
28
A. V. Kajava, A. C. Steven, “β‐rolls, β‐helices, and other β‐solenoid proteins” in Advances in Protein Chemistry (Academic Press, 2006), vol. 73, pp. 55–96.
29
S. Mesdaghi, R. M. Price, J. Madine, D. J. Rigden, Deep learning-based structure modelling illuminates structure and function in uncharted regions of β-solenoid fold space. J. Struct. Biol. 215, 108010 (2023).
30
B. Chakrabarty, N. Parekh, DbStRiPs: Database of structural repeats in proteins. Protein Sci. 31, 23–36 (2022).
31
M. Laitaoja, J. Valjakka, J. Jänis, Zinc coordination spheres in protein structures. Inorg. Chem. 52, 10983–10991 (2013).
32
T. Li, H. L. Bonkovsky, J.-T. Guo, Structural analysis of heme proteins: Implications for design and prediction. BMC Struct. Biol. 11, 13 (2011).
33
G. Apic, J. Gough, S. A. Teichmann, Domain combinations in archaeal, eubacterial and eukaryotic proteomes. J. Mol. Biol. 310, 311–325 (2001).
34
S. Batey, A. A. Nickson, J. Clarke, Studying the folding of multidomain proteins. HFSP J. 2, 365–377 (2008).
35
X. Zhou, J. Hu, C. Zhang, G. Zhang, Y. Zhang, Assembling multidomain protein structures through analogous global structural alignments. Proc. Natl. Acad. Sci. U.S.A. 116, 15930–15938 (2019).
36
Y. Hou, T. Xie, L. He, L. Tao, J. Huang, Topological links in predicted protein complex structures reveal limitations of AlphaFold. Commun. Biol. 6, 1098 (2023).
37
D. T. Jones, J. M. Thornton, The impact of AlphaFold2 one year on. Nat. Methods 19, 15–20 (2022).
38
H. K. Wayment-Steele, A. Ojoawo, R. Otten, J. M. Apitz, W. Pitsawong, M. Hömberger, S. Ovchinnikov, L. Colwell, D. Kern, Predicting multiple conformations via sequence clustering and AlphaFold2. Nature 625, 832–839 (2024).
39
D. del Alamo, D. Sala, H. S. Mchaourab, J. Meiler, Sampling alternative conformational states of transporters and receptors with AlphaFold2. eLife 11, e75751 (2022).
40
M. Mirdita, K. Schütze, Y. Moriwaki, L. Heo, S. Ovchinnikov, M. Steinegger, ColabFold: Making protein folding accessible to all. Nat. Methods 19, 679–682 (2022).
41
N. Bordin, I. Sillitoe, J. G. Lees, C. Orengo, Tracing evolution through protein structures: Nature captured in a few thousand folds. Front. Mol. Biosci. 8, 668184 (2021).
42
C. Chothia, One Thousand Families for the Molecular Biologist (Nature Publishing Group UK, 1992), 10.1038/357543a0.
43
Y. Zhang, J. Skolnick, TM-align: A protein structure alignment algorithm based on the TM-score. Nucleic Acids Res. 33, 2302–2309 (2005).
44
W. R. Taylor, Protein structural domain identification. Protein Eng. 12, 203–216 (1999).
45
S. J. Littler, S. J. Hubbard, Conservation of orientation and sequence in protein domain–domain interactions. J. Mol. Biol. 345, 1265–1279 (2005).
46
D. Holten, Hierarchical edge bundles: Visualization of adjacency relations in hierarchical data. IEEE Trans. Vis. Comput. Graph. 12, 741–748 (2006).
48
D. Ekman, S. Light, Å. K. Björklund, A. Elofsson, What properties characterize the hub proteins of the protein-protein interaction network of Saccharomyces cerevisiae? Genome Biol. 7, R45 (2006).
49
T. E. Lewis, I. Sillitoe, J. G. Lees, cath-resolve-hits: A new tool that resolves domain matches suspiciously quickly. Bioinformatics 35, 1766–1767 (2019).
50
D. Frishman, P. Argos, Knowledge-based protein secondary structure assignment. Proteins 23, 566–579 (1995).
51
N. Zhou, Y. Jiang, T. R. Bergquist, A. J. Lee, B. Z. Kacsoh, A. W. Crocker, K. A. Lewis, G. Georghiou, H. N. Nguyen, M. N. Hamid, L. Davis, T. Dogan, V. Atalay, A. S. Rifaioglu, A. Dalkıran, R. Cetin Atalay, C. Zhang, R. L. Hurto, P. L. Freddolino, Y. Zhang, P. Bhat, F. Supek, J. M. Fernández, B. Gemovic, V. R. Perovic, R. S. Davidović, N. Sumonja, N. Veljkovic, E. Asgari, M. R. K. Mofrad, G. Profiti, C. Savojardo, P. L. Martelli, R. Casadio, F. Boecker, H. Schoof, I. Kahanda, N. Thurlby, A. C. McHardy, A. Renaux, R. Saidi, J. Gough, A. A. Freitas, M. Antczak, F. Fabris, M. N. Wass, J. Hou, J. Cheng, Z. Wang, A. E. Romero, A. Paccanaro, H. Yang, T. Goldberg, C. Zhao, L. Holm, P. Törönen, A. J. Medlar, E. Zosa, I. Borukhov, I. Novikov, A. Wilkins, O. Lichtarge, P.-H. Chi, W.-C. Tseng, M. Linial, P. W. Rose, C. Dessimoz, V. Vidulin, S. Dzeroski, I. Sillitoe, S. Das, J. G. Lees, D. T. Jones, C. Wan, D. Cozzetto, R. Fa, M. Torres, A. Warwick Vesztrocy, J. M. Rodriguez, M. L. Tress, M. Frasca, M. Notaro, G. Grossi, A. Petrini, M. Re, G. Valentini, M. Mesiti, D. B. Roche, J. Reeb, D. W. Ritchie, S. Aridhi, S. Z. Alborzi, M.-D. Devignes, D. C. E. Koo, R. Bonneau, V. Gligorijević, M. Barot, H. Fang, S. Toppo, E. Lavezzo, M. Falda, M. Berselli, S. C. E. Tosatto, M. Carraro, D. Piovesan, H. Ur Rehman, Q. Mao, S. Zhang, S. Vucetic, G. S. Black, D. Jo, E. Suh, J. B. Dayton, D. J. Larsen, A. R. Omdahl, L. J. McGuffin, D. A. Brackenridge, P. C. Babbitt, J. M. Yunes, P. Fontana, F. Zhang, S. Zhu, R. You, Z. Zhang, S. Dai, S. Yao, W. Tian, R. Cao, C. Chandler, M. Amezola, D. Johnson, J.-M. Chang, W.-H. Liao, Y.-W. Liu, S. Pascarelli, Y. Frank, R. Hoehndorf, M. Kulmanov, I. Boudellioua, G. Politano, S. Di Carlo, A. Benso, K. Hakala, F. Ginter, F. Mehryary, S. Kaewphan, J. Björne, H. Moen, M. E. E. Tolvanen, T. Salakoski, D. Kihara, A. Jain, T. Šmuc, A. Altenhoff, A. Ben-Hur, B. Rost, S. E. Brenner, C. A. Orengo, C. J. Jeffery, G. Bosco, D. A. Hogan, M. J. Martin, C. O’Donovan, S. D. Mooney, C. S. Greene, P. Radivojac, I. Friedberg, The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol. 20, 244 (2019).
52
A. Bateman, M.-J. Martin, S. Orchard, M. Magrane, S. Ahmad, E. Alpi, E. H. Bowler-Barnett, R. Britto, H. Bye-A-Jee, A. Cukura, P. Denny, T. Dogan, T. G. Ebenezer, J. Fan, P. Garmiri, L. J. da Costa Gonzales, E. Hatton-Ellis, A. Hussein, A. Ignatchenko, G. Insana, R. Ishtiaq, V. Joshi, D. Jyothi, S. Kandasaamy, A. Lock, A. Luciani, M. Lugaric, J. Luo, Y. Lussi, A. MacDougall, F. Madeira, M. Mahmoudy, A. Mishra, K. Moulang, A. Nightingale, S. Pundir, G. Qi, S. Raj, P. Raposo, D. L. Rice, R. Saidi, R. Santos, E. Speretta, J. Stephenson, P. Totoo, E. Turner, N. Tyagi, P. Vasudev, K. Warner, X. Watkins, R. Zaru, H. Zellner, A. J. Bridge, L. Aimo, G. Argoud-Puy, A. H. Auchincloss, K. B. Axelsen, P. Bansal, D. Baratin, T. M. Batista Neto, M.-C. Blatter, J. T. Bolleman, E. Boutet, L. Breuza, B. C. Gil, C. Casals-Casas, K. C. Echioukh, E. Coudert, B. Cuche, E. de Castro, A. Estreicher, M. L. Famiglietti, M. Feuermann, E. Gasteiger, P. Gaudet, S. Gehant, V. Gerritsen, A. Gos, N. Gruaz, C. Hulo, N. Hyka-Nouspikel, F. Jungo, A. Kerhornou, P. Le Mercier, D. Lieberherr, P. Masson, A. Morgat, V. Muthukrishnan, S. Paesano, I. Pedruzzi, S. Pilbout, L. Pourcel, S. Poux, M. Pozzato, M. Pruess, N. Redaschi, C. Rivoire, C. J. A. Sigrist, K. Sonesson, S. Sundaram, C. H. Wu, C. N. Arighi, L. Arminski, C. Chen, Y. Chen, H. Huang, K. Laiho, P. McGarvey, D. A. Natale, K. Ross, C. R. Vinayaka, Q. Wang, Y. Wang, J. Zhang; UniProt Consortium, UniProt: The Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51, D523–D531 (2023).
53
Data for: A. Lau, N. Bordin, S. M. Kandathil, I. Sillitoe, V. P. Waman, J. Wells, C. A. Orengo, D. T. Jones, Exploring structural diversity across the protein universe with The Encyclopedia of Domains, version 4, Zenodo (2024); https://doi.org/10.5281/zenodo.13369203.
54
C. A. Orengo, W. R. Taylor, SSAP: Sequential structure alignment program for protein structure comparison. Methods Enzymol. 266, 617–635 (1996).
55
R. A. Laskowski, J. Jabłońska, L. Pravda, R. S. Vařeková, J. M. Thornton, PDBsum: Structural summaries of PDB entries. Protein Sci. 27, 129–134 (2018).
56
W. S. J. Valdar, Scoring residue conservation. Proteins 48, 227–241 (2002).