- Access by Xinjiang University
Density estimation for ordinal biological sequences and its applications
Phys. Rev. E 110, 044408 – Published 30 October, 2024
DOI: https://doi.org/10.1103/PhysRevE.110.044408
Abstract
Biological sequences do not come at random. Instead, they appear with particular frequencies that reflect properties of the associated system or phenomenon. Knowing how biological sequences are distributed in sequence space is thus a natural first step toward understanding the underlying mechanisms. Here we propose a method for inferring the probability distribution from which a sample of biological sequences were drawn for the case where the sequences are composed of elements that admit a natural ordering. Our method is based on Bayesian field theory, a physics-based machine learning approach, and can be regarded as a nonparametric extension of the traditional maximum entropy estimate. As an example, we use it to analyze the aneuploidy data pertaining to gliomas from The Cancer Genome Atlas project. In addition, we demonstrate two follow-up analyses that can be performed with the resulting probability distribution. One of them is to investigate the associations among the sequence sites. This provides a way to infer the governing biological grammar. The other is to study the global geometry of the probability landscape, which allows us to look at the problem from an evolutionary point of view. It can be seen that this methodology enables us to learn from a sample of sequences about how a biological system or phenomenon in the real world works.
Physics Subject Headings (PhySH)
Article Text
References (55)
- F. Sanger, S. Nicklen, and A. R. Coulson, DNA sequencing with chain-terminating inhibitors, Proc. Natl. Acad. Sci. USA 74, 5463 (1977).
- J. Shendure, S. Balasubramanian, G. M. Church, W. Gilbert, J. Rogers, J. A. Schloss, and R. H. Waterston, DNA sequencing at 40: Past, present and future, Nature (London) 550, 345 (2017).
- D. S. Johnson, A. Mortazavi, R. M. Myers, and B. Wold, Genome-wide mapping of in vivo protein-DNA interactions, Science 316, 1497 (2007).
- N. Cloonan et al., Stem cell transcriptome profiling via massive-scale mRNA sequencing, Nat. Methods 5, 613 (2008).
- R. Lister, R. C. O'Malley, J. Tonti-Filippini, B. D. Gregory, C. C. Berry, A. H. Millar, and J. R. Ecker, Highly integrated single-base resolution maps of the epigenome in Arabidopsis, Cell 133, 523 (2008).
- A. Mortazavi, B. A. Williams, K. McCue, L. Schaeffer, and B. Wold, Mapping and quantifying mammalian transcriptomes by RNA-seq, Nat. Methods 5, 621 (2008).
- U. Nagalakshmi, Z. Wang, K. Waern, C. Shou, D. Raha, M. Gerstein, and M. Snyder, The transcriptional landscape of the yeast genome defined by RNA sequencing, Science 320, 1344 (2008).
- B. T. Wilhelm, S. Marguerat, S. Watt, F. Schubert, V. Wood, I. Goodhead, C. J. Penkett, J. Rogers, and Jürg Bähler, Dynamic repertoire of a eukaryotic transcriptome surveyed at single-nucleotide resolution, Nature (London) 453, 1239 (2008).
- R. Durbin, S. R. Eddy, A. Krogh, and G. Mitchison, Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids (Cambridge University Press, Cambridge, England, 1998).
- E. Schneidman, M. J. Berry, II, R. Segev, and W. Bialek, Weak pairwise correlations imply strongly correlated network states in a neural population, Nature (London) 440, 1007 (2006).
- J. Humplik and G. Tkačik, Probabilistic models for neural populations that naturally capture global coupling and criticality, PLoS Comput. Biol. 13, e1005763 (2017).
- S. Cocco, C. Feinauer, M. Figliuzzi, R. Monasson, and M. Weigt, Inverse statistical physics of protein sequences: A key issues review, Rep. Prog. Phys. 81, 032601 (2018).
- A. J. Riesselman, J. B. Ingraham, and D. S. Marks, Deep generative models of genetic variation capture the effects of mutations, Nat. Methods 15, 816 (2018).
- W. C. Chen, J. Zhou, J. M. Sheltzer, J. B. Kinney, and D. M. McCandlish, Field-theoretic density estimation for biological sequence space with applications to splice site diversity and aneuploidy in cancer, Proc. Natl. Acad. Sci. USA 118, e2025782118 (2021).
- W. Bialek, C. G. Callan, and S. P. Strong, Field theories for learning probability distributions, Phys. Rev. Lett. 77, 4693 (1996).
- I. Nemenman and W. Bialek, Occam factors and model independent Bayesian learning of continuous distributions, Phys. Rev. E 65, 026137 (2002).
- T. A. Enßlin, M. Frommert, and F. S. Kitaura, Information field theory for cosmological perturbation reconstruction and nonlinear signal analysis, Phys. Rev. D 80, 105005 (2009).
- T. A. Enßlin, Information field theory, AIP Conf. Proc. 1553, 184 (2013).
- J. B. Kinney, Estimation of probability densities using scale-free field theories, Phys. Rev. E 90, 011301(R) (2014).
- J. B. Kinney, Unification of field theory and maximum entropy methods for learning probability densities, Phys. Rev. E 92, 032107 (2015).
- W. C. Chen, A. Tareen, and J. B. Kinney, Density estimation on small data sets, Phys. Rev. Lett. 121, 160605 (2018).
- R. Staden, Computer methods to locate signals in nucleic acid sequences, Nucleic Acids Res. 12, 505 (1984).
- G. D. Stormo, Modeling the specificity of protein-DNA interactions, Quant. Biol. 1, 115 (2013).
- A. S. Lapedes, B. Giraud, L. Liu, and G. D. Stormo, in Statistics in Molecular Biology and Genetics, edited by F. Seillier-Moiseiwitsch (Institute of Mathematical Statistics, 1999), pp. 236–256.
- G. Yeo and C. B. Burge, Maximum entropy modeling of short sequence motifs with applications to RNA splicing signals, J. Comput. Biol. 11, 377 (2004).
- W. Bialek and R. Ranganathan, Rediscovering the power of pairwise interactions, arXiv:0712.4397.
- M. Weigt, R. A. White, H. Szurmant, J. A. Hoch, and T. Hwa, Identification of direct residue contacts in protein-protein interaction by message passing, Proc. Natl. Acad. Sci. USA 106, 67 (2009).
- T. Mora, A. M. Walczak, W. Bialek, and C. G. Callan, Jr., Maximum entropy models for antibody diversity, Proc. Natl. Acad. Sci. USA 107, 5405 (2010).
- S. Balakrishnan, H. Kamisetty, J. G. Carbonell, S. I. Lee, and C. J. Langmead, Learning generative models for protein fold families, Proteins 79, 1061 (2011).
- F. Morcos, A. Pagnani, B. Lunt, A. Bertolino, D. S. Marks, C. Sander, R. Zecchina, J. N. Onuchic, T. Hwa, and M. Weigt, Direct-coupling analysis of residue coevolution captures native contacts across many protein families, Proc. Natl. Acad. Sci. USA 108, E1293 (2011).
- M. Ekeberg, C. Lövkvist, Y. Lan, M. Weigt, and E. Aurell, Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models, Phys. Rev. E 87, 012707 (2013).
- E. van Nimwegen, Inferring contacting residues within and between proteins: What do the probabilities mean? PLoS Comput. Biol. 12, e1004726 (2016).
- R. M. Levy, A. Haldane, and W. F. Flynn, Potts Hamiltonian models of protein co-variation, free energy landscape, and evolutionary fitness, Curr. Opin. Struct. Biol. 43, 55 (2017).
- A. M. Taylor et al., Genomic and functional approaches to understanding cancer aneuploidy, Cancer Cell 33, 676 (2018).
- B. Nogrady, How cancer genomics is transforming diagnosis and treatment, Nature (London) 579, S10 (2020).
- B. T. Whitfield and J. T. Huse, Classification of adult-type diffuse gliomas: Impact of the World Health Organization 2021 update, Brain Pathol. 32, e13062 (2022).
- C. Hutter and J. C. Zenklusen, The Cancer Genome Atlas: Creating lasting value beyond its data, Cell 173, 283 (2018).
- Y. M. M. Bishop, S. E. Fienberg, and P. W. Holland, Discrete Multivaruate Analysis: Theory and Practice (MIT Press, Cambridge, MA, 1975).
- T. M. Cover and J. A. Thomas, in Elements of Information Theory, 2nd ed. (John Wiley & Sons, Hoboken, NJ, 2006), pp. 409–425.
- A. Vasudevan, K. M. Schukken, E. L. Sausville, V. Girish, O. A. Adebambo, J. M. Sheltzer, Aneuploidy as a promoter and suppressor of malignant growth, Nat. Rev. Cancer 21, 89 (2021).
- The Cancer Genome Atlas Research Network, Comprehensive genomic characterization defines human glioblastoma genes and core pathways, Nature (London) 455, 1061 (2008).
- The Cancer Genome Atlas Research Network, The somatic genomic landscape of glioblastoma, Cell 155, 462 (2013).
- The Cancer Genome Atlas Research Network, Comprehensive, integrative genomic analysis of diffuse lower-grade gliomas, N. Engl. J. Med. 372, 2481 (2015).
- M. Krzywinski, J. Schein, I. Birol, J. Connors, R. Gascoyne, D. Horsman, S. J. Jones, and M. A. Marra, Circos: An information aesthetic for comparative genomics, Genome Res. 19, 1639 (2009).
- A. Tareen and J. B. Kinney, Logomaker: Beautiful sequence logos in Python, Bioinformatics 36, 2272 (2020).
- C. Geisenberger et al., Molecular profiling of long-term survivors identifies a subgroup of glioblastoma characterized by chromosome 19/20 co-gain, Acta Neuropathol. 130, 419 (2015).
- A. Cohen, M. Sato, K. Aldape, C. C. Mason, K. Alfaro-Munoz, L. Heathcock, S. T. South, L. M. Abegglen, J. D. Schiffman and H. Colman, DNA copy number analysis of Grade II–III and Grade IV gliomas reveals differences in molecular ontogeny including chromothripsis associated with IDH mutation status, Acta Neuropathol. Commun. 3, 34 (2015).
- D. M. McCandlish, Visualizing fitness landscapes, Evolution 65, 1544 (2011).
- S. H. Almal and H. Padh, Implications of gene copy-number variation in health and diseases, J. Hum. Genet. 57, 6 (2012).
- G. Romeo and V. A. McKusick, Phenotypic diversity, allelic series and modifier genes, Nat. Genet. 7, 451 (1994).
- J. Podani, Braun-Blanquet's legacy and data analysis in vegetation science, J. Veg. Sci. 17, 113 (2006).
- W. C. Chen, J. Zhou, and D. M. McCandlish, OrdinalSeqDEFT (2024), GitHub repository, https://github.com/wcchen-ccu/OrdinalSeqDEFT.
- M. Kimura, On the probability of fixation of mutant genes in a population, Genetics 47, 713 (1962).
- R. A. Fisher, The Genetical Theory of Natural Selection (Oxford University Press, Oxford, England, 1930).
- S. Wright, Evolution in Mendelian populations, Genetics 16, 97 (1931).