PubMed Health⌕ Search

Biomedical subjects

Bonnie Berger

Publications and source records attributed to Bonnie Berger.

16 recordsLinked to original sources

Foundation model reveals the shared organization of transcription and topologically associating domains.

The three-dimensional organization of chromatin into topologically associating domains (TADs) may impact gene regulation by bringing distant genes into contact. However, studies of TADs' function and their influence on transcription have been constrained by ambiguities in TAD boundary definitions and challenges in directly measuring their regulatory effects. We overcome these limitations by developing species-level consensus TAD maps for human and mouse by using a bag-of-genes approach that exposes an emergent regulatory structure. To quantify TAD-mediated relationships, we use a foundation model trained on 33 million transcriptomes to define a contextual similarity metric that captures higher-order relationships missed by co-expression. We find that TADs are regions of elevated co-regulation, with our framework yielding testable hypotheses about chromatin organization across cellular states. This TAD-linked enhancement is strongest during early development and declines with aging, while cancer cells show distinct TAD usage that shifts with chemotherapy. Together, these findings suggest that chromatin organization acts through probabilistic rather than deterministic mechanisms.

Humans↗

Decoding the Functional Interactome of Non-Model Organisms with PHILHARMONIC.

Despite the widespread availability of genome sequencing pipelines, many genes remain part of the genome's "dark matter," where existing inference tools cannot even begin to guess the biological function of their proteins from sequence alone. This challenge is especially pronounced in organisms that are highly evolutionarily distant from well-studied models, where homology-based methods break down. Here, we describe PHILHARMONIC, a computational method that combines deep learning-based de novo protein interaction network inference with robust unsupervised spectral clustering and remote homology to illuminate functional organization in any non-model organism. From only a sequenced proteome, we show PHILHARMONIC predicts protein functions, functional communities, and higher-order network structure with high accuracy. We validate its performance using experimental gene expression and pathway data in D. melanogaster, and we demonstrate its broad utility by analyzing temperature sensing and stress response pathways in the reef-building coral P. damicornis and its algal symbiont C. goreaui. PHILHARMONIC provides a general-purpose engine for functional discovery and biological hypothesis generation in non-model organisms, enabling systems-level insights across the full diversity of life.

Journal Article↗

Predicting transmembrane beta-barrels and interstrand residue interactions from sequence.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria, and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins (omps) makes them an important protein class. At the present time, very few nonhomologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane proteins. A novel method using pairwise interstrand residue statistical potentials derived from globular (nonouter membrane) proteins is introduced to predict the supersecondary structure of transmembrane beta-barrel proteins. The algorithm transFold employs a generalized hidden Markov model (i.e., multitape S-attribute grammar) to describe potential beta-barrel supersecondary structures and then computes by dynamic programming the minimum free energy beta-barrel structure. Hence, the approach can be viewed as a "wrapping" component that may capture folding processes with an initiation stage followed by progressive interaction of the sequence with the already-formed motifs. This approach differs significantly from others, which use traditional machine learning to solve this problem, because it does not require a training phase on known TMB structures and is the first to explicitly capture and predict long-range interactions. TransFold outperforms previous programs for predicting TMBs on smaller (<or=200 residues) proteins and matches their performance for straightforward recognition of longer proteins. An exception is for multimeric porins where the algorithm does perform well when an important functional motif in loops is initially identified. We verify our simulations of the folding process by comparing them with experimental data on the functional folding of TMBs. A Web server running transFold is available and outputs contact predictions and locations for sequences predicted to form TMBs.

Amino Acid Sequence↗

transFold: a web server for predicting the structure and residue contacts of transmembrane beta-barrels.

Transmembrane beta-barrel (TMB) proteins are embedded in the outer membrane of Gram-negative bacteria, mitochondria and chloroplasts. The cellular location and functional diversity of beta-barrel outer membrane proteins makes them an important protein class. At the present time, very few non-homologous TMB structures have been determined by X-ray diffraction because of the experimental difficulty encountered in crystallizing transmembrane (TM) proteins. The transFold web server uses pairwise inter-strand residue statistical potentials derived from globular (non-outer-membrane) proteins to predict the supersecondary structure of TMB. Unlike all previous approaches, transFold does not use machine learning methods such as hidden Markov models or neural networks; instead, transFold employs multi-tape S-attribute grammars to describe all potential conformations, and then applies dynamic programming to determine the global minimum energy supersecondary structure. The transFold web server not only predicts secondary structure and TMB topology, but is the only method which additionally predicts the side-chain orientation of transmembrane beta-strand residues, inter-strand residue contacts and TM beta-strand inclination with respect to the membrane. The program transFold currently outperforms all other methods for accuracy of beta-barrel structure prediction. Available at http://bioinformatics.bc.edu/clotelab/transFold.

Amino Acids↗

Fold recognition and accurate sequence-structure alignment of sequences directing beta-sheet proteins.

The ability to predict structure from sequence is particularly important for toxins, virulence factors, allergens, cytokines, and other proteins of public health importance. Many such functions are represented in the parallel beta-helix and beta-trefoil families. A method using pairwise beta-strand interaction probabilities coupled with evolutionary information represented by sequence profiles is developed to tackle these problems for the beta-helix and beta-trefoil folds. The algorithm BetaWrapPro employs a "wrapping" component that may capture folding processes with an initiation stage followed by processive interaction of the sequence with the already-formed motifs. BetaWrapPro outperforms all previous motif recognition programs for these folds, recognizing the beta-helix with 100% sensitivity and 99.7% specificity and the beta-trefoil with 100% sensitivity and 92.5% specificity, in crossvalidation on a database of all nonredundant known positive and negative examples of these fold classes in the PDB. It additionally aligns 88% of residues for the beta-helices and 86% for the beta-trefoils accurately (within four residues of the exact position) to the structural template, which is then used with the side-chain packing program SCWRL to produce 3D structure predictions. One striking result has been the prediction of an unexpected parallel beta-helix structure for a pollen allergen, and its recent confirmation through solution of its structure. A Web server running BetaWrapPro is available and outputs putative PDB-style coordinates for sequences predicted to form the target folds.

Algorithms↗

Pertactin beta-helix folding mechanism suggests common themes for the secretion and folding of autotransporter proteins.

Many virulence factors secreted from pathogenic Gram-negative bacteria are autotransporter proteins. The final step of autotransporter secretion is C --> N-terminal threading of the passenger domain through the outer membrane (OM), mediated by a cotranslated C-terminal porin domain. The native structure is formed only after this final secretion step, which requires neither ATP nor a proton gradient. Sequence analysis reveals that, despite size, sequence, and functional diversity among autotransporter passenger domains, >97% are predicted to form parallel beta-helices, indicating this structural topology may be important for secretion. We report the folding behavior of pertactin, an autotransporter passenger domain from Bordetella pertussis. The pertactin beta-helix folds reversibly in isolation, but folding is much slower than expected based on size and native-state topology. Surprisingly, pertactin is not prone to aggregation during folding, even though folding is extremely slow. Interestingly, equilibrium denaturation results in the formation of a partially folded structure, a stable core comprising the C-terminal half of the protein. Examination of the pertactin crystal structure does not reveal any obvious reason for the enhanced stability of the C terminus. In vivo, slow folding would prevent premature folding of the passenger domain in the periplasm, before OM secretion. Moreover, the extra stability of the C-terminal rungs of the beta-helix might serve as a template for the formation of native protein during OM secretion; hence, vectorial folding of the beta-helix could contribute to the energy-independent translocation mechanism. Coupled with the sequence analysis, the results presented here suggest a general mechanism for autotransporter secretion.

Bacterial Outer Membrane Proteins↗

The association between mood states and physical activity in postmenopausal, obese, sedentary women.

Mood states influence evaluative judgments that can affect the decision to exercise or to continue to exercise. This study examined how mood associated with graded exercise testing (GXT) in sedentary, obese, postmenopausal women (N = 25) was associated with physical activity and predicted VO2max during and after a behavioral weight-loss program (BWLP). Measures of physical activity included planned exercise, calories from physical activity, leisure-time physical activity, and predicted VO2max. Mood before and after pre-BWLP GXT was assessed using the Profile of Mood States. Mood before and after the GXT was more strongly associated with planned exercise than other forms of physical activity, and this effect became stronger over time. Mood enhancement in response to exercise was not related to physical activity. Mood before and after exercise might yield important clinical information that can be used to promote physical activity in sedentary adults.

Behavior Therapy↗

Struct2net: integrating structure into protein-protein interaction prediction.

UNLABELLED: This paper presents a framework for predicting protein-protein interactions (PPI) that integrates structure-based information with other functional annotations, e.g. GO, co-expression and co-localization, etc., Given two protein sequences, the structure-based interaction prediction technique threads these two sequences to all the protein complexes in the PDB and then chooses the best potential match. Based on this match, structural information is incorporated into logistic regression to evaluate the probability of these two proteins interacting. This paper also describes a random forest classifier which can effectively combine the structure-based prediction results and other functional annotations together to predict protein interactions. Experimental results indicate that the predictive power of the structure-based method is better than many other information sources. Also, combining the structure-based method with other information sources allows us to achieve a better performance than when structure information is not used. We also tested our method on a set of approximately 1000 yeast genes and, interestingly, the predicted interaction network is a scale-free network. Our method predicted some potential interactions involving yeast homologs of human disease-related proteins. SUPPLEMENTARY INFORMATION: http://theory.csail.mit.edu/struct2net

Algorithms↗

Herpesviral protein networks and their interaction with the human proteome.

The comprehensive yeast two-hybrid analysis of intraviral protein interactions in two members of the herpesvirus family, Kaposi sarcoma-associated herpesvirus (KSHV) and varicella-zoster virus (VZV), revealed 123 and 173 interactions, respectively. Viral protein interaction networks resemble single, highly coupled modules, whereas cellular networks are organized in separate functional submodules. Predicted and experimentally verified interactions between KSHV and human proteins were used to connect the viral interactome into a prototypical human interactome and to simulate infection. The analysis of the combined system showed that the viral network adopts cellular network features and that protein networks of herpesviruses and possibly other intracellular pathogens have distinguishing topologies.

Cell Line↗

A tree-decomposition approach to protein structure prediction.

This paper proposes a tree decomposition of protein structures, which can be used to efficiently solve two key subproblems of protein structure prediction: protein threading for backbone prediction and protein side-chain prediction. To develop a unified tree-decomposition based approach to these two subproblems, we model them as a geometric neighborhood graph labeling problem. Theoretically, we can have a low-degree polynomial time algorithm to decompose a geometric neighborhood graph G = (V, E) into components with size O(|V|((2/3))log|V|). The computational complexity of the tree-decomposition based graph labeling algorithms is O(|V|Delta(tw+1)) where Delta is the average number of possible labels for each vertex and tw( = O(|V|((2/3))log|V|)) the tree width of G. Empirically, tw is very small and the tree-decomposition method can solve these two problems very efficiently. This paper also compares the computational efficiency of the tree-decomposition approach with the linear programming approach to these two problems and identifies the condition under which the tree-decomposition approach is more efficient than the linear programming approach. Experimental result indicates that the tree-decomposition approach is more efficient most of the time.

Algorithms↗

MSARI: multiple sequence alignments for statistical detection of RNA secondary structure.

We present a highly accurate method for identifying genes with conserved RNA secondary structure by searching multiple sequence alignments of a large set of candidate orthologs for correlated arrangements of reverse-complementary regions. This approach is growing increasingly feasible as the genomes of ever more organisms are sequenced. A program called msari implements this method and is significantly more accurate than existing methods in the context of automatically generated alignments, making it particularly applicable to high-throughput scans. In our tests, it discerned clustalw-generated multiple sequence alignments of signal recognition particle or RNaseP orthologs from controls with 89.1% sensitivity at 97.5% specificity and with 74.4% sensitivity with no false positives in 494 controls. We used msari to conduct a comprehensive scan for secondary structure in mRNAs of coding genes, and we found many genes with known mRNA secondary structure and compelling evidence for secondary structure in other genes. msari uses a method for coping with sequence redundancy that is likely to have applications in a large set of other comparison-based search methods. The program is available for download from http://theory.csail.mit.edu/MSARi.

Algorithms↗

Methods in comparative genomics: genome correspondence, gene identification and regulatory motif discovery.

In Kellis et al. (2003), we reported the genome sequences of S. paradoxus, S. mikatae, and S. bayanus and compared these three yeast species to their close relative, S. cerevisiae. Genomewide comparative analysis allowed the identification of functionally important sequences, both coding and noncoding. In this companion paper we describe the mathematical and algorithmic results underpinning the analysis of these genomes. (1) We present methods for the automatic determination of genome correspondence. The algorithms enabled the automatic identification of orthologs for more than 90% of genes and intergenic regions across the four species despite the large number of duplicated genes in the yeast genome. The remaining ambiguities in the gene correspondence revealed recent gene family expansions in regions of rapid genomic change. (2) We present methods for the identification of protein-coding genes based on their patterns of nucleotide conservation across related species. We observed the pressure to conserve the reading frame of functional proteins and developed a test for gene identification with high sensitivity and specificity. We used this test to revisit the genome of S. cerevisiae, reducing the overall gene count by 500 genes (10% of previously annotated genes) and refining the gene structure of hundreds of genes. (3) We present novel methods for the systematic de novo identification of regulatory motifs. The methods do not rely on previous knowledge of gene function and in that way differ from the current literature on computational motif discovery. Based on genomewide conservation patterns of known motifs, we developed three conservation criteria that we used to discover novel motifs. We used an enumeration approach to select strongly conserved motif cores, which we extended and collapsed into a small number of candidate regulatory motifs. These include most previously known regulatory motifs as well as several noteworthy novel motifs. The majority of discovered motifs are enriched in functionally related genes, allowing us to infer a candidate function for novel motifs. Our results demonstrate the power of comparative genomics to further our understanding of any species. Our methods are validated by the extensive experimental knowledge in yeast and will be invaluable in the study of complex genomes like that of the human.

Algorithms↗

TRILOGY: Discovery of sequence-structure patterns across diverse proteins.

We describe a new computer program, trilogy, for the automated discovery of sequence-structure patterns in proteins. trilogy implements a pattern discovery algorithm that begins with an exhaustive analysis of flexible three-residue patterns; a subset of these patterns are selected as seeds for an extension process in which longer patterns are identified. A key feature of the method is explicit treatment of both the sequence and structure components of these motifs: each trilogy pattern is a pair consisting of a sequence pattern and a structure pattern. Matches to both these component patterns are identified independently, allowing the program to assign a significance score to each sequence-structure pattern that assesses the degree of correlation between the corresponding sequence and structure motifs. trilogy identifies several thousand high-scoring patterns that occur across protein families. These include both previously identified and potentially novel motifs. We expect that these sequence-structure patterns will be useful in predicting protein structure from sequence, annotating newly determined protein structures, and identifying novel motifs of potential functional or structural significance. Further details on 7,768 significant patterns identified by trilogy can be found at http://theory.lcs.mit.edu/trilogy.

Algorithms↗

Predicting the beta-helix fold from protein sequence data.

A method is presented that uses beta-strand interactions to predict the parallel right-handed beta-helix super-secondary structural motif in protein sequences. A program called BetaWrap implements this method and is shown to score known beta-helices above non-beta-helices in the Protein Data Bank in cross-validation. It is demonstrated that BetaWrap learns each of the seven known SCOP beta-helix families, when trained primarily on beta-structures that are not beta-helices, together with structural features of known beta-helices from outside the family. BetaWrap also predicts many bacterial proteins of unknown structure to be beta-helices; in particular, these proteins serve as virulence factors, adhesins, and toxins in bacterial pathogenesis and include cell surface proteins from Chlamydia and the intestinal bacterium Helicobacter pylori. The computational method used here may generalize to other beta-structures for which strand topology and profiles of residue accessibility are well conserved.

Algorithms↗

ARACHNE: a whole-genome shotgun assembler.

We describe a new computer system, called ARACHNE, for assembling genome sequence using paired-end whole-genome shotgun reads. ARACHNE has several key features, including an efficient and sensitive procedure for finding read overlaps, a procedure for scoring overlaps that achieves high accuracy by correcting errors before assembly, read merger based on forward-reverse links, and detection of repeat contigs by forward-reverse link inconsistency. To test ARACHNE, we created simulated reads providing approximately 10-fold coverage of the genomes of H. influenzae, S. cerevisiae, and D. melanogaster, as well as human chromosomes 21 and 22. The assemblies of these simulated reads yielded nearly complete coverage of the respective genomes, with a small number of contigs joined into a smaller number of supercontigs (or scaffolds). For example, analysis of the D. melanogaster genome yielded approximately 98% coverage with an N50 contig length of 324 kb and an N50 supercontig length of 5143 kb. The assembly accuracy was high, although not perfect: small errors occurred at a frequency of roughly 1 per 1 Mb (typically, deletion of approximately 1 kb in size), with a very small number of other misassemblies. The assembly was rapid: the Drosophila assembly required only 21 hours on a single 667 MHz processor and used 8.4 Gb of memory.

Algorithms↗

Wrap-and-Pack: a new paradigm for beta structural motif recognition with application to recognizing beta trefoils.

A method is presented that uses beta-strand interactions at both the sequence and the atomic level, to predict beta-structural motifs of protein sequences. A program called Wrap-and- Pack implements this method and is shown to recognize beta-trefoils, an important class of globular beta-structures, in the Protein Data Bank with 92% specificity and 92.3% sensitivity in cross-validation. It is demonstrated that Wrap-and-Pack learns each of the ten known SCOP beta-trefoil families, when trained primarily on beta-structures that are not beta-trefoils, together with three-dimensional structures of known beta-trefoils from outside the family. Wrap-and-Pack also predicts many proteins of unknown structure to be beta-trefoils. The computational method used here may generalize to other beta-structures for which strand topology and profiles of residue accessibility are well conserved.

Computational Biology↗