PubMed HealthSearch

Biomedical subjects

S Brunak

Publications and source records attributed to S Brunak.

12 recordsLinked to original sources

Analysis of eukaryotic promoter sequences reveals a systematically occurring CT-signal.

A general data study of eukaryotic promoter sequences from widely different species is presented. Mammalian promoters with known transcription initiation sites represented the largest subclass of the data, and for this group neural network algorithms were trained to predict the location of the initiation site in a test set. The prediction accuracy of this local method was higher than what could be expected from the known non-local structure of eukaryotic promoters. Subsequent analysis revealed, besides the consensus of the two known important subregions: the TATA-box TATAAA and the Cap-signal CA, a CT-signal positioned on the average seven nucleotides downstream of the transcription initiation site. The consensus of the CT-signal is CTNCNG. The details of this core promoter element were disclosed using multiple alignment and have earlier only been described in a few isolated examples.

Algorithms

Periodic sequence patterns in human exons.

We analyse the sequential structure of human exons and their flanking introns by hidden Markov models. Together, models of donor site regions, acceptor site regions and flanked internal exons, show that exons--besides the reading frame--hold a specific periodic pattern. The pattern, which has the consensus: non-T(A/T)G and a minimal periodicity of roughly 10 nucleotides, is not a consequence of the nucleotide statistics in the three codon positions, nor of the well known nucleosome positioning signal. We discuss the relation between the pattern and other known sequence elements responsible for the intrinsic bending or curvature of DNA.

Base Sequence

Neural network model of the genetic code is strongly correlated to the GES scale of amino acid transfer free energies.

A neural network trained to classify the 61 nucleotide triplets of the genetic code into 20 amino acid categories develops in its internal representation a pattern matching the relative cost of transferring amino acids with satisfied backbone hydrogen bonds from water to an environment of dielectric constant of roughly 2.0. Such environments are typically found in lipid membranes or in the interior of proteins. In learning the mapping between the codons and the categories, the network groups the amino acids according to the scale of transfer free energies developed by Engelman, Goldman and Steitz. Several other scales based on internal preference statistics also agree reasonably well with the network grouping. The network is able to relate the structure of the genetic code to quantifications of amino acid hydrophobicity-hydrophilicity more systematically than the numerous attempts made earlier. Due to its inherent non-linearity, the code is also shown to impose decisive constraints on algorithmic analysis of the protein coding potential of DNA.

Amino Acid Sequence

Protein structures from distance inequalities.

A computer method for folding protein backbones from distance inequalities is presented. It involves an algorithm that uses a novel approach for handling inequalities through the minimization of a continuous energy function. Tests of the folding algorithm have been carried out on a small protein, the 6PTI (bovine pancreatic trypsin inhibitor) with 56 amino acid residues, and on a medium-size protein, the 1TRM (rat trypsin) with 223 amino acid residues. Reconstructions based on a real-valued distance matrix led to folded three-dimensional structures with root-mean-square values of 0.04 A when compared with the crystallographic data. The obtained root-mean-square measures were of the order of 1 A, when distance inequalities were used for the reconstruction. Subsequently, the folding approach has been applied to distance inequalities predicted by neural network techniques that use the amino acid sequence as the only input. The inaccuracy in the inequalities predicted by the neural network was the reason for the root-mean-square value of 5.2 A. An error analysis of the method for reconstruction was performed and showed that no more than 3% inaccurate distance inequalities could be corrected for. Finally, a simple technique for root-mean-square comparisons of different protein structures is discussed.

Algorithms

G+C-rich tract in 5' end of human introns.

Analysis of an artificial neural network trained to classify DNA as coding or non-coding revealed compositional differences between sequence parts translated into protein and those that were not. The 5' end of human introns was found to have a base composition that was non-random to an extent matching the non-randomness in the 3' end that contains the polypyrimidine tract. The prevailing nucleotides in the initial 50 nucleotides of human introns are guanine and cytosine, the trinucleotide GGG was found to occur almost four times as frequently as it would in sequences with a uniform distribution of the nucleotides. The initial part of terminal exons and their associated terminal introns were shown to have a very special base composition deviating strongly from the normal picture in other exons and introns.

Base Composition

Multiple alignment using simulated annealing: branch point definition in human mRNA splicing.

A method for the simultaneous alignment of a very large number of sequences using simulated annealing is presented. The total running time of the algorithm does not depend explicitly on the number of sequences treated. The method has been used for the simultaneous alignment of 1462 human intron sequences upstream of the intron-exon boundary. The consensus sequence of the aligned set together with a calculation of the Shannon information clearly shows that several sequence motives are conserved: (i) a previously undetected guanosine rich region, (ii) the branch point and (iii) the polypyrimidine tract. The nucleotide frequencies at each position of the branch point consensus sequence qualitatively reproduce the frequencies of the experimentally determined branch points.

Algorithms

Prediction of human mRNA donor and acceptor sites from the DNA sequence.

Artificial neural networks have been applied to the prediction of splice site location in human pre-mRNA. A joint prediction scheme where prediction of transition regions between introns and exons regulates a cutoff level for splice site assignment was able to predict splice site locations with confidence levels far better than previously reported in the literature. The problem of predicting donor and acceptor sites in human genes is hampered by the presence of numerous amounts of false positives: here, the distribution of these false splice sites is examined and linked to a possible scenario for the splicing mechanism in vivo. When the presented method detects 95% of the true donor and acceptor sites, it makes less than 0.1% false donor site assignments and less than 0.4% false acceptor site assignments. For the large data set used in this study, this means that on average there are one and a half false donor sites per true donor site and six false acceptor sites per true acceptor site. With the joint assignment method, more than a fifth of the true donor sites and around one fourth of the true acceptor sites could be detected without accompaniment of any false positive predictions. Highly confident splice sites could not be isolated with a widely used weight matrix method or by separate splice site networks. A complementary relation between the confidence levels of the coding/non-coding and the separate splice site networks was observed, with many weak splice sites having sharp transitions in the coding/non-coding signal and many stronger splice sites having more ill-defined transitions between coding and non-coding.

Base Sequence

Neural network detects errors in the assignment of mRNA splice sites.

The use of databanks in genetic research assumes reliability of the information they contain. Currently, error-detection in the manually or electronically entered data contained in the nucleotide sequence databanks at EMBL, Heidelberg and GenBank at Los Alamos is limited. We have used a subset of sequences from these databanks to train neural networks to recognize pre-mRNA splicing signals in human genes. During the training on 33 human genes from the EMBL databank seven genes appeared to disturb the learning process. Subsequent investigation revealed discrepancies from the original published papers, for three genes. In four genes, we found wrongly assigned splicing frames of introns. We believe this to be a reflection of the fact that splicing frames cannot always be unambiguously assigned on the basis of experimental data. Thus incorrect assignment appear both due to mere typographical misprints as well as erroneous interpretation of experiments. Training on 241 human sequences from GenBank revealed nine new errors. We propose that such errors could be detected by computer algorithms designed to check the consistency of data prior to their incorporation in databanks.

Algorithms

Analysis of the secondary structure of the human immunodeficiency virus (HIV) proteins p17, gp120, and gp41 by computer modeling based on neural network methods.

A neural network computer program, trained to predict secondary structure of proteins by exposing it to matching sets of primary and secondary structures from a database, was used to analyze the human immunodeficiency virus (HIV) proteins p17, gp120, and gp41 from their amino acid sequences. The results are compared to those obtained by the Chou-Fasman analysis. Two alpha-helical sequences corresponding to the putative fusigenic domain and to the transmembrane domain of gp41 could be predicted, as well as a possible binding site between p17 and gp41. On the basis of the secondary structure predictions, a three-dimensional model of p17 was constructed. This model was found to represent a stable conformation by an analysis using an energy-minimization program. The model predicts that p17 is attached to the membrane only by the acylated N-terminus, in analogy with the N-terminus of the gag protein of other retroviruses and also with the src oncogene protein p60src. The intracellular C-terminal part of gp41 may act as a receptor by electrostatic interaction with p17.

Algorithms

Protein secondary structure and homology by neural networks. The alpha-helices in rhodopsin.

Neural networks provide a basis for semiempirical studies of pattern matching between the primary and secondary structures of proteins. Networks of the perceptron class have been trained to classify the amino-acid residues into two categories for each of three types of secondary feature: alpha-helix or not, beta-sheet or not, and random coil or not. The explicit prediction for the helices in rhodopsin is compared with both electron microscopy results and those of the Chou-Fasman method. A new measure of homology between proteins is provided by the network approach, which thereby leads to quantification of the differences between the primary structures of proteins.

Amino Acid Sequence