PubMed HealthSearch

Biomedical subjects

M Attimonelli

Publications and source records attributed to M Attimonelli.

7 recordsLinked to original sources

WORDUP: an efficient algorithm for discovering statistically significant patterns in DNA sequences.

We present here a fast and sensitive method designed to isolate short nucleotide sequences which have non-random statistical properties and may thus be biologically active. It is based on a first order Markov analysis and allows us to detect statistically significant sequence motifs from six to ten nucleotides long which are significantly shared (or avoided) in the sequences under investigation. This method has been tested on a set of 521 sequences extracted from the Eukaryotic Promoter Database (2). Our results demonstrate the accuracy and the efficiency of the method in that the sequence motifs which are known to act as eukaryotic promoters, such as the TATA-box and the CAAT-box, were clearly identified. In addition we have found other statistically significant motifs, the biological roles of which are yet to be clarified.

Algorithms

A simple method for global sequence comparison.

A simple method of sequence comparison, based on a correlation analysis of oligonucleotide frequency distributions, is here shown to be a reliable test of overall sequence similarity. The method does not involve sequence alignment procedures and permits the rapid screening of large amounts of sequence data. It identifies those sequences which deserve more careful analysis of sequence similarity at the level of resolution of the single nucleotide. It uses observed quantities only and does not involve the adoption of any theoretical model.

Algorithms

A statistical method for detecting regions with different evolutionary dynamics in multialigned sequences.

We describe a stochastic method for tracing the evolutionary pattern of multialigned sequences. This method allows us to detect gene regions with distinct evolutionary dynamics, e.g., regions that significantly deviate from the expected behavior. Accurate detection of hypervariable or hyperconstrained regions may provide useful information on the structure/function relationship of biosequences. This information can help localize functional constraints. In addition, the selection of distinct evolutionary dynamics may assist in the correct use of biosequences as reliable molecular clocks.

Animals

Reorganization and merging of the EMBL and GenBank keyword indexes in a tree structure for more efficient retrieval of nucleic acid sequences.

EMBL and GenBank keyword indexes have no hierarchical structure. In this paper we present a method for merging and reorganizing them in a tree structure whose primary roots are the keywords 'protein', 'DNA', 'RNA', and 'unclassified'. Synonymous keywords have been grouped together and erroneous keywords have been corrected. This taxonomic organization of keywords results in a more extensive and efficient retrieval which is further aided by "synonyms declaration". The tree has been produced using the computer programs GENPOINT and CREANET.

Abstracting and Indexing

Estimation of protein secondary structure from circular dichroism spectra: a critical examination of the CONTIN program.

The computer program CONTIN uses the Provencher and Glöckner procedure to calculate protein secondary structure from circular dichroism spectra. We have tested this program with peptides and proteins in which unfolding was either induced by denaturing treatment or was already present. Results indicate that the program does not clearly discriminate between the ordered and the unordered states of a protein.

Circular Dichroism

Structural elements highly preserved during the evolution of the D-loop-containing region in vertebrate mitochondrial DNA.

A detailed comparative study of the regions surrounding the origin of replication in vertebrate mitochondrial DNA (mtDNA) has revealed a number of interesting properties. This region, called the D-loop-containing region, can be divided into three domains. The left (L) and right (R) domains, which have a low G content and contain the 5' and the 3' D-loop ends, respectively, are highly variable for both base sequence and length. They, however, contain thermodynamically stable secondary structures which include the conserved sequence blocks called CSB-1 and TAS which are associated with the start and stop sites, respectively, for D-loop strand synthesis. We have found that a "mirror symmetry" exists between the CSB-1 and TAS elements, which suggests that they can act as specific recognition sites for regulatory, probably dimeric, proteins. Long, statistically significant repeats are found in the L and R domains. Between the L and R domains we observed in all mtDNA sequences a region with a higher G content which was apparently free of complex secondary structure. This central domain, well preserved in mammals, contains an open reading frame of variable length in the organisms considered. The identification of common features well preserved in evolution despite the high primary structural divergence of the D-loop-containing region of vertebrate mtDNA suggests that these properties are of prime importance for the mitochondrial processes that occur in this region and may be useful for singling out the sites on which one should operate experimentally in order to discover functionally important elements.

Amino Acid Sequence

Multisequence comparisons in protein coding genes. Search for functional constraints.

A very powerful method for detecting functional constraints operative in biological macromolecules is presented. This method entails performing a base permanence analysis of protein coding genes at each codon position simultaneously in different species. It calculates the degree of permanence of subregions of the gene by dividing it into segments, c codons long, counting how many sites remain unchanged in each segment among all species compared. By comparing the base permanence among several sequences with the expectations based on a stochastic evolutionary process, gene regions showing different degrees of conservation can be selected. This means that wherever the permanence deviates significantly from the expected value generated by the simulation, the corresponding regions are considered "constrained" or "hypervariable". The constrained regions are of two types: alpha and beta. The alpha regions result from constraints at the amino acid level, whereas the beta regions are those probably involved in "control" processing. The method has been applied to mitochondrial genes coding for subunit 6 of the ATPase and subunit 1 of the cytochrome oxidase in four mammalian species: human, rat, mouse, and cow. In the two mitochondrial genes a few regions that are highly conserved in all codon positions have been identified. Among these regions a sequence, common to both genes, that is complementary to a strongly conserved region of 12S rRNA has been found. This method can also be of great help in studying molecular evolution mechanisms.

Amino Acid Sequence