PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Thermodynamic prediction of conserved secondary structure: application to the RRE element of HIV, the tRNA-like element of CMV and the mRNA of prion protein.

An algorithm for prediction of conserved secondary structure of single-stranded RNA is presented. For each RNA of a set of homologous RNAs optimal and suboptimal secondary structures are calculated and stored in a base-pair probability matrix. A multiple sequence alignment is performed for the set of RNAs. The resulting gaps are introduced into the individual probability matrices. These homologous probability matrices are summed to give a consensus probability matrix emphasizing the conserved secondary structure elements of the RNA set. Thus the algorithm combines the advantages of thermodynamic structure prediction by energy minimization with the information obtained from phylogenetic alignment of sequences. The algorithm is applied to three examples. The REV-responsive element of HIV, the structure of which is well known from the literature, was chosen to test the algorithm. The second example is the 3' terminal segment of genomic single-stranded RNAs of cucumber mosaic viruses; a structure similar to that of the related brome mosaic virus was expected and was confirmed. The third example is the prion-protein mRNA from different organisms; the structure of this mRNA is not known. By application of the algorithm highly conserved hairpins were found in the prion-protein mRNA.

Algorithms↗

Molecular phylogeny of some European heteronemertean (Nemertea) species and the monophyletic status of Riseriellus, Lineus, and Micrura.

The 16S rRNA mitochondrial gene was used to reconstruct the relationships among 10 heteronemertean species (subclass Heteronemertea, phylum Nemertea); Lineus ruber and L. viridis are represented by more than one specimen to assess intraspecific variation in these enigmatic species, and the analysis includes in total 14 terminal taxa incorporating one palaeonemertean species (Tubulanus annulatus) for outgroup rooting. The aligned sequences were subjected to maximum parsimony, maximum-likelihood, and neighbor-joining analyses to estimate the phylogenetic relationship of the species. The results were concordant from all analyses and indicate that neither Lineus nor Micrura are monophyletic taxa, and that there is no support from a phylogenetic point of view to establish the monotypic genus Riseriellus.

Animals↗

Evolution of ABCA4 proteins in vertebrates.

The ABCA4 (ABCR) gene encodes a retinal-specific ATP-binding cassette transporter. Mutations in ABCA4 are responsible for several recessive macular dystrophies and susceptibility to age related macular degeneration (AMD). The protein appears to function as a flippase of all-trans-retinaldehyde and/or its derivatives across the membrane of outer segment disks and is a potentially important element in recycling visual cycle metabolites. However, the understanding of ABCA4's role in the visual cycle is limited due to the lack of a direct functional assay. An evolutionary analysis of ABCA4 may aid in the identification of conserved elements, the preservation of which implies functional importance. To date, only human, murine, and bovine ABCA4 genes are described. We have identified ABCA4 genes from African (Xenopus laevis) and Western (Silurana tropicalis) clawed frogs. A comparative analysis describing the evolutionary relationships between the frog ABCA4s, annotated T. rubripes ABCA4, and mammalian ABCA4 proteins was carried out. Several segments are conserved in both intradiscal loop (IL) domains, in addition to the transmembrane and ATP-binding domains. Nonconserved segments were found in the IL and cytoplasmic linker domains. Maximum likelihood analyses of the aligned sequences strongly suggest that ABCA4 was subject to purifying selection. Collectively, these data corroborate the current evolutionary model where two distinct ABCA half-transporter progenitors were combined to form a full ABCA4 progenitor in ancestral chordates. We speculate that evolutionary alterations may increase the retinoid metabolite recycling capacity of ABCA4 and may improve dark adaptation.

ATP-Binding Cassette Transporters↗

Photoactivated adenylyl cyclase (PAC) genes in the flagellate Euglena gracilis mutant strains.

The unicellular, green flagellate wild-type Euglena gracilis(strain Z) and its colorless phototaxis-mutant strains as well as the non-photosynthetic close relative, Astasia longa, possess several genes of the photoactivated adenylyl cyclase (PAC) family. The corresponding gene products were found to be responsible for step-up (but not step-down) photophobic responses as well as both positive and negative phototaxis. The proteins consist of two PACalpha(M(r) 105 kDa) and two PACbeta(90 kDa) subunits. While the proteins were first believed all to be located in the paraxonemal body (PAB), confocal microscopy revealed that Astasia longa as well as some of the mutant strains do not contain a PAB. Immunofluorescence using PAC antibodies showed that the PAC proteins are also located along the total length of the flagellum at least in some of the strains. In order to determine if the genes responsible for the PAC proteins in the PAB and flagella are identical, sequences of all PAC proteins were analyzed in the Euglena and Astasia strains studied for PAC protein location. Full sequence analysis using PCR and 3' and 5' RACE indicated a substantial divergence between strains with a homology between strains of between 45 and 100%. Sequence alignment and sequence tree construction for the main functional groups (BLUF domain, which binds FAD, and adenylyl cyclase) showed that the pacalpha and the pacbeta gene products form clusters each with some of the mutants being closely related while others show a substantial degree of genetic diversity. The conclusion of these results is that there is a family of very dissimilar PAC proteins located in the PAB and the flagellum where they serve different functions in phototaxis and step-up photophobic reactions.

Adenylyl Cyclases↗

The evolution of extracellular hemoglobins of annelids, vestimentiferans, and pogonophorans.

The evolution of extracellular hemoglobins of annelids, vestimentiferans, and pogonophorans was investigated by applying cladistic and distance-based approaches to reconstruct the phylogenetic relationships of this group of respiratory pigments. We performed this study using the aligned sequences of globin and linker chains that are the constituents of these complex molecules. Three novel globin and two novel linker chains of Sabella spallanzanii described in an accompanying paper (Pallavicini, A., Negrisolo, E., Barbato, R., Dewilde, S., Ghiretti-Magaldi, A., Moens, L., and Lanfranchi, G. (2001) J. Biol. Chem. 276, 26384--26390) were also included. Our results allowed us to test previous hypotheses on the evolutionary pathways of these proteins and to formulate a new most parsimonious model of molecular evolution. According to this novel model, the genes coding for the polypeptides forming these composite molecules were already present in the common ancestor of annelids, vestimentiferans, and pogonophorans.

Amino Acid Sequence↗

Statistical modeling, phylogenetic analysis and structure prediction of a protein splicing domain common to inteins and hedgehog proteins.

Inteins, introns spliced at the protein level, and the hedgehog family of proteins involved in eucaryotic development both undergo autocatalytic proteolysis. Here, a specific and sensitive hidden Markov model (HMM) of protein splicing domain shared by inteins and the hedgehog proteins has been trained and employed for further analysis. The HMM characterizes the common features of this domain including the position where a site-specific DNA endonuclease domain is inserted in the majority of the inteins. The HMM was used to identify several new putative inteins, such as that in the Methanococcus jannaschii klbA protein, and to generate a multiple sequence alignment of sequences possessing this domain. Phylogenetic analysis suggests that hedgehog proteins evolved from inteins. Secondary and tertiary structure predictions suggest that the domain has a structure similar to a beta-sandwich. Similarities between the serine protease cleavage mechanism and the protein splicing reaction mechanism are discussed. Examination of the locations of inteins indicates that they are not inserted randomly in an extein, but are often inserted at functionally important positions in the host proteins. A specific and sensitive HMM for a domain present in klbA proteins identified several additional bacterial and archaeal family members, and analysis of the site of insertion of the intein suggests residues that may be functionally important. This domain may play a role in formation of surface-associated protein complexes.

Algorithms↗

Optimizing substitution matrices by separating score distributions.

MOTIVATION: Homology search is one of the most fundamental tools in Bioinformatics. Typical alignment algorithms use substitution matrices and gap costs. Thus, the improvement of substitution matrices increases accuracy of homology searches. Generally, substitution matrices are derived from aligned sequences whose relationships are known, and gap costs are determined by trial and error. To discriminate relationships more clearly, we are encouraged to optimize the substitution matrices from statistical viewpoints using both positive and negative examples utilizing Bayesian decision theory. RESULTS: Using Cluster of Orthologous Group (COG) database, we optimized substitution matrices. The classification accuracy of the obtained matrix is better than that of conventional substitution matrices to COG database. It also achieves good performance in classifying with other databases.

Algorithms↗

Phylogenetic reconstruction of vertebrate Hox cluster duplications.

In vertebrates and the cephalochordate, amphioxus, the closest vertebrate relative, Hox genes are linked in a single cluster. Accompanying the emergence of higher vertebrates, the Hox gene cluster duplicated in either a single step or multiple steps, resulting in the four-cluster state present in teleosts and tetrapods. Mammalian Hox clusters (designated A, B, C, and D) extend over 100 kb and are located on four different chromosomes. Reconstructing the history of the duplications and its relation to vertebrate evolution has been problematic due to the lack of alignable sequence information. In this study, the problem was approached by conducting a statistical analysis of sequences from the fibrillar-type collagens (I, II, III, and IV), genes closely linked to each Hox cluster which likely share the same duplication history as the Hox genes. We find statistical support for the hypothesis that the cluster duplication occurred as multiple distinct events and that the four-cluster situation arose by a three-step sequential process.

Animals↗

Amino acid-amino acid contacts at the cooperativity interface of the bacteriophage lambda and P22 repressors.

The bacteriophage lambda repressor and its relatives bind cooperatively to adjacent as well as artificially separated operator sites. This cooperativity is mediated by a protein-protein interaction between the DNA-bound dimers. Here we use a genetic approach to identify two pairs of amino acids that interact at the dimer-dimer interface. One of these pairs is nonconserved in the aligned sequences of the lambda and P22 repressors; we show that a lambda repressor variant bearing the P22 residues at these two positions interacts specifically with the P22 repressor. The other pair consists of a conserved ion pair; we reverse the charges at these two positions and demonstrate that, whereas the individual substitutions abolish the interaction of the DNA-bound dimers, these changes in combination restore the interaction of both lambdacI and P22c2 dimers.

Amino Acid Sequence↗

Searching for potential drug targets in two-component and phosphorelay signal-transduction systems using three-dimensional cluster analysis.

Two-component and phosphorelay signal transduction systems are central components in the virulence and antimicrobial resistance responses of a number of bacterial and fungal pathogens; in some cases, these systems are essential for bacterial growth and viability. Herein, we analyze in detail the conserved surface residue clusters in the phosphotransferase domain of histidine kinases and the regulatory domain of response regulators by using complex structure-based three-dimensional cluster analysis. We also investigate the protein-protein interactions that these residue clusters participate in. The Spo0B-Spo0F complex structure was used as the reference structure, and the multiple aligned sequences of phosphotransferases and response regulators were paired correspondingly. The results show that a contiguous conserved residue cluster is formed around the active site, which crosses the interface of histidine kinases and response regulators. The conserved residue clusters of phosphotransferase and the regulatory domains are directly involved in the functional implementation of two-component signal transduction systems and are good targets for the development of novel antimicrobial agents.

Binding Sites↗

Ancestral maximum likelihood of evolutionary trees is hard.

Maximum likelihood (ML) (Neyman, 1971) is an increasingly popular optimality criterion for selecting evolutionary trees. Finding optimal ML trees appears to be a very hard computational task--in particular, algorithms and heuristics for ML take longer to run than algorithms and heuristics for maximum parsimony (MP). However, while MP has been known to be NP-complete for over 20 years, no such hardness result has been obtained so far for ML. In this work we make a first step in this direction by proving that ancestral maximum likelihood (AML) is NP-complete. The input to this problem is a set of aligned sequences of equal length and the goal is to find a tree and an assignment of ancestral sequences for all of that tree's internal vertices such that the likelihood of generating both the ancestral and contemporary sequences is maximized. Our NP-hardness proof follows that for MP given in (Day, Johnson and Sankoff, 1986) in that we use the same reduction from Vertex Cover; however, the proof of correctness for this reduction relative to AML is different and substantially more involved.

Algorithms↗

Amplicon: software for designing PCR primers on aligned DNA sequences.

SUMMARY: Amplicon is a program for designing PCR primers on aligned groups of DNA sequences. The most important application for Amplicon is the design of 'group-specific' PCR primer sets that amplify a DNA region from a given taxonomic group but do not amplify orthologous regions from other taxonomic groups. AVAILABILITY: Amplicon is freely available as a script that will run on any platform with Python 2.3 installed (http://www.python.org). It is also available as a Windows executable. Free downloads that do not require registration can be found at http://www.aad.gov.au/amplicon

Algorithms↗

Fast small molecule similarity searching with multiple alignment profiles of molecules represented in one-dimension.

Multiple sequence alignment has proven to be a powerful method for creating protein and DNA sequence alignment profiles. These profiles of protein families are useful tools for identifying conserved motifs, such as the catalytic triad of the serine protease family or the seven transmembrane helices of the G-protein coupled receptor family. Ultimately, the understanding of the critical motifs within a family is useful for identifying new members of the family. Due to the complexity of protein-ligand recognition, no universally accepted method exists for clustering small molecules into families with the same or similar biological activity. A combination of the concept of multiple sequence alignment and the 1-dimensional molecular representation described earlier offers a new method for profiling sets of small molecules with the same biological activity. These small molecule profiles can isolate key commonalities within the set of bioactive compounds much like a multiple sequence alignment can isolate critical motifs within a protein family. The small molecule profiles then make useful tools for searching small molecule databases for new compounds with the same biological activity. The technique is demonstrated here using the human ether-a-go-go potassium channel and the kinase SRC.

Algorithms↗

A polynomial-time algorithm for a class of protein threading problems.

This paper presents an algorithm for constructing an optimal alignment between a three-dimensional protein structure template and an amino acid sequence. A protein structure template is given as a sequence of amino acid residue positions in three-dimensional space, along with an array of physical properties attached to each position; these residue positions are sequentially grouped into a series of core secondary structures (central helices and beta sheets). In addition to match scores and gap penalties, as in a traditional sequence-sequence alignment problem, the quality of a structure-sequence alignment is also determined by interaction preferences among amino acids aligned with structure positions that are spatially close (we call these 'long-range interactions'). Although it is known that constructing such a structure-sequence alignment in the most general form is NP-hard, our algorithm runs in polynomial time when restricted to structures with a 'modest' number of long-range amino acid interactions. In the current work, long-range interactions are limited to interactions between amino acids from different core secondary structures. Dividing the series of core secondary structures into two subseries creates a cut set of long-range interactions. If we use N, M and C to represent the size of an amino acid sequence, the size of a structure template, and the maximum cut size of long-range interactions, respectively, the algorithm finds an optimal structure-sequence alignment in O(21C NM) time, a polynomial function of N and M when C = O(log(N + M)). When running on structure-sequence alignment problems without long-range intersections, i.e. C = 0, the algorithm achieves the same asymptotic computational complexity of the Smith-Waterman sequence-sequence alignment algorithm.

Algorithms↗

Bayesian adaptive alignment and inference.

Sequence alignment without the specification of gap penalties or a scoring matrix is attained by using Bayesian inference and a recursive algorithm. This procedure's recursive algorithm sums over all possible alignments on the forward step to obtain normalizing constants essential to Bayesian inferences, and samples from the exact posterior distribution on the backward step. Since both terminal and intervening unrelated subsequences will often be excluded from an alignment, the resulting alignments may be seen as extensions of local alignments. An alignment's significance is assessed using the Bayesian evidence. A shuffling simulation shows that Bayesian evidence against the null hypothesis tends to be a conservative measure of significance compared to classical p-values. An application to proteins from the GTPase superfamily shows that the posterior distribution of the number of gaps is often flat and that the posterior distribution of the evolutionary distance is often flat and sometimes bimodal. An alignment of 1GIA with 1ETU shows good correspondence with a structural alignment.

Algorithms↗

Parallelized multiple alignment.

UNLABELLED: Multiple sequence alignment is a frequently used technique for analyzing sequence relationships. Compilation of large alignments is computationally expensive, but processing time can be considerably reduced when the computational load is distributed over many processors. Parallel processing functionality in the form of single-instruction multiple-data (SIMD) technology was implemented into the multiple alignment program Praline by using 'message passing interface' (MPI) routines. Over the alignments tested here, the parallelized program performed up to ten times faster on 25 processors compared to the single processor version. AVAILABILITY: Example program code for parallelizing pairwise alignment loops is available from http://mathbio.nimr.mrc.ac.uk/~jkleinj/tools/mpicode. The 'message passing interface' package (MPICH) is available from http:/www.unix.mcs.anl.gov/mpi/mpich. CONTACT: jhering@nimr.mrc.ac.uk SUPPLEMENTARY INFORMATION: Praline is accessible at http://mathbio.nimr.mrc.ac.uk/praline.

Algorithms↗

Phylogenetic analysis of 5'-UTR and P1 protein of Indian common strain of potato virus Y reveals its possible introduction in India.

The 5' untranslated region (UTR) and P1 region of the Indian strain of potato virus Y ordinary strain (PVYO) was cloned and sequenced for the first time. Database searches and multiple sequence alignment showed the highest sequence similarity with the PVYO strains of European origin. Based on the phylogenetic analysis and multiple sequence alignment, the possible evolution of PVYN from PVYO is predicted. PVYO strains from China and India were perhaps introduced into these countries from a similar geographical location. All major PVY strains available in the database can be classified into two major subgroups of North American and European origin. The Chinese and Indian PVYO strains fall within the European union subgroup suggesting a long association since potato was introduced from Europe into these countries by two separate independent events. The possible function of P1 protein in plant virus replication is suggested due to in-silico prediction of nuclear localization signal (NLS) and other phosphorylation regulatory domains at the vicinity of the NLS.

5' Untranslated Regions↗

Analysis of missense variation in human BRCA1 in the context of interspecific sequence variation.

INTRODUCTION: Interpretation of results from mutation screening of tumour suppressor genes known to harbour high risk susceptibility mutations, such as APC, BRCA1, BRCA2, MLH1, MSH2, TP53, and PTEN, is becoming an increasingly important part of clinical practice. Interpretation of truncating mutations, gene rearrangements, and obvious splice junction mutations, is generally straightforward. However, classification of missense variants often presents a difficult problem. From a series of 20,000 full sequence tests of BRCA1 carried out at Myriad Genetic Laboratories, a total of 314 different missense changes and eight in-frame deletions were observed. Before this study, only 21 of these missense changes were classified as deleterious or suspected deleterious and 14 as neutral or of little clinical significance. METHODS: We have used a combination of a multiple sequence alignment of orthologous BRCA1 sequences and a measure of the chemical difference between the amino acids present at individual residues in the sequence alignment to classify missense variants and in-frame deletions detected during mutation screening of BRCA1. RESULTS: In the present analysis we were able to classify an additional 50 missense variants and two in-frame deletions as probably deleterious and 92 missense variants as probably neutral. Thus we have tentatively classified about 50% of the unclassified missense variants observed during clinical testing of BRCA1. DISCUSSION: An internal test of the analysis is consistent with our classification of the variants designated probably deleterious; however, we must stress that this classification is tentative and does not have sufficient independent confirmation to serve as a clinically applicable stand alone method.

Amino Acid Sequence↗