PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Stretch coding and block coding: two new strategies to represent questionably aligned DNA sequences.

Most coding strategies that address the problem of questionable alignment (elision, case sensitive, missing, polymorphic, gaps as presence/absence matrix) conflict with phylogenetic principles, particularly those relating to the concept of homology (shared similiarity explained by common ancestry). In some cases, the test of conjunction is failed. In other cases, characters that are coded ambiguously can lead to character-state optimization in the terminal taxa that conflicts with the original observations. Only data exclusion and contraction avoid these pitfalls. In highly dissimilar sequences additional character states can represent the available information. Two new methods that accomplish this-block and stretch coding-are introduced here. These two new coding strategies are not in conflict with the test of conjunction and do not contradict the original observations. They are comparable to coding practices with morphological data once the intrinsic differences due to character-state identity and topographical identity have been taken into account. It is suggested that, of the three recoding methods, the one is selected that preserves the maximum potential phylogenetic information as measured with the minimum number of steps required for the particular part of the data matrix.

Animals↗

Identification of consensus RNA secondary structures using suffix arrays.

BACKGROUND: The identification of a consensus RNA motif often consists in finding a conserved secondary structure with minimum free energy in an ensemble of aligned sequences. However, an alignment is often difficult to obtain without prior structural information. Thus the need for tools to automate this process. RESULTS: We present an algorithm called Seed to identify all the conserved RNA secondary structure motifs in a set of unaligned sequences. The search space is defined as the set of all the secondary structure motifs inducible from a seed sequence. A general-to-specific search allows finding all the motifs that are conserved. Suffix arrays are used to enumerate efficiently all the biological palindromes as well as for the matching of RNA secondary structure expressions. We assessed the ability of this approach to uncover known structures using four datasets. The enumeration of the motifs relies only on the secondary structure definition and conservation only, therefore allowing for the independent evaluation of scoring schemes. Twelve simple objective functions based on free energy were evaluated for their potential to discriminate native folds from the rest. CONCLUSION: Our evaluation shows that 1) support and exclusion constraints are sufficient to make an exhaustive search of the secondary structure space feasible. 2) The search space induced from a seed sequence contains known motifs. 3) Simple objective functions, consisting of a combination of the free energy of matching sequences, can generally identify motifs with high positive predictive value and sensitivity to known motifs.

Algorithms↗

Frequency of gaps observed in a structurally aligned protein pair database suggests a simple gap penalty function.

Gap penalty is an important component of the scoring scheme that is needed when searching for homologous proteins and for accurate alignment of protein sequences. Most homology search and sequence alignment algorithms employ a heuristic 'affine gap penalty' scheme q + r x n, in which q is the penalty for opening a gap, r the penalty for extending it and n the gap length. In order to devise a more rational scoring scheme, we examined the pattern of gaps that occur in a database of structurally aligned protein domain pairs. We find that the logarithm of the frequency of gaps varies linearly with the length of the gap, but with a break at a gap of length 3, and is well approximated by two linear regression lines with R2 values of 1.0 and 0.99. The bilinear behavior is retained when gaps are categorized by secondary structures of the two residues flanking the gap. Similar results were obtained when another, totally independent, structurally aligned protein pair database was used. These results suggest a modification of the affine gap penalty function.

Computational Biology↗

Grouping together highly diverged PD-(D/E)XK nucleases and identification of novel superfamily members using structure-guided alignment of sequence profiles.

The PD-(D/E)XK nuclease domains, initially identified in type II restriction enzymes, serve as models for studying aspects of protein-DNA interactions, mechanisms of phosphodiester hydrolysis, and provide indispensable tools for techniques in genetic engineering and molecular medicine. However, the low degree of amino acid conservation hampers the possibility of identification of PD-(D/E)XK superfamily members based solely on sequence comparisons. In several proteins implicated in DNA recombination and repair the restriction enzyme-like nuclease domain has been found only after the corresponding structures were determined experimentally. Here, we identified highly diverged variants of the PD-(D/E)XK domain in many proteins and open reading frames using iterative database searches and progressive, structure-guided alignment of sequence profiles. We predicted the possible cellular function for many hypothetical proteins based on their relative similarity to characterized nucleases or observed presence of additional domains. We also identified the nuclease domain in genuine recombinases and restriction enzymes, whose homology to other PD-(D/E)XK enzymes has not been demonstrated previously. The first superfamily-wide comparative analysis, not limited to nucleases of known structure, will guide cloning and characterization of novel enzymes and planning new experiments to better understand those already studied.

Amino Acid Sequence↗

Biological and molecular characterization of chicken anemia virus isolates from Slovenia.

The presence of chicken anemia virus (CAV) in Slovenia was confirmed by inoculation of 1-day-old chickens without antibodies against CAV and isolation of the virus on the Marek's disease chicken cell-MSB1 line and by polymerase chain reaction (PCR). Experimental inoculation of 1-day-old chickens resulted in lower hematocrit values, atrophy of the thymus, and atrophy of bone marrow. CAV was confirmed by PCR in the thymus, bone marrow, bursa of Fabricius, liver, spleen, ileocecal tonsils, duodenum, and proventriculus. The nucleotide sequence of the whole viral protein (VP)1 gene was determined by direct sequencing. Alignment of VP1 nucleotide sequences of Slovenian CAV isolates (CAV-69/00, CAV-469/01, and CAV-130/03) showed 99.4% to 99.9% homology. The VP1 nucleotide sequence alignment of Slovenian isolates with 19 other CAV strains demonstrated 94.4% to 99.4% homology. Slovenian isolates shared highest homology with the BD-3 isolate from Bangladesh. Alignment of the deduced VP1 amino acids showed that the Slovenian isolates shared 100% homology and had an amino acid sequence most similar to the BD-3 strain from Bangladesh (99.6%) and were 99.1% similar to the G6 strain from Japan and the L-028 strain from the United States. The Slovenian isolates were least similar (96.6%) to the 82-2 strain from Japan. A phylogeneric analysis on the basis of the alignment of the VP1 amino acids showed that CAV isolates used in the study formed three groups that indicated the possible existence of genetic groups among CAV strains. The CAV isolates were grouped together independent of their geographic origin and pathogenicity.

Amino Acid Sequence↗

Envelope protein sequences of dengue virus isolates TH-36 and TH-Sman, and identification of a type-specific genetic marker for dengue and tick-borne flaviviruses.

Complementary DNAs were synthesized from the envelope protein genes of two isolates of dengue virus (TH-36 and TH-Sman, previously suggested as possible dengue virus type 5 and dengue virus type 6 respectively) and amplified by the polymerase chain reaction using sense and antisense primers designed from conserved dengue virus gene sequences. The amplified cDNA clones were sequenced in both directions by double-stranded dideoxynucleotide sequencing. Alignment with published dengue virus sequences enabled us to assign these viruses accurately to classified serotypes, confirming that TH-36 and TH-Sman are strains of dengue virus type 2 and dengue virus type 1 respectively. Amino acid changes between the proteins encoded by these two isolates and strains of their respective serotypes may account for the significant antigenic differences observed during previous serological typing of these viruses. Moreover, sequence alignment of flavivirus envelope proteins revealed a hypervariable region, within which members of the dengue and tick-borne virus antigenic complexes show unique peptide sequences. This type-specific hypervariable domain may be useful as a genetic marker for typing dengue and tick-borne flaviviruses.

Amino Acid Sequence↗

Investigating phylogenetic relationships within the Apicomplexa using sequence data: the search for homology.

Whether stated explicitly or not, all molecular studies that seek to infer "homologies" among sequences or that attempt to determine the "relatedness" of taxa based on sequence comparisons are evolutionary studies. The generation of a reliable evolutionary hypothesis based on molecular sequences is dependent almost exclusively on the ability to align sequences such that bases or amino acids in the same position of two sequences are positionally homologous (i.e., they share the same position in the gene under study). The selection of suitable gene targets (commonly 18S small subunit rRNA gene sequences in the Apicomplexa) and appropriate ingroup and outgroup taxa will affect the ability to align sequences unambiguously. Mathematically derived alignments based on local sequence similarity have been shown to be less reliable than alignments based on conserved secondary structures coupled with an analysis of compensatory base changes. Use of staggered sequence alignments through hypervariable regions of 18S small subunit rRNA gene sequences in which subsets of taxa are aligned independently may permit inclusion of more of the primary sequences with an associated increase in information content in the data set. The use of these highly variable regions is critical for determining the branching order of closely related terminal taxa in the phylum Apicomplexa.

Animals↗

HomologyPlot: searching for homology to a family of proteins using a database of unique conserved patterns.

A new database of conserved amino acid residues is derived from the multiple sequence alignment of over 84 families of protein sequences that have been reported in the literature. This database contains sequences of conserved hydrophobic core patterns which are probably important for structure and function, since they are conserved for most sequences in that family. This database differs from other single-motif or signature databases reported previously, since it contains multiple patterns for each family. The new database is used to align a new sequence with the conserved regions of a family. This is analogous to reports in the literature where multiple sequence alignments are used to improve a sequence alignment. A program called HomologyPlot (suitable for IBM or compatible computers) uses this database to find homology of a new sequence to a family of protein sequences. There are several advantages to using multiple patterns. First, the program correctly identifies a new sequence as a member of a known family. Second, the search of the entire database is rapid and requires less than one minute. This is similar to performing a multiple sequence alignment of a new sequence to all of the known protein family sequences. Third, the alignment of a new sequence to family members is reliable and can reproduce the alignment of conserved regions already described in the literature. The speed and efficiency of this method is enhanced, since there is no need to score for insertions or deletions as is done in the more commonly used sequence alignment methods. In this method only the patterns are aligned. HomologyPlot also provides general information on each family, as well as a listing of patterns in a family.

Amino Acid Sequence↗

Evolutionarily conserved regions and hydrophobic contacts at the superfamily level: The case of the fold-type I, pyridoxal-5'-phosphate-dependent enzymes.

The wealth of biological information provided by structural and genomic projects opens new prospects of understanding life and evolution at the molecular level. In this work, it is shown how computational approaches can be exploited to pinpoint protein structural features that remain invariant upon long evolutionary periods in the fold-type I, PLP-dependent enzymes. A nonredundant set of 23 superposed crystallographic structures belonging to this superfamily was built. Members of this family typically display high-structural conservation despite low-sequence identity. For each structure, a multiple-sequence alignment of orthologous sequences was obtained, and the 23 alignments were merged using the structural information to obtain a comprehensive multiple alignment of 921 sequences of fold-type I enzymes. The structurally conserved regions (SCRs), the evolutionarily conserved residues, and the conserved hydrophobic contacts (CHCs) were extracted from this data set, using both sequence and structural information. The results of this study identified a structural pattern of hydrophobic contacts shared by all of the superfamily members of fold-type I enzymes and involved in native interactions. This profile highlights the presence of a nucleus for this fold, in which residues participating in the most conserved native interactions exhibit preferential evolutionary conservation, that correlates significantly (r = 0.70) with the extent of mean hydrophobic contact value of their apolar fraction.

Conserved Sequence↗

Statistical significance in biological sequence analysis.

One of the major goals of computational sequence analysis is to find sequence similarities, which could serve as evidence of structural and functional conservation, as well as of evolutionary relations among the sequences. Since the degree of similarity is usually assessed by the sequence alignment score, it is necessary to know if a score is high enough to indicate a biologically interesting alignment. A powerful approach to defining score cutoffs is based on the evaluation of the statistical significance of alignments. The statistical significance of an alignment score is frequently assessed by its P-value, which is the probability that this score or a higher one can occur simply by chance, given the probabilistic models for the sequences. In this review we discuss the general role of P-value estimation in sequence analysis, and give a description of theoretical methods and computational approaches to the estimation of statistical signifiance for important classes of sequence analysis problems. In particular, we concentrate on the P-value estimation techniques for single sequence studies (both score-based and score-free), global and local pairwise sequence alignments, multiple alignments, sequence-to-profile alignments and alignments built with hidden Markov models. We anticipate that the review will be useful both to researchers professionally working in bioinformatics as well as to biomedical scientists interested in using contemporary methods of DNA and protein sequence analysis.

Computational Biology↗

A novel randomized iterative strategy for aligning multiple protein sequences.

The rigorous alignment of multiple protein sequences becomes impractical even with a modest number of sequences, since computer memory and time requirements increase as the product of the lengths of the sequences. We have devised a strategy to approach such an optimal alignment, which modifies the intensive computer storage and time requirements of dynamic programming. Our algorithm randomly divides a group of unaligned sequences into two subgroups, between which an optimal alignment is then obtained by a Needleman-Wunsch style of algorithm. Our algorithm uses a matrix with dimensions corresponding to the lengths of the two aligned sequence subgroups. The pairwise alignment process is repeated using different random divisions of the whole group into two subgroups. Compared with the rigorous approach of solving the n-dimensional lattice by dynamic programming, our iterative algorithm results in alignments that match or are close to the optimal solution, on a limited set of test problems. We have implemented this algorithm in a computer program that runs on the IBM PC class of machines, together with a user-friendly environment for interactively selecting sequences or groups of sequences to be aligned either simultaneously or progressively.

Amino Acid Sequence↗

Structural clues in the sequences of the aquaporins.

The large number of sequences available for the aquaporin family represents a valuable source of information to incorporate into three-dimensional structure determination. Phylogenetic analysis was used to define type sequences to avoid extreme over-representation of some subfamilies, and as a measure of the quality of multiple sequence alignment. Inspection of the sequence alignment suggested eight conserved segments that define the core architecture of six transmembrane helices and two functional loops, B and E, projecting into the plane of the membrane. The sum of the core segments and the minimum lengths of the interlinking loops constitute the 208 residues necessary to satisfy the aquaporin architecture. Analysis of hydrophobic and conservation periodicity and of correlated mutations across the alignment indicated the likely assignment and orientation of the helices in the bilayer. This assignment is examined with respect to the structure of the erythrocyte aquaporin 1 determined by electron crystallography. The aquaporin 1 tetramer is described as three rings of helices, each ring with a different exposure to the lipid environment. The sequence analysis clearly suggests that two helices are exposed along their whole lengths, two helices are exposed only at their N termini, and two helices are not exposed to lipid. It is further proposed that, besides loops B and E, the highly conserved motifs on helices 1 and 4, ExxxTxxF/L, could line the water channel.

Amino Acid Sequence↗

Generation of an integrated transcription map of the BRCA2 region on chromosome 13q12-q13.

An integrated approach involving physical mapping, identification of transcribed sequences, and computational analysis of genomic sequence was used to generate a detailed transcription map of the 1. 0-Mb region containing the breast cancer susceptibility locus BRCA2 on chromosome 13q12-q13. This region is included in the genetic interval bounded by D13S1444 and D13S310. Retrieved sequences from exon amplification or hybrid selection procedures were grouped into physical intervals and subsequently grouped into transcription units by clone overlap. Overlap was established by direct hybridization, cDNA library screening, PCR cDNA linking (island hopping), and/or sequence alignment. Extensive genomic sequencing was performed in an effort to understand transcription unit organization. In total, approximately 500 kb of genomic sequence was completed. The transcription units were further characterized by hybridization to RNA from a series of human tissues. Evidence for seven genes, two putative pseudogenes, and nine additional putative transcription units was obtained. One of the transcription units was recently identified as BRCA2 but all others are novel genes of unknown function as only limited alignment to sequences in public databases was observed. One large gene with a transcript size of 10.7 kb showed significant similarity to a gene predicted by the Caenorhabditis elegans genome and the Saccharomyces cerevisiae genome sequencing efforts, while another contained a motif sequence similar to the human 2',3' cyclic nucleotide 3' phosphodiesterase gene. Several retrieved transcribed sequences were not aligned into transcription units because no corresponding cDNAs were obtained when screening libraries or because of a lack of definitive evidence for splicing signals or putative coding sequence based on computational analysis. However, the presence of additional genes in the BRCA2 interval is suggested as groups of putative exons and hybrid selected clones that were transcribed in consistent orientations could be localized to common physical intervals.

BRCA2 Protein↗

Detecting recombination with MCMC.

MOTIVATION: We present a statistical method for detecting recombination, whose objective is to accurately locate the recombinant breakpoints in DNA sequence alignments of small numbers of taxa (4 or 5). Our approach explicitly models the sequence of phylogenetic tree topologies along a multiple sequence alignment. Inference under this model is done in a Bayesian way, using Markov chain Monte Carlo (MCMC). The algorithm returns the site-dependent posterior probability of each tree topology, which is used for detecting recombinant regions and locating their breakpoints. RESULTS: The method was tested on a synthetic and three real DNA sequence alignments, where it was found to outperform the established detection methods PLATO, RECPARS, and TOPAL.

Algorithms↗

Testing homology with Contact Accepted mutatiOn (CAO): a contact-based Markov model of protein evolution.

Point Accepted Mutation (PAM) is the Markov model of amino acid replacements in proteins introduced by Dayhoff and her co-workers (Dayhoff et al., 1978). The PAM matrices and other matrices based on the PAM model have been widely accepted as the standard scoring system of protein sequence similarity in protein sequence alignment tools. Here, we present Contact Accepted mutatiOn (CAO), a Markov model of protein residue contact mutations. The CAO model simulates the interchanging of structurally defined side-chain contacts, and introduces additional structural information into protein sequence alignments. Therefore, similarities between structurally conserved sequences can be detected even without apparent sequence similarity. CAO has been benchmarked on the HOMSTRAD database and a subset of the CATH database, by comparing sequence alignments with reference alignments derived from structural superposition. CAO yields scores that reflect coherently the structural quality of sequence alignments, which has implications particularly for homology modelling and threading techniques.

Amino Acid Sequence↗

RALEE--RNA ALignment editor in Emacs.

UNLABELLED: Production of high quality multiple sequence alignments of structured RNAs relies on an iterative combination of manual editing and structure prediction. An essential feature of an RNA alignment editor is the facility to mark-up the alignment based on how it matches a given secondary structure prediction, but few available alignment editors offer such a feature. The RALEE (RNA ALignment Editor in Emacs) tool provides a simple environment for RNA multiple sequence alignment editing, including structure-specific colour schemes, utilizing helper applications for structure prediction and many more conventional editing functions. This is accomplished by extending the commonly used text editor, Emacs, which is available for Linux, most UNIX systems, Windows and Mac OS. AVAILABILITY: The ELISP source code for RALEE is freely available from http://www.sanger.ac.uk/Users/sgj/ralee/ along with documentation and examples. CONTACT: sgj@sanger.ac.uk

Algorithms↗

Eukaryotic translation elongation factor 1 gamma contains a glutathione transferase domain--study of a diverse, ancient protein superfamily using motif search and structural modeling.

Using computer methods for multiple alignment, sequence motif search, and tertiary structure modeling, we show that eukaryotic translation elongation factor 1 gamma (EF1 gamma) contains an N-terminal domain related to class theta glutathione S-transferases (GST). GST-like proteins related to class theta comprise a large group including, in addition to typical GSTs and EF1 gamma, stress-induced proteins from bacteria and plants, bacterial reductive dehalogenases and beta-etherases, and several uncharacterized proteins. These proteins share 2 conserved sequence motifs with GSTs of other classes (alpha, mu, and pi). Tertiary structure modeling showed that in spite of the relatively low sequence similarity, the GST-related domain of EF1 gamma is likely to form a fold very similar to that in the known structures of class alpha, mu, and pi GSTs. One of the conserved motifs is implicated in glutathione binding, whereas the other motif probably is involved in maintaining the proper conformation of the GST domain. We predict that the GST-like domain in EF1 gamma is enzymatically active and that to exhibit GST activity, EF1 gamma has to form homodimers. The GST activity may be involved in the regulation of the assembly of multisubunit complexes containing EF1 and aminoacyl-tRNA synthetases by shifting the balance between glutathione, disulfide glutathione, thiol groups of cysteines, and protein disulfide bonds. The GST domain is a widespread, conserved enzymatic module that may be covalently or noncovalently complexed with other proteins. Regulation of protein assembly and folding may be 1 of the functions of GST.

Amino Acid Sequence↗

Comparison of three classes of snake neurotoxins by homology modeling and computer simulation graphics.

We present a systematic structure comparison of three major classes of postsynaptic snake toxins, which include short and long chain alpha-type neurotoxins plus one angusticeps-type toxin of black mamba snake family. Two novel alpha-type neurotoxins isolated from Taiwan cobra (Naja naja atra) possessing distinct primary sequences and different postsynaptic neurotoxicities were taken as exemplars for short and long chain neurotoxins and compared with the major lethal short-chain neurotoxin in the same venom, i.e., cobrotoxin, based on the derived three-dimensional structure of this toxin in solution by NMR spectroscopy. A structure comparison among these two alpha-neurotoxins and angusticeps-type toxin (denoted as FS2) was carried out by the secondary-structure prediction together with computer homology-modeling based on multiple sequence alignment of their primary sequences and established NMR structures of cobrotoxin and FS2. It is of interest to find that upon pairwise superpositions of these modeled three-dimensional polypeptide chains, distinct differences in the overall peptide flexibility and interior microenvironment between these toxins can be detected along the three constituting polypeptide loops, which may reflect some intrinsic differences in the surface hydrophobicity of several hydrophobic peptide segments present on the surface loops of these toxin molecules as revealed by hydropathy profiles. Construction of a phylogenetic tree for these structurally related and functionally distinct toxins corroborates that all long and short toxins present in diverse snake families are evolutionarily related to each other, supposedly derived from an ancestral polypeptide by gene duplication and subsequent mutational substitutions leading to divergence of multiple three-loop toxin peptides.

Amino Acid Sequence↗