PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

CoSMoS: Conserved Sequence Motif Search in the proteome.

BACKGROUND: With the ever-increasing number of gene sequences in the public databases, generating and analyzing multiple sequence alignments becomes increasingly time consuming. Nevertheless it is a task performed on a regular basis by researchers in many labs. RESULTS: We have now created a database called CoSMoS to find the occurrences and at the same time evaluate the significance of sequence motifs and amino acids encoded in the whole genome of the model organism Escherichia coli K12. We provide a precomputed set of multiple sequence alignments for each individual E. coli protein with all of its homologues in the RefSeq database. The alignments themselves, information about the occurrence of sequence motifs together with information on the conservation of each of the more than 1.3 million amino acids encoded in the E. coli genome can be accessed via the web interface of CoSMoS. CONCLUSION: CoSMoS is a valuable tool to identify highly conserved sequence motifs, to find regions suitable for mutational studies in functional analyses and to predict important structural features in E. coli proteins.

Amino Acid Motifs↗

Parametric alignment of ordered trees.

MOTIVATION: Computing the similarity between two ordered trees has applications in RNA secondary structure comparison, genetics and chemical structure analysis. Alignment of tree is one of the proposed measures. Similar to pair-wise sequence comparison, there is often disagreement about how to weight matches, mismatches, indels and gaps when we compare two trees. For sequence comparison, the parametric sequence alignment tools have been developed. The users are allowed to see explicitly and completely the effect of parameter choices on the optimal sequence alignments. A similar tool for aligning two ordered trees is required in practice. RESULTS: We develop a parametric tool for aligning two ordered trees that allow users to see the effect of parameter choices on the optimal alignment of trees. Our contributions include: (1) develop a parametric tool for aligning two ordered trees; (2) design an efficient algorithm for aligning two ordered trees with gap penalties that runs in O(n(2)deg(2)) time, where n is the number of nodes in the trees and deg is the degree of the trees; and (3) reduce the space of the algorithm from O(n(2)deg(2)) to O(n log n. deg(2)). AVAILABILITY: The software is available at http://www.cs.cityu.edu.hk/~lwang/software/ParaTree

Algorithms↗

A phylogenetic survey of recombination frequency in plant RNA viruses.

The severe economic consequences of emerging plant viruses highlights the importance of studies of plant virus evolution. One question of particular relevance is the extent to which the genomes of plant viruses are shaped by recombination. To this end we conducted a phylogenetic survey of recombination frequency in a wide range of positive-sense RNA plant viruses, utilizing 975 capsid gene sequences and 157 complete genome sequences. In total, 12 of the 36 RNA virus species analyzed showed evidence for recombination, comprising 17% of the capsid gene sequence alignments and 44% of the genome sequence alignments. Given the conservative nature of our analysis, we propose that recombination is a relatively common process in some plant RNA viruses, most notably the potyviruses.

Computational Biology↗

PromoLign: a database for upstream region analysis and SNPs.

The study of transcriptional regulation at the genomic level has been hindered by the lack of functional annotation in the putative regulatory regions. Phylogenetic footprinting, in which cross-species sequence alignment among orthologous genes is applied to locate conserved sequence blocks, is an effective strategy to attack this problem. Single nucleotide polymorphisms (SNPs) in transcription factor (TF) binding sites contribute to the heterogeneity of TF binding sites and might disrupt or enhance their regulatory activity. The correlation of SNPs with the TF sites will not only help in functional evaluation of SNPs, but will also help in the study of transcription regulation by focusing attention on specific TF sites. PromoLign (http://polly.wustl.edu/promolign/main.html) is an online database application that presents SNPs and TF binding profiles in the context of human-mouse orthologous sequence alignment with a hyperlinked graphical interface. PromoLign could be applied to a variety of SNPs and transcription related studies, including association genetics, population genetics, and pharmacogenetics.

Animals↗

Differential signatures of bacterial and mammalian IMP dehydrogenase enzymes.

IMP dehydrogenase (IMPDH) is an essential enzyme of de novo guanine nucleotide synthesis. IMPDH inhibitors have clinical utility as antiviral, anticancer or immunosuppressive agents. The essential nature of this enzyme suggests its therapeutic applications may be extended to the development of antimicrobial agents. Bacterial IMPDH enzymes show biochemical and kinetic characteristics that are different than the mammalian IMPDH enzymes, suggesting IMPDH may be an attractive target for the development of antimicrobial agents. We suggest that the biochemical and kinetic differences between bacterial and mammalian enzymes are a consequence of the variance of specific, identifiable amino acid residues. Identification of these residues or combination of residues that impart this mammalian or bacterial enzyme signature is a prerequisite for the rational identification of agents that specifically target the bacterial enzyme. We used sequence alignments of IMPDH proteins to identify sequence signatures associated with bacterial or eukaryotic IMPDH enzymes. These selections were further refined to discern those likely to have a role in catalysis using information derived from the bacterial and mammalian IMPDH crystal structures and site-specific mutagenesis. Candidate bacterial sequence signatures identified by this process include regions involved in subunit interactions, the active site flap and the NAD binding region. Analysis of sequence alignments in these regions indicates a pattern of catalytic residues conserved in all enzymes and a secondary pattern of amino acid conservation associated with the major phylogenetic groups. Elucidation of the basis for this mammalian/bacterial IMPDH signature will provide insight into the catalytic mechanism of this enzyme and the foundation for the development of highly specific inhibitors.

Amino Acid Sequence↗

Molecular evolution and secondary structural conservation in the B-cell lymphoma leukemia 2 (bcl-2) family of proto-oncogene products.

The nature of the bcl-2 family of proto-oncogenes was analyzed by sequence alignment, secondary structure prediction, and phylogenetic techniques. Phylogenies were inferred from both the nucleic acid and amino acid sequences of the human, murine, rat, and chicken sequences for BCL-2 and BCL-X, human MCL1, murine A1, the nematode Caenorhabditis elegans and Caenorhabditis briggsiae ced-9 proteins, and the sequences BHRF1 from Epstein-Barr and LMW5-HL from African swine fever viruses. Both sequence alignment and secondary structure prediction techniques supported the conservation of both the overall secondary structure and the carboxy-terminal transmembrane domain in all members of the family. All the treeing methods employed (distance matrix, maximum likelihood, and parsimony) supported a tree in which the proapoptotic proteins BCL-2 and BCL-X represent the most recent additions to the group. All the trees also indicated that the viral proteins BHRF1 and LMW-HL arose from a common ancestor, an ancestor they shared in common with the pro-apoptotic control protein BAX, indicating that this function of BAX evolved only recently. The most ancient branches are represented by the nematode ced-9 protein and by the control genes MCL1 and A1, which in the treeing methods employed represent separate lineages within the most ancient grouping. These results demonstrate the evolution of a highly conserved family of developmental control genes from nematode to man--genes that encode proteins essential for normal development but which are highly conserved in terms of predicted structure and possible cellular localization. The evolutionary analysis also indicates that the family may be even larger than originally predicted and that other members are waiting to be discovered.

Amino Acid Sequence↗

A fast homology program for aligning biological sequences.

The algorithm of Gotoh computes in two passes of MN steps the alignment of a pair of sequences of lengths M and N, subject to a constraint on the form of the gap weighting function. This compares with the previous algorithm of Waterman et al. which runs in M2N steps. Gotoh also gave a method using two passes of (L+2)MN steps in the case where gap weights remain constant for gaps of length greater than L. Here we describe a procedure for computing the alignment (evolutionary distance and optimal path) in a single pass of MN steps for both cases.

Amino Acid Sequence↗

Aligning multiple genomic sequences with the threaded blockset aligner.

We define a "threaded blockset," which is a novel generalization of the classic notion of a multiple alignment. A new computer program called TBA (for "threaded blockset aligner") builds a threaded blockset under the assumption that all matching segments occur in the same order and orientation in the given sequences; inversions and duplications are not addressed. TBA is designed to be appropriate for aligning many, but by no means all, megabase-sized regions of multiple mammalian genomes. The output of TBA can be projected onto any genome chosen as a reference, thus guaranteeing that different projections present consistent predictions of which genomic positions are orthologous. This capability is illustrated using a new visualization tool to view TBA-generated alignments of vertebrate Hox clusters from both the mammalian and fish perspectives. Experimental evaluation of alignment quality, using a program that simulates evolutionary change in genomic sequences, indicates that TBA is more accurate than earlier programs. To perform the dynamic-programming alignment step, TBA runs a stand-alone program called MULTIZ, which can be used to align highly rearranged or incompletely sequenced genomes. We describe our use of MULTIZ to produce the whole-genome multiple alignments at the Santa Cruz Genome Browser.

Animals↗

Prediction of protein solvent accessibility using support vector machines.

A Support Vector Machine learning system has been trained to predict protein solvent accessibility from the primary structure. Different kernel functions and sliding window sizes have been explored to find how they affect the prediction performance. Using a cut-off threshold of 15% that splits the dataset evenly (an equal number of exposed and buried residues), this method was able to achieve a prediction accuracy of 70.1% for single sequence input and 73.9% for multiple alignment sequence input, respectively. The prediction of three and more states of solvent accessibility was also studied and compared with other methods. The prediction accuracies are better than, or comparable to, those obtained by other methods such as neural networks, Bayesian classification, multiple linear regression, and information theory. In addition, our results further suggest that this system may be combined with other prediction methods to achieve more reliable results, and that the Support Vector Machine method is a very useful tool for biological sequence analysis.

Bayes Theorem↗

Structural relationships of homologous proteins as a fundamental principle in homology modeling.

Protein structure prediction is based mainly on the modeling of proteins by homology to known structures; this knowledge-based approach is the most promising method to date. Although it is used in the whole area of protein research, no general rules concerning the quality and applicability of concepts and procedures used in homology modeling have been put forward yet. Therefore, the main goal of the present work is to provide tools for the assessment of accuracy of modeling at a given level of sequence homology. A large set of known structures from different conformational and functional classes, but various degrees of homology was selected. Pairwise structure superpositions were performed. Starting with the definition of the structurally conserved regions and determination of topologically correct sequence alignments, we correlated geometrical properties with sequence homology (defined by the 250 PAM Dayhoff Matrix) and identity. It is shown that both the topological differences of the protein backbones and the relative positions of corresponding side chains diverge with decreasing sequence identity. Below 50% identity, the deviation in regions that are structurally not conserved continually increases, thus implying that with decreasing sequence identity modeling has to take into account more and more structurally diverging loop regions that are difficult to predict.

Bacterial Proteins↗

Typing of Mycoplasma pneumoniae by nucleic acid sequence-based amplification, NASBA.

Nucleic acid sequence-based amplification, NASBA, is an isothermal amplification technique for nucleic acids and was used for typing a collection of 24 Mycoplasma pneumoniae strains. A set of primers was chosen from the 16S rRNA sequence alignment of Mycoplasma species. The nucleotide sequences of the (-)RNA amplicons were determined for M. pneumoniae strains M15/83 (type 1) and FH (type 2), and revealed a one-point difference at the 16S rRNA level between the two types. Based on this result, two type-specific probes were constructed. The probes were hybridized in solution with the amplified nucleic acids of 24 M. pneumoniae strains in an enzyme-linked gel assay (ELGA). The results obtained by NASBA-based typing are in agreement with the classification of the 24 M. pneumoniae strains into two types by other typing methods, confirming the reliability of this technique.

Bacterial Typing Techniques↗

Theory and practice of parallel direct optimization.

Our ability to collect and distribute genomic and other biological data is growing at a staggering rate (Pagel, 1999). However, the synthesis of these data into knowledge of evolution is incomplete. Phylogenetic systematics provides a unifying intellectual approach to understanding evolution but presents formidable computational challenges. A fundamental goal of systematics, the generation of evolutionary trees, is typically approached as two distinct NP-complete problems: multiple sequence alignment and phylogenetic tree search. The number of cells in a multiple alignment matrix are exponentially related to sequence length. In addition, the number of evolutionary trees expands combinatorially with respect to the number of organisms or sequences to be examined. Biologically interesting datasets are currently comprised of hundreds of taxa and thousands of nucleotides and morphological characters. This standard will continue to grow with the advent of highly automated sequencing and development of character databases. Three areas of innovation are changing how evolutionary computation can be addressed: (1) novel concepts for determination of sequence homology, (2) heuristics and shortcuts in tree-search algorithms, and (3) parallel computing. In this paper and the online software documentation we describe the basic usage of parallel direct optimization as implemented in the software POY (ftp://ftp.amnh.org/pub/molecular/poy).

Animals↗

Sequence and functional conservation of the intergenic region between the head-to-head genes encoding the small heat shock proteins alphaB-crystallin and HspB2 in the mammalian lineage.

An unexpected feature of the large mammalian genome is the frequent occurrence of closely linked head-to-head gene pairs. Close apposition of such gene pairs has been suggested to be due to sharing of regulatory elements. We show here that the head-to-head gene pair encoding two small heat shock proteins, alphaB-crystallin and HspB2, is closely linked in all major mammalian clades, suggesting that this close linkage is of selective advantage. Yet alphaB-crystallin is abundantly expressed in lens and muscle and in response to a heat shock, while HspB2 is abundant only in muscle and not upregulated by a heat shock. The intergenic distance between the genes for these two proteins in mammals ranges from 645 bp (platypus) to 1069 bp (opossum), with an average of about 900 bp; in chicken the distance was the same as in duck (1.6 kb). Phylogenetic footprinting and sequence alignment identified a number of conserved sequence elements close to the HspB2 promoter and two farther upstream. All known regulatory elements of the mouse alphaB-crystallin promoter are conserved, except in platypus and birds. The lens-specific region 1 (LSR1) and the heat shock elements (HSEs) lack in birds; in platypus the LSR1 is reduced to a Pax-6 site, while the Pax-6 site in LSR2 and a HSE are absent. Most likely the primordial mammalian alphaB-crystallin promoter had two LSRs and two HSEs. In transfection experiments the platypus alphaB-crystallin promoter retained heat shock responsiveness and lens expression. It also directed lens expression in Xenopus laevis transgenes, as did the HspB2 promoter of rat or blind mole rat. Deletion of the middle of the intergenic region including the upstream enhancer affected the activity of both the rat alphaB-crystallin and the HspB2 promoters, suggesting sharing of the enhancer region by the two promoters.

Animals↗

Classification of spider neurotoxins using structural motifs by primary structure features. Single residue distribution analysis and pattern analysis techniques.

In recent years the data on the novel structures of spider toxins have been greatly increasing. The sequence data should be classified. We introduced two primary structure analysis techniques-single residue distribution analysis (SRDA) and pattern analysis for classifying spider polypeptide toxins with molecular weight less than 10kDa. For multiple sequence alignment, we also introduced three novel sequence representation formats named as a simple record, motif record and a pattern record, which can be useful for large-scale analysis of structures. About 300 sequences of spider toxins were analyzed and nine primary structure motifs were identified. New classification of spider toxins was proposed on the basis of previously described principal structural motif (PSM) and extra structural motif (ESM) [Kozlov, S.A., Malyavka, A.A., McCutchen, B., Lu, A., Schepers, E., Herrmann, R., Grishin, E.V., 2005. A novel strategy for the identification of toxin-like structures in spider venom. Proteins 59 (1), 131-140]. Five main structural classes were revealed, and for putative ion channel inhibitors from the most numerous classes 1, 2, and 3, five-digital personal ID numbers were introduced. A reference table with simple, motif and pattern representation sequence formats was created for all analyzed structures.

Amino Acid Motifs↗

Evidence for distinct prototype sequences within the Plasmodium falciparum Pf60 multigene family.

Using oligonucleotides derived from Pf60.1, a member of the Plasmodium falciparum Pf60 multigene family, numerous fragments were amplified from genomic and cDNA from the 3D7 P. falciparum clone. DNA sequencing showed that the various fragments presented considerable diversity, indicating that the 3D7 repertoire contains at least 20 distinct versions of the region analysed. The various sequences aligned with either of two prototype sequences. Characteristic of the A-type was the presence of a 21 bp motif, present in variable copy number, as well as a sequence homologous to the Babesia sp. RAP-1 consensus. The B prototype sequence did not present such features and substantially differed from the A-type, due to accumulation of point mutations and numerous triplet deletions. Consistent with the marked differences between both sub-families, individual members from each sub-family did not cross-hybridise, produced distinct multiple band patterns on Southern blots and distinct chromosome profiles. Numerous hybrid sequences were observed. Interestingly, most var genes and var-related unspliced cDNAs described so far are of A/B hybrid type. These data suggest that the family has evolved by successive amplifications from two ancestral copies, with accumulation of mutations, as well as recombination and/or gene conversion events.

Amino Acid Sequence↗

Identification of an achaete-scute homolog, Fash1, from Fugu rubripes.

Proteins of the achaete-scute family of transcription factors play important roles in neurogenesis in both invertebrates and vertebrates. Here, we report the cloning and characterization of a Japanese pufferfish, Fugu rubripes achaete-scute homolog 1, Fash1. Sequence alignment of the predicted amino acid sequence of Fash1 with other vertebrate homologs of the achaete scute homolog 1 subclass shows that the carboxyl 2/3 of the protein, including the basic helix-loop-helix, a putative nuclear localization signal and several consensus phosphorylation sites, is highly conserved. Strikingly, the similarity in this region between eight vertebrate species is close to 90%.

Amino Acid Sequence↗

Identification of the N-linked oligosaccharide sites in chick corneal lumican and keratocan that receive keratan sulfate.

Corneal proteoglycans have chondroitin/dermatan and keratan sulfate (KS) chains and belong to the leucine-rich proteoglycan gene family. Corneal KS is N-linked to Asn of an NX(S/T) site through a complex oligosaccharide linkage region. Only some sites receive KS, whereas others remain in a high mannose form. To determine whether the attachment of KS was biased toward specific sites, we isolated trypsin-digested KS-containing fragments of chick corneal proteoglycans and sequenced the peptides. Results showed that all of the peptides sequenced aligned to the deduced amino acid sequence of either chick lumican or chick keratocan at the first, third, and fourth potential N-linked sites. Sites 1 and 4 in lumican and keratocan are in a homologous location. By analogy with the structure of ribonuclease inhibitor (a Leu-rich repeat containing protein), the KS chains would extend outward on the outer face of a horseshoe-like structure. The amino acid sequences surrounding the potential N-linked sites were also compared. Sites receiving KS tend to have a higher occurrence of aromatic residues, in particular Phe, located within 3 amino acids of NX(S/T). These conserved Phe residues may have a role in the conversion of high mannose N-linked oligosaccharides to polylactosamine and/or keratan sulfate.

Amino Acid Sequence↗

Crystal structure and functional analysis of lipoamide dehydrogenase from Mycobacterium tuberculosis.

We report the 2.4 A crystal structure for lipoamide dehydrogenase encoded by lpdC from Mycobacterium tuberculosis. Based on the Lpd structure and sequence alignment between bacterial and eukaryotic Lpd sequences, we generated single point mutations in Lpd and assayed the resulting proteins for their ability to catalyze lipoamide reduction/oxidation alone and in complex with other proteins that participate in pyruvate dehydrogenase and peroxidase activities. The results suggest that amino acid residues conserved in mycobacterial species but not conserved in eukaryotic Lpd family members modulate either or both activities and include Arg-93, His-98, Lys-103, and His-386. In addition, Arg-93 and His-386 are involved in forming both "open" and "closed" active site conformations, suggesting that these residues play a role in dynamically regulating Lpd function. Taken together, these data suggest protein surfaces that should be considered while developing strategies for inhibiting this enzyme.

Amino Acid Sequence↗