PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

CASTp: computed atlas of surface topography of proteins with structural and topographical mapping of functionally annotated residues.

Cavities on a proteins surface as well as specific amino acid positioning within it create the physicochemical properties needed for a protein to perform its function. CASTp (http://cast.engr.uic.edu) is an online tool that locates and measures pockets and voids on 3D protein structures. This new version of CASTp includes annotated functional information of specific residues on the protein structure. The annotations are derived from the Protein Data Bank (PDB), Swiss-Prot, as well as Online Mendelian Inheritance in Man (OMIM), the latter contains information on the variant single nucleotide polymorphisms (SNPs) that are known to cause disease. These annotated residues are mapped to surface pockets, interior voids or other regions of the PDB structures. We use a semi-global pair-wise sequence alignment method to obtain sequence mapping between entries in Swiss-Prot, OMIM and entries in PDB. The updated CASTp web server can be used to study surface features, functional regions and specific roles of key residues of proteins.

Amino Acids↗

Proposed structure for the DNA-binding domain of the helix-loop-helix family of eukaryotic gene regulatory proteins.

A modelled tertiary structure for the dimeric HLH domain of the E47 protein is presented. Structural information was obtained from the aligned sequences of > 40 members of the HLH family. The information was used to model each monomer as an alpha-helical hairpin, with knobs-into-holes packing of side-chains as found in antiparallel coiled-coil. The dimer forms a four-helix bundle with additional knobs-into-holes packing at the dimer interface. The size and electrostatic properties of core-forming residues are all accounted for in the model. The model does not violate any known properties of protein structure. The monomers are related by two-fold rotational symmetry, in agreement with the observed DNA-binding sites which are imperfect inverted repeats. The N-terminal basic region, in which DNA binding and base specificity reside, forms the first part of helix 1. A prediction based on the model structure is that the HLH domains do not bind to DNA in its B form but require a partially unwound conformation in order to enter the major groove.

Amino Acid Sequence↗

Enteroviral infection causing fatal myocarditis and subclinical myopathy.

UNLABELLED: Enteroviral RNA detection in myocarditis and dilated cardiomyopathy is rare. Enteroviral particles and RNA have recently been identified in patient's skeletal muscle, suggesting that skeletal more than heart muscle hosts the virus in chronic infection. Enteroviral RNA and virus-like particles were found in the myocardium and in the skeletal muscle of two patients with fatal myocarditis: a 39 year old man who died five days after the onset of febrile flu; and a 49 year old woman, assisted for 50 days with a left ventricular assist device, who then died from cerebral haemorrhage. Automated sequencing, alignment, and sequence comparison confirmed the enteroviral origin of polymerase chain reaction products and excluded contamination. These findings agree with prior observations of enteroviral localisation in the skeletal muscle of patients with dilated cardiomyopathy, and further support the hypothesis that skeletal rather than heart muscle may host the virus and serve as a reservoir in cardiomyopathies related to chronic infection. KEYWORDS: enterovirus; myocarditis; viral particles; skeletal muscle

Adult↗

The molecular diversity of the 5S rRNA gene in barley (Hordeum vulgare).

The 5S rRNA genes from several accessions of cultivated barley, Hordeum vulgare L., were amplified by the polymerase chain reaction, cloned, and sequenced. Analysis of the aligned sequences, followed by principal coordinate analysis, support the recognition of at least two distinct classes of 5S rDNA genes. The short repeat class corresponds to the 300-bp tandem repeat defined by E.V. Ananiev as containing several TAG repeating units. The long repeat class contains long tandem repeats and lacks the TAG repeating unit. Sequences in each class can be further subdivided, with the long repeat class containing two groups and the short repeat class containing two and possibly three groups. These results suggest that in cultivated barley the sequence diversity found within the 5S rDNA nontranscribed spacer region may be encoded by three or more loci and may be useful for phylogenetic analyses provided that orthology can be established.

Base Sequence↗

Exact algorithms for computing pairwise alignments and 3-medians from structure-annotated sequences (extended abstract).

Given the problem of mutation saturation in ancient molecular sequences, there is great interest in inferring phylogenies from higher-order types of molecular data that change more slowly, such as genomic organization and the secondary and tertiary structures of ribosomal RNA and proteins. In this paper, we define edit distances based on two representations of RNA secondary structure, arc annotation and hierarchical string annotation, and give algorithms for computing these distances on pairs of annotated sequences, aligning pairs of annotated sequences, and computing 3-median annotated sequences from triples of annotated sequences. The 3-median algorithms can be used as part of a well-known iterative heuristic for inferring phylogenies. All given algorithms are adapted from algorithms for computing longest common annotated subsequences of pairs of annotated sequences.

Algorithms↗

Pseudo-likelihood for non-reversible nucleotide substitution models with neighbour dependent rates.

In the field of molecular evolution genome substitution models with neighbour dependent substitution rates have recently received much attention. It is well-known that substitution of nucleotides does not occur independently of neighbouring nucleotides, but there has been less focus on the phenomenon that this substitution process is also not time-reversible. In this paper I construct a pseudo-likelihood type method for inference in non-reversible substitution models with neighbour dependent substitution rates. I also construct an EM-algorithm for maximising the pseudo-likelihood. For human-mouse aligned sequence data a number of different models are investigated, where I show that strand-symmetric models are appropriate, and that overlapping di-nucleotide models do not fit the data well.

Algorithms↗

The fortuitous cloning of retroelement-like sequences from wheat and rye as by-products of a specific polymerase chain reaction.

Cloning of by-products of a specific PCR reaction, directed to the Em genes of wheat and rye, has resulted in the identification of ten sequences with homology to the known Tyl-copia-like retroelements WIS 2-1A from wheat and BARE-1 from barley. These sequences were amplified by only one of the primers due to the presence of an inverted repeat. Nine sequences are ca. 740 bp long and contain part of the left LTR, the adjacent primer-binding site and part of the leader sequence, whereas one shorter sequence (535 bp) consists of part of the leader sequence only. The dendrogram, constructed from the multiple sequence alignment, classified the isolated sequences into two narrowly related groups that belong to the WIS-2 family of cereal retroelements.

Base Sequence↗

Direct immunosuppressive effects of EBV-encoded latent membrane protein 1.

In neoplastic cells of EBV-positive lymphoid malignancies latent membrane protein (LMP1) is expressed. Because no adequate cellular immune response can be detected against LMP1, we investigated whether LMP1 had a direct effect on T lymphocyte activation. In this study we show that nanogram amounts of purified recombinant LMP1 (rLMP1) strongly suppresses activation of T cells. By sequence alignment two sequences (LALLFWL and LLLLAL) in the first transmembrane domain of LMP1 were identified showing strong homology to the immunosuppressive domain (LDLLFL) of the retrovirus-encoded transmembrane protein p15E. The effects of rLMP1 and LMP1-derived peptides were tested in T cell proliferation and NK cytotoxicity assays and an Ag-induced IFN-gamma release enzyme-linked immunospot assay. LMP1 derived LALLFWL peptides showed strong inhibition of T cell proliferation and NK cytotoxicity, while acetylated LALLFWL peptides had an even stronger effect. In addition, Ag-specific IFN-gamma release was severely inhibited. To exert immunosuppressive effects in vivo, LMP1 has to be excreted from the cells. Indeed, LMP1 was detected in supernatant of EBV-positive B cell lines (LCL), and differential centrifugation in combination with Western blot analysis of the pellets indicated that LMP1 is probably secreted by LCL in the form of exosomes. The amount of secreted LMP1 in B cell cultures is well below the immunosuppressive level observed with rLMP1. Our results demonstrate direct immunosuppressive properties of LMP1 (fragments) and suggest that EBV-positive tumor cells may actively secrete LMP1 and thus mediate immunosuppressive effects on tumor-infiltrating lymphocytes. Moreover, we demonstrate, for the first time, that transmembrane protein-mediated immunosuppression is not solely restricted to RNA tumor viruses, but can also be found in DNA tumor viruses.

Amino Acid Sequence↗

New strategy to detect single nucleotide polymorphisms.

A great effort has been made to identify and map a large set of single nucleotide polymorphisms. The goal is to determine human DNA variants that contribute most significantly to population variation in each trait. Different algorithms and software packages, such as PolyBayes and PolyPhred, have been developed to address this problem. We present strategies to detect single nucleotide polymorphisms, using chromatogram analysis and consensi of multiple aligned sequences. The algorithms were tested using HIV datasets, and the results were compared with those produced by PolyBayes and PolyPhred using the same dataset. Our algorithms produced significantly better results than these two software packages.

Algorithms↗

Ribosomal RNA sequences of Clostridium piliforme isolated from rodent and rabbit: re-examining the phylogeny of the Tyzzer's disease agent and development of a diagnostic polymerase chain reaction assay.

We used polymerase chain reaction (PCR) technology to amplify the 16S rRNA gene, the intergenic spacer, and most of the 23S rRNA gene from 6 isolates (2 mice, 1 hamster, 1 rat, and 2 rabbit isolates) of the Tyzzer's disease agent (Clostridium piliforme) and C. colinum. Sequence similarity searches of GenBank identified 45 closely related bacteria, which we used for phylogenetic analysis by parsimony and maximum-likelihood methods using Escherichia coli to root the resulting phylogram. Microorganisms identified as C. piliforme form 3 clusters within a single clade; the nearest related distinguishable species is C. colinum. Other bacterial clades closely related to C. piliforme are clostridia previously identified by molecular methods in the bovine, porcine, and human gastrointestinal tracts. DNA sequence alignment highlighting sequence differences were used to design a rodent and rabbit C. piliforme-specific PCR assay, which targets a 639-basepair region at the 3' end of the 16S rRNA gene and the 5' end of the intergenic spacer. We used this PCR assay to examine 4 rat fecal samples from C. piliformeseropositive rats and reexamine 2 rabbit fecal samples previously identified as containing DNA sequences consistent with C. piliforme infection by 16S PCR assay. Our new assay did not detect the presence of C. piliforme DNA sequences in either the rat or rabbit fecal DNA samples, consistent with the absence of clinical disease in the colonies evaluated.

Animals↗

Isolation and characterization of a skate retinal GABA transporter cDNA.

PURPOSE: The inhibitory neurotransmitter gamma-aminobutyric acid (GABA) is believed to play a crucial role in the processing of information within the vertebrate retina. Extracellular concentrations of GABA are thought to be tightly regulated by carrier-mediated transport proteins in neurons and glial cells. The purpose of this work was to isolate the gene that encodes one of these transport proteins in the skate retina. METHODS: cDNA clones were isolated from a skate retinal cDNA library using a mouse retinal GABA transporter (GAT1) cDNA as a probe. The PCR technique was used to fill sequence gaps, and 5' and 3' RACE were employed to amplify the 5' and 3' untranslated regions. The amplified fragments were subcloned into a T-vector. Blots containing RNA from 10 different tissues were probed to determine the size of the transcript and the tissue distribution. RESULTS: Sequence analysis revealed that the skate retinal GABA transporter cDNA shared 72% identity with the mouse GABA transporter-1 at the DNA level and 80% identity at the amino acid level. Multiple sequence alignments showed that our sequence is closest to the Torpedo GABA transporter-1. Two transcripts, 4.5 and 7 kb, were detected in retina and possibly brain by RNA blot analysis. Fourteen introns were detected in the skate GABA transporter gene. CONCLUSIONS: We successfully isolated a full length GABA transporter cDNA from the retina of the skate. The size of the full length sequence of the skate retinal GABA transporter is in agreement with the size of the smaller transcript detected on RNA blots. The larger transcript observed on the RNA blot may be the result of either alternative splicing or utilization of a downstream poly A signal.

Animals↗

Alignment and comparison of multiple related DNA sequences can be improved by consideration of known patterns of mutation.

The rate at which bases are replaced in mammalian DNA is influenced by the sequence involved. I have modified a multiple sequence alignment and comparison program to take account of the observed patterns of base replacement in mammals, and used the program to analyse 100 human Alu sequences. The results show that using a sequence matching matrix based on the actual rate of base replacement, as opposed to arbitrary base matching matrix, gives significant differences in the way that sequences are aligned and compared with each other. Use of the observed matching matrix gives significantly better discrimination between sub-families of Alu sequences than conventional methods, and shows that comparison systems which take account of the mechanism of mutation give more biologically realistic results.

Animals↗

Recent developments in linear-space alignment methods: a survey.

A dynamic-programming strategy for sequence alignment first proposed in 1975 by Dan Hirschberg can be adapted to yield a number of extremely space-efficient algorithms. Specifically, these algorithms align two sequences using only "linear space," i.e., an amount of computer memory that is proportional to the sum of the lengths of the two sequences being aligned. This paper begins by reviewing the basic idea, as it applies to the global (i.e., end-to-end) alignment of two DNA or protein sequences. Three of our recent extensions of the technique are then outlined. The first extension computes an optimal alignment subject to the constraint that each position, i, of the first sequence must be aligned somewhere between positions L[i] and U[i] of the second sequence, for given values of L and U. The second finds all aligned position pairs (i.e., potential columns of the alignment) that occur in an alignment whose score exceeds a given threshold. The third treats the case where each of the two sequences is allowed to be an alignment (e.g., a sequence of aligned pairs), using a sensitive scoring scheme. We also describe two linear-space methods for computing k best local (i.e., involving only a part of each sequence) alignments, where k > or = 1. One is a linear-space version of the algorithm of Waterman and Eggert (1987), and the other is based on the strategy proposed by Wilbur and Lipman (1983). Finally, we describe programs that implement various combinations of these techniques to provide a multisequence alignment method that is especially suited to handling a few very long sequences. The utility of these programs is illustrated by analysis of the locus control region of the beta-like globin gene cluster of several mammals.

Algorithms↗

Estimation of P-values for global alignments of protein sequences.

MOTIVATION: The global alignment of protein sequence pairs is often used in the classification and analysis of full-length sequences. The calculation of a Z-score for the comparison gives a length and composition corrected measure of the similarity between the sequences. However, the Z-score alone, does not indicate the likely biological significance of the similarity. In this paper, all pairs of domains from 250 sequences belonging to different SCOP folds were aligned and Z-scores calculated. The distribution of Z-scores was fitted with a peak distribution from which the probability of obtaining a given Z-score from the global alignment of two protein sequences of unrelated fold was calculated. A similar analysis was applied to subsequence pairs found by the Smith-Waterman algorithm. These analyses allow the probability that two protein sequences share the same fold to be estimated by global sequence alignment. RESULTS: The relationship between Z-score and probability varied little over the matrix/gap penalty combinations examined. However, an average shift of +4.7 was observed for Z-scores derived from global alignment of locally-aligned subsequences compared to global alignment of the full-length sequences. This shift was shown to be the result of pre-selection by local alignment, rather than any structural similarity in the subsequences. The search ability of both methods was benchmarked against the SCOP superfamily classification and showed that global alignment Z-scores generated from the entire sequence are as effective as SSEARCH at low error rates and more effective at higher error rates. However, global alignment Z-scores generated from the best locally-aligned subsequence were significantly less effective than SSEARCH. The method of estimating statistical significance described here was shown to give similar values to SSEARCH and BLAST, providing confidence in the significance estimation. AVAILABILITY: Software to apply the statistics to global alignments is available from http://barton.ebi.ac.uk. CONTACT: geoff@ebi.ac.uk

Databases, Protein↗

Detection and analysis of alternative splicing in the silkworm by aligning expressed sequence tags with the genomic sequence.

We identified 277 alternative splice forms in silkworm genes based on aligning expressed sequence tags with genomic sequences, using a transcript assembly program. A large fraction (74%) of these alternative splices are located in protein-coding regions and alter protein products, whereas only 26% are in untranslated regions. From the alternative splices located in protein-coding regions, some (43%) affect protein domains that bind various biological molecules. The vast majority of the detected alternative forms in this study appear to be novel, and potentially affect biologically meaningful control of function in silkworm genes. Our results indicate that alternative splicing in silkworm largely produces protein diversity and functional diversity, and is a widely used mechanism for regulating gene expression.

Alternative Splicing↗

MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.

Analysis of multiple sequence alignments can generate important, testable hypotheses about the phylogenetic history and cellular function of genomic sequences. We describe the MultiPipMaker server, which aligns multiple, long genomic DNA sequences quickly and with good sensitivity (available at http://bio.cse.psu.edu/ since May 2001). Alignments are computed between a contiguous reference sequence and one or more secondary sequences, which can be finished or draft sequence. The outputs include a stacked set of percent identity plots, called a MultiPip, comparing the reference sequence with subsequent sequences, and a nucleotide-level multiple alignment. New tools are provided to search MultiPipMaker output for conserved matches to a user-specified pattern and for conserved matches to position weight matrices that describe transcription factor binding sites (singly and in clusters). We illustrate the use of MultiPipMaker to identify candidate regulatory regions in WNT2 and then demonstrate by transfection assays that they are functional. Analysis of the alignments also confirms the phylogenetic inference that horses are more closely related to cats than to cows.

Algorithms↗

Detecting overlapping coding sequences with pairwise alignments.

MOTIVATION: Overlapping gene coding sequences (CDSs) are particularly common in viruses but also occur in more complex genomes. Detecting such genes with conventional gene-finding algorithms can be difficult for several reasons. If an overlapping CDS is on the same read-strand as a known CDS, then there may not be a distinct promoter or mRNA. Furthermore, the constraints imposed by double-coding can result in atypical codon biases. However, these same constraints lead to particular mutation patterns that may be detectable in sequence alignments. RESULTS: In this paper, we investigate several statistics for detecting double-coding sequences with pairwise alignments--including a new maximum-likelihood method. We also develop a model for double-coding sequence evolution. Using simulated sequences generated with the model, we characterize the distribution of each statistic as a function of sequence composition, length, divergence time and double-coding frame. Using these results, we develop several algorithms for detecting overlapping CDSs. The algorithms were tested on known overlapping CDSs and other overlapping open reading frames (ORFs) in the hepatitis B virus (HBV), Escherichia coli and Salmonella typhimurium genomes. The algorithms should prove useful for detecting novel overlapping genes--especially short coding ORFs in viruses. AVAILABILITY: Programs may be obtained from the authors. SUPPLEMENTARY INFORMATION: http://biochem.otago.ac.nz/double.html.

Algorithms↗

Identification of catalytically relevant amino acids of the extracellular Serratia marcescens endonuclease by alignment-guided mutagenesis.

By sequence alignment of the extracellular Serratia marcescens nuclease with three related nucleases we have identified seven charged amino acid residues which are conserved in all four sequences. Six of these residues together with four other partially conserved His or Asp residues were changed to alanine by site-directed PCR-mediated mutagenesis using a variant of the nuclease gene in which the coding sequence of the signal peptide was replaced by the coding sequence for an N-terminal affinity tag [Met(His)6GlySer]. Four of the mutant proteins showed almost no reduction in nuclease activity but five displayed a 10- to 1000-fold reduction in activity and one (His110Ala) was inactive. Based upon these results it is suggested that the S.marcescens nuclease employs a mechanism in which His110 acts in concert with a Mg2+ ion and three carboxylates (Asp107, Glu148 and Glu232) as well as one or two basic amino acid residues (Arg108, Arg152).

Amino Acid Sequence↗