PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Comparative ab initio prediction of gene structures using pair HMMs.

We present a novel comparative method for the ab initio prediction of protein coding genes in eukaryotic genomes. The method simultaneously predicts the gene structures of two un-annotated input DNA sequences which are homologous to each other and retrieves the subsequences which are conserved between the two DNA sequences. It is capable of predicting partial, complete and multiple genes and can align pairs of genes which differ by events of exon-fusion or exon-splitting. The method employs a probabilistic pair hidden Markov model. We generate annotations using our model with two different algorithms: the Viterbi algorithm in its linear memory implementation and a new heuristic algorithm, called the stepping stone, for which both memory and time requirements scale linearly with the sequence length. We have implemented the model in a computer program called DOUBLESCAN. In this article, we introduce the method and confirm the validity of the approach on a test set of 80 pairs of orthologous DNA sequences from mouse and human. More information can be found at: http://www.sanger.ac.uk/Software/analysis/doublescan/

Algorithms↗

Studies of chromophores in model membranes by polarized light-absorption spectroscopy. The orientation and binding of tetracaine and procaine.

A method for studying the orientation and binding of chromophores in macroscopically aligned membranes by polarized light absorption spectroscopy is described. Here tetracaine and procaine solubilized in a lamellar phase of octanoyl-1-glyceride (monooctanoin) and water have been investigated. Tetracaine is found to be located in the lipid region with a preferential orientation of the molecular long axis parallel to the hydrocarbon chains. The orientation of procaine, mainly residing in the water region, is very small.

Binding Sites↗

Trypanothione reductase of Trypanosoma congolense: gene isolation, primary sequence determination, and comparison to glutathione reductase.

The gene encoding trypanothione reductase, the redox disulfide-containing flavoenzyme that is unique to the parasitic trypanosomatids (Shames et al., 1986), has been isolated from the cattle pathogen Trypanosoma congolense. Library screening was carried out with inosine-containing oligonucleotide probes encoding sequences determined from two active site peptides isolated from the purified Crithidia fasciculata enzyme. The nucleotide sequence of the gene was determined according to the dideoxy chain termination method of Sanger. The structural gene is 1476 nucleotides long and encodes 492 amino acids. We have identified the active site peptide containing the redox-active disulfide, a peptide corresponding to the histidine-467 region of human erythrocyte glutathione reductase, as well as the flavin binding domain that is highly conserved in all disulfide-containing flavoprotein reductase enzymes. Alignment of five tryptic peptides (80 residues) isolated from the C. fasciculata trypanothione reductase with the primary sequence of the T. congolense enzyme showed 88% homology with 76% identity. Additionally, a sequence comparison of the glutathione reductase from Escherichia coli or human erythrocytes to T. congolense trypanothione reductase reveals greater than 50% homology. A search for the amino acid residues in the primary sequence of trypanothione reductase functionally active in binding/catalysis in human erythrocyte glutathione reductase shows that only the two arginine residues (Arg-37 and Arg-347), shown by X-ray crystallographic data to hydrogen bond to the GS1 glutathione glycyl carboxylate, are absent.

Amino Acid Sequence↗

[Hemoglobins, XXV. Hemoglobin (erythrocruorin) CTT III from Chironomus thummi thummi (Diptera). Primary structure and relationship to other heme proteins (author's transl)].

The amino acid sequence analysis of hemoglobin (erythrocruorin) CTT III from Chironomus thummi th. (Diptera) has been checked with automatic methods and completed. The protein chain consists of 136 amino acids and contains a neutral exchange isoleucine/threonine in position 57. The molecular weight of the heme protein (Thr) is 15400. The primary structure gives the chemical basis for the refinement of the X-ray structure and the understanding of the mechanism of the Bohr effect in this monomeric hemoglobin. A homologous alignment to vertebrate globins is reported. The resulting data for the phylogeny of proto-and deuterostomian animals and the function of this hemoglobin are discussed.

Amino Acid Sequence↗

Isolation and sequence analysis of the small subunit ribosomal RNA gene from the euryhaline yeast Debaryomyces hansenii.

The small subunit ribosomal RNA gene (SSU rDNA) from the euryhaline yeast Debaryomyces hansenii has been isolated and sequenced. After appropriate alignment of this sequence with SSU rDNA sequences from 30 other taxa, phylogenetic reconstruction using distance matrix and maximum parsimony methods indicates that D. hansenii is most closely affiliated with Candida albicans, and occurs in the cluster of the yeasts Saccharomyces cerevisiae, Torulaspora delbruekii, Candida glabrata, and Kluyveromyces lactis. It appears that the capacity to tolerate high salt is independent of phylogenetic affiliations based on SSU rDNA analyses.

Base Sequence↗

Cyclospora cayetanensis: a review of an emerging parasitic coccidian.

Cyclospora cayetanensis is a sporulating parasitic protozoan that infects the upper small intestinal tract. It has been identified as both a food and waterborne pathogen endemic in many developing countries. It is an important agent of Traveller's Diarrohea in developed countries and was responsible for numerous foodborne outbreaks in the United States and Canada in the late 1990s. Like Cryptosporidium, infection has been associated with a variety of sequelae such as Guillain-Barré syndrome, reactive arthritis syndrome (formally Reiter syndrome) and acalculous cholecystitis. There has been much debate as to where to place C. cayetanensis taxonomically due to its homology with Eimeria species. To date, the only genomic DNA sequences available are the ribosomal DNA of C. cayetanensis and three other species; within these a high degree of homology has been observed. This homology and the lack of sequence data from other Cyclospora species have hindered identification methods.

Animals↗

The role of pattern databases in sequence analysis.

In the wake of the numerous now-fruitful genome projects, we are entering an era rich in biological data. The field of bioinformatics is poised to exploit this information in increasingly powerful ways, but the abundance and growing complexity both of the data and of the tools and resources required to analyse them are threatening to overwhelm us. Databases and their search tools are now an essential part of the research environment. However, the rate of sequence generation and the haphazard proliferation of databases have made it difficult to keep pace with developments. In an age of information overload, researchers want rapid, easy-to-use, reliable tools for functional characterisation of newly determined sequences. But what are those tools? How do we access them? Which should we use? This review focuses on a particular type of database that is increasingly used in the task of routine sequence analysis--the so-called pattern database. The paper aims to provide an overview of the current status of pattern databases in common use, outlining the methods behind them and giving pointers on their diagnostic strengths and weaknesses.

Amino Acid Motifs↗

The rrs (16S)-rrl (23S) ribosomal intergenic spacer region as a target for the detection of Haemophilus ducreyi by a heminested-PCR assay.

The intergenic spacer region between the rrs and rrl ribosomal RNA genes of Haemophilus ducreyi was analysed and the DNA sequence was used for the selection of specific PCR primers. A highly sensitive and specific heminested-PCR assay for the identification of H. ducreyi was developed. The assay showed a sensitivity of 96% on genital ulcer specimens from patients with clinically diagnosed chancroid, compared with a sensitivity of 56% for culture methods. These results indicate that this PCR assay has the potential to become an accurate and easy reference method for the detection of H. ducreyi.

Base Sequence↗

Detection of neuron-specific gamma-enolase messenger ribonucleic acid in normal human leukocytes by polymerase chain reaction amplification with nested primers.

BACKGROUND: NNE (non-neuronal alpha-enolase) is a glycolytic enzyme detected in most tissues. NSE (neuron-specific gamma-enolase) is detected in normal neurons and tumors such as neuroblastoma. Staining with antibodies against NSE is therefore used to detect neuroblastoma cells invading bone marrow. Since staining of normal leukocytes has been reported we asked whether bona fide NSE is in fact expressed in normal blood and marrow. EXPERIMENTAL DESIGN: We designed nested coding region specific primers for NSE and NNE and, after reverse transcription of mRNA, we amplified the coding region between these primers in a semi-nested polymerase chain reaction. In order to distinguish both iso-mRNAs from each other, we amplified a long (1,047 bp) template in a first round of 30 cycles with primers specific for NNE or NSE. One percent of this product was used in a second round of 30 cycles in which both sense primers and two nested anti-sense primers of alternate specificities yielding shorter products of discernible sizes (768 bp or 619 bp) were added together in the same reaction tube. With this combination of four primers, only that shorter product was amplified to visibility, the specificity of which was homologous to the template produced in the first 30 cycles. Restriction enzyme digestion of the amplified products was used to verify this polymerase chain reaction-based approach for the distinction of isoforms of RNA. RESULTS: This semi-nested polymerase chain reaction clearly allows for the distinction of mRNA for NNE or NSE and shows the presence of transcripts for NSE in normal human leukocytes from blood and bone marrow. CONCLUSIONS: This method exploiting short stretches of nucleotide differences in the coding regions for priming can more generally be applied to the distinction of all isoforms of RNA where nested specific primers can be designed. However, the presence of NSE specific transcripts in normal human leukocytes invalidates the use of this highly sensitive method as a disease marker in neuroblastoma.

Base Sequence↗

Molecular epidemiological study of dengue virus type 1 in Taiwan.

Taiwan has experienced several major outbreaks of dengue (DEN) virus since 1981. The predominant virus type involved has been dengue virus type one (DEN-1), which first appeared in 1987. To understand the molecular epidemiology of this virus, 15 strains of DEN-1 isolated during 1987-1991 and 1994-1995, including 11 epidemic strains, two sporadic strains, and two imported strains have been studied. Fragments of 490 nucleotides (nt) from the E/NS1 junction were amplified by reverse transcription-polymerase chain reaction and the nt sequences were determined. Of the 490 nt of the E/NS1 junction, 240 nt (nt 2282-2521) were aligned and compared. Nucleotide substitutions were found at 54 positions among 15 isolates. Most nt changes were synonymous substitutions, and only three amino acid changes were found. A total of 61 strains isolated worldwide were analyzed by the Neighbor-joining method, and separated phylogenetically into three distinct genotypes, I-III. Genotype I comprised isolates from Japan and Hawaii collected in the 1940s. Genotype II included most strains isolated from Asia in 1977-1995. Genotype III consisted of isolates from three continents in 1964-1995: Asia, the Americas, and Africa. Genotype III was divided further into two subgenotypes, IIIA and IIIB. Most recent isolates from Taiwan, except for the sporadic strain isolated in 1995, were similar genetically and have been classified as Genotype II.

Amino Acid Sequence↗

The primary structure of high density apolipoprotein-glutamine-I.

The major protein constituent of human plasma high density lipoproteins has been isolated and its complete amino-acid sequence determined. The protein, designated apolipoprotein-glutamine-I by the presence of carboxyl-terminal glutamine, is a single polypeptide chain of 245 amino-acid residues, including three residues of methionine. The protein is devoid of cysteine, cystine, and isoleucine. Cleavage of apolipoprotein-glutamine-I with cyanogen bromide yields four fragments with 94, 90, 36, and 25 amino acids. The amino-acid sequence of each fragment was determined by conventional methods, with proteolytic digestion with trypsin, chymotrypsin, and thermolysin. The alignment of the cyanogen bromide fragments was determined by the isolation of the methionine-containing tryptic peptides from apolipoprotein-glutamine-I. Inspection of the sequence of apolipoprotein-glutamine-I suggests an interesting distribution of amino acids that may account for its helical structure and its ability to bind and transport lipid.

Amino Acid Sequence↗

Structural studies on the coat protein of alfalfa mosaic virus. The complete primary structure.

The complete amino acid sequence of the coat protein of alfalfa mosaic virus (strain 425) is reported. Sequence determinations were mainly performed on peptides obtained from fragmentation by cyanogen bromide and trypsin. Both manual and automatic sequence methods were used. Some refinements of the solid-phase Edman degradation were introduced. The final alignment of the peptides was established by means of alternative cleavage methods, such as limited tryptic digestion of intact virus particles, tryptic digestion after blockage of lysine residues and chymotryptic digestion. The coat protein consists of 220 amino acid residues corresponding to a molecular weight of 24252. A remarkable clustering of basic residues occurs in the N-terminal part of the protein chain. Several internal hydrophobic clusters and a strongly acidic site at the C-terminus can be observed. Two regions of sequence homology (12 residues) were found. Some features of the secondary structure are predicted.

Amides↗

Secondary structure prediction and unrefined tertiary structure prediction for cyclin A, B, and D.

We present heuristic-based predictions of the secondary and tertiary structures of cyclins A, B, and D, representatives of the cyclin superfamily. The list of suggested constraints for tertiary structure assembly was left unrefined in order to submit this report before an announced crystal structure for cyclin A becomes available. To predict these constraints, a master sequence alignment over 270 positions of cyclin types A, B, and D was adjusted based on individual secondary structure predictions for each type. We used new heuristics for predicting aromatic residues at protein-protein interfaces and to identify sequentially distinct regions in the protein chain that cluster in the folded structure. The boundaries of two conjectured domains in the cyclin fold were predicted based on experimental data in the literature. The domain that is important for interaction of the cyclins with cyclin-dependent kinases (CDKs) is predicted to contain six helices; the second domain in the consensus model contains both helices and a beta-sheet that is formed by sequentially distant regions in the protein chain. A plausible phosphorylation site is identified. This work represents a blinded test of the method for prediction of secondary and, to a lesser extent, tertiary structure from a set of homologous protein sequences. Evaluation of our predictions will become possible with the publication of the announced crystal structure.

Amino Acid Sequence↗

Peptide sequences binding to MHC class I proteins.

Motifs for peptides which bind specifically to the human class I major histocompatibility complex molecules HLA-A2 and B7 were determined by sequence analysis of class I-bound peptides selected from a random synthetic library of nonamers. Thirteen individual peptides were sequenced for HLA-A2, twelve individual and nine pooled peptides were sequenced for HLA-B7. Analysis of sequence alignment implicated four peptide positions in potential contact with the class I HLA-A2 molecule and three positions for the HLA-B7 molecule. The results demonstrate that a synthetic peptide library can be used to identify allele-specific motifs for class I molecules, providing information comparable to the results obtained from sequencing endogenous peptides. This method utilizes denatured class I heavy chains, and similar results were obtained using a class I protein purified from mammalian cells or by expression in Escherichia coli. This method has the potential to detect peptides which may not be generated physiologically, but due to their binding properties, may be valuable to predict or engineer immunomodulatory T cell epitopes.

Amino Acid Sequence↗

Identification and characterization of subfamily-specific signatures in a large protein superfamily by a hidden Markov model approach.

BACKGROUND: Most profile and motif databases strive to classify protein sequences into a broad spectrum of protein families. The next step of such database studies should include the development of classification systems capable of distinguishing between subfamilies within a structurally and functionally diverse superfamily. This would be helpful in elucidating sequence-structure-function relationships of proteins. RESULTS: Here, we present a method to diagnose sequences into subfamilies by employing hidden Markov models (HMMs) to find windows of residues that are distinct among subfamilies (called signatures). The method starts with a multiple sequence alignment (MSA) of the subfamily. Then, we build a HMM database representing all sliding windows of the MSA of a fixed size. Finally, we construct a HMM histogram of the matches of each sliding window in the entire superfamily. To illustrate the efficacy of the method, we have applied the analysis to find subfamily signatures in two well-studied superfamilies: the cadherin and the EF-hand protein superfamilies. As a corollary, the HMM histograms of the analyzed subfamilies revealed information about their Ca2+ binding sites and loops. CONCLUSIONS: The method is used to create HMM databases to diagnose subfamilies of protein superfamilies that complement broad profile and motif databases such as BLOCKS, PROSITE, Pfam, SMART, PRINTS and InterPro.

Binding Sites↗

Discovering new genes with advanced homology detection.

Most genome annotation protocols combine ab initio predictions with transcription and homology analyses to produce reliable gene predictions but they often fail to detect many actual genes. Alternative approaches involving more sensitive homology recognition methods are playing an increasingly important role in the next stage of gene discovery. The hunt for new genes is far from over.

Database Management Systems↗

Rapid identification of clinically relevant Nocardia species to genus level by 16S rRNA gene PCR.

Two regions of the gene coding for 16S rRNA in Nocardia species were selected as genus-specific primer sequences for a PCR assay. The PCR protocol was tested with 60 strains of clinically relevant Nocardia isolates and type strains. It gave positive results for all strains tested. Conversely, the PCR assay was negative for all tested species belonging to the most closely related genera, including Dietzia, Gordona, Mycobacterium, Rhodococcus, Streptomyces, and Tsukamurella. Besides, unlike the latter group of isolates, all Nocardia strains exhibited one MlnI recognition site but no SacI restriction site. This assay offers a specific and rapid alternative to chemotaxonomic methods for the identification of Nocardia spp. isolated from pathogenic samples.

Bacteriological Techniques↗