PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Phylogenetic position of symbiotic protist Dinenympha [correction of Dinemympha] exilis in the hindgut of the termite Reticulitermes speratus inferred from the protein phylogeny of elongation factor 1 alpha.

The phylogenetic position of the symbiotic oxymonad Dinenympha exilis, found in the hindgut of the lower termite Reticulitermes speratus, was determined by analysis of translation elongation factor 1 alpha (EF-1 alpha). cDNA corresponding to a major part of the amino acid coding region of EF-1 alpha mRNA was amplified by the reverse transcription polymerase chain reaction (RT-PCR) method from total mRNA of termite hindgut microorganisms without cultivation. The product was cloned into a plasmid vector, pGEM-T, and the clones were isolated and sequenced. One of the EF-1 alpha clones isolated was assigned to the protist D. exilis by whole-cell in-situ hybridization using a specific oligonucleotide probe with enzymatic signal amplification. The deduced amino acid sequence was aligned with those of other eukaryotic and archaeabacterial EF-1 alpha s, and the phylogenetic relationships among early branching eukaryotes were inferred by using the distance matrix method and the maximum parsimony method. The phylogenetic analysis indicated that the D. exilis offshoot occurred before mitochondria-containing organisms and D. exilis branched out after the diplomonads clade. These results indicate that the oxymonad D. exilis is one of the early branching organisms and suggest that the oxymonads form a lineage independent of other early branching organisms.

Amino Acid Sequence↗

Use of the DNA sequence of variable regions of the 16S rRNA gene for rapid and accurate identification of bacteria in the Lactobacillus acidophilus complex.

The Lactobacillus acidophilus complex includes Lact. acidophilus, Lactobacillus amylovorus, Lactobacillus crispatus, Lactobacillus gallinarum, Lactobacillus gasseri and Lactobacillus johnsonii. The objective of this work was to develop a rapid and definitive DNA sequence-based identification system for unknown isolates of the Lact. acidophilus complex. A approximately = 500 bp region of the 16S rRNA gene, which contained the V1 and V2 variable regions, was amplified from the isolates by the polymerase chain reaction. The sequence of this region of the 16S rRNA gene from the type strains of the Lact. acidophilus complex was sufficiently variable to allow for clear differentiation amongst each of the strains. As an initial step in the characterization of potentially probiotic strains, this technique was successfully used to identify a variety of unknown human intestinal isolates. The approach described here represents a rapid and definitive method for the identification of Lact. acidophilus complex members.

Base Sequence↗

Homology modeling of an RNP domain from a human RNA-binding protein: Homology-constrained energy optimization provides a criterion for distinguishing potential sequence alignments.

We have recently described an automated approach for homology modeling using restrained molecular dynamics and simulated annealing procedures (Li et al, Protein Sci., 6:956-970,1997). We have employed this approach for constructing a homology model of the putative RNA-binding domain of the human RNA-binding protein with multiple splice sites (RBP-MS). The regions of RBP-MS which are homologous to the template protein snRNP U1A were constrained by "homology distance constraints," while the conformation of the non-homologous regions were defined only by a potential energy function. A full energy function without explicit solvent was employed to ensure that the calculated structures have good conformational energies and are physically reasonable. The effects of mis-alignment of the unknown and the template sequences were also explored in order to determine the feasibility of this homology modeling method for distinguishing possible sequence alignments based on considerations of the resulting conformational energies of modeled structures. Differences in the alignments of the unknown and the template sequences result in significant differences in the conformational energies of the calculated homology models. These results suggest that conformational energies and residual constraint violations in these homology-constrained simulated annealing calculations can be used as criteria to distinguish between correct and incorrect sequence alignments and chain folds.

Algorithms↗

Detection of Ralstonia solanacearum, which causes brown rot of potato, by fluorescent in situ hybridization with 23S rRNA-targeted probes.

During the past few years, Ralstonia (Pseudomonas) solanacearum race 3, biovar 2, was repeatedly found in potatoes in Western Europe. To detect this bacterium in potato tissue samples, we developed a method based on fluorescent in situ hybridization (FISH). The nearly complete genes encoding 23S rRNA of five R. solanacearum strains and one Ralstonia pickettii strain were PCR amplified, sequenced, and analyzed by sequence alignment. This resulted in the construction of an unrooted tree and supported previous conclusions based on 16S rRNA sequence comparison in which R. solanacearum strains are subdivided into two clusters. Based on the alignments, two specific probes, RSOLA and RSOLB, were designed for R. solanacearum and the closely related Ralstonia syzygii and blood disease bacterium. The specificity of the probes was demonstrated by dot blot hybridization with RNA extracted from 88 bacterial strains. Probe RSOLB was successfully applied in FISH detection with pure cultures and potato tissue samples, showing a strong fluorescent signal. Unexpectedly, probe RSOLA gave a less intense signal with target cells. Potato samples are currently screened by indirect immunofluorescence (IIF). By simultaneously applying IIF and the developed specific FISH, two independent targets for identification of R. solanacearum are combined, resulting in a rapid (1-day), accurate identification of the undesired pathogen. The significance of the method was validated by detecting the pathogen in soil and water samples and root tissue of the weed host Solanum dulcamara (bittersweet) in contaminated areas.

Base Sequence↗

A major component approach to presenting consensus sequences.

MOTIVATION: Summarizing and displaying the information contained in a set of aligned sequences is an important aid to identifying patterns within the sequences. A variety of forms of consensus sequences have been used previously to provide this information. However, these methods can cause a loss of information or introduce ambiguities into the consensus sequence, and some graphical approaches may become difficult to interpret due to visual distortion. RESULTS: We have developed a method to present a more precise and graphically clear view of a consensus sequence by using an approach based on defining the major components at each position in a sequence set. The major components are given in an ordered list and their frequencies are shown as histograms which can be colour coded to reflect conservative groupings. Minor components, a one-line character-based consensus sequence and information statistics can also be presented. As well as identifying the dominant sources of variation and conservation in the sequence set, the method also enables similarities and differences between subgroups of a sequence set to be readily assessed. AVAILIABILITY: On request from the authors. CONTACT: bcsmith@usthk.ust.hk, hxue@usthk. ust.hk

Amino Acid Sequence↗

Partial sequence identification of grapevine-leafroll-associated virus-1 and development of a highly sensitive IC-RT-PCR detection method.

Using immunocapture reverse transcription PCR (IC-RT-PCR) a specific PCR product from GLRaV-1 infected vine samples was amplified with the help of degenerate primers deduced from the conserved HSP70 region of closteroviruses. 511 basepairs of the 5'end of GLRaV-1 HSP70 gene were identified. Within this region, putative GLRaV-1 specific primers were designed and an IC-RT-PCR detection procedure was developed which is about 125 times more sensitive than the established ELISA method. No PCR product was amplified in GLRaV-2,-3 and -4 infected plants which indicates the specificity of the primers. This procedure may serve as an alternative method for GLRaV-1 detection where the sensitivity of ELISA is insufficient.

Base Sequence↗

Strain characterization and classification of oxyphotobacteria in clone cultures on the basis of 16S rRNA sequences from the variable regions V6, V7, and V8.

A major problem in development of a polyphasic taxonomy is that the identification of oxyphotobacterial strains (cyanobacteria and prochlorophytes) in culture collections may be incorrect. We have therefore developed a diagnostic system using the DNA sequence polymorphism in the 16S rRNA regions V6 to V8 for individual strain characterization and identification. PCR primers amplifying V6 to V8 from oxyphotobacteria in unialgal cultures were constructed. Direct solid-phase or cyclic sequencing was used to determine the sequences from the amplified DNA. This survey includes 10 strains of Nostoc/Anabaena/Aphanizomenon (Nostoc category), 5 strains of Microcystis (Microcystis category), and 4 strains of Planktothrix (Planktothrix category). Fifteen additional strains of cyanobacteria and two strains of prochlorophytes were included such that the major phyletic groups were represented. One of the strains, Phormidium sp. NIVA-CYA 203, contained an 11-nucleotide insertion with no homology to other known 16S rRNA sequences. Based on parsimony and neighbor-joining trees, the phyletic relationships of the strains were investigated. Thirteen major branches were found, with Pseudanabaena limnetica NIVA-CYA 276/6 as the most divergent strain. The strain categories Nostoc, Planktothrix, and Microcystis were all monophyletic. The sequence polymorphism within Nostoc was higher than that in Planktothrix and Microcystis. Based on the sequence and phyletic information, group-specific PCR primers for the categories Nostoc, Planktothrix, and Microcystis were constructed. For the strains included in this work, the amplifications were specific for the relevant groups. By combination of magnetic solid-phase DNA isolation and group-specific PCR amplifications, an accurate method for characterization, classification and identification of oxyphotobacterial clone cultures has been developed.

Base Sequence↗

Differentiation of phylogenetically related slowly growing mycobacteria by their gyrB sequences.

The conventional methods for identifying mycobacterial species are based on their phenotypic characterization. Since some problematic species are slow growers, their taxonomy takes several weeks or months to identify. The ribosomal DNA (rDNA) sequence-based identification strategy has been adopted to solve this problem. More recently, the gyrB sequences have been shown to be useful phylogenetic markers for the identification of species. We determined the gyrB sequences of 43 slowly growing strains belonging to 15 species in the genus Mycobacterium. The frequencies of base substitutions in the gyrB sequences were comparable to those in the 16S-23S rDNA internal transcribed spacer (ITS) sequences. The ITS sequences of four species belonging to the M. tuberculosis complex (M. tuberculosis, M. bovis, M. africanum, and M. microti) were 100% identical, while four synonymous substitutions were found in the gyrB sequences of these strains. Based on the differences found in the gyrB sequences, we developed PCR and PCR-restriction fragment length polymorphism methods to discriminate these species.

Animals↗

Development of amplified consensus genetic markers (ACGM) in Brassica napus from Arabidopsis thaliana sequences of known biological function.

A method for the development of consensus genetic markers between species of the same taxonomic family is described in this paper. It is based on the conservation of the peptide sequences and on the potential polymorphism within non-coding sequences. Six loci sequenced from Arabidopsis thaliana, AG, LFY3, AP3, FAD7, FAD3, and ADH, were analysed for one ecotype of A. thaliana, four lines of Brassica napus, and one line for each parental species, Brassica oleracea and Brassica rapa. Positive amplifications with the degenerate primers showed one band for A. thaliana, two to four bands in rapeseed, and one to two bands in the parental species. Direct sequencing of the PCR products confirms their peptide similarity with the "mother" sequence. By comparison of intron sequences, the correspondence between each rapeseed gene and its homologue in one of the parental species can be determined without ambiguity. Another important result is the presence of a polymorphism inside these fragments between the rapeseed lines. This variability could generally be detected by differences of electrophoretic migration on long non-denaturing polyacrylamide gels. This method enables a quick and easy shuttle between A. thaliana and Brassica species without cloning.

Amino Acid Sequence↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

Protein fold recognition by total alignment probability.

We present a protein fold-recognition method that uses a comprehensive statistical interpretation of structural Hidden Markov Models (HMMs). The structure/fold recognition is done by summing the probabilities of all sequence-to-structure alignments. The optimal alignment can be defined as the most probable, but suboptimal alignments may have comparable probabilities. These suboptimal alignments can be interpreted as optimal alignments to the "other" structures from the ensemble or optimal alignments under minor fluctuations in the scoring function. Summing probabilities for all alignments gives a complete estimate of sequence-model compatibility. In the case of HMMs that produce a sequence, this reflects the fact that due to our indifference to exactly how the HMM produced the sequence, we should sum over all possibilities. We have built a set of structural HMMs for 188 protein structures and have compared two methods for identifying the structure compatible with a sequence: by the optimal alignment probability and by the total probability. Fold recognition by total probability was 40% more accurate than fold recognition by the optimal alignment probability. Proteins 2000;40:451-462.

Algorithms↗

Identification and assessment of known and novel human papillomaviruses by polymerase chain reaction amplification, restriction fragment length polymorphisms, nucleotide sequence, and phylogenetic algorithms.

The identification and taxonomy of papillomaviruses has become increasingly complex, as approximately 70 human papillomavirus (HPV) types have been described and novel HPV genomes continue to be identified. Methods and corresponding DNA sequence data bases were designed for the reliable identification of mucosal HPV genomes from clinical specimens. HPVs are identified by the amplification of a fragment of the L1 region by consensus primer polymerase chain reaction (PCR) and subsequent hybridization or restriction fragment length polymorphism analysis. L1 PCR fragments may be further characterized by nucleotide sequencing. Conservation of 30 (of 151) predicted amino acids identifies HPV genomic fragments, and nucleotide sequence alignments allow calculation of their phylogenetic relatedness. Sequence differences > 10% from any known HPV type suggest a novel HPV type. Phylogenetic relationships with known HPV types may permit predictions of biology. With these criteria, 10 PCR fragments were identified that would qualify as new genital HPV types after complete genomic isolation.

Amino Acid Sequence↗

The complete amino acid sequence of momordin-a, a ribosome-inactivating protein from the seeds of bitter gourd (Momordica charantia).

The complete amino acid sequence of momordin-a, a ribosome-inactivating protein from the seeds of bitter gourd, has been analyzed. Twenty-two peptides were isolated from the tryptic digest of momordin-a and sequenced by the DABITC/PITC double coupling method. The alignment of these tryptic peptides was done by analyzing the amino acid sequences of the peptides derived from chymotryptic digestion and cyanogen bromide cleavage of momordin-a as well as V8 protease-digestion of the CNBr fragment. Momordin-a consisted of 250 amino acid residues and carbohydrate residues attached to Asn227, and its molecular mass was calculated to be 28,690 Da. The sequence comparison with ricin A-chain shows that 33% of the residues of momordin-a are identical to those of ricin A-chain and that the residues involved in the catalytic site of the ricin A-chain are conserved in momordin-a.

Amino Acid Sequence↗

Molecular analysis of ependymins from the cerebrospinal fluid of the orders Clupeiformes and Salmoniformes: no indication for the existence of an euteleost infradivision.

Ependymins represent the predominant protein constituents in the cerebrospinal fluid of many teleost fish and they are synthesized in meningeal fibroblasts. Here, we present the ependymin sequences from the herring (Clupea harengus) and the pike (Esox lucius). A comparison of ependymin homologous sequences from three different orders of teleost fish (Salmoniformes, Cypriniformes, and Clupeiformes) revealed the highest similarity between Clupeiformes and Cypriniformes. This result is unexpected because it does not reflect current systematics, in which Clupeiformes belong to a separate infradivision (Clupeomorpha) than Salmoniformes and Cypriniformes (Euteleostei). Furthermore, in Salmoniformes the evolutionary rate of ependymins seems to be accelerated mainly on the protein level. However, considering these inconstant rates, neither neighbor-joining trees nor DNA parsimony methods gave any indication that a separate euteleost infradivision exists.

Amino Acid Sequence↗

Partial DNA cloning and sequencing of a canine parvovirus vaccine strain: application of nucleic acid hybridization to the diagnosis of canine parvovirus disease.

The cloning and sequencing of an Eco RI-PstI fragment derived from the replicative form of a canine parvovirus (CPV) vaccine strain are reported. The variability of the 5' end of NS 1 protein gene in the genome is confirmed by comparison with previously determined DNA sequences. A 15 nucleotide deletion was also observed in this vaccine strain. In order to improve CPV diagnosis, radioactively labelled RNA or DNA and biotin labelled DNA obtained by random priming of the recombinant plasmid were used as probes mainly on gut or stool samples from naturally infected dogs. Results of filter hybridization correlated well with histopathological diagnosis of parvovirus infection and with hemagglutination tests performed on dog faeces. We propose that nucleic acid hybridization may be an alternative diagnostic method to ascertain the presence of CPV, especially in frozen samples.

Amino Acid Sequence↗

Microbial gene identification using interpolated Markov models.

This paper describes a new system, GLIMMER, for finding genes in microbial genomes. In a series of tests on Haemophilus influenzae , Helicobacter pylori and other complete microbial genomes, this system has proven to be very accurate at locating virtually all the genes in these sequences, outperforming previous methods. A conservative estimate based on experiments on H.pylori and H. influenzae is that the system finds >97% of all genes. GLIMMER uses interpolated Markov models (IMMs) as a framework for capturing dependencies between nearby nucleotides in a DNA sequence. An IMM-based method makes predictions based on a variable context; i.e., a variable-length oligomer in a DNA sequence. The context used by GLIMMER changes depending on the local composition of the sequence. As a result, GLIMMER is more flexible and more powerful than fixed-order Markov methods, which have previously been the primary content-based technique for finding genes in microbial DNA.

Algorithms↗

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence↗

Homology modelling by distance geometry.

BACKGROUND: Unknown protein structures can be predicted from known structures (the scaffolds) with sequences sufficiently homologous to that of the target, based on the observation that similar sequences usually adopt the same fold. When structural equivalences between residues in the scaffold and target proteins are expressed in terms of conserved interatomic distances, the resulting 'distance geometry' representation provides an elegant mechanism for simultaneous restraint satisfaction and bias-free conformation space exploration. RESULTS: We present a homology modelling algorithm based on distance geometry that relies on the gradual projection of simple model chain coordinates into Euclidean spaces with decreasing dimensionality. The similarity between the unknown target structure and the scaffold proteins with known structures was described by mapping secondary structure assignments and specific distance restraints between C alpha atoms onto the model through a multiple alignment. This information was complemented by additional restraints derived from stereochemical considerations and other general aspects of protein structure such as hydrophobic core formation or the absence of tangled mainchains. CONCLUSIONS: The method was capable of quickly locating the correct fold even from an alignment with modest average conservation indicating that it could serve as a fast tool for obtaining correct low-resolution starting conformations for detailed refinement.

Amino Acid Sequence↗