PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

A combined approach for locating box H/ACA snoRNAs in the human genome.

A novel combined method for locating box H/ACA small nucleolar RNAs (snoRNAs) is described, together with a software tool. The method adopts both a probabilistic hidden Markov model (HMM) and a minimum free energy (MFE) rule, and filters possible candidate box H/ACA snoRNAs obtained from genomic DNA sequences. With our novel method 12 known box H/ACA snoRNAs, and one strong candidate were identified in 30 nucleolar protein genomic sequences.

Algorithms↗

Group II nucleopolyhedrovirus subgroups revealed by phylogenetic analysis of polyhedrin and DNA polymerase gene sequences.

Two major clades, designated Groups I and II, of nucleopolyhedroviruses (NPVs) from lepidopteran hosts have been previously identified. To reveal more detailed relationships, a series of DNA polymerase nucleotide sequences from the taxa MbMNPV, SeMNPV, HzSNPV, HearNPV, SpltNPV, BusuNPV, and OranNPV have been determined using a polymerase chain reaction (PCR)-based approach. This technique enabled gene sequence determination using microliter samples of NPV-infected insect cadavers. Polyhedrin genes from HearNPV, OranNPV, SeMNPV, and SpltNPV were also isolated and sequenced using a similar approach. These sequences, together with other database entries, were aligned for positional homology of peptide sequences. Phylogenetic analysis of DNA polymerase molecular sequence alignments supports LdMNPV as a taxon of Group II and three Group II subclades, designated A, B, and C. Comparison of DNA polymerase trees with those estimated from occlusion protein molecular sequences enabled identification of three subclades of Group II. These are Subgroup II-A [MbMNPV, LeseNPV, MacoNPV, PaflNPV, SeMNPV, SpltNPV (India isolate), SfMNPV]; Subgroup II-B [SpliNPV, SpltNPV (Japan isolate), SpltNPV (Queensland isolate), and possibly HzSNPV, HearNPV, and ManeNPV], and Subgroup II-C [OpSNPV, OranNPV (S-type), BusuNPV (S-type), and possibly EcobNPV (S-type)]. Notably, all Subgroup II-A taxa are from noctuid hosts. Correlations of virus and host evolution within Group II taxa are discussed. The methods and data developed in this study will allow rapid sequencing of NPV DNA polymerase genes.

Amino Acid Sequence↗

Systematic and fully automated identification of protein sequence patterns.

We present an efficient algorithm to systematically and automatically identify patterns in protein sequence families. The procedure is based on the Splash deterministic pattern discovery algorithm and on a framework to assess the statistical significance of patterns. We demonstrate its application to the fully automated discovery of patterns in 974 PROSITE families (the complete subset of PROSITE families which are defined by patterns and contain DR records). Splash generates patterns with better specificity and undiminished sensitivity, or vice versa, in 28% of the families; identical statistics were obtained in 48% of the families, worse statistics in 15%, and mixed behavior in the remaining 9%. In about 75% of the cases, Splash patterns identify sequence sites that overlap more than 50% with the corresponding PROSITE pattern. The procedure is sufficiently rapid to enable its use for daily curation of existing motif and profile databases. Third, our results show that the statistical significance of discovered patterns correlates well with their biological significance. The trypsin subfamily of serine proteases is used to illustrate this method's ability to exhaustively discover all motifs in a family that are statistically and biologically significant. Finally, we discuss applications of sequence patterns to multiple sequence alignment and the training of more sensitive score-based motif models, akin to the procedure used by PSI-BLAST. All results are available at httpl//www.research.ibm.com/spat/.

Algorithms↗

Hidden Markov models for detecting remote protein homologies.

MOTIVATION: A new hidden Markov model method (SAM-T98) for finding remote homologs of protein sequences is described and evaluated. The method begins with a single target sequence and iteratively builds a hidden Markov model (HMM) from the sequence and homologs found using the HMM for database search. SAM-T98 is also used to construct model libraries automatically from sequences in structural databases. METHODS: We evaluate the SAM-T98 method with four datasets. Three of the test sets are fold-recognition tests, where the correct answers are determined by structural similarity. The fourth uses a curated database. The method is compared against WU-BLASTP and against DOUBLE-BLAST, a two-step method similar to ISS, but using BLAST instead of FASTA. RESULTS: SAM-T98 had the fewest errors in all tests-dramatically so for the fold-recognition tests. At the minimum-error point on the SCOP (Structural Classification of Proteins)-domains test, SAM-T98 got 880 true positives and 68 false positives, DOUBLE-BLAST got 533 true positives with 71 false positives, and WU-BLASTP got 353 true positives with 24 false positives. The method is optimized to recognize superfamilies, and would require parameter adjustment to be used to find family or fold relationships. One key to the performance of the HMM method is a new score-normalization technique that compares the score to the score with a reversed model rather than to a uniform null model. AVAILABILITY: A World Wide Web server, as well as information on obtaining the Sequence Alignment and Modeling (SAM) software suite, can be found at http://www.cse.ucsc.edu/research/compbi o/ CONTACT: karplus@cse.ucsc.edu; http://www.cse.ucsc.edu/karplus

Algorithms↗

Rate matrices for analyzing large families of protein sequences.

We propose and study a new approach for the analysis of families of protein sequences. This method is related to the LogDet distances used in phylogenetic reconstructions; it can be viewed as an attempt to embed these distances into a multidimensional framework. The proposed method starts by associating a Markov matrix to each pairwise alignment deduced from a given multiple alignment. The central objects under consideration here are matrix-valued logarithms L of these Markov matrices, which exist under conditions that are compatible with fairly large divergence between the sequences. These logarithms allow us to compare data from a family of aligned proteins with simple models (in particular, continuous reversible Markov models) and to test the adequacy of such models. If one neglects fluctuations arising from the finite length of sequences, any continuous reversible Markov model with a single rate matrix Q over an arbitrary tree predicts that all the observed matrices L are multiples of Q. Our method exploits this fact, without relying on any tree estimation. We test this prediction on a family of proteins encoded by the mitochondrial genome of 26 multicellular animals, which include vertebrates, arthropods, echinoderms, molluscs, and nematodes. A principal component analysis of the observed matrices L shows that a single rate model can be used as a rough approximation to the data, but that systematic deviations from any such model are unmistakable and related to the evolutionary history of the species under consideration.

Computational Biology↗

Amplification of ribosomal DNA of Anoplocephalidae: Anoplocephala perfoliata diagnosis by PCR as a possible alternative to coprological methods.

The diagnosis of tapeworm infections in horses relies on copro-diagnostic methods, which are time-consuming and of limited sensitivity for determination of the exact prevalence. The development of serological tests has slightly improved the detection of tapeworm infections, but more sensitive methods are still required. A polymerase chain reaction (PCR)-based approach may constitute a valuable tool to improve tapeworm diagnosis. Nuclear ribosomal DNA (rDNA) is a useful target for species and/or strain markers. Partial 18S, the internal transcribed spacer 1 (ITS-1), the 5.8S, the internal transcribed spacer 2 (ITS-2), and partial 28S rDNA of the equine tapeworms Anoplocephala perfoliata and Anoplocephaloides mamillana were amplified and sequenced. The lengths and GC contents of the regions sequenced were 2087-2091bp and 49.35-49.69% for A. perfoliata, and 2110-2119bp and 49.15-49.32% for A. mamillana, respectively. Sequence alignment and comparison of both taxa showed 79.3-80.2% identity. The lowest identities were found in the ITS regions with 39.9-43.5% for the ITS-1 and 59.5-61.2% for the ITS-2. No matches of the ITS-2 of A. perfoliata and A. mamillana were found with other species by BLAST search. For this reason, ITS-2 sequences seemed appropriate as accurate species markers and A. perfoliata ITS-2 primers were developed. The ITS-2 PCR enabled the detection of genomic DNA as low as 0.5 pgs. First efforts on the practical application of the PCR-based approach were made. A 6-mg fragment of a tapeworm proglottid was detected in 0.5 and 1g of faeces.

Animals↗

flaA-like sequences containing internal termination codons (TAG) in urease-positive thermophilic Campylobacter isolated in Japan.

AIMS: To demonstrate two flaA-like sequences containing two internal termination codons (TAG) in two Japanese strains of urease-positive thermophilic Campylobacter (UPTC). METHODS AND RESULTS: A primer pair of A1 and A2, which ought to generate a product of approx. 1700 bp of the flaA gene for Campylobacter jejuni, was used to amplify products of approx. 1450 bp for two Japanese strains of UPTC, CF89-12 and CF89-14. After molecular cloning and sequencing, the nucleotide sequences of the amplicons from the two strains were found to be 1461 bp in length and to have nucleotide sequence differences in relation to each other at four nucleotide positions, respectively. CONCLUSIONS: Nucleotide and amino acid sequence alignment and homology analysis demonstrated that the polymerase chain reaction (PCR) amplicons from the two Japanese strains have approx. 83% nucleotide and 80% amino acid sequence homology to the possible open reading frame of the flaA gene of UPTC NCTC 12892. SIGNIFICANCE AND IMPACT OF THE STUDY: Surprisingly, both PCR amplicons from the Japanese UPTC have two internal termination codons (TAG) at nucleotide positions from 775 to 777 and 817 to 819, respectively.

Amino Acid Sequence↗

Molecular phylogenetic studies on Brugia filariae using Hha I repeat sequences.

This paper is the first molecular phylogenetic study on Brugia parasites (family Onchocercidae) which includes 6 of the 10 species of this genus: B. beaveri Ash et Little, 1964; B. buckleyi Dissanaike et Paramananthan, 1961: B. malayi (Brug, 1927) Buckley, 1960; B. pahangi (Buckley et Edeson, 1956) Buckley, 1960; B. patei (Buckley, Nelson er Heisch, 1958) Buckley, 1960 and B. timori Partono et al., 1977. Hha l repeat sequences are 322 nucleotides long, highly repeated, tandemly arranged and unique to the nuclear genomes of the genus Brugia. Hha l repeat sequence data was collected by PCR, cloning and dideoxy sequencing. The Hha l repeat sequences were aligned and analyzed by maximum parsimony algorithms, distance methods and maximum likelihood methods to construct phylogenetic trees. Bootstrap analysis was used to test the robustness of the different phylogenetic reconstructions. The data indicated that the Hha l repeat sequences are highly conserved within species yet differ significantly between species. The various tree-building methods gave identical results. Bootstrap analyses on the Hha l repeat sequence data set identified at least two clades: the B. pahangi-B. beaveri clade and the B. malayi-B. timori-B. buckleyi clade; the first clade includes parasites of carnivores from Asia and America; the second includes species from primates and lagomorphs from Asiatic region. It was also noted that the Hha l repeat sequences obtained from B. malayi were identical to those obtained from B. timori, indicating very recent speciation.

Algorithms↗

Rapid differentiation of Fusarium oxysporum isolates using PCR-SSCP with the combination of pH-variable electrophoretic medium and low temperature.

Differentiation of Fusarium oxysporum is significantly important for unraveling the pathogenetic mechanism of Fusaria wilts. In this study, isolates of F. oxysporum were screened from the soils in the rhizosphere of watermelon plant by Komada medium and differentiated by SSCP approach with the combination of pH-variable electrophoretic medium (Tris-MES-EDTA (TME), pH 6.1) and low temperature (9 degrees C). We found that TME was a good electrophoretic medium and its pH value was variable over the course of electrophoresis in our apparatus. The pH-variable electrophoretic medium made more contribution for the better differentiation of F. oxysporum isolates than low temperature. The combination of TME pH 6.1 and low temperature showed an improved effect on resolution of ssDNAs. Leaving partial nondenatured dsDNA for SSCP was advantageous for differentiation of F. oxysporum isolates. The SSCP patterns of F. oxysporum isolates proved to be highly reproducible. Sequencing data confirmed that this SSCP method could detect one single base change within the 550 bp PCR fragment from the ribosomal internal transcribed spacer region of F. oxysporum.

Base Sequence↗

Protein topology recognition from secondary structure sequences: application of the hidden Markov models to the alpha class proteins.

The three-dimensional fold of a protein is described by the organization of its secondary structure elements in 3D space, i.e. its "topology". We find that the protein topology can be recognized from the ID sequence of secondary structure states of the residues alone. Automated recognition is facilitated by use of hidden Markov models (HMMs) to represent topology families of proteins. Such models can be trained on the experimentally observed secondary structure sequences of family members using well established algorithms. Here, we model various topology groups in the alpha class of proteins and identify, from a large database, those proteins having the topology described by each model. The correct topology family for protein secondary structure sequences could be recognized 12 out of 14 times. When the observed secondary structure sequences are replaced with predicted sequences recognition is still achievable 8 out of 14 times. The success rate for observed sequences indicates that our approach will become increasingly useful as the accuracy of secondary prediction algorithms is improved. Our study indicates that the HMMs are useful for protein topology recognition even when no detectable primary amino acid sequence similarity is present. To illustrate the potential utility of our method, protein topology recognition is attempted on leptin, the obese gene product, and the human interleukin-6 sequence, for which fold predictions have been previously published.

Algorithms↗

Sequence of 18S rDNA of actinorhizal Alnus glutinosa (Betulaceae).

The small subunit ribosomal DNA for a woody actinorhizal, Alnus glutinosa, was isolated by the PCR method. Amplification products were cloned into the Bluescript SK- vector. Full sequence, 1698 bp, was obtained with NS1 to NS8 primers. Sequence alignments were made by UWGCG sequence data analysis computer programs. 18S rDNA sequence of A. glutinosa was compared to analogous segments of four other angiosperms, tomato, rice, maize and soybean. Sequence homologies are discussed and application for the technique is suggested.

Base Sequence↗

Molecular genotyping of the murine H-2K MHC class I allele.

An increasing number of genetically modified mouse mutants are employed in immunological research. Many experiments require the crossbreeding of various mouse lines and screening for the genetic background of the offspring. Analysis for various immunological phenotypes such as the MHC class I haplotype is usually performed by FACS analysis of peripheral blood lymphocytes. Here, we report a PCR based technique for the differentiation of the MHC class I H-2K(k) and the H-2K(b) haplotype. Advantages of this method are cost efficiency and rapidity in particular when multiple genetic variables such as the disruption of a gene, transgenesis or the MHC haplotype are to be tested.

Alleles↗

Modeling amino acid replacement.

The estimation of amino acid replacement frequencies during molecular evolution is crucial for many applications in sequence analysis. Score matrices for database search programs or phylogenetic analysis rely on such models of protein evolution. Pioneering work was done by Dayhoff et al. (1978) who formulated a Markov model of evolution and derived the famous PAM score matrices. Her estimation procedure for amino acid exchange frequencies is restricted to pairs of proteins that have a constant and small degree of divergence. Here we present an improved estimator, called the resolvent method, that is not subject to these limitations. This extension of Dayhoff's approach enables us to estimate an amino acid substitution model from alignments of varying degree of divergence. Extensive simulations show the capability of the new estimator to recover accurately the exchange frequencies among amino acids. Based on the SYSTERS database of aligned protein families (Krause and Vingron, 1998) we recompute a series of score matrices.

Amino Acid Substitution↗

Identification and classification of protein fold families.

We have developed a method for identifying fold families in the protein structure data bank. Pairwise sequence alignments are first performed to extract families of homologous proteins having 35% or more sequence identity. Representatives are selected with the best resolution and R-factor to give a nonhomologous data set. Subsequent structure comparisons between all members of this set detect homologous folds with low sequence identity but highly conserved structures. By softening the requirement on structural similarity, families of analogous proteins are obtained that have related folds but more diverse structures. Representatives are selected to give a non-analogous data set. Starting with 1410 chains from the Brookhaven Data Bank, we generate a set of 150 nonhomologous folds and a set of 112 non-analogous folds. Analysis of sequence and structure conservation within the larger families shows the globins to be the most highly conserved family and the TIM barrels the most weakly conserved.

Classification↗

Specific detection of Plesiomonas shigelloides isolated from aquatic environments, animals and human diarrhoeal cases by PCR based on 23S rRNA gene.

Twenty-five strains of Plesiomonas shigelloides isolated from aquatic environment, 10 strains from human cases of diarrhoea and five strains from animals were identified by the polymerase chain reaction technique based on 23S rRNA gene. For this purpose, two primers targeted against part of the 5' half of the 23S rRNA gene of P. shigelloides (Escherichia coli number C-912, G-1195; Plesiomonas number C-906, G-1189) were designed. Results from our study indicated that this method might serve as a tool for a rapid and sensitive identification of P. shigelloides from different environmental and clinical sources.

Animals↗

A real-time PCR-based method to independently sample single simian immunodeficiency virus genomes from macaques with a range of viral loads.

The generation of a diverse population of viral variants is a hallmark of simian immunodeficiency virus (SIV) infection. In order to address what role this diversity plays in disease progression, accurate sampling of the viral population is necessary. However, traditional PCR-based methods often rely on amplification of multiple genomes in one reaction, leading to resampling of viral genomes and potential errors in the estimations of viral diversity, especially when sequences from only one or a small number of PCRs are examined and/or viral copy number is low. Here we describe a method to amplify one viral envelope gene per PCR, thereby avoiding resampling. For this purpose we developed a highly accurate real-time PCR method to quantify SIV copy number, then used a single SIV template in a sensitive, high-fidelity full-length envelope PCR. Using this method, we have estimated the intra-animal viral diversity for a cohort of five pig-tailed macaques (Macaca nemestrina) infected with SIVMne variants, which displayed a broad range of viral loads at setpoint.

Amino Acid Sequence↗

Cloning of the koi herpesvirus (KHV) gene encoding thymidine kinase and its use for a highly sensitive PCR based diagnosis.

BACKGROUND: Outbreaks with mass mortality among common carp Cyprinus carpio carpio and koi Cyprinus carpio koi have occurred worldwide since 1998. The herpes-like virus isolated from diseased fish is different from Herpesvirus cyprini and channel catfish virus and was accordingly designated koi herpesvirus (KHV). Diagnosis of KHV infection based on viral isolation and current PCR assays has a limited sensitivity and therefore new tools for the diagnosis of KHV infections are necessary. RESULTS: A robust and sensitive PCR assay based on a defined gene sequence of KHV was developed to improve the diagnosis of KHV infection. From a KHV genomic library, a hypothetical thymidine kinase gene (TK) was identified, subcloned and expressed as a recombinant protein. Preliminary characterization of the recombinant TK showed that it has a kinase activity using dTTP but not dCTP as a substrate. A PCR assay based on primers selected from the defined DNA sequence of the TK gene was developed and resulted in a 409 bp amplified fragment. The TK based PCR assay did not amplify the DNAs of other fish herpesviruses such as Herpesvirus cyprini (CHV) and the channel catfish virus (CCV). The TK based PCR assay was specific for the detection of KHV and was able to detect as little as 10 fentograms of KHV DNA corresponding to 30 virions. The TK based PCR was compared to previously described PCR assays and to viral culture in diseased fish and was shown to be the most sensitive method of diagnosis of KHV infection. CONCLUSION: The TK based PCR assay developed in this work was shown to be specific for the detection of KHV. The TK based PCR assay was more sensitive for the detection of KHV than previously described PCR assays; it was as sensitive as virus isolation which is the golden standard method for KHV diagnosis and was able to detect as little as 10 fentograms of KHV DNA corresponding to 30 virions.

Amino Acid Sequence↗

The ITS2 ribosomal DNA of Anopheles beklemishevi and further remarks on the phylogenetic relationships within the Anopheles maculipennis group of species (Diptera: Culicidae).

Anopheles beklemishevi specimens from Russia were analysed by their ITS2 ribosomal DNA sequence to amend and to specify the phylogenetic tree of the Anopheles maculipennis species complex. Surprisingly, with 638 base pairs, the ITS2 regions of all the 34 An beklemishevi specimens examined were considerably longer than those of all their sibling species. Sequence alignment with GenBank derived sequences of the other siblings was only possible in the beginning (for approx. 335 bp) and at the end (for approx. 150 bp) of the PCR-amplified DNA fragment, whereas in the middle, the An beklemishevi DNA sequence found no counterpart in sequences of the other siblings. Closer analysis of this intermediate part suggests a duplicated insertion of about 140 bp that has undergone subsequent mutational changes. Due to this large putative insertion, computerized phylogenetic analysis by the Bayesian inference method locates An beklemishevi in a closer relationship to the nearctic than to the palaearctic sibling species. However, when only ITS2 regions are compared, that have corresponding sequences in the other siblings, An beklemishevi forms a lineage with the palaearctic species although it is still most remotely related. It is hypothesized that during the evolution An beklemishevi separated first from the common ancestor of the palaearctic species, which had presumably made its way from the Nearctic to the Palaearctic.

Animals↗