PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Identification of Salmonella enterica serovar Typhimurium using specific PCR primers obtained by comparative genomics in Salmonella serovars.

Salmonella enterica serovar Typhimurium is a major foodborne pathogen throughout the world. Until now, the specific target genes for the detection and identification of serovar Typhimurium have not been developed. To determine the specific probes for serovar Typhimurium, the genes of serovar Typhimurium LT2 that were expected to be unique were selected with the BLAST (Basic Local Alignment Search Tool) program within GenBank. The selected genes were compared with 11 genomic sequences of various Salmonella serovars by BLAST. Of these selected genes, 10 were expected to be specific to serovar Typhimurium and were not related to virulence factor genes of Salmonella pathogenicity island or to genes of the O and H antigens of Salmonella. Primers for the 10 selected genes were constructed, and PCRs were evaluated with various genomic DNAs of Salmonella and non-Salmonella strains for the specific identification of Salmonella serovar Typhimurium. Among all the primer sets for the 10 genes, STM4497 showed the highest degree of specificity to serovar Typhimurium. In this study, a specific primer set for Salmonella serovar Typhimurium was developed on the basis of the comparison of genomic sequences between Salmonella serovars and was validated with PCR. This method of comparative genomics to select target genes or sequences can be applied to the specific detection of microorganisms.

Base Sequence↗

Molecular analysis of ependymins from the cerebrospinal fluid of the orders Clupeiformes and Salmoniformes: no indication for the existence of an euteleost infradivision.

Ependymins represent the predominant protein constituents in the cerebrospinal fluid of many teleost fish and they are synthesized in meningeal fibroblasts. Here, we present the ependymin sequences from the herring (Clupea harengus) and the pike (Esox lucius). A comparison of ependymin homologous sequences from three different orders of teleost fish (Salmoniformes, Cypriniformes, and Clupeiformes) revealed the highest similarity between Clupeiformes and Cypriniformes. This result is unexpected because it does not reflect current systematics, in which Clupeiformes belong to a separate infradivision (Clupeomorpha) than Salmoniformes and Cypriniformes (Euteleostei). Furthermore, in Salmoniformes the evolutionary rate of ependymins seems to be accelerated mainly on the protein level. However, considering these inconstant rates, neither neighbor-joining trees nor DNA parsimony methods gave any indication that a separate euteleost infradivision exists.

Amino Acid Sequence↗

Partial DNA cloning and sequencing of a canine parvovirus vaccine strain: application of nucleic acid hybridization to the diagnosis of canine parvovirus disease.

The cloning and sequencing of an Eco RI-PstI fragment derived from the replicative form of a canine parvovirus (CPV) vaccine strain are reported. The variability of the 5' end of NS 1 protein gene in the genome is confirmed by comparison with previously determined DNA sequences. A 15 nucleotide deletion was also observed in this vaccine strain. In order to improve CPV diagnosis, radioactively labelled RNA or DNA and biotin labelled DNA obtained by random priming of the recombinant plasmid were used as probes mainly on gut or stool samples from naturally infected dogs. Results of filter hybridization correlated well with histopathological diagnosis of parvovirus infection and with hemagglutination tests performed on dog faeces. We propose that nucleic acid hybridization may be an alternative diagnostic method to ascertain the presence of CPV, especially in frozen samples.

Amino Acid Sequence↗

A molecular method to detect Bacillus cereus from a coffee concentrate sample used in industrial preparations.

AIMS: The aim of this work was to develop specific primers which are able to detect Bacillus cereus in a coffee concentrate sample. METHODS AND RESULTS: A pre-PCR step to clean the DNA, used for PCR, was developed to avoid PCR inhibition by Maillard products. The combination of centrifugation and washing the pellet, employing EDTA and water, before DNA extraction improved the detection of low numbers of B. cereus cells (10 cells ml-1). The development of specific primers enabled to detect low numbers of B. cereus without the need of a pre-enrichment step. CONCLUSIONS: The data obtained demonstrated the specificity and the sensitivity of the primers that could be used to check the presence of B. cereus in different food products, avoiding the need for labourious and time-consuming culture-based techniques. SIGNIFICANCE AND IMPACT OF THE STUDY: The method could help food microbiologists to check food samples quickly for the presence of B. cereus.

Bacillus cereus↗

seq++: analyzing biological sequences with a range of Markov-related models.

SUMMARY: The seq++ package offers a reference set of programs and an extensible library to biologists and developers working on sequence statistics. Its generality arises from the ability to handle sequences described with any alphabet (nucleotides, amino acids, codons and others). seq++ enables sequence modelling with various types of Markov models, including variable length Markov models and the newly developed parsimonious Markov models, all of them potentially phased. Simulation modules are supplied for Monte Carlo methods. Hence, this toolbox allows the study of any biological process which can be described by a series of states taken from a finite set.

Algorithms↗

Microbial gene identification using interpolated Markov models.

This paper describes a new system, GLIMMER, for finding genes in microbial genomes. In a series of tests on Haemophilus influenzae , Helicobacter pylori and other complete microbial genomes, this system has proven to be very accurate at locating virtually all the genes in these sequences, outperforming previous methods. A conservative estimate based on experiments on H.pylori and H. influenzae is that the system finds >97% of all genes. GLIMMER uses interpolated Markov models (IMMs) as a framework for capturing dependencies between nearby nucleotides in a DNA sequence. An IMM-based method makes predictions based on a variable context; i.e., a variable-length oligomer in a DNA sequence. The context used by GLIMMER changes depending on the local composition of the sequence. As a result, GLIMMER is more flexible and more powerful than fixed-order Markov methods, which have previously been the primary content-based technique for finding genes in microbial DNA.

Algorithms↗

Characterization of infectious bronchitis virus isolates by slot blot hybridization.

We used slot blot hybridization of the hypervariable regions of the S1 subunit of spike peplomer gene to identify and characterize infectious bronchitis virus (IBV) strains. Template DNA was created from six reference strain IBVs of different serotypes and immobilized on a nitrocellulose membrane. We synthesized digoxigenin-labeled probes from reference and unknown field viruses and hybridized them to template DNA. All reference strains could be distinguished and isolates identified by serotype if they were at least 95% identical to a reference strain. This slot blot hybridization procedure was specific and reproducible, and strain typing was consistent with the S1 sequencing of the IBV genome. This study thus provides a simple and rapid method for typing of IBV.

Animals↗

Subfamily hmms in functional genomics.

The limitations of homology-based methods for prediction of protein molecular function are well known; differences in domain structure, gene duplication events and errors in existing database annotations complicate this process. In this paper we present a method to detect and model protein subfamilies, which can be used in high-throughput, genome-scale phylogenomic inference of protein function. We demonstrate the method on a set of nine PFAM families, and show that subfamily HMMs provide greater separation of homologs and non-homologs than is possible with a single HMM for each family. We also show that subfamily HMMs can be used for functional classification with a very low expected error rate. The BETE method for identifying functional subfamilies is illustrated on a set of serotonin receptors.

Animals↗

Playing with blocks: some pitfalls of forcing multiple alignments.

Block alignments of multiple amino acid sequences are useful representations of regions thought to share common ancestry and function. Often the block alignments are motivated by the expectation that a protein of interest is similar in function to members of a family of proteins. However, when alignments are forced by using ad hoc methods, it is often difficult to decide whether the proposed relationship is valid. Visual examination can be deceptive, especially when alignments are not carried out in the context of controls subjected to similar procedures. Even computer-aided methods can be misleading when biases are introduced. To illustrate some of the problems that can arise, a few examples from the literature are analyzed. It is concluded that when standard methods fail to find an interesting block alignment unaided by human intervention, then the result should be regarded with caution.

Amino Acid Sequence↗

State of the art: refinement of multiple sequence alignments.

BACKGROUND: Accurate multiple sequence alignments of proteins are very important in computational biology today. Despite the numerous efforts made in this field, all alignment strategies have certain shortcomings resulting in alignments that are not always correct. Refinement of existing alignment can prove to be an intelligent choice considering the increasing importance of high quality alignments in large scale high-throughput analysis. RESULTS: We provide an extensive comparison of the performance of the alignment refinement algorithms. The accuracy and efficiency of the refinement programs are compared using the 3D structure-based alignments in the BAliBASE benchmark database as well as manually curated high quality alignments from Conserved Domain Database (CDD). CONCLUSION: Comparison of performance for refined alignments revealed that despite the absence of dramatic improvements, our refinement method, REFINER, which uses conserved regions as constraints performs better in improving the alignments generated by different alignment algorithms. In most cases REFINER produces a higher-scoring, modestly improved alignment that does not deteriorate the well-conserved regions of the original alignment.

Algorithms↗

Homology modelling by distance geometry.

BACKGROUND: Unknown protein structures can be predicted from known structures (the scaffolds) with sequences sufficiently homologous to that of the target, based on the observation that similar sequences usually adopt the same fold. When structural equivalences between residues in the scaffold and target proteins are expressed in terms of conserved interatomic distances, the resulting 'distance geometry' representation provides an elegant mechanism for simultaneous restraint satisfaction and bias-free conformation space exploration. RESULTS: We present a homology modelling algorithm based on distance geometry that relies on the gradual projection of simple model chain coordinates into Euclidean spaces with decreasing dimensionality. The similarity between the unknown target structure and the scaffold proteins with known structures was described by mapping secondary structure assignments and specific distance restraints between C alpha atoms onto the model through a multiple alignment. This information was complemented by additional restraints derived from stereochemical considerations and other general aspects of protein structure such as hydrophobic core formation or the absence of tangled mainchains. CONCLUSIONS: The method was capable of quickly locating the correct fold even from an alignment with modest average conservation indicating that it could serve as a fast tool for obtaining correct low-resolution starting conformations for detailed refinement.

Amino Acid Sequence↗

Disulfide connectivity prediction using recursive neural networks and evolutionary information.

MOTIVATION: We focus on the prediction of disulfide bridges in proteins starting from their amino acid sequence and from the knowledge of the disulfide bonding state of each cysteine. The location of disulfide bridges is a structural feature that conveys important information about the protein main chain conformation and can therefore help towards the solution of the folding problem. Existing approaches based on weighted graph matching algorithms do not take advantage of evolutionary information. Recursive neural networks (RNN), on the other hand, can handle in a natural way complex data structures such as graphs whose vertices are labeled by real vectors, allowing us to incorporate multiple alignment profiles in the graphical representation of disulfide connectivity patterns. RESULTS: The core of the method is the use of machine learning tools to rank alternative disulfide connectivity patterns. We develop an ad-hoc RNN architecture for scoring labeled undirected graphs that represent connectivity patterns. In order to compare our algorithm with previous methods, we report experimental results on the SWISS-PROT 39 dataset. We find that using multiple alignment profiles allows us to obtain significant prediction accuracy improvements, clearly demonstrating the important role played by evolutionary information. AVAILABILITY: The Web interface of the predictor is available at http://neural.dsi.unifi.it/cysteines

Algorithms↗

Detection and identification of human pathogenic Leishmania and Trypanosoma species by hybridization of PCR-amplified mini-exon repeats.

A single pair of PCR primers within a conserved region of the mini-exon repeat was used to amplify the repeats from 10 species of pathogenic Leishmania belonging to four major clinical groups and also from three species of Trypanosoma. Oligonucleotide hybridization probes for the detection and identification of the PCR-amplified repeats were constructed from alignments of mini-exon intron and intergenic sequences. The probes generated from mini-exon intergenic regions of the L. (V.) braziliensis, L. (L.) donovani, and L. (L.) mexicana species hybridized specifically to their cognate groups without discriminating between the species within the groups. The probes for L. (L.) major and L. (L.) aethiopica were species-specific, while the L. (L.) tropica probe also hybridized with the L. (L.) aethiopica mini-exon repeat. The mini-exon intron-derived probes for T. cruzi, T. rangeli, and T. brucei were species-specific. This method involving the detection of specific PCR-amplified products produced using a single primer set represents a novel sensitive and specific assay for multiple trypanosomatid species and groups.

Animals↗

A quantitative structure-activity relationship and pharmacophore modeling investigation of aryl-X and heterocyclic bisphosphonates as bone resorption agents.

We have used quantitative structure-activity relationship (QSAR) techniques, together with pharmacophore modeling, to investigate the relationships between the structures of a wide variety of geminal bisphosphonates and their activity in inhibiting osteoclastic bone resorption. For aryl-X (X = alkyl, oxyalkyl, and sulfanylalkyl) derivatives of pamidronate and one alendronate, a molecular field analysis (MFA) yielded an R(2) value of 0.900 and an F-test of 54 for a training set of 29 compounds. Using reduced training sets, the activities of 20 such compounds were predicted with an average error of 2.1 over a 4000x range in activity. Such good results were only obtained when using the X-ray crystallographic structure of farnesyl pyrophosphate (FPP) bound to the target enzyme, farnesyl pyrophosphate synthase (FPP synthase), to guide the initial molecular alignment. For a series of heterocyclic bisphosphonates, use of the MFA method yielded an R(2) of 0.873 and an F-test of 36 for a training set of 26 compounds. Using a reduced training set, the activities of 20 compounds were predicted with an average error of 2.5 over a 2000x range in activity. With the heterocyclic compounds, test calculations indicated the importance of correct choice of protonation of the heterocyclic rings. For example, thiazoles, pyrazoles, and triazoles have low ( approximately 2-3) pK(a) values and the derived bisphosphonates are inactive in bone resorption since they cannot readily be side chain protonated and are thus poor carbocation reactive intermediate analogues. On the other hand, aminothiazoles, imidazoles, pyridyl, and aminopyridyl species typically have pK(a) values in the range approximately 5-9 and, in the absence of unfavorable steric interactions, the corresponding bisphosphonates are generally good inhibitors. However, aminoimidazole bisphosphonates are generally less active, since their pK(a)s ( approximately 11) are so high, due to guanidinium-like resonance, that they cannot readily be deprotonated, which we propose results in poor cellular uptake. The results of pharmacophore modeling using the Catalyst program revealed the importance of two negative ionizable and one positive charge feature for both aryl-X and heterocyclic pharmacophores, together with the presence of a distal hydrophobic feature in the aryl bisphosphonate and a more proximal aromatic feature in the heterocyclic bisphosphonate pharmacophores. When taken together, these results show that it is now possible to predict the activity, within a factor of about 2.3, of a wide range of aryl-X and heterocyclic bisphosphonates. The results emphasize the importance of utilizing crystallographic structural information to guide the initial alignment of extended bisphosphonates, and in the case of heterocyclic bisphosphonates, the importance of side chain protonation state. These simple ideas may facilitate the design of other, novel bisphosphonates, of use in bone resorption therapy, and as antiparasitic and immunotherapeutic agents.

Alendronate↗

Robust allele-specific polymerase chain reaction markers developed for single nucleotide polymorphisms in expressed barley sequences.

Many methods have been developed to assay for single nucleotide polymorphisms (SNPs), but generally these depend on access to specialised equipment. Allele-specific polymerase chain reaction (AS-PCR) is a method that does not require specialised equipment (other than a thermocycler), but there is a common perception that AS-PCR markers can be unreliable. We have utilised a three primer AS-PCR method comprising of two flanking-primers combined with an internal allele-specific primer. We show here that this method produces a high proportion of robust markers (from candidate allele specific primers). Forty-nine inter-varietal SNP sites in 31 barley (Hordeum vulgare L.) genes were targeted for the development of AS-PCR assays. The SNP sites were found by aligning barley expressed sequence tags from public databases. The targeted genes correspond to cDNAs that have been used as restriction fragment length polymorphic probes for linkage mapping in barley. Two approaches were adopted in developing the markers. In the first approach, designed to maximise the successful development of markers to a SNP site, markers were developed for 18 sites from 19 targeted (95% success rate). With the second approach, designed to maximise the number of markers developed per primer synthesised, markers were developed for 18 SNP sites from 30 that were targeted (a 60% success rate). The robustness of markers was assessed from the range of annealing temperatures over which the PCR assay was allele-specific. The results indicate that this form of AS-PCR is highly successful for the development of robust SNP markers.

Alleles↗

Stimulation of astaxanthin formation in the yeast Xanthophyllomyces dendrorhous by the fungus Epicoccum nigrum.

A fungal contaminant on an agar plate containing colonies of Xanthophyllomyces dendrorhous markedly increased carotenoid production by yeast colonies near to the fungal growth. Spent-culture filtrate from growth of the fungus in yeast-malt medium also stimulated carotenoid production by X. dendrorhous. Four X. dendrorhous strains including the wild-type UCD 67-385 (ATCC 24230), AF-1 (albino mutant, ATCC 96816), Yan-1 (beta-carotene mutant, ATCC 96815) and CAX (astaxanthin overproducer mutant) exposed to fungal concentrate extract enhanced astaxanthin up to approximately 40% per unit dry cell weight in the wild-type strain and in CAX. Interestingly, the fungal extract restored astaxanthin biosynthesis in non-astaxanthin-producing mutants previously isolated in our laboratory, including the albino and the beta-carotene mutant. The fungus was identified as Epicoccum nigrum by morphology of sporulating cultures, and the identity confirmed by genetic characterization including rDNA sequencing analysis of the large-subunit (LSU), the internal transcribed spacer, and the D1/D2 region of the LSU. These E. nigrum rDNA sequences were deposited in GenBank under accesssion numbers AF338443, AY093413 and AY093414. Systematic rDNA homology alignments were performed to identify fungi related to E. nigrum. Stimulation of carotenogenesis by E. nigrum and potentially other fungi could provide a novel method to enhance astaxanthin formation in industrial fermentations of X. dendrorhous and Phaffia rhodozyma.

Base Sequence↗

On single and multiple models of protein families for the detection of remote sequence relationships.

BACKGROUND: The detection of relationships between a protein sequence of unknown function and a sequence whose function has been characterised enables the transfer of functional annotation. However in many cases these relationships can not be identified easily from direct comparison of the two sequences. Methods which compare sequence profiles have been shown to improve the detection of these remote sequence relationships. However, the best method for building a profile of a known set of sequences has not been established. Here we examine how the type of profile built affects its performance, both in detecting remote homologs and in the resulting alignment accuracy. In particular, we consider whether it is better to model a protein superfamily using a single structure-based alignment that is representative of all known cases of the superfamily, or to use multiple sequence-based profiles each representing an individual member of the superfamily. RESULTS: Using profile-profile methods for remote homolog detection we benchmark the performance of single structure-based superfamily models and multiple domain models. On average, over all superfamilies, using a truncated receiver operator characteristic (ROC5) we find that multiple domain models outperform single superfamily models, except at low error rates where the two models behave in a similar way. However there is a wide range of performance depending on the superfamily. For 12% of all superfamilies the ROC5 value for superfamily models is greater than 0.2 above the domain models and for 10% of superfamilies the domain models show a similar improvement in performance over the superfamily models. CONCLUSION: Using a sensitive profile-profile method we have investigated the performance of single structure-based models and multiple sequence models (domain models) in detecting remote superfamily members. We find that overall, multiple models perform better in recognition although single structure-based models display better alignment accuracy.

Amino Acid Sequence↗

Multiple alignment using hidden Markov models.

A simulated annealing method is described for training hidden Markov models and producing multiple sequence alignments from initially unaligned protein or DNA sequences. Simulated annealing in turn uses a dynamic programming algorithm for correctly sampling suboptimal multiple alignments according to their probability and a Boltzmann temperature factor. The quality of simulated annealing alignments is evaluated on structural alignments of ten different protein families, and compared to the performance of other HMM training methods and the ClustalW program. Simulated annealing is better able to find near-global optima in the multiple alignment probability landscape than the other tested HMM training methods. Neither ClustalW nor simulated annealing produce consistently better alignments compared to each other. Examination of the specific cases in which ClustalW outperforms simulated annealing, and vice versa, provides insight into the strengths and weaknesses of current hidden Markov model approaches.

Algorithms↗