PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Phylogeny of Photorhabdus and Xenorhabdus species and strains as determined by comparison of partial 16S rRNA gene sequences.

Partial 16S rRNA gene sequences of 16 strains of the genera Photorhabdus and Xenorhabdus were determined by direct sequencing of PCR products. Aligned sequences were subjected to phylogenetic analysis by maximum-likelihood and maximum-parsimony methods. Distance matrix and phylogenetic analysis did not separate the genera unambiguously. Taxonomic grouping of the bacteria closely paralleled taxonomic grouping of their nematode associates and their geographic origins. We found at least two well-supported taxonomic groups in Photorhabdus species, which suggests that the genus Photorhabdus is coevolving with the nematodes and may be polyspecific.

Animals↗

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals↗

Phylogeny of the Mycoplasma mycoides cluster as shown by sequencing of a putative membrane protein gene.

The Mycoplasma mycoides cluster is made of six species that are closely related both genetically and phenotypically. Two are of particular importance, M. mycoides subsp. mycoides SC causing contagious bovine pleuropneumonia and M. capricolum subsp. capripneumoniae causing contagious caprine pleuropneumonia. The sequences of a putative membrane protein gene and partial flanking open reading frames have been obtained from various strains in this cluster, including all reference strains. Sequence analysis showed this locus is present and fully conserved in all strains of M. mycoides subsp. mycoides SC isolated from geographically most distant places worldwide. In M. capricolum subsp. capripneumoniae polymorphism in this locus has been found at seven positions and revealed that they can be used as epidemiological markers. Conserved regions were used to define a primer pair that enables the amplification by PCR of two fragments 302 and 1298bp long, respectively. The 302bp long fragment contains an intergenic sequence that can be used for phylogenetic studies or for identification purposes. Parsimony analysis on an alignment of 49 DNA sequences show a subdivision of the M. mycoides cluster into two subgroups that is in accordance with results obtained by phenotypic methods. Two lineages exist within the capricolum subgroup, one of them clustering strains identified as M. capricolum subsp. capricolum, M. capricolum subsp. capricolum and M. sp Bovine Group 7. However M. capricolum subsp. capripneumoniae strains can readily be identified by three specific nucleotide positions or by sequencing the 1298bp long fragment. There is no clear subdivision within the mycoides subgroup, supporting the idea that M. mycoides subsp. mycoides LC and M. mycoides subsp. capri should not be separated into two subspecies. Mycoplasma mycoides subsp. mycoides SC strains can easily be distinguished as they bear an insertion sequence 15bp downstream from the stop codon of the membrane protein gene.

Animals↗

Density functional studies of oxidized and reduced methane monooxygenase. Optimized geometries and exchange coupling of active site clusters.

The conflicting protein crystallography data for the oxidized form (MMOH(ox)) of methane monooxygenase present a dilemma regarding the identity of the solvent-derived bridging ligands within the active site: do they comprise a diiron unit bridged by 1H2O and 1OH(-) as postulated for Methylococcus capsulatus or 2OH(-) ligands as suggested for Methylosinus trichosporium? Using models derived explicitly from the M. capsulatus and M. trichosporium protein data, spin-unrestricted density functional methods have been used to study two structurally characterized forms of the hydroxylase component of methane monooxygenase. The active site geometries of the oxidized (MMOH(ox)) and two-electron-reduced (MMOH(red)) states have been geometry optimized using several quantum cluster models which take into account the antiferromagnetic (AF) and ferromagnetic (F) coupling of electron spins. Trends in cluster geometries, energetics, and Heisenberg J values have been evaluated. For the majority of models, calculated geometries are in good agreement with the X-ray analyses and appear relatively insensitive to the F or AF alignment of electron spins on adjacent Fe sites. Discrepancies between calculation and experiment appear in the orientation of the coordinated His and Glu amino acid side chains for both MMOH(ox) and MMOH(red) and also in unexpected intramolecular proton transfer in the MMOH(ox) cluster models. There is additional dispersion between (and among) calculated and experimental Fe(3+)-OH(-) distances with relevance to the correct protonation state of the solvent-derived ligands. In an accompanying paper (Lovell, T.; Li, J.; Noodleman, L. Inorg. Chem. 2001, 40, 5267), a comparison of the related energetics of the active site models examined herein is further evaluated in the full protein and solvent environment.

Binding Sites↗

Scorpion toxins specific for Na+-channels.

Na+-channel specific scorpion toxins are peptides of 60-76 amino acid residues in length, tightly bound by four disulfide bridges. The complete amino acid sequence of 85 distinct peptides are presently known. For some toxins, the three-dimensional structure has been solved by X-ray diffraction and NMR spectroscopy. A constant structural motif has been found in all of them, consisting of one or two short segments of alpha-helix plus a triple-stranded beta-sheet, connected by variable regions forming loops (turns). Physiological experiments have shown that these toxins are modifiers of the gating mechanism of the Na+-channel function, affecting either the inactivation (alpha-toxins) or the activation (beta-toxins) kinetics of the channels. Many functional variations of these peptides have been demonstrated, which include not only the classical alpha- and beta-types, but also the species specificity of their action. There are peptides that bind or affect the function of Na+-channels from different species (mammals, insects or crustaceans) or are toxic to more than one group of animals. Based on functional and structural features of the known toxins, a classification containing 10 different groups of toxins is proposed in this review. Attempts have been made to correlate the presence of certain amino acid residues or 'active sites' of these peptides with Na+-channel functions. Segments containing positively charged residues in special locations, such as the five-residue turn, the turn between the second and the third beta-strands, the C-terminal residues and a segment of the N-terminal region from residues 2-11, seems to be implicated in the activity of these toxins. However, the uncertainty, and the limited success obtained in the search for the site through which these peptides bind to the channels, are mainly due to the lack of an easy method for expression of cloned genes to produce a well-folded, active peptide. Many scorpion toxin coding genes have been obtained from cDNA libraries and from polymerase chain reactions using fragments of scorpion DNAs, as templates. The presence of an intron at the DNA level, situated in the middle of the signal peptide, has been demonstrated.

Amino Acid Sequence↗

A database analysis method identifies an endogenous trans-acting short-interfering RNA that targets the Arabidopsis ARF2, ARF3, and ARF4 genes.

Two classes of small RNAs, microRNAs and short-interfering RNA (siRNAs), have been extensively studied in plants and animals. In Arabidopsis, the capacity to uncover previously uncharacterized small RNAs by means of conventional strategies seems to be reaching its limits. To discover new plant small RNAs, we developed a protocol to mine an Arabidopsis nonannotated, noncoding EST database. Using this approach, we identified an endogenous small RNA, trans-acting short-interfering RNA-auxin response factor (tasiR-ARF), that shares a 21- and 22-nt region of sequence similarity with members of the ARF gene family. tasiR-ARF has characteristics of both short-interfering RNA and microRNA, recently defined as tasiRNA. Accumulation of trans-acting siRNA depends on DICER-LIKE1 and RNA-DEPENDENT RNA POLYMERASE6 but not RNA-DEPENDENT RNA POLYMERASE2. We demonstrate that tasiR-ARF targets three ARF genes, ARF2, ARF3/ETT, and ARF4, and that both the tasiR-ARF precursor and its target genes are evolutionarily conserved. The identification of tasiRNA-ARF as a low-abundance, previously uncharacterized small RNA species proves our method to be a useful tool to uncover additional small regulatory RNAs.

Arabidopsis↗

Prediction of contact maps by GIOHMMs and recurrent neural networks using lateral propagation from all four cardinal corners.

MOTIVATION: Accurate prediction of protein contact maps is an important step in computational structural proteomics. Because contact maps provide a translation and rotation invariant topological representation of a protein, they can be used as a fundamental intermediary step in protein structure prediction. RESULTS: We develop a new set of flexible machine learning architectures for the prediction of contact maps, as well as other information processing and pattern recognition tasks. The architectures can be viewed as recurrent neural network implemantations of a class of Bayesian networks we call generalized input-output HMMs (GIOHMMs). For the specific case of contact maps, contextual information is propagated laterally through four hidden planes, one for each cardinal corner. We show that these architectures can be trained from examples and yield contact map predictors that outperform previously reported methods. While several extensions and improvements are in progress, the current version can accurately predict 60.5% of contacts at a distance cutoff of 8 A and 45% of distant contacts at 10 A, for proteins of length up to 300.

Algorithms↗

Identification and interrogation of highly informative single nucleotide polymorphism sets defined by bacterial multilocus sequence typing databases.

A unified, bioinformatics-driven, single nucleotide polymorphism (SNP)-based approach to microbial genotyping has been developed. Multilocus sequence typing (MLST) databases consist of known variants of standardized housekeeping genes. Normally, seven fragments are defined; a sequence type (ST) consists of the variants of these fragments that are found in a particular isolate. A computer program that can identify highly informative sets of SNPs in entire MLST databases has been constructed. The SNPs either define a particular user-specified ST or provide a high value for Simpson's index of diversity (D), and may thus be generally applicable to that species. SNP sets that are diagnostic for Neisseria meningitidis ST-11 and ST-42, and high-D SNP sets for N. meningitidis and Staphylococcus aureus, were identified and real-time PCR methods to interrogate these SNPs were demonstrated. High-D SNP sets were also identified in other MLST databases. This widely applicable approach allows rapid genetic fingerprinting of infectious agents.

Algorithms↗

Protein structure comparison using the markov transition model of evolution.

A number of automatic protein structure comparison methods have been proposed; however, their similarity score functions are often decided by the researchers' intuition and trial-and-error, and not by theoretical background. We propose a novel theory to evaluate protein structure similarity, which is based on the Markov transition model of evolution. Our similarity score between structures i and j is defined as log P(j --> i)/P(i), where P(j --> i) is the probability that structure j changes to structure i during the evolutionary process, and P(i) is the probability that structure i appears by chance. This is a reasonable definition of structure similarity, especially for finding evolutionarily related (homologous) similarity. The probability P(j --> i) is estimated by the Markov transition model, which is similar to the Dayhoff's substitution model between amino acids. To estimate the parameters of the model, homologous protein structure pairs are collected using sequence similarity, and the numbers of structure transitions within the pairs are counted. Next these numbers are transformed to a transition probability matrix of the Markov transition. Transition probabilities for longer time are obtained by multiplying the probability matrix by itself several times. In this study, we generated three types of structure similarity scores: an environment score, a residue-residue distance score, and a secondary structure elements (SSE) score. Using these scores, we developed the structure comparison program, Matras (MArkovian TRAnsition of protein Structure). It employs a hierarchical alignment algorithm, in which a rough alignment is first obtained by SSEs, and then is improved with more detailed functions. We attempted an all-versus-all comparison of the SCOP database, and evaluated its ability to recognize a superfamily relationship, which was manually assigned to be homologous in the SCOP database. A comparison with the FSSP database shows that our program can recognize more homologous similarity than FSSP. We also discuss the reliability of our method, by studying the disagreement between structural classifications by Matras and SCOP.

Databases, Factual↗

A strategy to retrieve the whole set of protein modules in microbial proteomes.

Protein homology is often limited to long structural segments that we have previously called modules. We describe here a suite of programs used to catalog the whole set of modules present in microbial proteomes. First, the Darwin AllAll program detects homologous segments using thresholds for evolutionary distance and alignment length, and another program classifies these modules. After assembling these homologous modules in families, we further group families which are related by a chain of neighboring unrelated homologous modules. With the automatic analysis of these groups of families sharing homologous modules in independent multimodular proteins, one can split into their component parts many fused modules and/or deduce by logic more distant modules. All detected and inferred modules are reassembled in refined families. These two last steps are made by a unique program. Eventually, the soundness of the data obtained by this experimental approach is checked using independent tests. To illustrate this modular approach, we compared four proteobacterial proteomes (Campylobacter jejuni, Escherichia coli, Haemophilus influenzae, and Helicobacter pylori). It appears that this method might retrieve from present-day proteins many of the modules which can help to trace back ancient events of gene duplication and/or fusion.

Bacterial Proteins↗

Designated primers targeted canine TP53 gene hotspot regions.

BACKGROUND: Tumor protein 53 gene (TP53) is a critical factor that controls different cell activities such as cell cycle, DNA repair mechanism, autophagy, apoptosis, and metabolism. The TP53 gene is the most commonly mutated gene, especially in the 4-8 exons region. This mutation enhances the development of many abnormalities, such as the initiation of different types of cancer. AIM: The main objective of this study was to design and evaluate the efficacy of three different primer sets that targeted the TP53 gene at the hotspot regions. METHODS: To do that, twelve blood samples were collected from dogs belonging to the German Shepherd breed/K9 aged between 8-12 years. Then, the DNA extraction and polymerase chain reaction (PCR) took place by using the three primer sets, which were designed using SnapGene. The primer sets, namely, first primer, the second and the third targeted exons 5-9 located in the canine TP53 gene. In the following step, all the PCR products were sent for Sanger sequencing and then phylogenetic analysis. RESULTS: Our findings indicated that the first primer set consistently showed higher amplification signal efficiency and reduced dimer formation compared with the second and third primer sets, respectively, with a 60ºC annealing temperature. In addition, all the sequenced samples aligned with the reference canine TP53 gene in the phylogenetic tree. CONCLUSION: This study offered the best TP53 primer design that targeted the hotspot regions of the canine TP53 gene for researchers who are interested in targeting such regions in this gene.

Animals↗

Assignment of homology to genome sequences using a library of hidden Markov models that represent all proteins of known structure.

Of the sequence comparison methods, profile-based methods perform with greater selectively than those that use pairwise comparisons. Of the profile methods, hidden Markov models (HMMs) are apparently the best. The first part of this paper describes calculations that (i) improve the performance of HMMs and (ii) determine a good procedure for creating HMMs for sequences of proteins of known structure. For a family of related proteins, more homologues are detected using multiple models built from diverse single seed sequences than from one model built from a good alignment of those sequences. A new procedure is described for detecting and correcting those errors that arise at the model-building stage of the procedure. These two improvements greatly increase selectivity and coverage. The second part of the paper describes the construction of a library of HMMs, called SUPERFAMILY, that represent essentially all proteins of known structure. The sequences of the domains in proteins of known structure, that have identities less than 95 %, are used as seeds to build the models. Using the current data, this gives a library with 4894 models. The third part of the paper describes the use of the SUPERFAMILY model library to annotate the sequences of over 50 genomes. The models match twice as many target sequences as are matched by pairwise sequence comparison methods. For each genome, close to half of the sequences are matched in all or in part and, overall, the matches cover 35 % of eukaryotic genomes and 45 % of bacterial genomes. On average roughly 15% of genome sequences are labelled as being hypothetical yet homologous to proteins of known structure. The annotations derived from these matches are available from a public web server at: http://stash.mrc-lmb.cam.ac.uk/SUPERFAMILY. This server also enables users to match their own sequences against the SUPERFAMILY model library.

Amino Acid Sequence↗

A facile analytical method for the identification of protease gene profiles from Bacillus thuringiensis strains.

Five pairs of degenerate universal primers have been designed to identify the general protease gene profiles from some distinct Bacillus thuringiensis strains. Based on the PCR amplification patterns and DNA sequences of the cloned fragments, it was noted that the protease gene profiles of the three distinct strains of B. thuringiensis subsp. kurstaki HD73, tenebrionis and israelensis T14001 are varied. Seven protease genes, neutral protease B (nprB), intracellular serine protease A (ispA), extracellular serine protease (vpr), envelope-associated protease (prtH), neutral protease F (nprF), thermostable alkaline serine protease and alkaline serine protease (aprS), with known functions were identified from three distinct B. thuringiensis strains. In addition, five DNA sequences with unknown functions were also identified by this facile analytical method. However, based on the alignment of the derived protein sequences with the protein domain database, it suggested that at least one of these unknown genes, yunA, might be highly protease-related. Thus, the proposed PCR-mediated amplification design could be a facile method for identifying the protease gene profiles as well as for detecting novel protease genes of the B. thuringiensis strains.

Amino Acid Sequence↗

Liquid-phase hybridization and capture of hepatitis B virus DNA with magnetic beads and fluorescence detection of PCR product.

The polymerase chain reaction (PCR) exceeds all hitherto known detection limits. This sensitivity could lead to false positive results. Every manipulation increases the risk of contamination via, for example, aerosols. Most protocols for the extraction of template nucleic acids are complicated and possible centrifugation steps do not reduce the risk of aerosols. In addition, most of the methods for analysis are time-consuming and cannot be applied to different template materials. An alternative extraction method has been developed. The fast chemical denaturation of template by guanidine thiocyanate was followed by liquid hybridization to biotinylated oligonucleotides. The template nucleic acid could be washed after binding to streptavidin-coated paramagnetic beads to reduce influence on the enzymatic amplification steps. PCR of hepatitis B virus deoxyribonucleic acid was used to demonstrate how easy, versatile, and time-saving this method is without centrifugation. The level of extracted nucleic acids was quantitated and the properties for sensitive extraction were evaluated. After PCR an additional step was developed which used fluorescent staining to detect positive amplifications. This is useful to identify positive results in predominantly negative samples.

Base Sequence↗

Evaluation of amplified ribosomal DNA restriction analysis (ARDRA) and species-specific PCR for identification of Bifidobacterium species.

Molecular biological methods based on genus-specific PCR, species-specific PCR, and amplified ribosomal DNA restriction analysis (ARDRA) of two PCR amplicons (523 and 914bp) using six restriction enzymes were used to differentiate among species of Bifidobacterium. The techniques were established using DNA from 16 type and reference strains of bifidobacteria of 11 species. The discrimination power of 914bp amplicon digestion was higher than that of 523bp amplicon digestion. The 914bp amplicon digestion by six restrictases provided unique patterns for nine species; B. catenulatum and B. pseudocatenulatum were not differentiated yet. The NciI digestion of the 914bp PCR product enabled to discriminate between each of B. animalis, B. lactis, and B. gallicum. The reference strain B. adolescentis CCM 3761 was reclassified as a member of the B. catenulatum/B. pseudocatenulatum group. The above-mentioned methods were applied for the identification of seven strains of Bifidobacterium spp. collected in the Culture Collection of Dairy Microorganisms (CCDM). The strains collected in CCDM were differentiated to the species level. Six strains were identified as B. lactis, one strain as B. adolescentis.

Base Sequence↗

Simple method to identify bacteriocin induction peptides and to auto-induce bacteriocin production at low cell density.

The production of some bacteriocins by lactic acid bacteria is regulated by induction peptides (IPs) that are secreted by a dedicated secretion system. The IP gene cbaX, for carnobacteriocin A production by Carnobacterium piscicola LV17A, and a presumptive IP gene (orf6), associated with the genetic locus for enterocin B production in Enterococcus faecium BFE 900, were fused to the signal peptide of the bacteriocin divergicin A from Carnobacterium divergens LV13 to access the general secretory pathway. The culture supernatants of C. piscicola UAL26 and Lactococcus lactis MG1363 containing either of these constructs were used to induce bacteriocin production by Bac(-) cultures of C. piscicola LV17A or E. faecium CTC492. The cbaX fusion product induced bacteriocin production by Bac(-) C. piscicola LV17A, but the orf6 fusion product did not induce bacteriocin production by E. faecium CTC492. This represents a relatively simple method of confirming the role of presumptive IPs. The transformation of C. piscicola LV17A with the CbaX gene under expression of the P32 promoter from L. lactis resulted in constitutive production of bacteriocin by either the dedicated transport apparatus or the general secretory pathway.

Amino Acid Sequence↗

PCR-based identification of hyperthermophilic archaea of the family Thermococcaceae.

A method for rapid detection and identification of hyperthermophilic archaea of the family Thermococcaceae based on PCR amplification of 16S rRNA gene fragments with primers TcPc 173F (5'-TCCCCCATAGGYCTGRGGTACTGGAAGGTC-3') and TcPc 589R (5'-GCCGTGRGATTTCGCCAGGGACTTACGGGC-3') was developed and used for identification of new isolates.

Bacteria↗

Molecular identification of food-borne and water-borne protozoa.

Cryptosporidium and Giardia can be transmitted to humans by contaminated food and water, resulting in large outbreaks of diarrheal disease. Sensitive methods for detecting these parasites are needed to control and prevent infection. However, this issue is complicated by the fact that there is still uncertainty about the role played by different species/genotypes with respect to human disease. We are in the process of collecting samples from clinical cases (both sporadic and outbreak-related human infections) and from the environment (tap and waste water samples from different geographic regions), to test the efficacy of methods for detection and genotyping. Concerning Cryptosporidium parvum, we have developed new genotyping methods based on highly polymorphic microsatellite markers. The use of microsatellite markers allows the route of transmission to be traced; these methods can also be used not only to distinguish between anthroponotic and zoonotic transmission but also to identify the source(s) of infection. Regarding Giardia, which was found very frequently in environmental water samples, we are testing the beta-giardin gene as a marker to discriminate among species/genotypes.

Animals↗