PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Characterization of IS900 loci in Mycobacterium avium subsp. paratuberculosis and development of multiplex PCR typing.

Mycobacterium avium subsp. paratuberculosis is a pathogen that causes chronic inflammation of the intestine in many animals, including primates, and is implicated in Crohn's disease in humans. It differs from other members of the M. avium complex in having 14-18 copies of IS900 inserted into conserved loci in its genome. In the present study, genomic DNA flanking 14 of these insertions was characterized and homologues in the Mycobacterium tuberculosis and M. avium subsp. avium genomes were identified. These included regions encoding a sigma factor (sigJ) at locus 3, a nitrate reductase (nirA) at locus 4, a transcription regulator (tetR) and polyketide synthase at locus 6, and a 6-O-methylguanine methyltransferase at locus 9. In addition, locus numbers were assigned to 9 of 15 RFLP bands previously described. IS900 insertion at 7 of the 14 characterized loci was into the RBS of a gene substituting an RBS encoded by IS900 sited two bases closer to the initiation codon. IS900 insertion at five loci interrupted an ORF at the target site, one of which encoded a homologue of the immunodominant mycobacterial DesA1 protein. Eleven of eighty-one M. avium subsp. paratuberculosis isolates lacked the insertion site at locus 6 together with flanking genomic DNA. This region was also absent from seven reference strains of M. avium subsp. avium, from one M. avium subsp. silvaticum and from six other mycobacterial species. A multiplex PCR of IS900 loci (MPIL) typing method was developed which was able to discriminate 10 different types of M. avium subsp. paratuberculosis from the panel of 81 isolates with consistent differences between those of bovine and ovine origin. Nine MPIL types corresponded with a single PstI/Bst:EII RFLP type, suggesting that this method may be applicable to typing of M. avium subsp. paratuberculosis directly from a sample without the need for culture. The remaining MPIL type corresponded with seven PstI/BstEII RFLP types. Further resolution of these may come from sequencing the remaining four uncharacterized IS900 loci.

Animals↗

[Identification of Amomum villosum, Amomum villosum var. xanthioides and Amomum longiligulare on ITS-1 sequence].

OBJECTIVE: To identify Amomum villosum Lour. and some their adulterants on molecular biology. METHOD: The DNA of Amomum villosum Lour. and some their adulterants were extracted, and amplified using ITS-1 primer. The amplificed DNA were purified and then sequenced by direct PCR sequencing method. RESULT: The ITS-sequence of all of the samples are 248 bp in size. But there are 7 bases in Amomum villosum Lour var. xanthioides (Wall.ex Bak) T.L. Wu et Senjen and 12 hases in Amomum longiligulare T.L. Wu. differing from Amomum villosum Lour. CONCLUSION: The ITS-1 sequence can be used to identify effectively Amomum villosum Lour. and their adulterants.

Amomum↗

Cloning of the structural gene for Clostridium botulinum type C1 toxin and whole nucleotide sequence of its light chain component.

The toxigenicity of Clostridium botulinum type C1 is mediated by specific bacteriophages. DNA was extracted from one of these phages. Two DNA fragments, 3 and 7.8 kb, which produced the protein reacting with antitoxin serum were cloned by using bacteriophage lambda gt11 and Escherichia coli. Both DNA fragments were then subcloned into pUC118 plasmids and transferred into E. coli cells. The nucleotide sequences of the cloned DNA fragments were analyzed by the dideoxy chain termination method, and their gene products were analyzed by Western immunoblot. The 7.8-kb fragment coded for the entire light chain component and the N terminus of the heavy chain component of the toxin, whereas the 3-kb fragment coded for the remaining heavy chain component. The entire nucleotide sequence for the light chain component was determined, and the derived amino acid sequence was compared with that of tetanus toxin. It was found that the light chain component of C1 toxin possessed several amino acid regions, in addition to the N terminus, that were homologous to tetanus toxin.

Amino Acid Sequence↗

Thrombin-like enzymes from venom gland of Deinagkistrodon acutus: cDNA cloning, mechanism of diversity and phylogenetic tree construction.

AIM: To clone cDNAs of thrombin-like enzymes (TLEs) from venom gland of Deinagkistrodon acutus and analyze the mechanisms by which their structural diversity arose. METHODS: Reverse transcription-polymerase chain reaction and gene cloning techniques were used, and the cloned sequences were analyzed by using bioinformatics tools. RESULTS: Novel cDNAs of snake venom TLEs were cloned. The possibilities of post-transcriptional recombination and horizontal gene transfer are discussed. A phylogenetic tree was constructed. CONCLUSION: The cDNAs of snake venom TLEs exhibit great diversification. There are several types of structural variations. These variations may be attributable to certain mechanisms including recombination.

Amino Acid Sequence↗

Implication by site-directed mutagenesis of Arg314 and Tyr316 in the coenzyme site of pig mitochondrial NADP-dependent isocitrate dehydrogenase.

Sequence alignment of pig mitochondrial NADP-dependent isocitrate dehydrogenase with eukaryotic (human, rat, and yeast) and Escherichia coli isocitrate dehydrogenases reveals that Tyr316 is completely conserved and is equivalent to the E. coli Tyr345, which interacts with the 2'-phosphate of NADP in the crystal structure [Hurley et al., Biochemistry 30 (1991) 8671-8678]. Lys321 is also completely conserved in the five isocitrate dehydrogenases. Either an arginine or lysine residue is found among the enzymes from other species at the position corresponding to the pig enzyme Arg314. While Arg323 is not conserved among all species, its proximity to the coenzyme site makes it a good candidate for investigation. The importance of these four amino acids to the function of pig mitochondrial NADP-isocitrate dehydrogenase was studied by site-directed mutagenesis. Mutants (R314Q, Y316F, Y316L, K321Q, and R323Q) were generated by a megaprimer polymerase chain reaction method. Wild-type and mutant enzymes were expressed in E. coli and purified to homogeneity. All mutant and wild-type enzymes exhibited comparable molecular weights indicative of the dimeric enzyme. Mutations do not cause an appreciable change in enzyme secondary structure as revealed by circular dichroism measurements. The kinetic parameters (V(max) and K(M) values) of K321Q and R323Q are similar to those of wild-type, indicating that Lys321 and Arg323 are not involved in enzyme function. R314Q exhibits a 10-fold increase in K(M) for NADP as compared to that of wild-type, while they have comparable V(max) values. These results suggest that Arg314 contributes to the affinity between the enzyme and NADP. The hydroxyl group of Tyr316 is not required for enzyme function since Y316F exhibits similar kinetic parameters to those of wild-type. Y316L shows a 4-fold increase in K(M) for NADP and a decrease in V(max) as compared to wild-type, suggesting that the aromatic ring of the Tyr of isocitrate dehydrogenase contributes to the affinity for coenzyme, as well as to catalysis. The K(i) for NAD of R314Q, Y316F, and Y316L is comparable to that of wild-type, indicating that the Arg314 and Tyr316 may be located near the 2'-phosphate of enzyme-bound NADP.

Amino Acid Sequence↗

Molecular cloning and characterization of full-length cDNAs encoding a novel high-molecular-weight Dermatophagoides pteronyssinus mite allergen, Der p 11.

BACKGROUND: Dermatophagoides pteronyssinus (Dp) and D. farinae (Df) mites are the most important source of indoor aeroallergens. Most Dp mite allergens identified to date have relatively low molecular weights (MWs). Identification of high-MW mite allergens is a crucial step in characterizing the complete spectrum of mite allergens and to provide appropriate tools for diagnostic and therapeutic application. METHODS: The full-length Der p 11 cDNA clone was isolated using cDNA library immunoscreening, the 5'-3' rapid amplification of cDNA ends (RACE) system and polymerase chain reactions (PCR). The whole cDNA insert and its PCR-derived DNA fragments (p1 to p4) were generated and expressed in the Escherichia coli expression system. The allergenicity of the recombinant protein and its peptide fragments was examined by IgE immunodot assays. The IgE-binding reactivity of rDer p 11 was analyzed in the serum of 50 asthmatic children with positive reactivity to Dp mite extract. Its recombinant peptide fragments were also examined by immunodot assays in 30 mite-allergic children. RESULTS: Der p 11 cDNA consists of a 2625-bp open reading frame encoding a 103-kDa protein with 875 amino acids. It exhibits significant homology with the paramyosin of other invertebrates. The protein sequence alignment of this newly identified Dp mite allergen (denominated as Der p 11) revealed over 89% identity with Der f 11 and Blo t 1. Among 50 Dp-sensitive asthmatic children, rDer p 11 showed positive IgE-binding reactivity to 39 patients (78%). Using immunodot assays, multiple human IgE-binding activities were demonstrated in all four fragments of Der p 11. Using immunoblot assays, the dominant IgG-binding epitope for monoclonal antibody (mAb642) was located in fragment p3 only. In immunoblot assays, cross-inhibition between rDer p 11 and rDer f 11 was up to 73-80% at concentrations of 100 microg/ml. CONCLUSIONS: This study confirms that the newly identified recombinant Der p 11 is a novel and important high-MW Dp mite allergen for asthmatic children. Our data also indicates that human IgE-binding major epitopes are scattered over the entire molecule of Der p 11.

Adolescent↗

Hybridization-ligation versus parallel overlap assembly: an experimental comparison of initial pool generation for direct-proportional length-based DNA computing.

Previously, direct-proportional length-based DNA computing (DPLB-DNAC) for solving weighted graph problems has been reported. The proposed DPLB-DNAC has been successfully applied to solve the shortest path problem, which is an instance of weighted graph problems. The design and development of DPLB-DNAC is important in order to extend the capability of DNA computing for solving numerical optimization problem. According to DPLB-DNAC, after the initial pool generation, the initial solution is subjected to amplification by polymerase chain reaction and, finally, the output of the computation is visualized by gel electrophoresis. In this paper, however, we give more attention to the initial pool generation of DPLB-DNAC. For this purpose, two kinds of initial pool generation methods, which are generally used for solving weighted graph problems, are evaluated. Those methods are hybridization-ligation and parallel overlap assembly (POA). It is found that for DPLB-DNAC, POA is better than that of the hybridization-ligation method, in terms of population size, generation time, material usage, and efficiency, as supported by the results of actual experiments.

Base Sequence↗

Analysis of the entire nucleotide sequence of the cryptic plasmid QpH1 from Coxiella burnetti.

The complete plasmid QpH1 from Coxiella burnetti, isolate 'Nine Mile', phase I, was cloned as NotI fragment with a size of 37329 bp. The entire plasmid was sequenced by the chain termination method after EcoRI subcloning. 37 open reading frames coding for polypeptides larger than 100 amino acid residues were determined. The predicted polypeptide products of the open reading frames were compared by computer analysis with reported protein sequences. Homologies of predicted polypeptide products to analogous proteins are described.

Amino Acids↗

Nuclear small subunit rRNA group I intron variation among Beauveria spp provide tools for strain identification and evidence of horizontal transfer.

An optional group I intron was characterized at a single insertion point in nuclear small subunit rRNA (nuSSU rRNA) genes of the imperfect entomopathogenic fungi, Beauveria bassiana and B. brongniartii. Insertion points were conserved among nuSSU rRNA genes from 35 Beauveria isolates. PCR-RFLP and DNA sequencing identified 12 group I intron variants and were applied to the identification of strains isolated from insect hosts. Alignment of 383-404-nt subgroup IB3 group I introns indicated that four insertion/deletion (indel) mutations were the main basis of fragment length variation. Phylogeny reconstruction using parsimony and neighbor-joining methods suggested six lineages may be present among nuSSU rRNA group I intron sequences from Beauveria and related ascomycete fungi. Terminal node placement of Beauveria introns conflicted with previously published phylogenies constructed from gene sequences, suggesting horizontal transfer of group I introns. PCR-RFLP among introns provided a means for the differentiation of Beauveria isolates.

Base Sequence↗

Cloning and characterization of a phosphopantetheinyl transferase from Streptomyces verticillus ATCC15003, the producer of the hybrid peptide-polyketide antitumor drug bleomycin.

BACKGROUND: Phosphopantetheinyl transferases (PPTases) catalyze the posttranslational modification of carrier proteins by the covalent attachment of the 4'-phosphopantetheine (P-pant) moiety of coenzyme A to a conserved serine residue, a reaction absolutely required for the biosynthesis of natural products including fatty acids, polyketides, and nonribosomal peptides. PPTases have been classified according to their carrier protein specificity. In organisms containing multiple P-pant-requiring pathways, each pathway has been suggested to have its own PPTase activity. However, sequence analysis of the bleomycin biosynthetic gene cluster in Streptomyces verticillus ATCC15003 failed to reveal an associated PPTase gene. RESULTS: A general approach for cloning PPTase genes by PCR was developed and applied to the cloning of the svp gene from S. verticillus. The svp gene is mapped to an independent locus not clustered with any of the known NRPS or PKS clusters. The Svp protein was overproduced in Escherichia coli, purified to homogeneity, and shown to be a monomer in solution. Svp is a PPTase capable of modifying both type I and type II acyl carrier proteins (ACPs) and peptidyl carrier proteins (PCPs) from either S. verticillus or other Streptomyces species. As compared to Sfp, the only 'promiscuous' PPTase known previously, Svp displays a similar catalytic efficiency (k(cat)/K(m)) for the BlmI PCP but a 346-fold increase in catalytic efficiency for the TcmM ACP. CONCLUSIONS: PPTases have recently been re-classified on a structural basis into two subfamilies: ACPS-type and Sfp-type. The development of a PCR method for cloning Sfp-type PPTases from actinomycetes, the recognition of the Sfp-type PPTases to be associated with secondary metabolism with a relaxed carrier protein specificity, and the availability of Svp, in addition to Sfp, should facilitate future endeavors in engineered biosynthesis of peptide, polyketide, and, in particular, hybrid peptide-polyketide natural products.

Amino Acid Sequence↗

Mitochondrial DNA phylogeography of European hedgehogs.

European hedgehog populations belonging to Erinaceus europaeus and E. concolor have been investigated by mitochondrial DNA analysis. A 383 bp fragment of the cytochrome b gene has been sequenced and maximum parsimony and neighbour-joining trees of Tamura-Nei genetic distance values have been constructed. Similar topologies have been produced by both methods, showing a deep divergence between E. europaeus and E. concolor and a further subdivision of each species into a western and an eastern clade. A comparison with previously published allozyme data is made, and concordant and discordant patterns are discussed. The influence of Pleistocene glaciations on the observed pattern of divergence is inferred.

Animals↗

Combined multiple sequence reduced protein model approach to predict the tertiary structure of small proteins.

By incorporating predicted secondary and tertiary restraints into ab initio folding simulations, low resolution tertiary structures of a test set of 20 nonhomologous proteins have been predicted. These proteins, which represent all secondary structural classes, contain from 37 to 100 residues. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Predicted tertiary restraints are obtained from multiple sequence alignments via a two-step process: First, "seed" side chain contacts are identified from a correlated mutation analysis, and then, the seed contacts are "expanded" by an inverse folding algorithm. These predicted restraints are then incorporated into a lattice based, reduced protein model. Depending upon fold complexity, the resulting nativelike topologies exhibit a coordinate root-mean-square deviation, cRMSD, from native between 3.1 and 6.7 A. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

Gene prediction in eukaryotes with a generalized hidden Markov model that uses hints from external sources.

BACKGROUND: In order to improve gene prediction, extrinsic evidence on the gene structure can be collected from various sources of information such as genome-genome comparisons and EST and protein alignments. However, such evidence is often incomplete and usually uncertain. The extrinsic evidence is usually not sufficient to recover the complete gene structure of all genes completely and the available evidence is often unreliable. Therefore extrinsic evidence is most valuable when it is balanced with sequence-intrinsic evidence. RESULTS: We present a fairly general method for integration of external information. Our method is based on the evaluation of hints to potentially protein-coding regions by means of a Generalized Hidden Markov Model (GHMM) that takes both intrinsic and extrinsic information into account. We used this method to extend the ab initio gene prediction program AUGUSTUS to a versatile tool that we call AUGUSTUS+. In this study, we focus on hints derived from matches to an EST or protein database, but our approach can be used to include arbitrary user-defined hints. Our method is only moderately effected by the length of a database match. Further, it exploits the information that can be derived from the absence of such matches. As a special case, AUGUSTUS+ can predict genes under user-defined constraints, e.g. if the positions of certain exons are known. With hints from EST and protein databases, our new approach was able to predict 89% of the exons in human chromosome 22 correctly. CONCLUSION: Sensitive probabilistic modeling of extrinsic evidence such as sequence database matches can increase gene prediction accuracy. When a match of a sequence interval to an EST or protein sequence is used it should be treated as compound information rather than as information about individual positions.

Algorithms↗

Stabilization centers in proteins: identification, characterization and predictions.

Methods are presented to locate residues, stabilization center elements, which are expected to stabilize protein structures by preventing their decay with their cooperative long range interactions. Artificial neural network-based algorithms were developed to predict these residues from the primary structure of single proteins and from the amino acid sequences of homologous proteins. The prediction accuracy using only single sequence information is 65%, but the incorporation of evolutionary information in the form of multiple alignments and conservation scores raises the efficiency by 3%. The composition, relative accessibility, number and type of interactions, conservation and the X-ray thermal factor of the identified stabilization center residues are different, not only from the whole data set but from the rest of the long range interacting residues as well. The most frequent stabilization center residues are usually found at buried positions and have a hydrophobic or aromatic side-chain, but some polar or charged residues also play an important role in the stabilization. The stabilization centers show significant difference in the composition and in the type of linked secondary structural elements compared with the rest of the residues. The performed structural and sequential conservation analysis showed the higher conservation of stabilization centers over protein families. The relation of the proposed stabilization centers to folding nuclei is also discussed.

Algorithms↗

Predicting ligand-binding function in families of bacterial receptors.

The three-dimensional fold of a new protein sequence can often be inferred directly from sequence homology to a protein of known structure. The function of a new protein sequence is more difficult to predict, however, since homologues can have different molecular and cellular functions. To develop and automate computational methods for determining molecular function, we have analyzed ligand-binding specificity in two related families of binding proteins. One of these families includes Escherichia coli lactose repressor and ribose-binding protein, and the other includes E. coli sulfate- and phosphate-binding proteins. These proteins have similar folds but varying specificity, binding many different small molecules, including mono- and disaccharides, purines, oxyanions, ferric iron, and polyamines. Starting from template structural alignments, alignments of over 90 sequences per family were generated by iterative database searches with hidden Markov models. Phylogenetic trees were made of full-length sequences and of subsets of residues lining the binding cleft, to determine whether subbranches of the trees correlate with ligand-binding preference. Automated analyses of residues in the binding pocket were also used to predict ligand-binding function for many uncharacterized database sequences and to identify specific side chain-ligand contacts in proteins without solved structures. Our results demonstrate the utility of anchoring functional annotation within a protein family context.

Amino Acid Sequence↗

Geometric invariant core for the CL and CH1 domains of immunoglobulin molecules.

A previously developed algorithmic method for identifying a geometric invariant of protein structures, termed geometrical core, is extended to the C(L) and C(H1) domains of immunoglobulin molecules. The method uses the matrix of C(alpha) - C(alpha) distances and does not require the usual superposition of structures. The result of applying the algorithm to 53 Immunoglobulin structures led to the identification of two geometrical core sets of C(alpha) atom positions for the C(L) and C(H1) domains.

Algorithms↗

Diversity of polyketide synthase gene sequences in Aspergillus species.

Fungal polyketide synthases are responsible for the biosynthesis of several mycotoxins and other secondary metabolites. The aim of our work was to investigate the diversity of polyketide synthases in Aspergillus species using two approaches: PCR amplification using oligonucleotide primers, and bioinformatics. Ketosynthase domain probes amplified DNA fragments of about 700 bp in each examined isolate. Sequences of these domains were aligned and analyzed by phylogenetic methods. The ketosynthase domain sequences were highly diverse indicating that they most probably represent polyketide synthases responsible for different functions. A. albertensis and A. niger ketosynthase domain sequences clustered together with sequences of genes required for pigment biosynthesis (wA) in A. nidulans and P. patulum, while the ketosynthase domain sequence of A. muricatus was most closely related to an A. parasiticus wA type domain sequence, and those of the A. ochraceus isolates formed a distinct clade on the tree. These sequences were highly homologous to an A. terreus naphthopyrone synthase gene. An Aspergillus fumigatus genomic database was also searched for ketosynthase domain sequences, which have been included in the phylogenetic analysis. Altogether 14 putative ketosynthase domain sequences were identified. Clustering of the ketosynthase domain sequences correlated well with the type of metabolites produced by the corresponding polyketide synthases. At least 8 clusters with putative ketosynthase domain sequences of unknown function have been identified. Further studies are in progress to clarify the role of some of the identified polyketide synthase genes.

Amino Acid Sequence↗

Analysis of the D1S80 locus by capillary electrophoresis.

We have demonstrated in a systematic manner that the capillary electrophoresis (CE) method of genotyping the human D1S80 locus is an effective replacement for the commonly used gel electrophoresis method. The CE method is fast, with a total run time of less than 22 min per sample. Separation has been optimized so that resolution for all pairs of alleles from 14 through 41 is greater than 1, providing unambiguous identification. The method utilizes run buffer containing 0.30% hydroxyethyl cellulose as the sieving polymer in a 50 microm internal diameter (ID) column coated with a 0.1 microm DB-17 film. Laser-induced fluorescence (LIF) with YO-PRO-1 intercalating dye is used for detection. Internal standards of 300 and 1000 bp bracket the D1S80 region in the electropherogram. Two typing procedures were evaluated: matching of the normalized migration times of sample alleles to ladder alleles; and software alignment of sample and ladder electropherograms using the internal standard peaks as references. Using the first procedure, 91.5% of the D1S80 genotypes were unambiguous, while 100% were unambiguous using the second procedure. Seventy-nine routine samples were analyzed in a side-by-side comparison of the CE and slab gel methods, with complete agreement of the results. Twenty-two samples were selected from a large database previously analyzed by the slab gel method to demonstrate that all alleles from 14 through 41 were typed correctly, including samples containing adjacent alleles. Additionally, 43 samples containing alleles greater than the 41 repeat number were typed correctly.

Automation↗