PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

PCR isolation of catechol 2,3-dioxygenase gene fragments from environmental samples and their assembly into functional genes.

A method was developed to isolate central segments of catechol 2, 3-dioxygenase (C23O) genes from environmental samples and to insert these C23O gene segments into nahH (the structural gene for C23O encoded by catabolic plasmid NAH7) by replacing the corresponding nahH sequence with the isolated segments. To PCR-amplify the central C23O gene segments, a pair of degenerate primers was designed from amino acid sequences conserved among C23Os. Using these primers, central regions of the C23O genes were amplified from DNA isolated from a mixed culture of phenol-degrading or crude oil-degrading bacteria. Both the 5' and 3' regions of nahH were also PCR-amplified by using appropriate primers. These three PCR products, the 5'-nahH and 3'-nahH segments and the central C23O gene segments, were mixed and PCR-amplified again. Since the primers for the amplification of the central C23O gene segments were designed so that the 20 nucleotides at both ends of the segments are identical to the 3' end of the 5'-nahH segment and the 5' end of the 3'-nahH segment, respectively, the central C23O gene segments could anneal to both the 5'- and 3'-nahH segments. After the second PCR, hybrid C23O genes in the form of (5'-nahH segment-central C23O gene segment-3'-nahH segment) were amplified to full length. The resulting products were cloned into a vector and used to transform Escherichia coli. This method enabled divergent C23O sequences to be readily isolated, and more than 90% of the hybrid plasmids expressed C23O activity. Thus, the present method is useful to create, without isolating bacteria, a library of functional hybrid genes.

Amino Acid Sequence↗

Decoding non-unique oligonucleotide hybridization experiments of targets related by a phylogenetic tree.

MOTIVATION: The reliable identification of presence or absence of biological agents ("targets"), such as viruses or bacteria, is crucial for many applications from health care to biodiversity. If genomic sequences of targets are known, hybridization reactions between oligonucleotide probes and targets performed on suitable DNA microarrays will allow to infer presence or absence from the observed pattern of hybridization. Targets, for example all known strains of HIV, are often closely related and finding unique probes becomes impossible. The use of non-unique oligonucleotides with more advanced decoding techniques from statistical group testing allows to detect known targets with great success. Of great relevance, however, is the problem of identifying the presence of previously unknown targets or of targets that evolve rapidly. RESULTS: We present the first approach to decode hybridization experiments using non-unique probes when targets are related by a phylogenetic tree. Using a Bayesian framework and a Markov chain Monte Carlo approach we are able to identify over 94% of known targets and assign up to 70% of unknown targets to their correct clade in hybridization simulations on biological and simulated data. AVAILABILITY: Software implementing the method described in this paper and datasets are available from http://algorithmics.molgen.mpg.de/probetrees.

Algorithms↗

Enterobacterial repetitive intergenic consensus (ERIC) sequences in Escherichia coli: Evolution and implications for ERIC-PCR.

Enterobacterial repetitive intergenic consensus (ERIC) sequences are 127-bp imperfect palindromes that occur in multiple copies in the genomes of enteric bacteria and vibrios. Here we investigate the distribution of these elements in the complete genome sequences of nine Escherichia coli (including Shigella species) strains. There is a significant tendency for copies to be adjacent to more highly expressed genes. There is considerable variation among strains with respect to the presence of an element in any particular intergenic region, but some copies appear to have been conserved since before the divergence of E. coli and Salmonella enterica. In comparisons of orthologous copies between these species, ERIC sequences are surprisingly conserved, implying that they have acquired some function, perhaps related to mRNA stability. The relationships among copies within E. coli are consistent with a master copy mode of generation. Insertion of new copies seems to occur at, and involve duplication of, the dinucleotide TA. Two classes of inserts of about 70 bp each occur at different specific sites within ERIC sequences; these inserts evolve independently of the ERIC sequences. The small number of ERIC sequences in E. coli genomes indicates that a widely used bacterial fingerprinting method using primers based on ERIC sequences (ERIC-PCR) does not rely on the presence of ERIC sequences.

Base Sequence↗

Extraction of hidden Markov model representations of signal patterns in DNA sequences.

We have developed a method to extract the signal patterns in DNA sequences. In this method, the Genetic Algorithm (GA) and Baum-Welch algorithm are used to obtain the best Hidden Markov Model (HMM) representations of the signal patterns in DNA sequences. The GA is used to search the best network shapes and the initial parameters of the HMMs. Baum-Welch algorithm is used to optimize the HMM parameters for the given network shapes. Akaike Information Criterion (AIC), which gives a criterion for the balance of adaptation and complexity of a model, is applied in the HMM evaluation. We have applied the method to the extraction of the signal patterns in human promoters and 5' ends of yeast introns. As a result, we obtained HMM representations of characteristic features in these sequences. To validate the efficiency of the method, we have performed promoter recognition using obtained HMMs. Two entries including nine promoters are selected from GenBank 76.0, and it is observed that the HMM can predicts eight promoters correctly. These results imply that the method is efficient to design preferable HMM networks, and provides reliable models for the recognition of the signal patterns.

Algorithms↗

Phylogeny of betanodaviruses and molecular evolution of their RNA polymerase and coat proteins.

The betanodaviruses are the causative agent of the disease viral nervous necrosis in fishes. Betanodavirus genome consists of two single-stranded positive-sense RNA molecules (RNA1 and RNA2). RNA1 gene encodes the RNA polymerase, named also protein A, while RNA2 encodes the coat protein precursor, the CPp protein. We investigated the evolutionary relationships among betanodaviruses working on partial sequences of both RNA1 and RNA2. Phylogenetic analyses were performed by applying a maximum likelihood approach. The phylogenetic relationships among the major betanodavirus clades SJNNV-IV, TPNNV-III, BFNNV-II and RGNNV-I were resolved differently in the trees obtained, respectively, from RNA1 and RNA2 multiple alignments. The alternative topologies were corroborated by strong bootstrap values. The molecular evolution of proteins A and CPp was also investigated. Protein A appeared to have evolved under strong purifying selection while the CPp protein was subject to both purifying and neutral selection in different amino acid residues. Intragenic recombination in RNA1 and RNA2 genes was investigated by applying several methods and was not detected. Conversely reassortment of RNA1 and RNA2 genes was demonstrated in some isolates. Finally RNA1 and RNA2 genes substitution rates do not follow a clock-like behavior thus impeding estimation of a possible origin time for Betanodavirus genus.

Base Sequence↗

A self consistent mean field approach to simultaneous gap closure and side-chain positioning in homology modelling.

A new computational procedure which simultaneously provides gap closure and side-chain positioning in homology modelling is described. It uses a database search scheme to generate fragments to model gaps, a rotamer library to define side-chain conformations, and iteratively refines a conformational matrix CM, such that its elements CM(i,j,o) and CM(i,j,k) give the probabilities that the backbone of residue i adopts the conformation described by fragment j and that its side-chain adopts the conformation of its possible rotamer k. Each residue experiences the average of all possible environments, weighted by their respective probabilities. The method converges, thereby deserving the name of 'self consistent mean field' approach.

Algorithms↗

Using hidden Markov models and observed evolution to annotate viral genomes.

MOTIVATION: ssRNA (single stranded) viral genomes are generally constrained in length and utilize overlapping reading frames to maximally exploit the coding potential within the genome length restrictions. This overlapping coding phenomenon leads to complex evolutionary constraints operating on the genome. In regions which code for more than one protein, silent mutations in one reading frame generally have a protein coding effect in another. To maximize coding flexibility in all reading frames, overlapping regions are often compositionally biased towards amino acids which are 6-fold degenerate with respect to the 64 codon alphabet. Previous methodologies have used this fact in an ad hoc manner to look for overlapping genes by motif matching. In this paper differentiated nucleotide compositional patterns in overlapping regions are incorporated into a probabilistic hidden Markov model (HMM) framework which is used to annotate ssRNA viral genomes. This work focuses on single sequence annotation and applies an HMM framework to ssRNA viral annotation. A description of how the HMM is parameterized, whilst annotating within a missing data framework is given. A Phylogenetic HMM (Phylo-HMM) extension, as applied to 14 aligned HIV2 sequences is also presented. This evolutionary extension serves as an illustration of the potential of the Phylo-HMM framework for ssRNA viral genomic annotation. RESULTS: The single sequence annotation procedure (SSA) is applied to 14 different strains of the HIV2 virus. Further results on alternative ssRNA viral genomes are presented to illustrate more generally the performance of the method. The results of the SSA method are encouraging however there is still room for improvement, and since there is overwhelming evidence to indicate that comparative methods can improve coding sequence (CDS) annotation, the SSA method is extended to a Phylo-HMM to incorporate evolutionary information. The Phylo-HMM extension is applied to the same set of 14 HIV2 sequences which are pre-aligned. The performance improvement that results from including the evolutionary information in the analysis is illustrated.

Algorithms↗

Structural homology of lens crystallins. III. Secondary structure estimation from circular dichroism and prediction from amino acid sequences.

Circular dichroism spectra (196-240 nm) of calf alpha-, beta H-, beta L- and gamma-crystallins were measured and analyzed over the entire wavelength range with five curve-fitting procedures for estimating protein secondary structure. For gamma-crystallin the estimates are in good agreement with the X-ray structure. For all four crystallins the estimates are very similar: 0-9% alpha-helix and 51-68% beta-sheet. This is in accordance with the three-dimensional homology of beta Bp- and gamma 2-crystallin polypeptide chains as postulated from their 30% sequence homology, and suggests that alpha A- and alpha B-crystallin chains may also have a corresponding structure. Secondary structure elements in the four amino acid sequences were predicted using two different comprehensive prediction methods. For gamma 2-crystallin the predictions of beta-sheet are in good agreement with the X-ray structure and with circular dichroism estimates. For beta Bp-crystallin only the C-terminal domain secondary structure predictions are considered satisfactory, which possibly relates to the proposed role of the N-terminal domain in subunit interactions. The combined predictions for alpha A- and alpha B-chains (3% helix, 49% sheet) are in excellent agreement with circular dichroism. Moreover, the good alignment of predicted beta-sheet segments in alpha-crystallin chains with known beta-sheet strands in gamma 2- (and presumably beta Bp-) crystallin strongly supports a similar 4-motif folding pattern in all four calf crystallin chains.

Amino Acid Sequence↗

Determination of polypeptide amino acid sequences from the carboxyl terminus using angiotensin I converting enzyme.

A method for sequence analysis of polypeptides starting at the carboxyl terminus is described that utilizes degradation of the polypeptide into dipeptides with angiotensin I converting enzyme. Dipeptides were identified by gas chromatography-mass spectroscopy. Dipeptide alignment was achieved by replicate digestion of the polypeptide after modification at the carboxyl terminus either by chemical or enzymatic removal of one residue or by addition of a single residue. The addition reaction involved coupling of L-alpha-aminobutyric acid under conditions described herein which yielded essentially complete conversions. Unlike sequence determination methods that commence from the polypeptide amino terminus, this procedure does not require that a polypeptide have a free amino terminus for successful application. A number of polypeptides with varying chain lengths (up to 49 residues), containing among them most of the common amino acids, have been successfully analyzed in amounts as low as 5 nmol.

Amino Acid Sequence↗

Crystallographic refinement of interleukin 1 beta at 2.0 A resolution.

The structure of human recombinant interleukin 1 beta (IL-1 beta) has been refined by a restrained least-squares method to a crystallographic R factor of 17.2% to 2.0 A resolution. One-hundred sixty-eight solvent molecules have been located, and isotropic temperature factors for each atom have been refined. The overall structure is composed of 12 beta-strands that can best be described as forming the four triangular faces of a tetrahedron with hydrogen bonding resembling normal antiparallel beta-sheets only at the vertices. The interior of this tetrahedron is filled by hydrophobic side chains. Analysis of sequence alignments with IL-1 beta from other mammalian species shows the interior to be very well conserved with the exterior residues markedly less so. There does not appear to be a clustering of invariant amino acid side chains on the surface of the molecule, suggesting an area of interaction with the IL-1 receptor. Comparison of the IL-1 beta structure with IL-1 alpha sequences indicates that IL-1 alpha probably has a similar overall folding as IL-1 beta but binds to the receptor in a different fashion. The three-dimensional structure of the IL-1 beta is analyzed in light of what has been suggested by previously published work on mutants and fragments of the molecule.

Amino Acid Sequence↗

A sensitive one-step real-time RT-PCR method for detecting Grapevine leafroll-associated virus 2 variants in grapevine.

Grapevine leafroll syndrome is caused by a complex of up to nine different Grapevine leafroll-associated viruses (GLRaV-1-9) with GLRaV-2 being reported as one of the most variable species of this group. Many methods, including indexing, serological and molecular procedures, have been developed for the detection of GLRaV-2. However, due to the low concentration of the virus in plants and the high variability of GLRaV-2, a method with improved sensitivity and with the capacity to detect of all known variants is required. Such improvement is essential for grapevine rootstocks, as these are suspected to harbour frequent GLRaV-2 infections difficult to detect, thus contributing to the spread of the leafroll disease. The development of new universal primers is described using a target sequence located in the 3' end of the virus genome. These primers were combined with a one-step SYBR Green real-time RT-PCR assay to achieve quantitative detection. All 43 GLRaV-2 isolates tested in this study were identified readily and reproducibly, regardless of their geographical origin or variety of grapevine. Using the procedure developed in this study, the sensitivity was increased 125 times compared to a conventional single-tube RT-PCR. This real-time method opens new perspectives for the sanitary selection of grapevine and in leafroll 2 disease monitoring.

3' Flanking Region↗

phlD-based genetic diversity and detection of genotypes of 2,4-diacetylphloroglucinol-producing Pseudomonas fluorescens.

Diversity within a worldwide collection of 2,4-diacetylphloroglucinol-producing Pseudomonas fluorescens strains was assessed by sequencing the phlD gene. Phylogenetic analyses based on the phlD sequences of 70 isolates supported the previous classification into 18 BOX-PCR genotypes (A-Q and T). Exploiting polymorphisms within the sequence of phlD, we designed and used allele-specific PCR primers with a PCR-based dilution endpoint assay to quantify the population sizes of A-, B-, D-, K-, L- and P-genotype strains grown individually or in pairs in vitro, in the rhizosphere of wheat and in bulk soil. Except for P. fluorescens Q8r1-96, which strongly inhibited the growth of P. fluorescens Q2-87, inhibition between pairs of strains grown in vitro did not affect the accuracy of the method. The allele-specific primer-based technique is a rapid method for studies of the interactions between genotypes of 2,4-diacetylphloroglucinol producers in natural environments.

Alleles↗

Phylogenetic analysis of Ara+ and Ara- Burkholderia pseudomallei isolates and development of a multiplex PCR procedure for rapid discrimination between the two biotypes.

A Burkholderia pseudomallei-like organism has recently been identified among some soil isolates of B. pseudomallei in an area with endemic melioidosis. This organism is almost identical to B. pseudomallei in terms of morphological and biochemical profiles, except that it differs in ability to assimilate L-arabinose. These Ara+ isolates are also less virulent than the Ara- isolates in animal models. In addition, clinical isolates of B. pseudomallei available to date are almost exclusively Ara-. These features suggested that these two organisms may belong to distinctive species. In this study, the 16S rRNA-encoding genes from five clinical (four Ara- and one Ara+) and nine soil isolates (five Ara- and four Ara+) of B. pseudomallei were sequenced. The nucleotide sequences and phylogenetic analysis indicated that the 16S rRNA-encoding gene of the Ara+ biotype was similar to but distinctively different from that of the Ara- soil isolates, which were identical to the classical clinical isolates of B. pseudomallei. The nucleotide sequence differences in the 16S rRNA-encoding gene appeared to be specific for the Ara+ or Ara- biotypes. The differences were, however, not sufficient for classification into a new species within the genus Burkholderia. A simple and rapid multiplex PCR procedure was developed to discriminate between Ara- and Ara+ B. pseudomallei isolates. This new method could also be incorporated into our previously reported nested PCR system for detecting B. pseudomallei in clinical specimens.

Arabinose↗

High fidelity SNP genotyping using sequence-specific primer elongation and fluorescence correlation spectroscopy.

Reliable, efficient and cost-effective modalities are urgently needed for mass screening of gene mutations. Previous reports have shown that SSCP or genechip methods require substantial time and monetary costs, thus limiting their appeal. Sequence Specific Primer Polymerase Chain Reaction (SSP-PCR) is a reliable and cost-effective method that utilizes the 3'-end discrimination properties of polymerase. However, the applicability of conventional SSP-PCR is limited due to the difficulties associated with determining optimal conditions and because mis-matched primers are amplified, resulting in signal noise during end-point assay. To overcome this problem, we eliminated the reverse primers from SSP-PCR, thus preventing amplification of mis-matched primers. We designated this method Sequence-Specific Primer Cycle Elongation (SSPCE). However, the detection of elongated sequence specific primers was difficult using conventional electrophoresis due to the small amounts of amplification product present. We therefore combined SSPCE and Fluorescence Correlation Spectroscopy, which is a novel technique used to determine the number and size of fluorophores at nano-molar concentrations, and designated the method SSPCE-FCS. We compared conventional SSP-PCR and SSPCE-FCS with regard to determining optimal conditions using two Mitochondrial SNPs (G --> A at position 1598, G --> A at position 12192). We were able to determine the optimal conditions for the SNP at position 1598 using either method. However, optimal conditions could only be determined for SSPCE-FCS with the 12192 mutation because non-specific amplification was observed at a wide range of annealing temperatures in SSP-PCR. We then applied this method to three other SNPs and the results were consistent with the results of sequencing data.

Base Sequence↗

[Development of "Amplisens-HCV-genotype" reagent set for identification of hepatitis C virus genotypes 1a, 1b, 2a and 3a].

Multiple alignments of 119 nucleotide sequences of isolates of hepatitis C virus (HCV) were carried out to choose the type-specific primers for the 5'-ultra-core fragment of viral genome for the purpose of detecting the HCV 1a, 1b, 2a, and 3a subtypes. A PCR kit of reagents was designed for the amplification of cDNA HCV with selected type-specific primers and for making the electrophoresis in agarous gel. The kit comprises the positive control samples, i.e. HCV genome fragments, subtypes 1a, 1b, 2a and 3a, cloned in the plasmid vector. 440 cDNAHCV samples were simultaneously tested by using the worked out reagents' set and according to the method of Ohno et al. The results were found to be concordant in 336 cases, and were discordant in 4 samples. A sequencing of the PCR products and phylogenetic analysis showed that 1 sample belonged to subtype 4a, 2 samples belonged to subtypes 2k and 1 sample--to subtype 31.

DNA Primers↗

Expanding allelic diversity of Helicobacter pylori vacA.

The diversity of the gene encoding the vacuolating cytotoxin (vacA) of Helicobacter pylori was analyzed in 98 isolates obtained from different geographic locations. The studies focused on variation in the previously defined s and m regions of vacA, as determined by PCR and direct sequencing. Phylogenetic analysis revealed the existence of four distinct types of s-region alleles: aside from the previously described s1a, s1b, and s2 allelic types, a novel subtype, designated s1c, was found. Subtype s1c was observed exclusively in isolates from East Asia and appears to be the major s1 allele in that part of the world. Three different allelic forms (m1, m2a, and m2b) were detected in the m region. On the basis of sequence alignments, universal PCR primers that allow effective amplification of the s and m regions from H. pylori isolates from all over the world were defined. Amplimers were subsequently analyzed by reverse hybridization onto a line probe assay (LiPA) that allows the simultaneous and highly specific hybridization of the different vacA s- and m-region alleles and tests for the presence of the cytotoxin-associated gene (cagA). This PCR-LiPA method permits rapid analysis of the vacA and cagA status of H. pylori strains for clinical and epidemiological studies and will facilitate identification of any further variations.

Amino Acid Sequence↗

A detailed consideration of a principal domain of vertebrate fibrinogen and its relatives.

Vertebrate fibrinogen is a complex multidomained protein, the structure of which has been inferred mainly from electron microscopy and amino acid sequence studies. Among its most prominent features are two terminal globules, moieties that are mostly composed of the carboxyl-terminal two-thirds of the beta and gamma chains. Sequences homologous to the latter segments are found in several other animal proteins, always as the carboxyl-terminal contributions. An alignment of 15 amino acid sequences from various fibrinogens and related proteins has been used to make judgments about secondary structure. The nature of amino acids at each position in the alignment was used to distinguish alpha helices and beta structure on the one hand from loops and turns on the other, and the resulting assignments compared with predictions of secondary structure by other methods. Additionally, constraints imposed by the locations of cystines, carbohydrate attachment residues, and proteinase-sensitive points provided further insights into the general organization of the postulated secondary structures. Other ancillary data, including the effects of bound calcium and the locations of labeled or variant residues, were also considered. An intriguing similarity to a portion of the recently reported structure of a calcium-dependent lectin is noted.

Amino Acid Sequence↗