PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

[Cloning of the mouse Doc-1R gene by genomic walking].

OBJECTIVE: To obtain the genomic sequences of the mouse Doc-1R gene. METHODS: Gene-specific primers were designed and synthesized based on the cDNA sequences of the mouse Doc-1R gene. With the use of genomic walking strategy, the mouse genomic walking library was amplified by the polymerase chain reaction(PCR). Mouse genomic library constructed with a special adaptor was utilized as a template to amplify the desired fragment by nested PCR. RESULTS: A desired fragment of 1.5 kb was obtained. Sequence analysis of the desired fragment confirmed that the genomic cloning of the Doc-1R gene was successful. This gene contains four exons and three introns. All of the splice donor/acceptor site sequences are in accordance with the consensus 'GT-AG' rule. CONCLUSION: The genomic walking strategy is simple, efficient and reliable; it is an ideal method of cloning genomic fragments.

Animals↗

Homology modeling of an RNP domain from a human RNA-binding protein: Homology-constrained energy optimization provides a criterion for distinguishing potential sequence alignments.

We have recently described an automated approach for homology modeling using restrained molecular dynamics and simulated annealing procedures (Li et al, Protein Sci., 6:956-970,1997). We have employed this approach for constructing a homology model of the putative RNA-binding domain of the human RNA-binding protein with multiple splice sites (RBP-MS). The regions of RBP-MS which are homologous to the template protein snRNP U1A were constrained by "homology distance constraints," while the conformation of the non-homologous regions were defined only by a potential energy function. A full energy function without explicit solvent was employed to ensure that the calculated structures have good conformational energies and are physically reasonable. The effects of mis-alignment of the unknown and the template sequences were also explored in order to determine the feasibility of this homology modeling method for distinguishing possible sequence alignments based on considerations of the resulting conformational energies of modeled structures. Differences in the alignments of the unknown and the template sequences result in significant differences in the conformational energies of the calculated homology models. These results suggest that conformational energies and residual constraint violations in these homology-constrained simulated annealing calculations can be used as criteria to distinguish between correct and incorrect sequence alignments and chain folds.

Algorithms↗

A new criterion and method for amino acid classification.

It is accepted that many evolutionary changes of amino acid sequence in proteins are conservative: the replacement of one amino acid by another residue has a far greater chance of being accepted if the two residues have similar properties. It is difficult, however, to identify relevant physicochemical properties that capture this similarity. In this paper we introduce a criterion that determines similarity from an evolutionary point of view. Our criterion is based on the description of protein evolution by a Markov process and the corresponding matrix of instantaneous replacement rates. It is inspired by the conductance, a quantity that reflects the strength of mixing in a Markov process. Furthermore we introduce a method to divide the 20 amino acid residues into subsets that achieve good scores with our criterion. The criterion has the time-invariance property that different time distances of the same amino acid replacement rate matrix lead to the same grouping; but different rate matrices lead to different groupings. Therefore it can be used as an automated method to compare matrices derived from consideration of different types of proteins, or from parts of proteins sharing different structural or functional features. We present the groupings resulting from two standard matrices used in sequence alignment and phylogenetic tree estimation.

Algorithms↗

Detection of Ralstonia solanacearum, which causes brown rot of potato, by fluorescent in situ hybridization with 23S rRNA-targeted probes.

During the past few years, Ralstonia (Pseudomonas) solanacearum race 3, biovar 2, was repeatedly found in potatoes in Western Europe. To detect this bacterium in potato tissue samples, we developed a method based on fluorescent in situ hybridization (FISH). The nearly complete genes encoding 23S rRNA of five R. solanacearum strains and one Ralstonia pickettii strain were PCR amplified, sequenced, and analyzed by sequence alignment. This resulted in the construction of an unrooted tree and supported previous conclusions based on 16S rRNA sequence comparison in which R. solanacearum strains are subdivided into two clusters. Based on the alignments, two specific probes, RSOLA and RSOLB, were designed for R. solanacearum and the closely related Ralstonia syzygii and blood disease bacterium. The specificity of the probes was demonstrated by dot blot hybridization with RNA extracted from 88 bacterial strains. Probe RSOLB was successfully applied in FISH detection with pure cultures and potato tissue samples, showing a strong fluorescent signal. Unexpectedly, probe RSOLA gave a less intense signal with target cells. Potato samples are currently screened by indirect immunofluorescence (IIF). By simultaneously applying IIF and the developed specific FISH, two independent targets for identification of R. solanacearum are combined, resulting in a rapid (1-day), accurate identification of the undesired pathogen. The significance of the method was validated by detecting the pathogen in soil and water samples and root tissue of the weed host Solanum dulcamara (bittersweet) in contaminated areas.

Base Sequence↗

Searching expressed sequence tag databases: discovery and confirmation of a common polymorphism in the thymidylate synthase gene.

Databases of expressed sequence tags (EST) can be used to screen rapidly for potential polymorphisms in candidate proteins. As part of this study, we screened the gene for the enzyme thymidylate synthase (TS). TS is important physiologically because it is essential for the synthesis of deoxythymidylate, a nucleotide required for DNA synthesis and repair. TS is also a major target for cancer chemotherapeutic drugs, especially the widely used 5-fluorouracil. Using sequence alignment of ESTs, we identified a candidate 6-bp variation at bp 1494 in the 3'-untranslated region of the TS mRNA. This sequence variation occurred in 21 of 34 aligned ESTs at this location, including ESTs from various tissue sources. The presence of this polymorphism was confirmed in a Caucasian population (n = 95) by polymerase chain restriction amplification/RFLP analysis. The allele frequency of the 6-bp deletion was found to be 0.29 (wildtype +6 bp/+6 bp, 48%; +6 bp/-6 bp, 44%; -6 bp/-6 bp, 7%). Although the function of this polymorphism has not yet been investigated, the 3'-untranslated region of a gene can play a role in mRNA stability and translation. This study illustrates an approach to polymorphism discovery in candidate enzymes of physiological interest by searches of publicly available sequence data, a rapid and inexpensive method. The potential functional relevance of the common 6-bp deletion in the TS gene needs to be investigated, because this enzyme is plausibly of major importance not only in cancer treatment but also in cancer prevention.

Databases, Factual↗

A major component approach to presenting consensus sequences.

MOTIVATION: Summarizing and displaying the information contained in a set of aligned sequences is an important aid to identifying patterns within the sequences. A variety of forms of consensus sequences have been used previously to provide this information. However, these methods can cause a loss of information or introduce ambiguities into the consensus sequence, and some graphical approaches may become difficult to interpret due to visual distortion. RESULTS: We have developed a method to present a more precise and graphically clear view of a consensus sequence by using an approach based on defining the major components at each position in a sequence set. The major components are given in an ordered list and their frequencies are shown as histograms which can be colour coded to reflect conservative groupings. Minor components, a one-line character-based consensus sequence and information statistics can also be presented. As well as identifying the dominant sources of variation and conservation in the sequence set, the method also enables similarities and differences between subgroups of a sequence set to be readily assessed. AVAILIABILITY: On request from the authors. CONTACT: bcsmith@usthk.ust.hk, hxue@usthk. ust.hk

Amino Acid Sequence↗

Peptide motif for the rat MHC class II molecule RT1.Da: similarities to the multiple sclerosis-associated HLA-DRB1*1501 molecule.

Experimental autoimmune encephalomyelitis induced with myelin proteins in DA and LEW.1AV1 rats is a model of multiple sclerosis (MS). It reproduces major aspects of this detrimental disease of the central nervous system. MS is associated with the HLA-DRB1*1501, DRB5*0101, and DQB1*0602 haplotype. DA and LEW.1AV1 rats share the RT1av1 haplotype. So far, no MHC class II peptide motif of RT1.Da molecules has been described. Sequence alignment of the beta chain of the rat MHC class II molecule RT1.Da with human HLA class II molecules revealed strong similarity in the peptide-binding groove of RT1.Da and HLA-DRB1*1501. According to the putative peptide-binding pockets of RT1.Da, after comparison with the pockets of HLA-DRB1*1501, we predicted the peptide motif of RT1.Da. To verify the predicted motif, naturally processed peptides were eluted by acidic treatment from immunoaffinity-purified RT1.Da molecules of lymphoid tissue of DA rats and subsequently analyzed by ESI tandem mass spectrometry. In addition, we performed binding studies with combinatorial nonapeptide libraries to purified RT1.Da molecules. Based on these studies we could define a peptide-binding motif for RT1.Da characterized by aliphatic amino acid residues (L, I, V, M) and of F for the peptide pocket P1, aromatic residues (F, Y, W) for P4, basic residues (K, R) for P6, aliphatic residues (I, L, V) for P7, and aromatic residues (F, Y, W) and L for P9. Both methods revealed similar binding characteristics for peptides to RT1.Da. This data will allow epitope predictions for analysis of peptides, relevant for experimental autoimmune diseases.

Amino Acid Motifs↗

Partial sequence identification of grapevine-leafroll-associated virus-1 and development of a highly sensitive IC-RT-PCR detection method.

Using immunocapture reverse transcription PCR (IC-RT-PCR) a specific PCR product from GLRaV-1 infected vine samples was amplified with the help of degenerate primers deduced from the conserved HSP70 region of closteroviruses. 511 basepairs of the 5'end of GLRaV-1 HSP70 gene were identified. Within this region, putative GLRaV-1 specific primers were designed and an IC-RT-PCR detection procedure was developed which is about 125 times more sensitive than the established ELISA method. No PCR product was amplified in GLRaV-2,-3 and -4 infected plants which indicates the specificity of the primers. This procedure may serve as an alternative method for GLRaV-1 detection where the sensitivity of ELISA is insufficient.

Base Sequence↗

Yersiniabactin production by Pseudomonas syringae and Escherichia coli, and description of a second yersiniabactin locus evolutionary group.

The siderophore and virulence factor yersiniabactin is produced by Pseudomonas syringae. Yersiniabactin was originally detected by high-pressure liquid chromatography (HPLC); commonly used PCR tests proved ineffective. Yersiniabactin production in P. syringae correlated with the possession of irp1 located in a predicted yersiniabactin locus. Three similarly divergent yersiniabactin locus groups were determined: the Yersinia pestis group, the P. syringae group, and the Photorhabdus luminescens group; yersiniabactin locus organization is similar in P. syringae and P. luminescens. In P. syringae pv. tomato DC3000, the locus has a high GC content (63.4% compared with 58.4% for the chromosome and 60.1% and 60.7% for adjacent regions) but it lacks high-pathogenicity-island features, such as the insertion in a tRNA locus, the integrase, and insertion sequence elements. In P. syringae pv. tomato DC3000 and pv. phaseolicola 1448A, the locus lies between homologues of Psyr_2284 and Psyr_2285 of P. syringae pv. syringae B728a, which lacks the locus. Among tested pseudomonads, a PCR test specific to two yersiniabactin locus groups detected a locus in genospecies 3, 7, and 8 of P. syringae, and DNA hybridization within P. syringae also detected a locus in the pathovars phaseolicola and glycinea. The PCR and HPLC methods enabled analysis of nonpathogenic Escherichia coli. HPLC-proven yersiniabactin-producing E. coli lacked modifications found in irp1 and irp2 in the human pathogen CFT073, and it is not clear whether CFT073 produces yersiniabactin. The study provides clues about the evolution and dispersion of yersiniabactin genes. It describes methods to detect and study yersiniabactin producers, even where genes have evolved.

Amino Acid Sequence↗

Strain characterization and classification of oxyphotobacteria in clone cultures on the basis of 16S rRNA sequences from the variable regions V6, V7, and V8.

A major problem in development of a polyphasic taxonomy is that the identification of oxyphotobacterial strains (cyanobacteria and prochlorophytes) in culture collections may be incorrect. We have therefore developed a diagnostic system using the DNA sequence polymorphism in the 16S rRNA regions V6 to V8 for individual strain characterization and identification. PCR primers amplifying V6 to V8 from oxyphotobacteria in unialgal cultures were constructed. Direct solid-phase or cyclic sequencing was used to determine the sequences from the amplified DNA. This survey includes 10 strains of Nostoc/Anabaena/Aphanizomenon (Nostoc category), 5 strains of Microcystis (Microcystis category), and 4 strains of Planktothrix (Planktothrix category). Fifteen additional strains of cyanobacteria and two strains of prochlorophytes were included such that the major phyletic groups were represented. One of the strains, Phormidium sp. NIVA-CYA 203, contained an 11-nucleotide insertion with no homology to other known 16S rRNA sequences. Based on parsimony and neighbor-joining trees, the phyletic relationships of the strains were investigated. Thirteen major branches were found, with Pseudanabaena limnetica NIVA-CYA 276/6 as the most divergent strain. The strain categories Nostoc, Planktothrix, and Microcystis were all monophyletic. The sequence polymorphism within Nostoc was higher than that in Planktothrix and Microcystis. Based on the sequence and phyletic information, group-specific PCR primers for the categories Nostoc, Planktothrix, and Microcystis were constructed. For the strains included in this work, the amplifications were specific for the relevant groups. By combination of magnetic solid-phase DNA isolation and group-specific PCR amplifications, an accurate method for characterization, classification and identification of oxyphotobacterial clone cultures has been developed.

Base Sequence↗

Differentiation of phylogenetically related slowly growing mycobacteria by their gyrB sequences.

The conventional methods for identifying mycobacterial species are based on their phenotypic characterization. Since some problematic species are slow growers, their taxonomy takes several weeks or months to identify. The ribosomal DNA (rDNA) sequence-based identification strategy has been adopted to solve this problem. More recently, the gyrB sequences have been shown to be useful phylogenetic markers for the identification of species. We determined the gyrB sequences of 43 slowly growing strains belonging to 15 species in the genus Mycobacterium. The frequencies of base substitutions in the gyrB sequences were comparable to those in the 16S-23S rDNA internal transcribed spacer (ITS) sequences. The ITS sequences of four species belonging to the M. tuberculosis complex (M. tuberculosis, M. bovis, M. africanum, and M. microti) were 100% identical, while four synonymous substitutions were found in the gyrB sequences of these strains. Based on the differences found in the gyrB sequences, we developed PCR and PCR-restriction fragment length polymorphism methods to discriminate these species.

Animals↗

Development of amplified consensus genetic markers (ACGM) in Brassica napus from Arabidopsis thaliana sequences of known biological function.

A method for the development of consensus genetic markers between species of the same taxonomic family is described in this paper. It is based on the conservation of the peptide sequences and on the potential polymorphism within non-coding sequences. Six loci sequenced from Arabidopsis thaliana, AG, LFY3, AP3, FAD7, FAD3, and ADH, were analysed for one ecotype of A. thaliana, four lines of Brassica napus, and one line for each parental species, Brassica oleracea and Brassica rapa. Positive amplifications with the degenerate primers showed one band for A. thaliana, two to four bands in rapeseed, and one to two bands in the parental species. Direct sequencing of the PCR products confirms their peptide similarity with the "mother" sequence. By comparison of intron sequences, the correspondence between each rapeseed gene and its homologue in one of the parental species can be determined without ambiguity. Another important result is the presence of a polymorphism inside these fragments between the rapeseed lines. This variability could generally be detected by differences of electrophoretic migration on long non-denaturing polyacrylamide gels. This method enables a quick and easy shuttle between A. thaliana and Brassica species without cloning.

Amino Acid Sequence↗

Application of a novel and fast information-theoretic method to the discovery of higher-order correlations in protein databases.

We present a fast, discrete data-mining approach to the problem of finding kappa-tuples of correlated amino acid residues in protein sequence data. When sets of sequence-distant sites display high mutual information, they may bespeak important structural or functional features. Our novel methodology overcomes the limitations of previous methods which examined only single-residue features or pairwise interactions.

AIDS Vaccines↗

Total variation in the penA gene of Neisseria meningitidis: correlation between susceptibility to beta-lactam antibiotics and penA gene heterogeneity.

In recent decades, the prevalence of Neisseria meningitidis isolates with reduced susceptibility to penicillins has increased. The intermediate resistance to penicillin (Pen(i)) for most strains is due mainly to mosaic structures in the penA gene, encoding penicillin-binding protein 2. In this study, susceptibility to beta-lactam antibiotics was determined for 60 Swedish clinical N. meningitidis isolates and 19 reference strains. The penA gene was sequenced and compared to 237 penA sequences from GenBank in order to explore the total identified variation of penA. The divergent mosaic alleles differed by 3% to 24% compared to those of the designated wild-type penA gene. By studying the final 1,143 to 1,149 bp of penA in a sequence alignment, 130 sequence variants were identified. In a 402-bp alignment of the most variable regions, 84 variants were recognized. Good correlation between elevated MICs and the presence of penA mosaic structures was found especially for penicillin G and ampicillin. The Pen(i) isolates comprised an MIC of >0.094 microg/ml for penicillin G and an MIC of >0.064 microg/ml for ampicillin. Ampicillin was the best antibiotic for precise categorization as Pen(s) or Pen(i). In comparison with the wild-type penA sequence, two specific Pen(i) sites were altered in all except two mosaic penA sequences, which were published in GenBank and no MICs of the corresponding isolates were described. In conclusion, monitoring the relationship between penA sequences and MICs to penicillins is crucial for developing fast and objective methods for susceptibility determination. By studying the penA gene, genotypical determination of susceptibility in culture-negative cases can also be accomplished.

Amino Acid Sequence↗

Protein fold recognition by total alignment probability.

We present a protein fold-recognition method that uses a comprehensive statistical interpretation of structural Hidden Markov Models (HMMs). The structure/fold recognition is done by summing the probabilities of all sequence-to-structure alignments. The optimal alignment can be defined as the most probable, but suboptimal alignments may have comparable probabilities. These suboptimal alignments can be interpreted as optimal alignments to the "other" structures from the ensemble or optimal alignments under minor fluctuations in the scoring function. Summing probabilities for all alignments gives a complete estimate of sequence-model compatibility. In the case of HMMs that produce a sequence, this reflects the fact that due to our indifference to exactly how the HMM produced the sequence, we should sum over all possibilities. We have built a set of structural HMMs for 188 protein structures and have compared two methods for identifying the structure compatible with a sequence: by the optimal alignment probability and by the total probability. Fold recognition by total probability was 40% more accurate than fold recognition by the optimal alignment probability. Proteins 2000;40:451-462.

Algorithms↗

Ribosomal RNA sequences of Clostridium piliforme isolated from rodent and rabbit: re-examining the phylogeny of the Tyzzer's disease agent and development of a diagnostic polymerase chain reaction assay.

We used polymerase chain reaction (PCR) technology to amplify the 16S rRNA gene, the intergenic spacer, and most of the 23S rRNA gene from 6 isolates (2 mice, 1 hamster, 1 rat, and 2 rabbit isolates) of the Tyzzer's disease agent (Clostridium piliforme) and C. colinum. Sequence similarity searches of GenBank identified 45 closely related bacteria, which we used for phylogenetic analysis by parsimony and maximum-likelihood methods using Escherichia coli to root the resulting phylogram. Microorganisms identified as C. piliforme form 3 clusters within a single clade; the nearest related distinguishable species is C. colinum. Other bacterial clades closely related to C. piliforme are clostridia previously identified by molecular methods in the bovine, porcine, and human gastrointestinal tracts. DNA sequence alignment highlighting sequence differences were used to design a rodent and rabbit C. piliforme-specific PCR assay, which targets a 639-basepair region at the 3' end of the 16S rRNA gene and the 5' end of the intergenic spacer. We used this PCR assay to examine 4 rat fecal samples from C. piliformeseropositive rats and reexamine 2 rabbit fecal samples previously identified as containing DNA sequences consistent with C. piliforme infection by 16S PCR assay. Our new assay did not detect the presence of C. piliforme DNA sequences in either the rat or rabbit fecal DNA samples, consistent with the absence of clinical disease in the colonies evaluated.

Animals↗

Identification and assessment of known and novel human papillomaviruses by polymerase chain reaction amplification, restriction fragment length polymorphisms, nucleotide sequence, and phylogenetic algorithms.

The identification and taxonomy of papillomaviruses has become increasingly complex, as approximately 70 human papillomavirus (HPV) types have been described and novel HPV genomes continue to be identified. Methods and corresponding DNA sequence data bases were designed for the reliable identification of mucosal HPV genomes from clinical specimens. HPVs are identified by the amplification of a fragment of the L1 region by consensus primer polymerase chain reaction (PCR) and subsequent hybridization or restriction fragment length polymorphism analysis. L1 PCR fragments may be further characterized by nucleotide sequencing. Conservation of 30 (of 151) predicted amino acids identifies HPV genomic fragments, and nucleotide sequence alignments allow calculation of their phylogenetic relatedness. Sequence differences > 10% from any known HPV type suggest a novel HPV type. Phylogenetic relationships with known HPV types may permit predictions of biology. With these criteria, 10 PCR fragments were identified that would qualify as new genital HPV types after complete genomic isolation.

Amino Acid Sequence↗

The complete amino acid sequence of momordin-a, a ribosome-inactivating protein from the seeds of bitter gourd (Momordica charantia).

The complete amino acid sequence of momordin-a, a ribosome-inactivating protein from the seeds of bitter gourd, has been analyzed. Twenty-two peptides were isolated from the tryptic digest of momordin-a and sequenced by the DABITC/PITC double coupling method. The alignment of these tryptic peptides was done by analyzing the amino acid sequences of the peptides derived from chymotryptic digestion and cyanogen bromide cleavage of momordin-a as well as V8 protease-digestion of the CNBr fragment. Momordin-a consisted of 250 amino acid residues and carbohydrate residues attached to Asn227, and its molecular mass was calculated to be 28,690 Da. The sequence comparison with ricin A-chain shows that 33% of the residues of momordin-a are identical to those of ricin A-chain and that the residues involved in the catalytic site of the ricin A-chain are conserved in momordin-a.

Amino Acid Sequence↗