PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Artificial neural networks for molecular sequence analysis.

Artificial neural networks provide a unique computing architecture whose potential has attracted interest from researchers across different disciplines. As a technique for computational analysis, neural network technology is very well suited for the analysis of molecular sequence data. It has been applied successfully to a variety of problems, ranging from gene identification, to protein structure prediction and sequence classification. This article provides an overview of major neural network paradigms, discusses design issues, and reviews current applications in DNA/RNA and protein sequence analysis.

Algorithms↗

Phylogenetic structure of the genera Flexibacter, Flexithrix, and Microscilla deduced from 16S rRNA sequence analysis.

The 16S rDNA sequences of 40 strains of 17 species in the genus Flexibacter, 5 strains of 4 species in the genus Microscilla, and 1 strain of Flexithrix dorotheae, including all type strains of approved and validated species in these genera, were determined to reveal their phylogenetic relationships. The 16S rRNA sequence analysis demonstrated the extreme heterogeneity of the genera Flexibacter and Microscilla. The strains examined diverged into 24 distinct lines of descent (1 group included both flexibacteria and flexithrix, and 1 group included both flexibacteria and microscilla) that were remote from each other at the genus level or higher. Flexibacter strains were scattered across the cytophaga-flavobacteria-bacteroides phylum and divided into 20 phylogenetic groups, and the genus Microscilla was separated into 5 groups. Flexibacter flexilis, the type species of the genus Flexibacter, and Microscilla marina, the type species of the genus Microscilla, were isolated from other organisms in their respective genera. This means that each genus should be restricted to only the type species. Flexithrix dorotheae, the type species of the genus Flexithrix, clustered with Flexibacter aggregans. The heterogeneity was found not only within genera but also within species. Flexibacter aggregans, Flexibacter aurantiacus, Flexibacter flexilis, Flexibacter roseolus, Flexibacter tractuosus, and "Microscilla sericea" each contained phylogenetically distant strains. The taxonomic concept of the genera Flexibacter, Flexithrix, and Microscilla should be reorganized in accordance with the natural relationships revealed in this study.

Bacteroides↗

GeneQuiz: a workbench for sequence analysis.

We present the prototype of a software system, called GeneQuiz, for large-scale biological sequence analysis. The system was designed to meet the needs that arise in computational sequence analysis and our past experience with the analysis of 171 protein sequences of yeast chromosome III. We explain the cognitive challenges associated with this particular research activity and present our model of the sequence analysis process. The prototype system consists of two parts: (i) the database update and search system (driven by perl programs and rdb, a simple relational database engine also written in perl) and (ii) the visualization and browsing system (developed under C++/ET++). The principal design requirement for the first part was the complete automation of all repetitive actions: database updates, efficient sequence similarity searches and sampling of results in a uniform fashion. The user is then presented with "hit-lists" that summarize the results from heterogeneous database searches. The expert's primary task now simply becomes the further analysis of the candidate entries, where the problem is to extract adequate information about functional characteristics of the query protein rapidly. This second task is tremendously accelerated by a simple combination of the heterogeneous output into uniform relational tables and the provision of browsing mechanisms that give access to database records, sequence entries and alignment views. Indexing of molecular sequence databases provides fast retrieval of individual entries with the use of unique identifiers as well as browsing through databases using pre-existing cross-references. The presentation here covers an overview of the architecture of the system prototype and our experiences on its applicability in sequence analysis.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals↗

Identification of pathogenic dematiaceous fungi and related taxa based on large subunit ribosomal DNA D1/D2 domain sequence analysis.

The nucleotide sequences of the D1/D2 domains of large subunit (26S) ribosomal DNA for 76 strains of 46 species of pathogenic dematiaceous fungi and related taxa were determined. Intra-species sequence diversity of medically important dematiaceous fungi including Phialophora verrucosa, Fonsecaea pedrosoi, Fonsecaea compacta, Cladophialophora carrionii, Cladophialophora bantiana, Exophiala dermatitidis, Exophiala jeanselmei, Exophiala spinifera, Exophiala moniliae, and Hortaea werneckii were extremely small; as few as 0 changes were detected in C. bantiana, Fonsecaea and Exophiala species, 1 bp in C. carrionii and H. werneckii, and 2 bp in P. verrucosa. Inter-species nucleotide diversity between most species was higher. These data suggested that the D1/D2 domain is sufficiently variable for identification of pathogenic dematiaceous fungi and relevant species. The phylogenetic trees constructed from the sequence data revealed that most human pathogenic species formed a single cluster and that Cladosporium and Phialophora species were distributed polyphyletically into several clusters.

Ascomycota↗

Determination of the base recognition positions of zinc fingers from sequence analysis.

The CC/HH zinc finger is a small independently folded DNA recognition motif found in many eukaryotic proteins, which ligates zinc through two cysteine and two histidine ligands. A database of 1340 zinc fingers from 221 proteins has been constructed and a program for analysis of aligned sequences written. This paper describes sequence analysis aimed at determining the amino acid positions that recognize the DNA bases, by comparing two types of sequence variation. Using the idea that long runs of adjacent zinc fingers have arisen from internal gene duplication, the conservation of each position of the finger within the runs was calculated. The conservation of each position of the finger between homologous proteins from different species was also noted. A correlation of the two types of conservation showed clusters of related amino acids. One cluster of three positions was found to be especially variable within long runs, but highly conserved between corresponding fingers of homologous proteins; these positions are predicted to be the base contact positions. They match the amino acid positions that contact the bases in the co-crystal structure determined by Pavletich and Pabo [Science, 240, 809-817 (1991)]. An adjacent cluster of four positions on the plot may also be associated with DNA binding. This analysis shows that the base recognition positions can be identified even in the absence of a known structure for a zinc finger. These results are applicable to zinc fingers where the structure of the complex is unknown, in particular suggesting that the individual finger--DNA interaction seen in the Zif268--DNA structure has been conserved in many zinc finger--DNA interactions.

Amino Acid Sequence↗

A histidine gene cluster of the hyperthermophile Thermotoga maritima: sequence analysis and evolutionary significance.

The sequences of histidine operon genes in hyperthermophiles are informative for understanding high protein thermostability and the evolution of metabolic pathways. Therefore, a cluster of eight his genes from the hyperthermophilic and phylogenetically early bacterium Thermotoga maritima was cloned and sequenced. The cluster has the gene order hisDCBdHAFI-E, lacking only hisG and hisBp, and does not contain intercistronic regions. This compact organization of his genes resembles the his operon of enterobacteria. Sequence analysis downstream of the stop codon of hisI-E identifies a region with a significantly higher cytosine over guanosine content, which is indicative of a rho-dependent termination of transcription of the his operon. Multiple sequence alignments of N1-((5'-phosphoribosyl)-formimino)-5-aminoimidazole-4-carboxyam ide ribonucleotide isomerase (HisA) and of the cycloligase moiety of imidazoleglycerol phosphate synthase (HisF) support the previous assignment of the (beta alpha)8-barrel fold to these proteins. The alignments also reveal a second phosphate-binding motif located in the first halves of both enzymes and thereby support the hypothesis that HisA and HisF have evolved by a sequence of two gene duplication events. Comparison of the amino acid compositions of HisA and HisF from mesophiles and thermophiles shows that the thermostable variants of both enzymes contain a significantly increased number of charged amino acid residues and may therefore be stabilized by additional salt bridges.

Aldose-Ketose Isomerases↗

Comparison of TP53 mutations identified by oligonucleotide microarray and conventional DNA sequence analysis.

As the rate of gene discovery accelerates, more efficient methods are needed to analyze genes in human tissues. To assess the efficiency, sensitivity, and specificity of different methods, alterations of TP53 were independently evaluated in 108 ovarian tumors by conventional DNA sequence analysis and oligonucleotide microarray (p53 GeneChip). All mutations identified by oligonucleotide microarray and all disagreements with conventional gel-based DNA sequence analysis were confirmed by re-analysis with manual and automated dideoxy DNA sequencing. A total of 77 ovarian cancers were identified as having TP53 mutations by one of the two approaches, 71 by microarray and 63 by gel-based DNA sequence analysis. The same mutation was identified in 57 ovarian cancers, and the same wild type TP53 sequence was observed in 31 ovarian cancers by both methods, for a concordance rate of 81%. Among the mutation analyses discordant by these methods for TP53 sequence were 14 cases identified as mutated by microarray but not by conventional DNA sequence analysis and 6 cases identified as mutated by conventional DNA sequence analysis but not by microarray. Overall, the oligonucleotide microarray demonstrated a 94% accuracy rate, a 92% sensitivity, and an 100% specificity. Conventional DNA sequence analysis demonstrated an 87% accuracy rate, 82% sensitivity, and a 100% specificity. Patients with TP53 mutations had significantly shorter overall survival than those with no mutation (P = 0.02). Women with mutations in loop2, loop3, or the loop-sheet-helix domain had shorter survival than women with other mutations or women with no mutations (P = 0.01). Although further refinement would be helpful to improve the detection of certain types of TPS3 alterations, oligonucleotide microarrays were shown to be a powerful and effective tool for TP53 mutation detection.

Female↗

Simultaneous identification of rifampin-resistant Mycobacterium tuberculosis and nontuberculous mycobacteria by polymerase chain reaction-single strand conformation polymorphism and sequence analysis of the RNA polymerase gene (rpoB).

Interspecies variations and mutations associated with rifampin resistance in rpoB of Mycobacterium allow for the simultaneous identification of rifampin-resistant Mycobacterium tuberculosis and nontuberculous mycobacteria by PCR-SSCP analysis and PCR- sequencing. One hundred and ten strains of rifampin-susceptible M. tuberculosis, 14 strains of rifampin-resistant M. tuberculosis, and four strains of the M. avium complex were easily identified by PCR-SSCP. Of another seven strains, which showed unique SSCP patterns, three were identified as rifampin-resistant M. tuberculosis and four as M. terrae complex by subsequent sequence analysis of their rpoB DNAs (306 bp). These results were concordant with those obtained by susceptibility testing, biochemical identification, and 16S rDNA sequencing.

Antitubercular Agents↗

Expression and sequence analysis of cDNAs induced during the early stages of tuberisation in different organs of the potato plant (Solanum tuberosum L.).

cDNA clones of two genes (TUB8 and TUB13) which show a 25-30-fold increase in transcript in the stolon tip during the early stages of tuberisation, have been isolated by differential screening. These genes are also expressed in leaves, stems and roots and the expression pattern in these organs changes on tuberisation. Southern analysis shows homologous sequences in the non-tuberising wild type potato species Solanum brevidens and in Lycopersicon esculentum (tomato). Sequence analysis reveals a high degree of similarity between the TUB13 cDNA, and a human S-adenosylmethionine decarboxylase gene. The predicted TUB8 peptide sequence shows several repeats of alanine, glutamate and proline which suggests a structural role for the encoded protein.

Adenosylmethionine Decarboxylase↗

[Amplification, cloning and sequence analysis of spider dragline silk cDNA].

Spider dragline silk is synthesized in special gland named major ampulate (MA) gland. The MA glands were dissected from the abdomen of the spiders Nephila clavata and the total RNA was extracted by the TRIZOL. The cDNA of dragline silk was amplificated by RT-PCR (reverse transcription polymerase chain reaction), multiplex PCR and cloned. PCR identification, restriction analysis and DNA sequence analysis were carried out to verify the recombinant plasmids. The codon usage frequencies of the cloned cDNA were added up, and the predicted amino acid sequence was compared with Spidroin2 of Nephila clavipes. Predicted secondary structure of the predicted amino-acid sequence was analysized by DNAStar software. All results showed that the cloned cDNA we got (GenBank Accession No. AF441245) was the very fragment of spider dragline silk Spidroin2 cDNA.

Amino Acid Sequence↗

Isolation and sequence analysis of cDNAs for the major potato tuber protein, patatin.

A cDNA library from membrane-bound poly(A)+ mRNA of developing potato tubers was constructed and two classes of essentially full-length cDNA clones for a precursor to the major tuber glycoprotein, patatin were isolated. Sequence analysis shows that the two classes are approximately 99% homologous and correspond to the two major species of patatin identified previously by NH2-terminal amino acid sequence analysis. Sequence analysis also predicts that patatin is synthesized with a 23 amino acid signal sequence. Northern blot analysis shows that patatin mRNAs are 1550 +/- 50 nucleotides in length and are normally not present in polyribosomal or total RNA from stems or leaves.

Amino Acid Sequence↗

Campylobacter spp. subtype analysis using gel-based repetitive extragenic palindromic-PCR discriminates in parallel fashion to flaA short variable region DNA sequence analysis.

AIMS: The repetitive extragenic palindromic-PCR (rep-PCR) subtyping technique, which targets repetitive extragenic DNA sequences in a PCR, was optimized for Campylobacter spp. These data were then used for comparison with the established genotyping method of flaA short variable region (SVR) DNA sequence analysis as a tool for molecular epidemiology. METHODS AND RESULTS: Uprime Dt, Uprime B1 or Uprime RI primers were utilized to generate gel-based fingerprints from a set of 50 Campylobacter spp. isolates recovered from a variety of epidemiological backgrounds and sources. Analysis and phenogram tree construction, using the unweighted pair group method with arithmetic mean, of the generated fingerprints demonstrated that the Uprime Dt primers were effective in providing reproducible patterns (100% typability, 99% reproducibility) and at placing isolates into epidemiological relevant groups. Genetic stability of the rep-PCR Uprime Dt patterns under nonselective, short-term transfer conditions revealed a Pearson's correlation approaching 99%. These same 50 Campylobacter spp. isolates were analysed by flaA SVR DNA sequence analysis to obtain phylogenetic relationships. CONCLUSIONS: The Uprime Dt primer-generated rep-PCR phenogram was compared with a phenogram generated from flaA SVR DNA sequence analysis of the same isolates. Comparison of the two sets of resulting genomic relationships revealed that both methods segregated isolates into similar groups. SIGNIFICANCE AND IMPACT OF THE STUDY: These results indicate that rep-PCR analysis performed using the Mo Bio Ultra Clean Microbial Genomic DNA Isolation Kit for DNA isolation and the Uprime DT primer set for amplification is a useful and effective tool for accurate differentiation of Campylobacter spp. for subtyping and epidemiological analyses.

Bacterial Typing Techniques↗

Microcomputer programs for DNA sequence analysis.

Computer programs are described which allow (a) analysis of DNA sequences to be performed on a laboratory microcomputer or (b) transfer of DNA sequences between a laboratory microcomputer and another computer system, such as a DNA library. The sequence analysis programs are interactive, do not require prior experience with computers and in many other respects resemble programs which have been written for larger computer systems (1-7). The user enters sequence data into a text file, accesses this file with the programs, and is then able to (a) search for restriction enzyme sites or other specified sequences, (b) translate in one or more reading frames in one or both directions in order to find open reading frames, or (c) determine codon usage in the sequence in one or more given reading frames. The results are given in table format and a restriction map is generated. The modem program permits collection of large amounts of data from a sequence library into a permanent file on the microcomputer disc system, or transfer of laboratory data in the reverse direction to a remote computer system.

Base Sequence↗

Molecular analysis of rifampin-resistant Mycobacterium tuberculosis isolated from Korea by polymerase chain reaction-single strand conformation polymorphism sequence analysis.

OBJECTIVE: To assess the molecular mechanism of rifampin (RMP) resistance in clinical strains of Mycobacterium tuberculosis. DESIGN: The molecular nature of a part of the rpoB gene in 77 M. tuberculosis clinical strains isolated in Korea was analyzed using polymerase chain reaction-single strand conformation polymorphism (PCR-SSCP) and PCR-sequence analysis. RESULTS: Among 67 RMP-resistant isolates, 50 showed SSCP profiles different from that of an RMP-sensitive control strain, M. tuberculosis H37Rv, indicating the possible existence of a sequence alteration in this region of the rpoB gene, while 17 resistant isolates displayed SSCP profiles indistinguishable from that of the sensitive control strain. Subsequently, 17 clinical isolates whose SSCP profiles were difficult to distinguish from the control strain were subjected to sequence analysis. The analysis revealed that all 17 isolates did indeed contain mutations in the 81 bp region of the rpoB gene, which is associated with RMP resistance. CONCLUSION: The results from our study clearly indicate that the molecular mechanism of RMP resistance in M. tuberculosis isolates from Korea involves alterations in the rpoB gene. In addition, this study suggests that PCR-direct sequence analysis works more efficiently and accurately than PCR-SSCP analysis for rapid screening of RMP-resistant M. tuberculosis clinical isolates.

Antibiotics, Antitubercular↗

GATA: a graphic alignment tool for comparative sequence analysis.

BACKGROUND: Several problems exist with current methods used to align DNA sequences for comparative sequence analysis. Most dynamic programming algorithms assume that conserved sequence elements are collinear. This assumption appears valid when comparing orthologous protein coding sequences. Functional constraints on proteins provide strong selective pressure against sequence inversions, and minimize sequence duplications and feature shuffling. For non-coding sequences this collinearity assumption is often invalid. For example, enhancers contain clusters of transcription factor binding sites that change in number, orientation, and spacing during evolution yet the enhancer retains its activity. Dot plot analysis is often used to estimate non-coding sequence relatedness. Yet dot plots do not actually align sequences and thus cannot account well for base insertions or deletions. Moreover, they lack an adequate statistical framework for comparing sequence relatedness and are limited to pairwise comparisons. Lastly, dot plots and dynamic programming text outputs fail to provide an intuitive means for visualizing DNA alignments. RESULTS: To address some of these issues, we created a stand alone, platform independent, graphic alignment tool for comparative sequence analysis (GATA http://gata.sourceforge.net/). GATA uses the NCBI-BLASTN program and extensive post-processing to identify all small sub-alignments above a low cut-off score. These are graphed as two shaded boxes, one for each sequence, connected by a line using the coordinate system of their parent sequence. Shading and colour are used to indicate score and orientation. A variety of options exist for querying, modifying and retrieving conserved sequence elements. Extensive gene annotation can be added to both sequences using a standardized General Feature Format (GFF) file. CONCLUSIONS: GATA uses the NCBI-BLASTN program in conjunction with post-processing to exhaustively align two DNA sequences. It provides researchers with a fine-grained alignment and visualization tool aptly suited for non-coding, 0-200 kb, pairwise, sequence analysis. It functions independent of sequence feature ordering or orientation, and readily visualizes both large and small sequence inversions, duplications, and segment shuffling. Since the alignment is visual and does not contain gaps, gene annotation can be added to both sequences to create a thoroughly descriptive picture of DNA conservation that is well suited for comparative sequence analysis.

Algorithms↗

Identification by sequence analysis of a second rat brain cDNA encoding the dopamine (D2) receptor.

A rat brain cDNA library constructed in lambda ZAP II was screened with three oligonucleotide probes based on the reported coding region of the D2 receptor gene, RGB-2. A complete cDNA clone, D2(8)-1, showing positive signals with the three probes was subsequently identified by restriction analysis and dideoxy sequence analysis to be a variant of the RGB-2 gene. Comparison of the two genes revealed almost complete homology except that D2(8)-1 contains an 87 bp insert within the protein coding region and 265 additional nucleotides 5' upstream from the 5' end reported for RGB-2. It is suggested that at least two mRNA species encoding for D2 receptors exist in rat brain, possibly resulting from alternative splicing of RNA.

Amino Acid Sequence↗

ANTHEPROT 2.0: a three-dimensional module fully coupled with protein sequence analysis methods.

ANTHEPROT is a fully interactive graphics program devoted to the analysis of the sequences and structures of proteins. This program, originally developed to facilitate the protein sequence analysis coupled with multiple alignments and predicted secondary structures of proteins, now comprises a powerful 3D module to display and handle macromolecular structures. All the methods that were previously integrated into ANTHEPROT are now directly coupled with a 3D window that provides the user all the classic features of a molecular modeling package. Indeed, it allows real-time rotation and translation of 3D structures with many kinds of models in depth-cueing mode (space filling, backbone, wire models, main chain, and ribbons), selections (atom type, residue type, segments, and chain), color-coding systems (amino acid properties, predicted or observed secondary structures, temperature B factor, and subunits), geometric calculations (Ramachandran plot, distances, and angles), and fitting molecules. Stereo views are possible as well as HPGL standard files. A module specifically devoted to the determination of 3D structures using nuclear magnetic resonance is also available. This major release of our program for IBM rs6000 workstations is available by anonymous ftp to ibcp.fr for academic institutions.

Antigens↗