PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Good-Enough RFLP Matcher (GERM) program.

A spreadsheet-based program (Good-Enough RFLP Matcher or GERM) is presented that matches unknown restriction fragment length polymorphism (RFLP) patterns of ectomycorrhizal fungi to a database of known ectomycorrhizal fungi. The program uses three simple methods to determine whether a sample matches a known: (1) Forward Matching: whether every band in the unknown is present in a known sample within a given error range; (2) Backward Matching: whether every band in the known sample is present in the unknown within a given error range; (3) Sum of Bands: whether the sum of all bands in the known and unknown are similar within a given error range. The program is available through the web page of this journal.

Database Management Systems↗

Codon usage bias amongst plant viruses.

An internet database (DPVweb) was established containing details of all sequences of viruses, viroids and satellites of plants that are complete or that contain at least one complete gene (n>4600). The start and end positions of each feature (genes, non-translated regions etc) were recorded and checked for accuracy. Client software was written to enable easy selection of sequences and features of a chosen virus and to analyse codon usage bias. Codon usage was analysed for each gene of one example of each fully-sequenced plant virus. There were large differences in codon preferences, related to the nucleotide composition of the genome, particularly the GC content of the third codon position. There was no effect of gene size on codon bias. Genes from the same genome usually had similar coding strategies except where constrained by the overlap of reading frames. Although some synonymous codons were consistently used with low frequency by both plants and viruses, viruses were not generally adapted to use (or avoid) those codons most frequently used by their host plants and there was no obvious association with the type of transmission. Mutational bias, rather than translational selection appears to account for the majority of the variation detected. The software is available at http://www.dpvweb.net/analysis/codons.php.

Codon↗

The core region of the coat protein gene is highly useful for establishing the provisional identification and classification of begomoviruses.

Polymerase chain reaction (PCR) was applied to detect and establish provisional identity of begomoviruses through amplification of a approximately 575 bp fragment of the begomoviral coat protein gene (CP), referred to as the 'core' region of the CP gene (core CP). The core CP fragment contains conserved and unique regions, and was hypothesized to constitute a sequence useful for begomovirus classification. Virus relationships were predicted by distance and parsimony analyses using the A component (bipartite viruses) or full genome (monopartite viruses), CP gene, core CP, or the 200 5'-nucleotides (nt) of the CP. Reconstructed trees and sequence divergence estimates yielded very similar conclusions for all sequence sets, while the CP 5'-200 nt was the best strain discriminator. Alignment of the core CP region for 52 field isolates with reference begomovirus sequences permitted provisional virus identification based on tree position and extent of sequence divergence. Geographic origin of field isolates was predictable based on phylogenetic separation of field isolates examined here. A 'closest match' or genus-level identification could be obtained for previously undescribed begomoviruses using the BLAST program to search a reference core CP database located at our website and/or in GenBank. Here, we describe an informative molecular marker that permits provisional begomovirus identification and classification using a begomoviral sequence that is smaller than the presently accepted, but less accessible CP sequence.

Capsid↗

Revealing gene transcription and translation initiation patterns in archaea, using an interactive clustering model.

An interactive clustering model based on positional weight matrices is described and results obtained using the model to analyze gene regulation patterns in archaea are presented. The 5' flanking sequences of ORFs identified in four archaea, Sulfolobus solfataricus, Pyrobaculum aerophilum, Halobacterium sp. NRC-1, and Pyrococcus abyssi, were clustered using the model. Three regular patterns of clusters were identified for most ORFs. One showed genes with only a ribosome-binding site; another showed genes with a transcriptional regulatory region located at a constant location with respect to the start codon. A third pattern combined the previous two. Both P. aerophilum and Halobacterium sp. NRC-1 exhibited clusters of genes that lacked any regular pattern. Halobacterium sp. NRC-1 also presented regular features not seen in the other organisms. This group of archaea seems to use a combination of eubacterial and eukaryotic regulatory features as well as some unique to individual species. Our results suggest that interactive clustering may be used to examine the divergence of the gene regulatory machinery in archaea and to identify the presence of archaea-specific gene regulation patterns.

Archaea↗

FlgM anti-sigma factors: identification of novel members of the family, evolutionary analysis, homology modeling, and analysis of sequence-structure-function relationships.

FlgM proteins, also known as Anti-sigma-28 factor (sigma28), are negative regulators of flagellin synthesis. Recently, a three-dimensional structure of the Aquifex aeolicus sigma28/FlgM complex (PDB code: 1rp3) was determined by X-ray crystallography at 2.3 A resolution. Furthermore, experimental data on bacterial FlgM, including site-directed mutagenesis and structural characterization by NMR are also available. However, an interpretation of the sequence-structure-function relationships combining X-ray and NMR data with the evolutionary information extracted from the increasing number of FlgM-related sequences annotated in databases is not available. In the present study, we combined database sequence searches and sequence-analysis tools to update the multiple sequence alignment of a previously characterized cluster of orthologs (COG2747) and the PFAM classification of protein domains (PF04316) for the FlgM family. A phylogenetic analysis of 77 protein sequences revealed the presence of at least three major sequence clades within the FlgM family. Besides, we predicted functional residues using a SequenceSpace method. We also generated homology models for Bacillus subtilis and Salmonella typhimurium FlgM proteins, for which sequence-structure-function relationship data are available, and used the docking program ClusPro to hypothesize about the dimer association between FlgM proteins. In conclusion, the analysis presented in this work will be useful in designing new experiments to understand better protein-protein interactions between FglM, sigma factors, and putative molecules from the flagellar export apparatus. Electronic Supplementary Material is available in the online version of this article at http://link.springer.de/

Bacterial Proteins↗

Catalog of 162 single nucleotide polymorphisms (SNPs) in a 4.7-kb region of the HLA-DP loci in southern Chinese ethnic groups.

HLA class-II proteins are cell-surface molecules that present antigens to T cells, and their expressional regulation is crucial to the immune reaction. Sequence variation at the regulatory region can directly affect the gene expression level. We cloned and sequenced a 4.7-kb region containing the regulatory region, exon1, and partial intron1 of both HLA-DPA1 and DPB1 genes in 25 variable sequences from southern Chinese ethnic groups and got a high-density map of 162 single nucleotide polymorphisms (SNPs): seven in 5'-flanking regions, four in 5'-untranslated regions, and four in the coding regions. By comparing these data with SNPs in dbSNP database in the NCBI, 145 SNPs (89.5%) were novel. In addition, eight genetic variations of insertion-deletion polymorphisms (INDELs) were discovered within the 4.7-kb region. These high-resolution maps can be used as resources of markers for association studies of complex diseases, assessment of individuals' predisposition to diseases, and tailoring of therapies, as well as research markers for population genetics and evolution.

Base Sequence↗

Identification of a novel polymorphism involving a CGG repeat in the PTCH gene and a genome-wide screening of CGG-containing genes.

Mutations in the human homologue of the Drosophila patched gene (PTCH) are responsible for the hereditary disorder called nevoid basal cell carcinoma syndrome (NBCCS). PTCH has a CGG triplet repeat located 4 bp upstream of the first methionine codon. Here we report a novel polymorphism involving the number of the CGG-repeat. The major allele (86.3%) contained a repeat size of seven, whereas the minor allele contained eight. No significant difference in the distributions of genotypes was observed between normal and NBCCS individuals. However, when the repeat was inserted between a heterologous promoter and the luciferase gene, the longer repeats tended to induce higher luciferase activities, suggesting that the repeat length potentially affects the levels of gene expression. A genome-wide screening revealed that 68 and 146 genes contained a CGG/CCG repeat in the coding region and in the 5'-untranslated region (5'-UTR), respectively. None of the genes had this repeat in 3'-UTR. Interestingly, the number of genes with a CGG repeat in the 5'-UTR was significantly higher than that with a CCG repeat in the 5'-UTR. The localization of a CGG/CCG repeat in PTCH is quite unique in that only four other genes have been found in which the repeat is localized up to 4 bp upstream of the first methionine.

Alleles↗

Gene expression in Florida red tide dinoflagellate Karenia brevis: analysis of an expressed sequence tag library and development of DNA microarray.

Karenia brevis (Davis) is the dinoflagellate responsible for nearly annual red tides in the Gulf of Mexico. Although the mechanisms regulating the growth and toxicity of this problematic organism are of considerable interest, little information is available on its molecular biology. We therefore constructed a complementary DNA library from which to gain insight into its expressed genome and to develop tools for studying its gene expression. Large-scale sequencing yielded 7001 high-quality expressed sequence tags (ESTs), which clustered into 5280 unique gene groups. The vast majority of genes expressed fell into a low-abundance class, with the highest expressed gene accounting for only 1% of the total ESTs. Approximately 29% of genes were found to have similarity to known sequences in other organisms after BLAST similarity comparisons to the GenBank public protein database using a cutoff of P < 10e(-4). We identified for the first time in a dinoflagellate a suite of conserved eukaryotic genes involved in cell cycle control, intracellular signaling, and the transcription and translation machinery. At least 40% of gene clusters displayed single nucleotide polymorphisms, suggesting the presence of multiple gene copies. The average GC content of ESTs was 51%, with a slight preference for G or C in the third codon position (53.5%). The ESTs were used to develop an oligonucleotide microarray containing 4629 unique features and 3462 replicate probes. Microarray labeling has been optimized, and the microarray has been validated for probe specificity and reproducibility. This is the first information to be developed on the expressed genome of K. brevis and provides the basis from which to begin functional genomic studies on this harmful algal bloom species.

Animals↗

Development of EST-SSR markers by data mining in three species of shrimp: Litopenaeus vannamei, Litopenaeus stylirostris, and Trachypenaeus birdy.

We report on the data mining of publicly available Litopenaeus vannamei expressed sequence tags (ESTs) to generate simple sequence repeat (SSRs) markers and on their transferability between related Penaeid shrimp species. Repeat motifs were found in 3.8% of the evaluated ESTs at a frequency of one repeat every 7.8 kb of sequence data. A total of 206 primer pairs were designed, and 112 loci were amplified with the highest success in L. vannamei. A high percentage (69%) of EST-SSRs were transferable within the genus Litopenaeus. More than half of the amplified products were polymorphic in a small testing panel of L. vannamei. Evaluation of those primers in a larger testing panel showed that 72% of the markers fit Hardy-Weinberg equilibrium, which shows their utility for population genetic analysis. Additionally, a set of 26 of the EST-SSRs were evaluated for Mendelian segregation. A high percentage of monomorphic markers (46%) proved to be polymorphic by singles-stranded conformational polymorphism analysis. Because of the high number of ESTs available in public databases, a data mining approach similar to the one outlined here might yield high numbers of SSR markers in many animal taxa.

Animals↗

Transposon tagging in maize.

Through recent government- and industry-sponsored efforts, several forward and reverse genetic screening programs have emerged over the past few years to aid in the genetic dissection of gene function in maize. Despite a US maize crop valued at $18.4 billion last year (http://www.ncga.com/03world/main/US_crop_value_2000.html) and rich genetic history, maize has taken a back seat to Arabidopsis thaliana as the model genetic system for plants over the past decade. With a fully sequenced genome, short generation time and small size, studies of Arabidopsis have provided plant scientists with a molecular framework for hormonal, developmental and environmental signaling pathways in plants. As investigations into Arabidopsis continue, our capacity to engineer biochemical pathways and alter plant physiological responses will become increasingly sophisticated. Nevertheless, approximately 130 million years have passed since monocot and higher eudicot lineages diverged. Thus, our ability to engineer agronomically important monocot grasses such as maize, rice and wheat will become increasingly limited by our lack of understanding of the physiological and morphological differences that have evolved in the monocots and higher eudicots. The sophisticated transposon collections now being generated for maize are but one of several recent projects (http://www.nsf.gov/bio/pubs/awards/genome01.htm) to provide grass researchers with essential tools for genome analysis. Because grain crops are such a closely related group, it is hoped that many of the findings made in one grass will be directly applicable to understanding the biology of another. The goal of this review is to highlight the recent developments in maize transposon-based gene characterization programs and provide a critical examination of the advantages and disadvantages each system offers.

DNA Transposable Elements↗

A comparative genomic analysis of ESTs from Ustilago maydis.

A large-scale comparative genomic analysis of unisequence sets obtained from an Ustilago maydis EST collection was performed against publicly available EST and genomic sequence datasets from 21 species. We annotated 70% of the collection based on similarity to known sequences and recognized protein signatures. Distinct grouping of the ESTs, defined by the presence or absence of similar sequences in the species examined, allowed the identification of U. maydis sequences present only (1) in fungal species, (2) in plants but not animals, (3) in animals but not plants, or (4) in all three eukaryotic lineages assessed. We also identified 215 U. maydis genes that are found in the ascomycete but not in the basidiomycete genome sequences searched. Candidate genes were identified for further functional characterization. These include 167 basidiomycete-specific sequences, 58 fungal pathogen-specific sequences (including 37 basidiomycete pathogen-specific sequences), and 18 plant pathogen-specific sequences, as well as two sequences present only in other plant pathogen and plant species.

Databases, Nucleic Acid↗

Ten years of bacterial genome sequencing: comparative-genomics-based discoveries.

It has been more than 10 years since the first bacterial genome sequence was published. Hundreds of bacterial genome sequences are now available for comparative genomics, and searching a given protein against more than a thousand genomes will soon be possible. The subject of this review will address a relatively straightforward question: "What have we learned from this vast amount of new genomic data?" Perhaps one of the most important lessons has been that genetic diversity, at the level of large-scale variation amongst even genomes of the same species, is far greater than was thought. The classical textbook view of evolution relying on the relatively slow accumulation of mutational events at the level of individual bases scattered throughout the genome has changed. One of the most obvious conclusions from examining the sequences from several hundred bacterial genomes is the enormous amount of diversity--even in different genomes from the same bacterial species. This diversity is generated by a variety of mechanisms, including mobile genetic elements and bacteriophages. An examination of the 20 Escherichia coli genomes sequenced so far dramatically illustrates this, with the genome size ranging from 4.6 to 5.5 Mbp; much of the variation appears to be of phage origin. This review also addresses mobile genetic elements, including pathogenicity islands and the structure of transposable elements. There are at least 20 different methods available to compare bacterial genomes. Metagenomics offers the chance to study genomic sequences found in ecosystems, including genomes of species that are difficult to culture. It has become clear that a genome sequence represents more than just a collection of gene sequences for an organism and that information concerning the environment and growth conditions for the organism are important for interpretation of the genomic data. The newly proposed Minimal Information about a Genome Sequence standard has been developed to obtain this information.

Bacterial Vaccines↗

Gene expression profiling analysis in nephrology: towards molecular definition of renal disease.

The increase in progressive kidney disease, resulting in a constantly rising prevalence of endstage renal disease (ESRD), urgently warrants the development of more effective strategies to diagnose, prevent, and intervene in renal disease. Histological information obtained by renal biopsies (RBx) is a cornerstone of the current management of kidney disease. Renal tissue can provide critical information on the disease process not available by nontissue-based approaches. However, insight gained by conventional histopathology remains limited and additional strategies to define renal disease on a molecular level are required. The sequencing of the human genome, together with recent advances in genome-wide profiling techniques, has provided the framework for a comprehensive analysis of renal disease-associated transcriptional programs. In this review, strategies to apply these technological advances towards the analysis of RBx will be described, with special emphasis on their potential impact on clinical management, but also on their inherent limitations. Finally, an outlook towards the emerging proteomic studies of renal disease will be given.

Antigens, CD20↗

Bias explorer: measurements of compositional bias in EMBL and GenBank sequence files.

A Windows application for compositional analysis of sequenced genomes (EMBL or GenBank flat files) is available as freeware. The application allows the user to quantify word bias using Markov chain analysis and it allows the user to generate sliding window data for GC-skew, AT-skew, purine excess, keto excess and discrete word counts. The mathematical routines reside in a dynamic link library (DLL), which can be used independently by other applications. The software is available for download at http://www.dfuni.dk/~anfu/Bioinformatics/Main.htm.

Bias↗

Analysis of organ-specific, expressed genes in Oncidium orchid by subtractive expressed sequence tags library.

The pseudobulb of Oncidium orchid plays a key role in water, carbohydrate, and other nutrition support during floral development, yet a large scale of gene expression analysis involved in the metabolisms have not been evaluated. By subtracting RsaI-digested cDNAs of leaf from those of psuedobulb, an efficient subtractive cDNA library was developed. In total, 1080 subtractive expressed sequence tags (ESTs) were obtained. Analysis revealed approximately 636 unique gene parts, 120 clusters and 516 singles. Of these sequences, 74.8% were annotated on the database of NCBI GenBank. Peroxidase, sodium/dicarboxylate cotransporter, and mannose-binding lectin were highly expressed. Some gene profiles were identified as related to carbohydrate metabolism involved in mannan, pectin, starch and sucrose biosynthesis. A large fraction of the ESTs (35%) were classified into transportation, stress-related, cell cycle, or regulatory functions. Most genes that were differentially expressed are important in early flowering development, carbohydrate metabolism and stress-response physiology. This efficient organ-specific EST library represented an explicit transcriptome profile of Oncidium pseudobulb.

Base Sequence↗

Identification of novel non-autonomous CemaT transposable elements and evidence of their mobility within the C. elegans genome.

We describe here two new transposable elements, CemaT4 and CemaT5, that were identified within the sequenced genome of Caenorhabditis elegans using homology based searches. Five variants of CemaT4 were found, all non-autonomous and sharing 26 bp inverted terminal repeats (ITRs) and segments (152-367 bp) of sequence with similarity to the CemaT1 transposon of C. elegans. Sixteen copies of a short, 30 bp repetitive sequence, comprised entirely of an inverted repeat of the first 15 bp of CemaT4's ITR, were also found, each flanked by TA dinucleotide duplications, which are hallmarks of target site duplications of mariner-Tc transposon transpositions. The CemaT5 transposable element had no similarity to maT elements, except for sharing identical ITR sequences with CemaT3. We provide evidence that CemaT5 and CemaT3 are capable of excising from the C. elegans genome, despite neither transposon being capable of encoding a functional transposase enzyme. Presumably, these two transposons are cross-mobilised by an autonomous transposon that recognises their shared ITRs. The excisions of these and other non-autonomous elements may provide opportunities for abortive gap repair to create internal deletions and/or insert novel sequence within these transposons. The influence of non-autonomous element mobility and structural diversity on genome variation is discussed.

Animals↗