PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

DNA microarray analysis of predominant human intestinal bacteria in fecal samples.

A microarray method was developed for the detection of 40 bacterial species reported in the literature to be predominant in the human gastrointestinal tract. The 40 species include seven species each of Bacteroides and Clostridium, six species of Ruminococcus, five species of Bifidobacterium, four species of Eubacterium, two species each of Fusobacterium, Lactobacillus and Enterococcus, and single species each of Collinsella, Eggerthella, Escherichia, Faecalibacterium and Finegoldia. Three 40-mer oligos specific for each bacterial species were designed based on comparison of the 16S rDNA sequences available in the GenBank database, and were used to make the DNA-array on epoxy slides. Using two universal primers, the 16S rRNA gene from bacteria present in fecal samples were amplified and labeled with Cyanine5-dCTP by PCR, and then hybridized to the DNA-array. After resolving some difficulties caused by sequence conflicts in GenBank and inaccurate reference strains, all 40 bacterial reference species gave positive results. The microarray method was used to screen fecal samples obtained from 11 healthy human volunteers for the presence of these intestinal bacteria. The results indicated that 25-37 of the 40 species could be detected in each fecal sample and that 33 of the species were found in a majority of the samples.

Bacteria↗

Selection for 3' end triplets for polymerase chain reaction primers.

The 3' end of a primer is a key component of PCR primer design. Many recommendations for the composition and sequence of the 3' end have been suggested based on theoretical considerations, but have not been verified experimentally. We analyzed 3' end triplets of PCR primer sequences obtained from refereed journal articles, to test those recommendations and to make empirical recommendations for primer design. The frequencies of the 64 possible 3'end triplets among 2137 PCR primers from the VirOligo database were not uniformly distributed. From the analysis, we found that unfavored and preferred 3' end triplets existed, and that the apparent preferences were not due to base compositions in viral genome sequences. Comparison of the sequences preferred by practitioners to those recommended, suggested that no single recommendation is entirely satisfactory. We suggest that recommendations be replaced with a scoring system incorporating empirical frequencies such as those reported here.

Base Composition↗

A study of spoligotyping-defined Mycobacterium tuberculosis clades in relation to the origin of peopling and the demographic history in Madagascar.

Despite well-developed tuberculosis (TB) control policies in Madagascar, the incidence of TB remains high and is estimated at about 100 new cases per 100000 inhabitants. This paper describes genetic characteristics of TB bacilli in Madagascar. Using an international spoligotyping database, SpolDB4, we also attempted to identify the origin of strains circulating in Madagascar. DNA polymorphism of 333 Mycobacterium tuberculosis complex isolates was assessed. A total of 301 isolates belonging to 60 spoligotyping-defined clusters were found, whereas 32 isolates harbored orphan patterns. By comparison with the international database, we identified a new genetic group of closely genetically related M. tuberculosis strains which we suggested to be specific from Madagascar. Most of them belonging to the East-African-Indian (EAI) superfamily of strains that are responsible for 14% of total TB cases (shared types ST1514-1525). These strains are closely related to the most prevalent shared type ST109, whose distribution is mainly confined to Madagascar. The observed distribution of genotypes shows that principal genetic group 1 strains (EAI, Beijing, CAS, Afri, "Manu") is high (35.4%) suggesting an ancient evolutionary history of tuberculosis in Madagascar, in relation to the origin of peopling and the demographic history.

Bacterial Typing Techniques↗

In silico mining in expressed sequences of Neurospora crassa for identification and abundance of microsatellites.

In the present study, 3217 UniGene sequences of Neurospora crassa downloaded from the National Center for Biotechnology Information (NCBI) were mined for the identification of microsatellites or simple sequence repeats (SSRs). A total of 287 SSRs detected gives density of 1SSR/14.6 kb of 4187.86 kb sequences mined suggests that only 250 (7.8%) of sequences contained SSRs. Depending on the repeat units, the length of SSRs ranged from 14 to 17 bp for mono-, 14 to 48 bp for di-, 18 to 90 bp for tri-, 24 to 48 bp for tetra-, 30 for penta- and 42 to 48 bp for hexa-nucleotide repeats. Tri-nucleotide repeats were the most frequent repeat type (88.8%) followed by di-nucleotide repeats (5.9%). An attempt was also made with the help of bioinformatics approach to find out primer pairs for identified SSRs and primers were found only for 239 sequences. But, this part needs experimental validation. Annotation of SSRs containing sequences was also carried out.

Computational Biology↗

The MitoDrome database annotates and compares the OXPHOS nuclear genes of Drosophila melanogaster, Drosophila pseudoobscura and Anopheles gambiae.

The oxidative phosphorylation (OXPHOS) is the primary energy-producing process of all aerobic organisms and the only cellular function under the dual control of both the mitochondrial and the nuclear genomes. Functional characterization and evolutionary study of the OXPHOS system is of great importance for the understanding of many as yet unclear aspects of nucleus-mitochondrion genomic co-evolution and co-regulation gene networks. The MitoDrome database is a web-based database which provides genomic annotations about nuclear genes of Drosophila melanogaster encoding for mitochondrial proteins. Recently, MitoDrome has included a new section annotating genomic information about OXPHOS genes in Drosophila pseudoobscura and Anopheles gambiae and their comparative analysis with their Drosophila melanogaster and human counterparts. The introduction of this new comparative annotation section into MitoDrome is expected to be a useful resource for both functional and structural genomics related to the OXPHOS system.

Animals↗

Widespread distribution of antisense transcripts in the Plasmodium falciparum genome.

The availability of the complete genome sequence of Plasmodium falciparum has facilitated high-throughput profiling of its complex life cycle, following the application of micro-array, proteomic, and serial analysis of gene expression (SAGE) technologies in this system. These, in turn, have yielded unprecedented insight into global gene expression, including the foremost demonstration of antisense transcription in the parasite. For example, owing to its inherent ability to sample novel ORFs and to predict transcript orientation, SAGE analysis in asexual forms led to the initial discovery of highly abundant antisense RNAs. To determine the extent of this phenomenon in P. falciparum, we have surveyed the distribution of both sense and antisense transcripts across the asexual transcriptome for the first time. To this end, a relational database integrating SAGE expression data with genome annotation information was constructed. This allowed the comprehensive annotation of a total of 17245 SAGE tags, extending over a 350-fold expression range. Transcripts from approximately 30% of the estimated 3D7 gene loci were present at detectable levels in mixed asexual stages, where loci involved in invasion and immune evasion; and carbohydrate metabolism were highly represented in the sense transcriptome. Approximately 12% of SAGE tags, however, were derived from the non-coding strand of nuclear-encoded ORFs, indicating that endogenous antisense RNAs are widespread in this system. Notably, these antisense transcripts were absent from the mitochondrial genome. Interestingly, we note that sense and antisense tag counts from single loci across the transcriptome were inversely related. Taken together, this data may provide first hints as to the possible function of antisense transcription in this system.

Animals↗

Exploration of neuroendocrine and immune gene expression in peripheral blood mononuclear cells.

As pathways of communication between the nervous, endocrine, and immune systems are identified, the importance of the interplay of these systems for health and well-being is increasingly recognized. In this study, we created a comprehensive database of 1622 genes likely to be involved in synthetic, biochemical, and regulatory psycho-neuroendocrine-immune (PNI) pathways. Expression of 1058 of these genes was detected in the peripheral blood by querying both a peripheral blood-specific expressed sequence tag (EST) database and a peripheral blood database generated from microarray evaluation of 30,000 genes. Several neural and endocrine genes were expressed in the peripheral blood including hormone receptors, a hormone-responsive transcription factor, and neurotransmitter receptors. These findings document the expression of nervous and endocrine genes in the peripheral blood that have previously only been characterized in the respective system tissues, and indicate that the blood is a rich source of information that should help in deciphering the communication between the mind and the body.

Databases, Nucleic Acid↗

Mitochondrial genomes of parasitic nematodes--progress and perspectives.

Mitochondria are subcellular organelles in which oxidative phosphorylation and other important biochemical functions take place within the cell. Within these organelles is a mitochondrial (mt) genome, which is distinct from, but cooperates with, the nuclear genome of the cell. Studying mt genomes has implications for various fundamental areas, including mt biochemistry, physiology and molecular biology. Importantly, the mt genome is a rich source of markers for population genetic and systematic studies. To date, more than 696 mt genomes have been sequenced for a range of metazoan organisms. However, few of these are from parasitic nematodes, despite their socioeconomic importance and the need for fundamental investigations into areas such as nematode genetics, systematics and ecology. In this article, we review knowledge and recent progress in mt genomics of parasitic nematodes, summarize applications of mt gene markers to the study of population genetics, systematics, epidemiology and evolution of key nematodes, and highlight some prospects and opportunities for future research.

Animals↗

Data mining of Mycobacterium tuberculosis complex genotyping results using mycobacterial interspersed repetitive units validates the clonal structure of spoligotyping-defined families.

Recently, a combination of spoligotyping and bioinformatics was proposed as a potential tool for defining major circulating clades of tuberculosis bacilli. In the present study, we attempted to validate the above mentioned classification using a new high-throughput marker, named mycobacterial interspersed repetitive units (MIRUs). Using 12 MIRU loci and spoligotyping, we performed data mining of results on clinical isolates of the Mycobacterium tuberculosis complex representative of global mycobacterial allelic diversity. Knowledge rules permitting automatic labeling of major M. tuberculosis families were defined. Using this strategy, MIRU 24 appeared to be most appropriate for classifying our dataset. The Bovis family was shown to be perfectly classified by a maximum of 3 MIRUs, followed by Africanum and East African Indian (EAI) families by 4 MIRUs, the Beijing family by 6 MIRUs, Haarlem and X families by 8 MIRUs, the T family by 9, and the Latin-American and Mediterranean (LAM) family by 10 MIRUs. Considering the hierarchy of family divergence, our results corroborate a recent suggestion that EAI is the ancestral family followed by Africanum and Bovis. On the other hand, T, X, LAM and Haarlem families appear to be of more recent evolution. These results indicate that data mining of MIRUs is a valuable new tool for analyzing the evolutionary dynamics of the M. tuberculosis complex, and for monitoring an infectious disease such as tuberculosis.

Bacterial Typing Techniques↗

Extracting novel information from gene expression data.

Data from high throughput technologies, such as DNA microarrays, necessitated the development of new computational methodologies for analyzing the high dimensional information contained within the gene expression data. Liao's group suggested the use of network component analysis to predict transcription factor activities by integrating gene expression data from Escherichia coli with known connectivity information between their genes and transcription factors. This introduces an approach for obtaining novel information from gene expression data.

Computational Biology↗

Meta-analysis for microarray studies of the genetics of complex traits.

In comparison to other complex disease traits, alcoholism and alcohol abuse are influenced by the combined effects of many genes that alter susceptibility, phenotypic expression and associated morbidity, respectively. Many genetic studies, in both animal models and humans, have identified genetic intervals containing genes that influence alcoholism or behavioral responses to ethanol. Concurrently, a growing number of microarray studies have identified gene expression differences related to ethanol drinking or other ethanol behaviors. However, concerns about the statistical power of these experiments, combined with the complexity of the underlying phenotypes, have greatly hampered the identification of candidate genes underlying ethanol behaviors. Meta-analysis approaches using recent compilations of large datasets of microarray, behavioral and genetic data promise improved statistical power for detecting the genes or gene networks affecting ethanol behaviors and other complex traits.

Alcohol Drinking↗

Enhancer sequence conservation between vertebrates is favoured in developmental regulator genes.

Sequence conservation has been used to find genes and to pinpoint functional non-coding sequences such as transcriptional regulatory elements. In this article, we analysed the conservation of 104 experimentally validated murine enhancer sequences between the mouse and zebrafish genomes. Surprisingly, only 10.5% of the mouse enhancers have homologues in zebrafish. All of the genes with conserved cis-elements have regulatory functions during embryonic development, perhaps reflecting substantial structural constraints on the integration of spatio-temporal signalling cues during the formation of the vertebrate body.

Animals↗

Computational discovery of DNA motifs associated with cell type-specific gene expression in Ciona.

Temporally and spatially co-expressed genes are expected to be regulated by common transcription factors and therefore to share cis-regulatory elements. In the ascidian Ciona intestinalis, the whole-genome sequences and genome-scale gene expression profiles allow the use of computational techniques to investigate cis-elements that control transcription. We collected 5' flanking sequences of 50 tissue-specific genes from genome databases of C. intestinalis and a closely related species Ciona savignyi. We searched for DNA motifs over-represented in upstream regions of a group of co-expressed genes. Several motifs were distributed predominantly in upstream regions of photoreceptor, pan-neuronal, or muscle-specific gene groups. One muscle-specific motif, M2, was distributed preferentially in regions from -200 to -100 bp relative to the translational start sites. Promoters of muscle-specific genes of C. intestinalis were isolated, connected with a green fluorescent protein gene (GFP), and introduced into C. intestinalis embryos. In muscle cells, these promoters specifically drove GFP expression, which mutations of the M2 sites greatly reduced. When M2 sites were located upstream of a basal promoter, the reporter GFP was specifically expressed in muscle cells. These results suggest the validity of our computational prediction of cis-regulatory elements. Thus, bioinformatics can help identify cis-regulatory elements involved in chordate development.

Animals↗

A database of mRNA expression patterns for the sea urchin embryo.

We present an initial characterization of a database that contains temporal expression profiles of sequences found in 35,282 gene predictions within the sea urchin genome. The relative RNA abundance for each sequence was determined at 5 key stages of development using high-density oligonucleotide microarrays that were hybridized with populations of polyA+ RNA sequence. These stages were two-cell, which represents maternal RNA, early blastula, the time at which major tissue territories are specified, early and late gastrula, during which important morphogenetic events occur, and the pluteus larva, which marks the culmination of pre-feeding embryogenesis. We provide evidence that the microarray reliably reports the temporal profiles for the large majority of predicted genes, as shown by comparison to data for many genes with known expression patterns. The sensitivity of this assay allows detection of mRNAs whose concentration is only several hundred copies/embryo. The temporal expression profiles indicate that 5% of the gene predictions encode mRNAs that are found only in the maternal population while 24% are embryo-specific. Further, we find that the concentration of >80% of different mRNAs is modulated by more than a factor of 3 during development. Along with the annotated sea urchin genome sequence and the whole-genome tiling array (the transcriptome, Samanta, M., Tongprasit, W., Istrrail, S., Cameron, R., Tu, Q., Davidson, E., Stolc, V., in press. A high-resolution transcriptome map of the sea urchin embryo. Science), this database proves a valuable resource for designing experiments to test the function of specific genes during development.

Animals↗

Characterization of the segmental duplication LCR7-20 in the human genome.

Our previous study described the amplification of a genomic sequence containing exon 9 of CFTR in the human genome. Here we report that this CFTR sequence is part of a large duplicated sequence unit, provisionally named LCR7-20. Through successive screening of two human chromosome 7-specific cosmid libraries to construct a cosmid contig, we assembled two sequenced BAC clones into a single contig containing a prototypic LCR7-20 unit. Subsequent searches of existing human genome sequences identified additional six copies of LCR7-20-like sequences with more than 90% sequence homology. Additional genomic clones containing LCR7-20-like sequences were then isolated from total genomic BAC and PAC libraries. Restriction fragment analysis and limited sequencing data indicated that there could be around 30 copies of LCR7-20-like sequences in the human genome and that the average region of homology could extend over 120 kb. As indicated by fluorescence in situ hybridization analysis, LCR7-20-like sequences are dispersed on different chromosomes, mainly in the centromeric and pericentromeric regions, and some may exist in tandem copies. Our study also indicates that many genomic regions containing LCR7-20's either have been misassembled or are missing in current versions of the human genome sequence.

Blotting, Southern↗

Isolation and analysis of candidate myeloid tumor suppressor genes from a commonly deleted segment of 7q22.

Monosomy 7 and deletions of 7q are recurring leukemia-associated cytogenetic abnormalities that correlate with adverse outcomes in children and adults. We describe a 2.52-Mb genomic DNA contig that spans a commonly deleted segment of chromosome band 7q22 identified in myeloid malignancies. This interval currently includes 14 genes, 19 predicted genes, and 5 predicted pseudogenes. We have extensively characterized the FBXL13, NAPE-PLD, and SVH genes as candidate myeloid tumor suppressors. FBXL13 encodes a novel F-box protein, SVHis a member of a gene family that contains Armadillo-like repeats, and NAPE-PLD encodes a phospholipase D-type phosphodiesterase. Analysis of a panel of leukemia specimens with monosomy 7 did not reveal mutations in these or in the candidate genes LRRC17, PRO1598, and SRPK2. This fully sequenced and annotated contig provides a resource for candidate myeloid tumor suppressor gene discovery.

Base Sequence↗

Comprehensive analysis of pathway or functionally related gene expression in the National Cancer Institute's anticancer screen.

We have analyzed the level of gene coregulation, using gene expression patterns measured across the National Cancer Institute's 60 tumor cell panels (NCI(60)), in the context of predefined pathways or functional categories annotated by KEGG (Kyoto Encyclopedia of Genes and Genomes), BioCarta, and GO (Gene Ontology). Statistical methods were used to evaluate the level of gene expression coherence (coordinated expression) by comparing intra- and interpathway gene-gene correlations. Our results show that gene expression in pathways, or groups of functionally related genes, has a significantly higher level of coherence than that of a randomly selected set of genes. Transcriptional-level gene regulation appears to be on a "need to be" basis, such that pathways comprising genes encoding closely interacting proteins and pathways responsible for vital cellular processes or processes that are related to growth or proliferation, specifically in cancer cells, such as those engaged in genetic information processing, cell cycle, energy metabolism, and nucleotide metabolism, tend to be more modular (lower degree of gene sharing) and to have genes significantly more coherently expressed than most signaling and regular metabolic pathways. Hierarchical clustering of pathways based on their differential gene expression in the NCI(60) further revealed interesting interpathway communications or interactions indicative of a higher level of pathway regulation. The knowledge of the nature of gene expression regulation and biological pathways can be applied to understanding the mechanism by which small drug molecules interfere with biological systems.

Algorithms↗

The presence of GC-AG introns in Neurospora crassa and other euascomycetes determined from analyses of complete genomes: implications for automated gene prediction.

A combination of experimental and computational approaches was employed to identify introns with noncanonical GC-AG splice sites (GC-AG introns) within euascomycete genomes. Evaluation of 2335 cDNA-confirmed introns from Neurospora crassa revealed 27 such introns (1.2%). A similar frequency (1.0%) of GC-AG introns was identified in Fusarium graminearum, in which 3 of 292 cDNA-confirmed introns contained GC-AG splice sites. Computational analyses of the N. crassa genome using a GC-AG intron consensus sequence identified an additional 20 probable GC-AG introns in this fungus. For 8 of the 47 GC-AG introns identified in N. crassa a GC donor site is also present in a homolog from Magnaporthe grisea, F. graminearum, or Aspergillus nidulans. In most cases, however, homologs in these fungi contain a GT-AG intron or no intron at the corresponding position. These findings have important implications for fungal genome annotation, as the automated annotations of euascomycete genomes incorrectly identified intron boundaries for all of the confirmed and probable GC-AG introns reported here.

Alternative Splicing↗