PubMed Health⌕ Search

Biomedical subjects

Dawood B Dudekula

Publications and source records attributed to Dawood B Dudekula.

10 recordsLinked to original sources

CisView: a browser and database of cis-regulatory modules predicted in the mouse genome.

To facilitate the analysis of gene regulatory regions of the mouse genome, we developed a CisView (http://lgsun.grc.nia.nih.gov/cisview), a browser and database of genome-wide potential transcription factor binding sites (TFBSs) that were identified using 134 position-weight matrices and 219 sequence patterns from various sources and were presented with the information about sequence conservation, neighboring genes and their structures, GO annotations, protein domains, DNA repeats and CpG islands. Analysis of the distribution of TFBSs revealed that many TFBSs (N = 145) were over-represented near transcription start sites. We also identified potential cis-regulatory modules (CRMs) defined as clusters of conserved TFBSs in the entire mouse genome. Out of 739 074 CRMs, 157 442 had a significantly higher regulatory potential score than semi-random sequences generated with a 3rd-order Markov process. The CisView browser provides a user-friendly computer environment for studying transcription regulation on a whole-genome scale and can also be used for interpreting microarray experiments and identifying putative targets of transcription factors.

Animals↗

Transcript copy number estimation using a mouse whole-genome oligonucleotide microarray.

The ability to quantitatively measure the expression of all genes in a given tissue or cell with a single assay is an exciting promise of gene-expression profiling technology. An in situ-synthesized 60-mer oligonucleotide microarray designed to detect transcripts from all mouse genes was validated, as well as a set of exogenous RNA controls derived from the yeast genome (made freely available without restriction), which allow quantitative estimation of absolute endogenous transcript abundance.

Animals↗

A web-based tool for principal component and significance analysis of microarray data.

UNLABELLED: We have developed a program for microarray data analysis, which features the false discovery rate for testing statistical significance and the principal component analysis using the singular value decomposition method for detecting the global trends of gene-expression patterns. Additional features include analysis of variance with multiple methods for error variance adjustment, correction of cross-channel correlation for two-color microarrays, identification of genes specific to each cluster of tissue samples, biplot of tissues and corresponding tissue-specific genes, clustering of genes that are correlated with each principal component (PC), three-dimensional graphics based on virtual reality modeling language and sharing of PC between different experiments. The software also supports parameter adjustment, gene search and graphical output of results. The software is implemented as a web tool and thus the speed of analysis does not depend on the power of a client computer. AVAILABILITY: The tool can be used on-line or downloaded at http://lgsun.grc.nia.nih.gov/ANOVA/

Algorithms↗

Genome-wide assembly and analysis of alternative transcripts in mouse.

To build a mouse gene index with the most comprehensive coverage of alternative transcription/splicing (ATS), we developed an algorithm and a fully automated computational pipeline for transcript assembly from expressed sequences aligned to the genome. We identified 191,946 genomic loci, which included 27,497 protein-coding genes and 11,906 additional gene candidates (e.g., nonprotein-coding, but multiexon). Comparison of the resulting gene index with TIGR, UniGene, DoTS, and ESTGenes databases revealed that it had a greater number of transcripts, a greater average number of exons and introns with proper splicing sites per gene, and longer ORFs. The 27,497 protein-coding genes had 77,138 transcripts, i.e., 2.8 transcripts per gene on average. Close examination of transcripts led to a combinatorial table of 23 types of ATS units, only nine of which were previously described, i.e., 14 types of alternative splicing, seven types of alternative starts, and two types of alternative termination. The 47%, 18%, and 14% of 20,323 multiexon protein-coding genes with proper splice sites had alternative splicings, alternative starts, and alternative terminations, respectively. The gene index with the comprehensive ATS will provide a useful platform for analyzing the nature and mechanism of ATS, as well as for designing the accurate exon-based DNA microarrays. The sequence data from this study have been submitted to GenBank under accession numbers: CK329321-CK334090; CF891695-CF906652; CF906741-CF916750; CK334091-CK347104; CK387035-CK393993; CN660032-CN690720; CN690721-CN725493.

Algorithms↗

Age-associated alteration of gene expression patterns in mouse oocytes.

Decreasing oocyte competence with maternal aging is a major factor in human infertility. To investigate the age-dependent molecular changes in a mouse model, we compared the expression profiles of metaphase II oocytes collected from 5- to 6-week-old mice with those collected from 42- to 45-week-old mice using the NIA 22K 60-mer oligo microarray. Among approximately 11,000 genes whose transcripts were detected in oocytes, about 5% (530) showed statistically significant expression changes, excluding the possibility of global decline in transcript abundance. Consistent with the generally accepted view of aging, the differentially expressed genes included ones involved in mitochondrial function and oxidative stress. However, the expression of other genes involved in chromatin structure, DNA methylation, genome stability and RNA helicases was also altered, suggesting the existence of additional mechanisms for aging. Among the transcripts decreased with aging, we identified and characterized a group of new oocyte-specific genes, members of the human NACHT, leucine-rich repeat and PYD-containing (NALP) gene family. These results have implications for aging research as well as for clinical ooplasmic donation to rejuvenate aging oocytes.

Aging↗

The status, quality, and expansion of the NIH full-length cDNA project: the Mammalian Gene Collection (MGC).

The National Institutes of Health's Mammalian Gene Collection (MGC) project was designed to generate and sequence a publicly accessible cDNA resource containing a complete open reading frame (ORF) for every human and mouse gene. The project initially used a random strategy to select clones from a large number of cDNA libraries from diverse tissues. Candidate clones were chosen based on 5'-EST sequences, and then fully sequenced to high accuracy and analyzed by algorithms developed for this project. Currently, more than 11,000 human and 10,000 mouse genes are represented in MGC by at least one clone with a full ORF. The random selection approach is now reaching a saturation point, and a transition to protocols targeted at the missing transcripts is now required to complete the mouse and human collections. Comparison of the sequence of the MGC clones to reference genome sequences reveals that most cDNA clones are of very high sequence quality, although it is likely that some cDNAs may carry missense variants as a consequence of experimental artifact, such as PCR, cloning, or reverse transcriptase errors. Recently, a rat cDNA component was added to the project, and ongoing frog (Xenopus) and zebrafish (Danio) cDNA projects were expanded to take advantage of the high-throughput MGC pipeline.

Animals↗

Transcriptome analysis of mouse stem cells and early embryos.

Understanding and harnessing cellular potency are fundamental in biology and are also critical to the future therapeutic use of stem cells. Transcriptome analysis of these pluripotent cells is a first step towards such goals. Starting with sources that include oocytes, blastocysts, and embryonic and adult stem cells, we obtained 249,200 high-quality EST sequences and clustered them with public sequences to produce an index of approximately 30,000 total mouse genes that includes 977 previously unidentified genes. Analysis of gene expression levels by EST frequency identifies genes that characterize preimplantation embryos, embryonic stem cells, and adult stem cells, thus providing potential markers as well as clues to the functional features of these cells. Principal component analysis identified a set of 88 genes whose average expression levels decrease from oocytes to blastocysts, stem cells, postimplantation embryos, and finally to newborn tissues. This can be a first step towards a possible definition of a molecular scale of cellular potency. The sequences and cDNA clones recovered in this work provide a comprehensive resource for genes functioning in early mouse embryos and stem cells. The nonrestricted community access to the resource can accelerate a wide range of research, particularly in reproductive and regenerative medicine.

Animals↗

In situ-synthesized novel microarray optimized for mouse stem cell and early developmental expression profiling.

Applications of microarray technologies to mouse embryology/genetics have been limited, due to the nonavailability of microarrays containing large numbers of embryonic genes and the gap between microgram quantities of RNA required by typical microarray methods and the miniscule amounts of tissue available to researchers. To overcome these problems, we have developed a microarray platform containing in situ-synthesized 60-mer oligonucleotide probes representing approximately 22,000 unique mouse transcripts, assembled primarily from sequences of stem cell and embryo cDNA libraries. We have optimized RNA labeling protocols and experimental designs to use as little as 2 ng total RNA reliably and reproducibly. At least 98% of the probes contained in the microarray correspond to clones in our publicly available collections, making cDNAs readily available for further experimentation on genes of interest. These characteristics, combined with the ability to profile very small samples, make this system a resource for stem cell and embryogenomics research.

Animals↗

Assembly, verification, and initial annotation of the NIA mouse 7.4K cDNA clone set.

A set of 7407 cDNA clones (NIA mouse 7.4K) was assembled from >20 cDNA libraries constructed mainly from early mouse embryos, including several stem cell libraries. The clone set was assembled from embryonic and newborn organ libraries consisting of ~120,000 cDNA clones, which were initially re-arrayed into a set of ~11,000 unique cDNA clones. A set of tubes was constructed from the racks in this set to prevent contamination and potential mishandling errors in all further re-arrays. Sequences from this set (11K) were analyzed further for quality and clone identity, and high-quality clones with verified identity were re-arrayed into the final set (7.4K). The set is freely available, and a corresponding database was built to provide comprehensive annotation for those clones with known identity or homology, and has been made available through an extensive Web site that includes many link-outs to external databases and analysis servers.

Animals↗

The NIA cDNA project in mouse stem cells and early embryos.

A catalog of mouse genes expressed in early embryos, embryonic and adult stem cells was assembled, including 250000 ESTs, representing approximately 39000 unique transcripts. The cDNA libraries, enriched in full-length clones, were condensed into the NIA 15 and 7.4K clone sets, freely distributed to the research community, providing a standard platform for expression studies using microarrays. They are essential tools for studying mammalian development and stem cell biology, and to provide hints about the differential nature of embryonic and adult stem cells.

Animals↗