PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Transcriptomic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Composition and dynamics of the Caenorhabditis elegans early embryonic transcriptome.

Temporal profiles of transcript abundance during embryonic development were obtained by whole-genome expression analysis from precisely staged C. elegans embryos. The result is a highly resolved time course that commences with the zygote and extends into mid-gastrulation, spanning the transition from maternal to embryonic control of development and including the presumptive specification of most major cell fates. Transcripts for nearly half (8890) of the predicted open reading frames are detected and expression levels for the majority of them (>70%) change over time. The transcriptome is stable up to the four-cell stage where it begins rapidly changing until the rate of change plateaus before gastrulation. At gastrulation temporal patterns of maternal degradation and embryonic expression intersect indicating a mid-blastula transition from maternal to embryonic control of development. In addition, we find that embryonic genes tend to be expressed transiently on a time scale consistent with developmental decisions being made with each cell cycle. Furthermore, overall rates of synthesis and degradation are matched such that the transcriptome maintains a steady-state frequency distribution. Finally, a versatile analytical platform based on cluster analysis and developmental classification of genes is provided.

Animals↗

Detection and evaluation of intron retention events in the human transcriptome.

Alternative splicing is a very frequent phenomenon in the human transcriptome. There are four major types of alternative splicing: exon skipping, alternative 3' splice site, alternative 5' splice site, and intron retention. Here we present a large-scale analysis of intron retention in a set of 21,106 known human genes. We observed that 14.8% of these genes showed evidence of at least one intron retention event. Most of the events are located within the untranslated regions (UTRs) of human transcripts. For those retained introns interrupting the coding region, the GC content, codon usage, and the frequency of stop codons suggest that these sequences are under selection for coding potential. Furthermore, 26% of the introns within the coding region participate in the coding of a protein domain. A comparison with mouse shows that at least 22% of all informative examples of retained introns in human are also present in the mouse transcriptome. We discuss that the data we present suggest that a significant fraction of the observed events is not spurious and might reflect biological significance. The analyses also allowed us to generate a reliable set of intron retention events that can be used for the identification of splicing regulatory elements.

Animals↗

Transcriptome analysis of acetate metabolism in Corynebacterium glutamicum using a newly developed metabolic array.

Following the determination of the whole-genome sequence of Corynebacterium glutamicum, we have developed a DNA array to extensively investigate gene expression and regulation relevant to carbon metabolism. For this purpose, a total of 120 C. glutamicum genes, including those in central metabolism and amino acid biosyntheses, were amplified by PCR and printed onto glass slides. The resulting array, designated a "metabolic array", was used for hybridization with fluorescently labeled cDNA probes generated by reverse transcription from total RNA samples. As the first demonstration of transcriptome analysis in this industrially important microorganism, we applied the metabolic array to study differential transcription profiles between cells grown on glucose and on acetate as the sole carbon source. The changes in gene expression observed for the known acetate-regulated genes (aceA, aceB, pta, and ack) were well consistent with the literature data of northern analyses and enzyme assays, indicating the utility of the metabolic array in transcriptome analysis of C. glutamicum. In addition to the known responses, many previously unrecognized co-regulated genes were identified. For example, several TCA cycle genes, such as gltA, sdhA, sdhB, fumH, and mdh, and the gluconeogenic gene pck were up-regulated in the acetate medium. On the other hand, a few genes involved in glycolysis and the pentose phosphate pathway, as well as many amino acid biosynthetic genes, were down-regulated in acetate. Furthermore, two gap genes, gapA and gapB, were found to be inversely regulated, suggesting the presence of a new regulatory step for carbon metabolism between glycolysis and gluconeogenesis.

Acetates↗

The transcriptome of the intraerythrocytic developmental cycle of Plasmodium falciparum.

Plasmodium falciparum is the causative agent of the most burdensome form of human malaria, affecting 200-300 million individuals per year worldwide. The recently sequenced genome of P. falciparum revealed over 5,400 genes, of which 60% encode proteins of unknown function. Insights into the biochemical function and regulation of these genes will provide the foundation for future drug and vaccine development efforts toward eradication of this disease. By analyzing the complete asexual intraerythrocytic developmental cycle (IDC) transcriptome of the HB3 strain of P. falciparum, we demonstrate that at least 60% of the genome is transcriptionally active during this stage. Our data demonstrate that this parasite has evolved an extremely specialized mode of transcriptional regulation that produces a continuous cascade of gene expression, beginning with genes corresponding to general cellular processes, such as protein synthesis, and ending with Plasmodium-specific functionalities, such as genes involved in erythrocyte invasion. The data reveal that genes contiguous along the chromosomes are rarely coregulated, while transcription from the plastid genome is highly coregulated and likely polycistronic. Comparative genomic hybridization between HB3 and the reference genome strain (3D7) was used to distinguish between genes not expressed during the IDC and genes not detected because of possible sequence variations. Genomic differences between these strains were found almost exclusively in the highly antigenic subtelomeric regions of chromosomes. The simple cascade of gene regulation that directs the asexual development of P. falciparum is unprecedented in eukaryotic biology. The transcriptome of the IDC resembles a "just-in-time" manufacturing process whereby induction of any given gene occurs once per cycle and only at a time when it is required. These data provide to our knowledge the first comprehensive view of the timing of transcription throughout the intraerythrocytic development of P. falciparum and provide a resource for the identification of new chemotherapeutic and vaccine candidates.

Animals↗

Yellow pages to the transcriptome.

Transcriptomics has become an important tool for the large-scale analysis of biological processes. This review aims to provide sufficient criteria to make an appropriate choice among the variety of 'closed' systems, represented by DNA microarrays, and 'open' systems like fragment display, tag sequencing and subtractive hybridization, depending on the biological system under investigation. The most important technologies currently available are presented, their strengths and weaknesses are discussed and companies active in the field are listed. The potential of transcriptomics in the pharmaceutical research and development process is highlighted by applications in oncology, research on neurological diseases, and predictive toxicology. Finally, a prognosis for future developments of the technologies is given.

Animals↗

A genomic view of estrogen actions in human breast cancer cells by expression profiling of the hormone-responsive transcriptome.

Estrogen controls key cellular functions of responsive cells including the ability to survive, replicate, communicate and adapt to the extracellular milieu. Changes in the expression of 8400 genes were monitored here by cDNA microarray analysis during the first 32 h of human breast cancer (BC) ZR-75.1 cell stimulation with a mitogenic dose of 17beta-estradiol, a timing which corresponds to completion of a full mitotic cycle in hormone-stimulated cells. Hierarchical clustering of 344 genes whose expression either increases or decreases significantly in response to estrogen reveals that the gene expression program activated by the hormone in these cells shows 8 main patterns of gene activation/inhibition. This newly identified estrogen-responsive transcriptome represents more than a simple cell cycle response, as only a few affected genes belong to the transcriptional program of the cell division cycle of eukaryotes, or showed a similar expression profile in other mitogen-stimulated human cells. Indeed, based on the functions assigned to the products of the genes they control, estrogen appears to affect several key features of BC cells, including their metabolic status, proliferation, survival, differentiation and resistance to stress and chemotherapy, as well as RNA and protein synthesis, maturation and turn-over rates. Interestingly, the estrogen-responsive transcriptome does not appear randomly interspersed in the genome. In chromosome 17, for example, a site particularly rich in genes activated by the hormone, physical association of co-regulated genes in clusters is evident in several instances, suggesting the likely existence of estrogen-responsive domains in the human genome.

Breast Neoplasms↗

Integrated transcriptome analysis and machine learning to construct a homeostatic model of acetylation for bladder cancer and validate the key gene CES1.

BACKGROUND: Bladder cancer (BLCA) is one of the most common malignant tumors of the urinary system. Protein acetylation (PA) plays a critical role in regulating multiple biological processes (BPs), cellular homeostasis, and cancer-related signaling pathways. This study aimed to construct a homeostatic model of acetylation for BLCA using integrated transcriptome analysis and machine learning and to validate the key gene CES1. METHODS: RNA sequencing (RNA-seq) and clinical data were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Acetylation-related differentially expressed genes (DEGs) in BLCA were screened using differential expression analysis (DEA). An acetylation homeostatic model was constructed via univariate, machine learning-based least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression analyses, followed by validation in multiple cohorts. Single-cell RNA-seq analysis was used to explore gene expression patterns in diverse cell types. Enrichment analysis (EA), immune infiltration, and drug sensitivity analysis (DSA) were performed to characterize molecular features of different risk groups. Finally, the biological function of CES1 as the key gene was verified by in vitro knockdown experiments. RESULTS: We established a robust acetylation homeostatic model consisting of five genes, which effectively predicted overall survival (OS) and served as an independent prognostic factor in BLCA. High-risk patients showed significantly poorer prognosis, distinct immune infiltration profiles, and differential drug sensitivity. CES1 was identified and validated as the key gene in this model, which was highly expressed in BLCA and associated with poor prognosis. Knockdown of CES1 markedly suppressed cell proliferation, invasion, and migration, and reduced intracellular coenzyme A (CoA) levels, thereby regulating PA homeostasis. CONCLUSIONS: We developed and validated a novel acetylation homeostatic model for survival stratification and personalized treatment guidance in BLCA, based on integrated transcriptome analysis and machine learning. CES1 is closely associated with intracellular CoA levels and the malignant progression of BLCA. Its potential association with PA homeostasis requires further mechanistic validation, and it may act as a candidate therapeutic biomarker for BLCA.

Bladder cancer (BLCA)↗

Transcriptome Analysis and Experimental Validation of Palmitoylation- Related Biomarkers in Atherosclerosis.

INTRODUCTION: Protein palmitoylation contributes to membrane localisation, signal transduction, and cell-fate regulation. It is closely associated with lipid metabolic dysfunction, immune inflammation, and vascular remodelling in atherosclerosis (AS). However, key palmitoylation-related transcriptomic markers and their potential causal associations with AS remain incompletely defined. METHODS: The Gene Expression Omnibus (GEO) dataset GSE100927 was used as the training cohort, and GSE43292 was used as an external validation cohort. Differentially expressed genes were identified using limma and intersected with palmitoylation-related genes to obtain palmitoylation-related differentially expressed genes (PRDEGs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were then performed using clusterProfiler. Two-sample Mendelian randomisation was used to evaluate potential causal relationships between characteristic genes and AS. Feature selection was conducted using random forest and support vector machine recursive feature elimination (SVM-RFE), and the overlapping genes selected by both methods were retained. Receiver operating characteristic (ROC) curves were used to assess diagnostic performance. A five-gene nomogram was constructed, and its clinical utility was evaluated using calibration curves and decision curve analysis (DCA). Gene set variation analysis (GSVA) was applied to compare pathway activity between high- and low-expression groups for each core gene. Single-cell analysis using Seurat and expression-based cell-cell communication analysis using CellChat were conducted with GSE159677, and upstream transcription factors were predicted using NetworkAnalyst. For in vivo validation, an AS model was established in ApoE⁸/⁸ mice fed a high-fat diet, and aortic gene and protein expression were assessed by RT-qPCR and western blotting. RESULTS: In GSE100927, 51 PRDEGs were identified. GO and KEGG enrichment analyses highlighted pathways associated with regulation of monoatomic ion transport, sarcomere and myofibril organisation, and immune inflammation. Mendelian randomisation suggested a potential protective causal association between SLC7A7 and AS. By integrating MR with random forest and SVM-RFE feature selection, we prioritised five core genes: PLCB2, GMIP, NEXN, PLN, and SLC7A7. These genes showed good diagnostic performance in GSE43292. The resulting nomogram was well calibrated and demonstrated stable net benefit in decision curve and clinical impact curve analyses. Single-gene GSVA identified consistently activated pathways across multiple genes, including innate and adaptive immune recognition, calcium signalling and myocardial contraction/cardiomyopathy, extracellular matrix-receptor interaction, cell junction pathways, autophagy-lysosome pathways, and several metabolic programmes. At the single-cell level, PLCB2 and GMIP were predominantly expressed in T cells and macrophages, NEXN and PLN were enriched in vascular smooth muscle cells, and SLC7A7 was mainly expressed in macrophages. CellChat analysis indicated increased signals for immune-related ligand-receptor interactions. In ApoE⁸/⁸ mice fed a high-fat diet, PLCB2, GMIP, and SLC7A7 were upregulated, whereas NEXN and PLN were downregulated; protein-level changes were concordant with the transcriptomic trends. DISCUSSION: These findings indicate that palmitoylation-related dysregulation in AS converges on immune inflammation, calcium signalling/contractile programmes, ECM remodelling, and autophagy-linked metabolism. The five-gene panel is supported by external validation, single-cell localisation to immune and vascular compartments, and concordant results in ApoE⁸/⁸ mice. CONCLUSION: This study identified and validated five palmitoylation-related genes associated with AS. SLC7A7 showed a potential protective causal signal in MR analysis. The enriched pathway patterns linked these genes to immune inflammation, calcium signalling-contraction coupling, ECM remodelling, cell adhesion, and autophagy- associated metabolic reprogramming. The five-gene nomogram showed potential utility for diagnostic classification and decision support, nominating candidate biomarkers and pathway targets for AS molecular subtyping, diagnosis, and mechanistic investigation.

Atherosclerosis (AS)↗

Application of transcriptome analysis to clinical pharmacology studies.

Clinical pharmacology is the investigation of drug effects in humans. This review discusses the basic tenets of clinical pharmacology research, including pharmacokinetic and pharmacodynamic analysis, therapeutic window, and clinical trial design, and the issues that may arise in the application of transcriptome analysis to clinical pharmacology studies. Examples of how transcriptome analysis can be applied to clinical pharmacology research are described, including a model system of endotoxin challenge (in vitro and in vivo), and an example of a cross-over drug study in normal volunteers. Various data display and analysis methods are also illustrated, including principal component analysis, hierarchical cluster analysis, and pathways analysis.

Animals↗

From traditional biomarkers to transcriptome analysis in drug development.

Traditional biomarkers have played an important role in drug development as well as patient care. A single traditional biomarker or surrogate endpoint is unlikely to either characterize the complete pathophysiology of a complex disease or capture all the therapeutic benefits or potential adverse effects that a drug will have in a diverse patient population. Transciptome analysis, on the other hand, can provide a large-scale survey of gene expression associated with the etiology of a human disease or pharmacological responses to a therapeutic intervention. The quantitative and qualitative readouts can provide increased power to identify novel drug targets or biomarkers indicative of drug safety or efficacy. Transcriptomics has positively impacted drug development and will continue to improve the medicines of the future. Here, we describe the increasingly important roles that traditional biomarkers and transcriptome analysis have played in various phases of drug discovery and development as well as the opportunities and challenges that they present to the pharmaceutical industry.

Biomarkers↗

Transcriptomic Association Between Poliovirus Receptor (PVR/CD155) and Claudin Signaling Pathways in Colorectal Cancer.

BACKGROUND/AIM: Enterotoxigenic Bacteroides fragilis promotes colorectal carcinogenesis through toxin-mediated cleavage of E-cadherin, a process facilitated by membrane-associated Claudin-4 (CLDN4). Separately, the poliovirus receptor (PVR/CD155) modulates tumor epithelial and immune dynamics. This study explored potential transcriptomic interactions and co-expression frameworks between PVR and claudin signaling pathways in colorectal cancer. MATERIALS AND METHODS: Transcriptomic and proteomic data from the The Cancer Genome Atlas-colon adenocarcinoma cohort (TCGA-COAD) were evaluated. An exploratory E-cadherin Cleavage Index was modeled to capture transcript-protein discordance. To control for tissue composition heterogeneity without mathematical circularity, a de-circularized, non-parametric partial rank residual model adjusted for independent CLDN4 expression was deployed within the stable microsatellite-stable (MSS) sub-cohort (N=473). RESULTS: Multivariable survival models showed no independent associations between overall survival and continuous PVR (p=0.79) or CLDN3 (p=0.56) expression. Robust linear modeling revealed no significant baseline interaction between PVR and CLDN4 regarding the exploratory Cleavage Index (p=0.82). However, de-circularized partial correlation analysis revealed a highly stable, positive co-expression between PVR and CLDN3 (rho=0.2459, p=3.23×10-7). Both epithelial markers retained modest inverse correlations with the infiltrating lymphocytic axis (TIGIT and CD96). CONCLUSION: Baseline PVR expression is coordinated with CLDN3 tissue programs independent of general epithelial cellularity but does not interact with the CLDN4 axis or impact overall survival in an unexposed cohort. Because TCGA lacks virome or active microbial exposure tracking, these findings serve as baseline benchmarks for future context-dependent mechanistic studies.

Bacteroides fragilis toxin↗

African Swine Fever Virus MGF 360-2L Disrupts Host Antiviral Immunity Based on Transcriptomic Analysis.

Background/Objectives: The African swine fever virus (ASFV) multi-gene family (MGF) 360 proteins play critical roles in immune evasion, replication regulation, and virulence determination. Despite substantial advances in this field, the functional roles of many members within this gene family remain to be fully characterized. Methods: In this study, Transcriptional kinetics analysis indicated that the expression profile of MGF 360-2L was consistent with that of the late marker gene B646L (p72). Transcriptomic profiling identified 13 and 171 differentially expressed genes (DEGs) at 12 and 24 h post-infection (hpi) with ΔMGF 360-2L, respectively. Results: Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses indicated that these DEGs were predominantly enriched in Type I interferon (IFN-I) signaling pathways. It is noteworthy that transcriptome analysis further demonstrates that the absence of MGF 360-2L specifically results in the dysregulation of expression of the replication-essential genes E199L and E301R. These findings indicate that MG F360-2L is essential for maintaining the stable expression of these proteins. Conclusions:MGF 360-2L is a late gene that contributes to the precise regulation of viral protein expression and modulates the host immune response during infection.

African swine fever virus↗

S100P as a Shared Biomarker in Inflammatory Bowel Disease, Colorectal Cancer, and Pancreatic Adenocarcinoma: An Integrated Transcriptomic Analysis.

Inflammatory bowel disease (IBD) is associated with an increased risk of colorectal cancer (CRC) and pancreatic adenocarcinoma (PAAD), yet the molecular features shared among these diseases remain incompletely understood. This study aimed to identify common genes and biological pathways associated with IBD, CRC, and PAAD through integrated transcriptomic analysis and experimental validation. Gene expression datasets for IBD, CRC, and PAAD were obtained from The Cancer Genome Atlas and Gene Expression Omnibus databases. Weighted gene co-expression network analysis and differential expression analysis were performed to identify disease-associated and shared genes. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes (analyses were used to explore enriched biological functions and pathways. Immune cell infiltration was evaluated using Cell-type Identification by Estimating Relative Subsets of RNA Transcripts. Receiver operating characteristic analysis was performed to assess the diagnostic performance of common genes. Single-cell RNA sequencing analysis was conducted to examine the cellular distribution of S100P. In addition, the effects of S100P downregulation were evaluated in lipopolysaccharide (LPS)-stimulated colonic epithelial cells. A total of 162 disease-associated genes and four common genes were identified. Functional enrichment analyses indicated significant enrichment of immune- and inflammation-related pathways, including the interleukin-17 signaling pathway. Immune infiltration analysis revealed similar trends in several immune cell populations across IBD, CRC, and PAAD. Single-cell analysis showed elevated S100P expression in epithelial cells from all three diseases. Downregulation of S100P restored the proliferative capacity of LPS-stimulated colonic epithelial cells and reduced inflammatory cytokine expression. Integrated transcriptomic analysis identified S100P as a biomarker associated with IBD, CRC, and PAAD and highlighted shared immune-related features across these diseases.

Humans↗

Parietal Cortex Transcriptomics Refines Parkinson Disease GWAS Nomination and Highlights STAT3 as a Putative Upstream Glial Regulator.

Parkinson disease (PD) affects more than 1.1 million individuals in the United States and around 12 million worldwide. Although Genome Wide Association Studies (GWAS) have substantially advanced our understanding of PD genetic architecture, the regulatory mechanisms linking PD risk loci to disease-relevant gene expression remain incompletely characterized, limiting our ability to infer disease mechanisms from genetic associations. Here, we integrated disease-state parietal cortex transcriptomics with the International Parkinson's Disease Genomics Consortium (iPDGC) locus prioritization to refine PD gene nomination and identify biologically plausible candidates missed by GWAS-only approaches. Using bulk RNA-seq from 99 neuropathologically confirmed PD cases and 30 neuropathologically confirmed controls, we prioritized candidate genes across 78 loci and classified them according to concordance between genetic evidence and differential expression in diseased cortices. This integrative approach recovered candidate genes not captured by external GWAS-based prioritization methods and highlighted synaptic, lysosomal, and proteostasis pathways as major components of PD risk biology. Network and transcription factor analyses further suggested coordinated regulation of these genes, with STAT3 emerging as a putative upstream glial regulator. Together, these findings suggest that integrating disease-state transcriptomics with genetic prioritization can refine PD risk-gene nomination and uncover regulatory programs that may be missed by GWAS alone.

Journal Article↗

Integrated Pan-Cancer, Single-Cell, and Spatial Transcriptomic Analyses Identify ZDHHC12 as a Biomarker Associated with Macrophage Infiltration and the Immune Landscape in Glioma.

BACKGROUND: The tumor immune microenvironment (TME) critically influences cancer progression and therapeutic response. However, the pan-cancer expression landscape, prognostic relevance, and spatial distribution of ZDHHC12 remain incompletely characterized. This study investigated the prognostic value of ZDHHC12 and its associations with immune microenvironmental features and drug sensitivity. METHODS: Data from The Cancer Genome Atlas (TCGA) and the Genotype-Tissue Expression (GTEx) datasets were used to evaluate ZDHHC12 expression and prognosis across cancer types. Immune infiltration analyses, single-cell RNA sequencing, and spatial transcriptomics were integrated to characterize the associations of ZDHHC12 with the cancer immunity cycle and the spatial architecture of glioma. Drug sensitivity and immunotherapy-related metrics were assessed using pharmacogenomic databases and computational prediction models. RESULTS: ZDHHC12 was aberrantly expressed across multiple tumors and was associated with patient prognosis. Its expression was broadly correlated with immune cell recruitment- and activation-related signatures. In glioma, single-cell and spatial transcriptomic analyses showed enrichment of ZDHHC12 in monocyte/macrophage populations and spatial co-localization with BAK1, CD68, and CD163. ZDHHC12 expression was also associated with predicted drug sensitivity and immunotherapy-related metrics. CONCLUSION: ZDHHC12 may serve as a candidate pan-cancer prognostic biomarker. In glioma, its expression is associated with macrophage-enriched and immunosuppressive microenvironmental features. Functional studies are required to establish causality and determine its therapeutic relevance.

GBM↗

A preliminary transcriptome map of non-small cell lung cancer.

We constructed a genome-wide transcriptome map of non-small cell lung carcinomas based on gene-expression profiles generated by serial analysis of gene expression (SAGE) using primary tumors and bronchial epithelial cells of the lung. Using the human genome working draft and the public databases, 25,135 nonredundant UniGene clusters were mapped onto unambiguous chromosomal positions. Of the 23,056 SAGE tags that appeared more than once among the nine SAGE libraries, 11,156 tags representing 7,097 UniGene clusters were positioned onto chromosomes. A total of 43 and 55 clusters of differentially expressed genes were observed in squamous cell carcinoma and adenocarcinoma, respectively. The number of genes in each cluster ranged from 18 to 78 in squamous cell carcinomas and from 20 to 165 in adenocarcinomas. The size of these clusters varied from 1.8 Mb to 65.5 Mb in squamous cell carcinomas and from 1.6 Mb to 98.1 Mb in adenocarcinomas. Overall, the clusters with genes over-represented in tumors had an average of 3-4-fold increase in gene expression compared with the normal control. In contrast, clusters of genes with reduced expression had about 50-65% of the gene expression level compared with the normal. Examination of clusters identified in squamous cell lung cancer suggested that 9 of 15 clusters with overexpressed genes and 13 of 28 clusters with underexpressed genes were concordant with previously reported cytogenetic, comparative genomic hybridization or loss of heterozygosity studies. Therefore, at least a portion of the gene clusters identified via the transcriptome map most likely represented the transcriptional or genetic alterations occurred in the tumors. Integrating chromosomal mapping information with gene expression profiles may help reveal novel molecular changes associated with human lung cancer.

Carcinoma, Non-Small-Cell Lung↗

Transcriptome analysis and related databases of Lactococcus lactis.

Several complete genome sequences of Lactococcus lactis and their annotations will become available in the near future, next to the already published genome sequence of L. lactis ssp. lactis IL 1403. This will allow intraspecies comparative genomics studies as well as functional genomics studies aimed at a better understanding of physiological processes and regulatory networks operating in lactococci. This paper describes the initial set-up of a DNA-microarray facility in our group, to enable transcriptome analysis of various Gram-positive bacteria, including a ssp. lactis and a ssp. cremoris strain of Lactococcus lactis. Moreover a global description will be given of the hardware and software requirements for such a set-up, highlighting the crucial integration of relevant bioinformatics tools and methods. This includes the development of MolGenIS, an information system for transcriptome data storage and retrieval, and LactococCye, a metabolic pathway/genome database of Lactococcus lactis.

Databases, Nucleic Acid↗

[Transcriptomes for serial analysis of gene expression].

The availability of the sequences for whole genomes is changing our understanding of cell biology. Functional genomics refers to the comprehensive analysis, at the protein level (proteome) and at the mRNA level (transcriptome) of all events associated with the expression of whole sets of genes. New methods have been developed for transcriptome analysis. Serial Analysis of Gene Expression (SAGE) is based on the massive sequential analysis of short cDNA sequence tags. Each tag is derived from a defined position within a transcript. Its size (14 bp) is sufficient to identify the corresponding gene and the number of times each tag is observed provides an accurate measurement of its expression level. Since tag populations can be widely amplified without altering their relative proportions, SAGE may be performed with minute amounts of biological extract. Dealing with the mass of data generated by SAGE necessitates computer analysis. A software is required to automatically detect and count tags from sequence files. Criterias allowing to assess the quality of experimental data can be included at this stage. To identify the corresponding genes, a database is created registering all virtual tags susceptible to be observed, based on the present status of the genome knowledge. By using currently available database functions, it is easy to match experimental and virtual tags, thus generating a new database registering identified tags, together with their expression levels. As an open system, SAGE is able to reveal new, yet unknown, transcripts. Their identification will become increasingly easier with the progress of genome annotation. However, their direct characterization can be attempted, since tag information may be sufficient to design primers allowing to extend unknown sequences. A major advantage of SAGE is that, by measuring expression levels without reference to an arbitrary standard, data are definitively acquired and cumulative. All publicly available data can thus be stored in a unique database, facilitating whole-genome analysis of differential expression between cell types, normal and diseased samples, or samples with and without drug treatment. SAGE data are readily amenable to statistical comparisons, allowing to determine the level of confidence of the observed variations. A major limitation of SAGE is that, because each analysis is obligatory performed on the whole set of expressed genes, it can hardly be performed on multiple samples, for example in kinetics studies or to compare the effects of large numbers of drugs. To overcome this limitation, high-throughput detection of a subset of mRNAs is more rapidly performed by parallel hybridization of mRNAs on arrays of nucleic acids immobilized on solid supports. From this point of view, a SAGE platform is a powerful instrument for selecting the most informative subset of genes, assembling them to design microarrays dedicated to a specific problem and calibrating measurement by comparison with a standard cell model for which SAGE data are available. This approach is an attractive alternative to strategies based exclusively on pangenomic arrays. A very large amount of SAGE data are already available and the problem is now to extract their biological meaning. Knowledge on metabolic pathways is already organized so that its successful integration in a SAGE platform can be undertaken. For other cell components and pathways, the problem lies on the lack of controlled vocabulary to describe gene activities, starting form a clear definition of the concept of biological function itself. Progress in gene and cell ontology is expected to facilitate computer-based extraction of biological knowledge from existing and forthcoming SAGE data.

Animals↗