PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “comparative transcriptomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,099 records · Page 61Linked to original sources

A multi-modal survival prediction framework with group-based batch training and structural consistency alignment.

OBJECTIVE: Integrating whole-slide images (WSIs) with transcriptomic profiles is pivotal for enhancing cancer survival prediction. However, the intrinsic gigapixel resolution and variable sequence lengths of WSIs create a fundamental trade-off between training efficiency and the preservation of data heterogeneity in existing frameworks. Furthermore, substantial statistical and structural discrepancies between histological and genomic modalities often impede effective cross-modal alignment and fusion, thereby limiting prognostic accuracy. METHODS: We propose PRISM, an efficient multi-modal learning framework for integrating WSIs with transcriptomic profiles. To reconcile training efficiency with full data heterogeneity, PRISM first stochastically partitions variable-length WSI sequences into a main subset and a complementary residual subset, both of which are packed into fixed-length groups for batch training. The main subset is processed in the main branch, utilizing isolation masking to maintain intra-group sequence independence. Simultaneously, the residual subset is consolidated into "hyperslides" within a residual branch that leverages tailored supervision, effectively capturing inter-slide correlations. Furthermore, PRISM integrates an Informative Token Aggregation (ITA) module to reduce redundancy in WSIs and employs Cross-batch Structural Consistency Alignment (CBSCA) mechanism to enhance inter-modal structural connectivity. Finally, efficient cross-modal feature interaction is achieved through a Low-rank Bilinear Gated Fusion (LBGF) module. Code is available at https://github.com/Alisa2080/PRISM. RESULTS: Compared with existing methods, PRISM achieves the best overall C-index across five TCGA cohorts. On the larger TCGA-BRCA dataset, PRISM requires only 6 hours of training time, substantially reducing computational cost relative to strong multimodal baselines. Furthermore, comprehensive evaluations demonstrate that PRISM achieves the best overall IBS ranking and favorable time-dependent AUC performance at 1, 3, and 5 years, thereby delivering a more favorable trade-off between prognostic performance and computational efficiency. CONCLUSION: PRISM provides a favorable balance between predictive performance, calibration quality, and computational efficiency, highlighting its potential for practical deployment in multimodal survival modeling for computational pathology.

Humans↗

Comparison of basal gene expression profiles and effects of hepatocarcinogens on gene expression in cultured primary human hepatocytes and HepG2 cells.

Toxicogenomics is a relatively new discipline of toxicology. Microarrays and bioinformatics tools are being used successfully to understand the effects of toxicants on in vivo and in vitro model systems, and to gain a better understanding of the relevance of in vitro models commonly used in toxicological studies. In this study, cDNA filter arrays were used to determine the basal expression patterns of human cultured primary hepatocytes from different male donors; compare the gene expression profile of HepG2 to that of primary hepatocytes; and analyze the effects of three genotoxic hepatocarcinogens; aflatoxin B(1) (AFB(1)), 2-acetylaminofluorene (2AAF), and dimethylnitrosamine (DMN), as well as one non-gentoxic hepatotoxin, acetaminophen (APAP) on gene expression in both in vitro systems. Real-time PCR was used to verify differential gene expression for selected genes. Of the approximately 31,000 genes screened, 3-6% were expressed in primary hepatocytes cultured on matrigel for 16 h. Of these genes, 867 were expressed in cultured hepatocytes from all donors. HepG2 cells expressed about 98% of the genes detectable in cultured primary hepatocytes, however, 31% of the HepG2 transcriptome was unique to the cell line. A number of these genes are expressed in human liver but expression is apparently lost during culture. There was considerable variability in the response to chemical carcinogen exposure in primary hepatocytes from different donors. The transcription factors, E2F1 and ID1 mRNA were increased three-fold and six-fold (P < 0.05, P < 0.01), respectively, in AFB(1) treated primary human hepatocytes but were not altered in HepG2. ID1 expression was also increased by dimethylnitrosamine, acetylaminofluorene and acetaminophen in both primary hepatocytes and HepG2. Identification of genes that are expressed in primary hepatocytes from most donors, as well as those genes with variable expression, will aid in understanding the variability in human reactions to drugs and chemicals. This study suggests that identification of biomarkers of exposure to some chemicals may be possible in the human through microarray analysis, despite the variability in responses.

Carcinogens↗

Identification of a necroptosis-related lncRNA prognostic signature and the hub RBP HNRNPK in esophageal squamous cell carcinoma.

ObjectiveEsophageal squamous cell carcinoma (ESCC) is a malignant tumor with poor prognosis. Necroptosis is important for tumor immunity, but its role in ESCC remains unclear. This retrospective bioinformatics study aimed to investigate the prognostic value of necroptosis-related long non-coding RNAs (lncRNAs) and to identify key lncRNA-binding proteins (RBPs) in ESCC patients.MethodsRNA transcriptome and clinical data of ESCC patients were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Necroptosis-related lncRNAs were identified through correlation analysis with necroptosis-related genes, subjected to consensus cluster analysis, and used to construct a prognostic risk model via least absolute shrinkage and selection operator (LASSO) regression. The hub RBP was experimentally validated by quantitative polymerase chain reaction (qPCR) using 30 pairs of ESCC and adjacent normal tissues from patients who underwent surgical resection.ResultsA total of 30 necroptosis-related lncRNAs were significantly correlated with overall survival (OS). The upregulated lncRNAs in the risk model were associated with high immune scores, innate immune cell infiltration, cluster 2 classification, and advanced T-stage disease (p&#x2009;<&#x2009;0.05). Three hub RBPs (HNRNPA1, HNRNPC, and HNRNPK) were identified through protein-protein interaction network analysis. qPCR confirmed that HNRNPK was significantly overexpressed in ESCC tissues compared to adjacent normal tissues (p&#x2009;<&#x2009;0.05).ConclusionsThe necroptosis-related lncRNA risk model is an independent prognostic factor for ESCC patients. HNRNPK was identified as a hub RBP significantly overexpressed in ESCC tissues. We hypothesize that HNRNPK may promote tumor progression through regulating proto-oncogene expression or modulating the immune microenvironment, though this requires further mechanistic validation.

Humans↗

Functional genomics studied by proteomics.

The human genome contains about 30,000 genes, each creating several transcripts per gene. Transcript structures and expression are studied by high-throughput transcriptomic techniques using microarrays. Generally, transcripts are not directly operating molecules, but are translated into functional proteins, post-translationally modified by proteolysis, glycosylation, phosphorylation, etc., sometimes with great functional impact. Proteins need to be analyzed by proteomic techniques, less suited for high-throughput. Two-dimensional polyacrylamide gel electrophoresis (2D-PAGE), separating thousands of proteins has developed slowly over the past quarter of a century. This technique is now quite reproducible and suitable for differential proteomics, comparing normal and diseased cells/tissues revealing differentially regulated proteins. 2D-PAGE is combined with protein-identification methods, currently mass spectrometry (MS), which has been significantly improved over the last decade. Other proteomic techniques studying protein-protein interactions are now either established or still being developed, such as peptide or protein arrays, phage display, and the yeast two-hybrid system. The strengths and weaknesses of these techniques are discussed.

Databases, Factual↗

Proteomic analysis of differentially expressed proteins in hepatocellular carcinoma developed in patients with chronic viral hepatitis C.

Hepatocellular carcinoma (HCC) is a major complication of chronic viral hepatitis C. Therapy for HCC is still disappointing. It is thus of great importance to identify novel HCC markers for early detection of the disease, and tumor-specific proteins as potential therapeutic targets. We have used a proteomic approach to identify new proteins involved in HCC development. Four cases of HCC developing from chronic viral hepatitis C were analyzed by two-dimensional electrophoresis (2-DE), and results were compared to those of paired adjacent non-tumorous liver tissues. For MS fingerprinting, protein spots with differential intensity between HCC and non-tumorous liver were directly cut out of gels and processed for MALDI-MS and nano-LC-ESI-MS/MS analysis. Approximately 850 spots were visualized in each gel. The comparative analysis of paired samples indicated that 345 protein spots showed significant differences in expression level between non-tumor and tumor tissue. Among the 345 protein spots analyzed, 238 spots corresponding to 155 different proteins were identified; 49 proteins were up-regulated, whereas 106 proteins were down-regulated. Among these 155 proteins, 91 proteins were regulated in at least three cases. Although 52 out of these 91 proteins have been already described by previous proteomic or transcriptomic studies, or are already known to be involved in hepatocarcinogenesis, this experiment revealed 39 new proteins differentially expressed in HCC developing from viral hepatitis C. Variations in protein accumulation were confirmed for two selected proteins (apolipoprotein E, chloride intracellular channel 1) by Western blotting in ten additional cases of HCC developing in patients with viral hepatitis C.

Aged↗

Co-expression of the Mammaglobin (SCGB2A2) Gene With hsa-miR-184 and hsa-miR-190b Indicates Its Possible Role in Oncogenic Pathways in Breast Cancer.

BACKGROUND/AIM: Breast cancer is the most common cancer in women worldwide, and early detection remains a significant challenge. Recent studies have identified increased expression of Mammaglobin A (Q13296, Gene: SCGB2A2) mRNA in breast cancer, suggesting its potential as a disease marker, although its function is not fully understood. To elucidate Mammaglobin's role, this study sought to identify co-expressed miRNAs and analyze the biological pathways they regulate. MATERIALS AND METHODS: Using TCGAbiolinks and Firebrowse, miRNA and gene expression data were collected from 86 patients, including tumor and normal tissue samples from the Cancer Genome Atlas (TCGA) Breast Cancer cohort. Transcriptomic data were analyzed with DESeq2, and a Spearman correlation was calculated for significant p-values, which were further explored using enrichment tools and target gene databases. RESULTS: DESeq2 was used to identify differential expression of miRNAs between normal and tumor breast tissues. Out of 782 miRNAs differentially expressed in breast cancer, hsa-mir-184 and hsa-mir-190b showed a significant positive correlation with SCGB2A expression. These markers were also upregulated in breast cancer tissues compared to normal tissues. Bioinformatics analysis revealed that hsa-mir-184 and hsa-mir-190b play important roles in cancer and cellular proliferation. These miRNAs target a wide range of genes, including sorting nexin 9 (SNX9) and annexin 6 (ANXA6), which are involved in membrane stability, vesicular trafficking, and cell mobility, and they contribute to cancer metastasis. CONCLUSION: The positive correlation among the expression of hsa-miR-184, hsa-miR-190b, and SCGB2A2 suggests that they may participate in shared biological pathways. These pathways govern critical cellular processes, such as membrane trafficking and cell signaling, which are frequently disrupted in cancer. Consequently, these findings enable a better understanding of the role of Mammaglobin in breast cancer signaling.

MicroRNAs (miRNAs)↗

Comparative host gene transcription by microarray analysis early after infection of the Huh7 cell line by severe acute respiratory syndrome coronavirus and human coronavirus 229E.

The pathogenesis of severe acute respiratory syndrome-associated coronavirus (SARS-CoV) at the cellular level is unclear. No human cell line was previously known to be susceptible to both SARS-CoV and other human coronaviruses. Huh7 cells were found to be susceptible to both SARS-CoV, associated with SARS, and human coronavirus 229E (HCoV-229E), usually associated with the common cold. Highly lytic and productive rates of infections within 48 h of inoculation were reproducible with both viruses. The early transcriptional profiles of host cell response to both types of infection at 2 and 4 h postinoculation were determined by using the Affymetrix HG-U133A microarray (about 22,000 genes). Much more perturbation of cellular gene transcription was observed after infection by SARS-CoV than after infection by HCoV-229E. Besides the upregulation of genes associated with apoptosis, which was exactly opposite to the previously reported effect of SARS-CoV in a colonic carcinoma cell line, genes related to inflammation, stress response, and procoagulation were also upregulated. These findings were confirmed by semiquantitative reverse transcription-PCR, reverse transcription-quantitative PCR for mRNA of genes, and immunoassays for some encoded proteins. These transcriptomal changes are compatible with the histological changes of pulmonary vasculitis and microvascular thrombosis in addition to the diffuse alveolar damage involving the pneumocytes.

Apoptosis↗

Gene for gene alignment between the Brassica and Arabidopsis genomes by direct transcriptome mapping.

We report a global gene for gene alignment of the genomes of Brassica oleracea and Arabidopsis thaliana by construction of a transcriptome map based on B. oleracea cDNAs obtained from leaf tissue. cDNAs were synthesized from total RNA extracted from individual F2s of a mapping population resulting from crossing double-haploids of broccoli and cauliflower. The map consisted of 247 cDNA markers obtained by the SRAP technique. After sequencing 190 of the polymorphic cDNA bands, FASTA detected 169 sequences with similarity to genes reported in Arabidopsis. There was extensive colinearity between the two genomes for chromosomal segments rather than for whole chromosomes, often showing inversions and deletions/insertions. Large-scale duplications were observed in the B. oleracea genome, but were unevenly distributed, arguing against ancient triplication of the entire genome. The most duplicated segments corresponded to those found on Arabidopsis chromosomes 1 and 5, whereas chromosomes 2 and 4 were the least represented in Brassica. Clear differences in the similarity score value of related sequences allowed the identification of orthologs. Transcriptome mapping is an efficient approach that allows gene-for-gene alignment between a fully sequenced and a poorly characterized genome.

Arabidopsis↗

Sex-dependent gene expression and evolution of the Drosophila transcriptome.

Comparison of the gene-expression profiles between adults of Drosophila melanogaster and Drosophila simulans has uncovered the evolution of genes that exhibit sex-dependent regulation. Approximately half the genes showed differences in expression between the species, and among these, approximately 83% involved a gain, loss, increase, decrease, or reversal of sex-biased expression. Most of the interspecific differences in messenger RNA abundance affect male-biased genes. Genes that differ in expression between the species showed functional clustering only if they were sex-biased. Our results suggest that sex-dependent selection may drive changes in expression of many of the most rapidly evolving genes in the Drosophila transcriptome.

Animals↗

Proteomic and transcriptomic analyses of differential stress/inflammatory responses in mandibular lymph nodes and oropharyngeal tonsils of European wild boars naturally infected with Mycobacterium bovis.

Differential stress/inflammatory responses were characterized at the mRNA and protein levels in mandibular lymph nodes (MLN) and oropharyngeal tonsils of European wild boars (Sus scrofa), naturally infected with Mycobacterium bovis. Suppression-subtractive hybridization combined with immunohistochemistry and/or quantitative real-time RT-PCR were used to identify and characterize abundant stress/inflammatory gene sequences differentially expressed in tuberculous (TB+) wild boars. Genes identified in MLN and tonsils corresponded to serum amyloid A, arginase I, osteopontin, lysozyme, annexin I, and heat shock proteins, respectively. Global protein patterns in MLN and tonsils were compared between TB+ and nontuberculous (TB-) boars by 2-DE and MALDI-TOF MS. Five proteins, including stress/inflammatory proteins annexin V, serum albumin, and apolipoprotein A1 were found at lower levels in MLN of TB+ boars. Manganese superoxide dismutase was found up-regulated in MLN of TB+ boars. Five proteins, including creatine kinase and MHC class II antigens were found up-regulated in tonsils of TB+ boars. These results demonstrated differential stress/inflammatory responses in wild boars naturally infected with M. bovis and suggest possible markers of tuberculosis in this species that may prove useful for future studies of host-pathogen interactions and for diagnostics and vaccine development.

Amino Acid Sequence↗

Comprehensive expression atlas of fibroblast growth factors and their receptors generated by a novel robotic in situ hybridization platform.

A recently developed robotic platform termed "Genepaint" can carry out large-scale nonradioactive in situ hybridization (ISH) on tissue sections. We report a series of experiments that validate this novel platform. Signal-to-noise ratio and mRNA detection limits were comparable to traditional ISH procedures, and hybridization was transcript-specific, even in cases in which probes could have hybridized to several transcripts of a multigene family. We established an atlas of expression patterns of fibroblast growth factors (Fgfs) and their receptors (Fgfrs) for the embryonic day 14.5 mouse embryo. This atlas provides a comprehensive overview of previously known as well as novel sites of expression for this important family of signaling molecules. The Fgf/Fgfr atlas was integrated into the transcriptome database (www.genepaint.org), where individual Fgf and Fgfr expression patterns can be interactively viewed at cellular resolution and where sites of expressions can be retrieved using an anatomy-based search.

Animals↗

Profiling estrogen-regulated gene expression changes in normal and malignant human ovarian surface epithelial cells.

Estrogens regulate normal ovarian surface epithelium (OSE) cell functions but also affect epithelial ovarian cancer (OCa) development. Little is known about how estrogens play such opposing roles. Transcriptional profiling using a cDNA microarray containing 2400 named genes identified 155 genes whose expression was altered by estradiol-17beta (E2) in three immortalized normal human ovarian surface epithelial (HOSE) cell lines and 315 genes whose expression was affected by the hormone in three established OCa (OVCA) cell lines. All but 19 of the genes in these two sets were different. Among the 19 overlapping genes, five were found to show discordant responses between HOSE and OVCA cell lines. The five genes are those that encode clone 5.1 RNA-binding protein (RNPS1), erythrocyte adducin alpha subunit (ADD1), plexin A3 (PLXNA3 or the SEX gene), nuclear protein SkiP (SKIIP), and Rap-2 (rap-2). RNPS1, ADD1, rap-2, and SKIIP were upregulated by E2 in HOSE cells but downregulated by estrogen in OVCA cells, whereas PLXNA3 showed the reverse pattern of regulation. The estrogen effects was observed within 6-18 h of treatment. In silicon analyses revealed presence of estrogen response elements in the proximal promoters of all five genes. RNPS1, ADD1, and PLXNA3 were underexpressed in OVCA cell lines compared to HOSE cell lines, while the opposite was true for rap-2 and SKIIP. Functional studies showed that RNPS1 and ADD1 exerted multiple antitumor actions in OVCA cells, while PLXNA3 only inhibited cell invasiveness. In contrast, rap-2 was found to cause significant oncogenic effects in OVCA cells, while SKIIP promotes only anchorage-independent growth. In sum, gene profiling data reveal that (1) E2 exerts different actions on HOSE cells than on OVCA cells by affecting two distinct transcriptomes with few overlapping genes and (2) among the overlapping genes, a set of putative oncogenes/tumor suppressors have been identified due to their differential responses to E2 between the two cell types. These findings may explain the paradoxical roles of estrogens in regulating normal and malignant OSE cell functions.

Cell Line↗

Bioinformatic screening of human ESTs for differentially expressed genes in normal and tumor tissues.

BACKGROUND: Owing to the explosion of information generated by human genomics, analysis of publicly available databases can help identify potential candidate genes relevant to the cancerous phenotype. The aim of this study was to scan for such genes by whole-genome in silico subtraction using Expressed Sequence Tag (EST) data. METHODS: Genes differentially expressed in normal versus tumor tissues were identified using a computer-based differential display strategy. Bcl-xL, an anti-apoptotic member of the Bcl-2 family, was selected for confirmation by western blot analysis. RESULTS: Our genome-wide expression analysis identified a set of genes whose differential expression may be attributed to the genetic alterations associated with tumor formation and malignant growth. We propose complete lists of genes that may serve as targets for projects seeking novel candidates for cancer diagnosis and therapy. Our validation result showed increased protein levels of Bcl-xL in two different liver cancer specimens compared to normal liver. Notably, our EST-based data mining procedure indicated that most of the changes in gene expression observed in cancer cells corresponded to gene inactivation patterns. Chromosomes and chromosomal regions most frequently associated with aberrant expression changes in cancer libraries were also determined. CONCLUSION: Through the description of several candidates (including genes encoding extracellular matrix and ribosomal components, cytoskeletal proteins, apoptotic regulators, and novel tissue-specific biomarkers), our study illustrates the utility of in silico transcriptomics to identify tumor cell signatures, tumor-related genes and chromosomal regions frequently associated with aberrant expression in cancer.

Algorithms↗

Algorithms and tools for data-driven omics integration to achieve multilayer biological insights: a narrative review.

Systems biology is a holistic approach to biological sciences that combines experimental and computational strategies, aimed at integrating information from different scales of biological processes to unravel pathophysiological mechanisms and behaviours. In this scenario, high-throughput technologies have been playing a major role in providing huge amounts of omics data, whose integration would offer unprecedented possibilities in gaining insights on diseases and identifying potential biomarkers. In the present review, we focus on strategies that have been applied in literature to integrate genomics, transcriptomics, proteomics, and metabolomics in the year range 2018-2024. Integration approaches were divided into three main categories: statistical-based approaches, multivariate methods, and machine learning/artificial intelligence techniques. Among them, statistical approaches (mainly based on correlation) were the ones with a slightly higher prevalence, followed by multivariate approaches, and machine learning techniques. Integrating multiple biological layers has shown great potential in uncovering molecular mechanisms, identifying putative biomarkers, and aid classification, most of the time resulting in better performances when compared to single omics analyses. However, significant challenges remain. The high-throughput nature of omics platforms introduces issues such as variable data quality, missing values, collinearity, and dimensionality. These challenges further increase when combining multiple omics datasets, as the complexity and heterogeneity of the data increase with integration. We report different strategies that have been found in literature to cope with these challenges, but some open issues still remain and should be addressed to disclose the full potential of omics integration.

Algorithms↗

Snail immunity to schistosomes: insights from omics studies.

Schistosomiasis is a serious public health concern, with transmission facilitated by a small number of freshwater snail intermediate host species. Infection outcomes vary greatly across the primary vector genera, Biomphalaria (for Schistosoma mansoni), Bulinus (for S. haematobium), and Oncomelania (for S. japonicum), even within species, ranging from full resistance to high compatibility. Omics methods have altered this field by correlating host genotype, baseline immunological status, and time-resolved responses to whether invading miracidia are eliminated or develop sporocysts. Evidence from genomes, transcriptomics, proteomics, and epigenomics suggests that resistance is frequently primed prior to exposure. However, the clearest divergence between resistant and susceptible trajectories occurs during a small early window (<12-48&#x202f;h) after penetration. During this time, recognition, hemocyte recruitment, and soluble effector deployment either come together quickly or are delayed and guided by parasite-derived modulators. Established infections cause the host to adapt to chronic conditions through immune regulation, metabolic reprogramming, tissue and neuroendocrine remodeling, microbiome modification, and parasite castration. Comparative genomics reveals that each vector genus has evolved its own immunogenomic profile, which includes lineage-specific expansions of recognition and effector gene families. Together, these findings can help with field surveillance and intervention by providing molecular compatibility markers, functional tools for testing candidate genes, and tactics that target parasite-derived immune modulators. Integrated multi-omics approaches are a top priority, yet they are still limited in snail vectors compared to other disease vector systems.

Animals↗

True and false discovery in DNA microarray experiments: transcriptome changes in the hippocampus of presenilin 1 mutant mice.

In transcriptome profiling experiments using DNA microarrays, it is critical to maximize putatively true data discovery while keeping the false discovery rate at acceptable levels. Using previously published and verified transcriptome datasets of mice with genetically altered PS1 physiology, we present a simple, robust, and system-specific assessment of type I and type II errors in two independent microarray experimental series. We provide evidence to suggest that for maximizing true discovery and minimizing false discovery, statistical criteria alone are inferior to statistical significance plus magnitude of change criteria. Furthermore, we found that, regardless of the exact criteria used for determining differential expression, different data extraction protocols give rise to different discovery and false discovery rates. In addition, a large proportion of expression differences were both dataset and analytical approach dependent. The data assessment methods presented and discussed in this manuscript can be easily carried out on any microarray dataset using basic spreadsheet functions as the only tool needed. Finally, we provide an in-depth analysis of the hippocampal transcriptome of DeltaE9 hPS1 transgenic mice and mice with a conditional ablation of the PS1 gene.

Animals↗

5'-end SAGE for the analysis of transcriptional start sites.

Identification of the mRNA start site is essential in establishing the full-length cDNA sequence of a gene and analyzing its promoter region, which regulates gene expression. Here we describe the development of a 5'-end serial analysis of gene expression (5' SAGE) that can be used to globally identify transcriptional start sites and the frequency of individual mRNAs. Of the 25,684 5' SAGE tags in the HEK293 human cell library, 19,893 matched to the human genome. Among 15,448 tags in one locus of the genome, 85.8%-96.1% of the 5' SAGE tags were assigned within -500 to +200 nt of mRNA start sites using the RefSeq, UniGene and DBTSS databases. This technique should facilitate 5'-end transcriptome analysis in a variety of cells and tissues.

5' Flanking Region↗

Complete sequencing and characterization of 21,243 full-length human cDNAs.

As a base for human transcriptome and functional genomics, we created the "full-length long Japan" (FLJ) collection of sequenced human cDNAs. We determined the entire sequence of 21,243 selected clones and found that 14,490 cDNAs (10,897 clusters) were unique to the FLJ collection. About half of them (5,416) seemed to be protein-coding. Of those, 1,999 clusters had not been predicted by computational methods. The distribution of GC content of nonpredicted cDNAs had a peak at approximately 58% compared with a peak at approximately 42%for predicted cDNAs. Thus, there seems to be a slight bias against GC-rich transcripts in current gene prediction procedures. The rest of the cDNAs unique to the FLJ collection (5,481) contained no obvious open reading frames (ORFs) and thus are candidate noncoding RNAs. About one-fourth of them (1,378) showed a clear pattern of splicing. The distribution of GC content of noncoding cDNAs was narrow and had a peak at approximately 42%, relatively low compared with that of protein-coding cDNAs.

Chromosomes, Human, 21-22 and Y↗