PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “comparative transcriptomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

Current topics in pharmacological research on bone metabolism: molecular basis of ectopic bone formation induced by mechanical stress.

Ectopic bone formation (EBF) is frequently found in various tissues and affects the prognosis of diseases accompanied by EBF. Although the mechanism of EBF remains unclear, several local factors that influence the progression of EBF have been proposed. We have been focusing on the role of mechanical stress as a local factor in EBF in spinal ligament tissues, that is, ossification of the posterior longitudinal ligament (OPLL), which causes serious neurological deficiencies. Transcriptome analyses revealed that the expressions of several marker genes related to bone remodeling were enhanced after exposure of ligament cells derived from OPLL patients (OPLL cells) to cyclic stretching as a type of mechanical stress. However, no significant alterations in gene expressions were detected after cyclic stretching of ligament cells derived from non-OPLL patients. OPLL cells exposed to cyclic stretching released several autocrine/paracrine factors that are known to mediate bone remodeling. These results suggest that OPLL cells have been transformed into cells that are highly sensitive to mechanical stress, which may induce the progression of OPLL. These observations provide information regarding the role of mechanical stress in the process of EBF.

Bone Morphogenetic Proteins↗

Microarray analysis of bacterial pathogenicity.

The DNA microarray, a surface that contains an ordered arrangement of each identified open reading frame of a sequenced genome, is the engine of functional genomics. Its output, the expression profile, provides a genome wide snap-shot of the transcriptome. Refined by array-specific statistical instruments and data-mined by clustering algorithms and metabolic pathway databases, the expression profile discloses, at the transcriptional level, how the microbe adapts to new conditions of growth--the regulatory networks that govern the adaptive response and the metabolic and biosynthetic pathways that effect the new phenotype. Adaptation to host microenvironments underlies the capacity of infectious agents to persist in and damage host tissues. While monitoring the whole genome transcriptional response of bacterial pathogens within infected tissues has not been achieved, it is likely that the complex, tissue-specific response is but the sum of individual responses of the bacteria to specific physicochemical features that characterize the host milieu. These are amenable to experimentation in vitro and whole-genome expression studies of this kind have defined the transcriptional response to iron starvation, low oxygen, acid pH, quorum-sensing pheromones and reactive oxygen intermediates. These have disclosed new information about even well-studied processes and provide a portrait of the adapting bacterium as a 'system', rather than the product of a few genes or even a few regulons. Amongst the regulated genes that compose this adaptive system are transcription factors. Expression profiling experiments of transcription factor mutants delineate the corresponding regulatory cascade. The genetic basis for pathogenicity can also be studied by using microarray-based comparative genomics to characterize and quantify the extent of genetic variability within natural populations at the gene level of resolution. Also identified are differences between pathogen and commensal that point to possible virulence determinants or disclose evolutionary history. The host vigorously engages the pathogen; expression studies using host genome microarrays and bacterially infected cell cultures show that the initial host reaction is dominated by the innate immune response. However, within the complex expression profile of the host cell are components mediated by pathogen-specific determinants. In the future, the combined use of bacterial and host microarrays to study the same infected tissue will reveal the dialogue between pathogen and host in a gene-by-gene and site- and time-specific manner. Translating this conversation will not be easy and will probably require a combination of powerful bioinformatic tools and traditional experimental approaches--and considerable effort and time.

Animals↗

Genetics of constant and severe pain in the NAPS2 cohort of recurrent acute and chronic pancreatitis patients.

Recurrent acute and chronic pancreatitis (RAP, CP) are complex, progressive inflammatory diseases with variable pain experiences impacting patient function and quality of life. The genetic variants and pain pathways in patients contributing to most severe pain experiences are unknown. We used previously genotyped individuals with RAP/CP from the North American Pancreatitis Study II (NAPS2) of European Ancestry for nested genome-wide associated study (GWAS) for pain-severity, chronicity, or both. Lead variants from GWAS were determined using FUMA. Loci with p<1e-5 were identified for post-hoc candidate identification. Transcriptome-wide association studies (TWAS) identified loci in cis and trans to the lead variants. Serum from phenotyped individuals with CP from the PROspective Evaluation of Chronic Pancreatitis for EpidEmiologic and Translational StuDies (PROCEED) was assessed for BDNF levels using Meso Scale Discovery Immunoassay. We identified four pain systems defined by candidate genes: 1) Pancreas-associated injury/stress mitigation genes include: REG gene cluster, CTRC, NEURL3 and HSF22. 2) Neural development and axon guidance tracing genes include: SNPO, RGMA, MAML1 and DOK6 (part of the RET complex). 3) Genes linked to psychiatric stress disorders include TMEM65, RBFOX1, and ZNF385D. 4) Genes in the dorsal horn pain-modulating BDNF/neuropathic pathway included SYNPR, NTF3 and RBFOX1. In an independent cohort BDNF was significantly elevated in patients with constant-severe pain. Extension and expansion of this exploratory study may identify pathway- and mechanism-dependent targets for individualized pain treatments in CP patients. PERSPECTIVE: Pain is the most distressing and debilitating feature of chronic pancreatitis. Yet many patients with chronic pancreatitis have little or no pain. The North American Pancreatitis Study II (NAPS2) includes over 1250 pancreatitis patients of all progressive stages with all clinical and phenotypic characteristics carefully recorded. Pain did not correlate well with disease stage, inflammation, fibrosis or other features. Here we spit the patients into groups with the most severe pain and/or chronic pain syndromes and compared them genetically with patients reporting mild or minimal pain. Although some genetic variants associated with pain were expressed in cells (1) of the pancreas, most genetic variants were linked to genes expressed in the nervous system cells associated with (2) neural development and axon guidance (as needed for the descending inhibition pathway), (3) psychiatric stress disorders, and (4) cells regulating sensory nerves associated with BDNF and neuropathic pain. Similar and overlapping genetic variants in systems 2 -4 are also seen in pain syndromes form other organs. The implications for treating pancreatic pain are great in that we can no longer focus on just the pancreas. Furthermore, new treatments designed for pain disorders in other tissues may be effective in some patient with pain syndromes from the pancreas. Further research is needed to replicate and extend these observations so that new, genetics-guided rational treatments can be developed and delivered.

Humans↗

The desmoplastic response to infiltrating breast carcinoma: gene expression at the site of primary invasion and implications for comparisons between tumor types.

The gene expression patterns of desmoplasia are becoming exposed through the application of global gene expression technologies such as cDNA microarrays or serial analysis of gene expression (SAGE). These patterns represent the sum of the many cellular components of the host stromal response to an infiltrating carcinoma. In studies of human neoplasms, it would be useful to identify those prototypical genes that characteristically indicate the recognizable forms of the responses to individual tumor types. Such genes may offer clues to better understand the process of invasion itself, the interactions between tumor and host cells, and tumor-specific differences in invasion. We used SAGE-defined genes and in situ transcript labeling to characterize the desmoplastic stroma induced by infiltrating ductal carcinomas of the breast. Principal component analysis identified 103 SAGE tags as specific for invasive breast carcinomas, in comparison with in situ duct carcinomas or normal breast epithelium. Of these, 68 tags corresponded to known genes. Six of the 68 genes from this breast cancer "invasion-specific" cluster were further characterized by in situ hybridization to breast cancer tissues. Results of in situ hybridization demonstrated that each gene was expressed within one of five distinct regions of the invasive tumors (neoplastic epithelium; angioendothelium; inflammatory, panstromal, and juxtatumoral stroma), reflecting a defined architectural structure to the transcriptome of invasive breast cancers. Two of these 6 genes were specifically expressed by the stromal cells within the invasive carcinoma; however, 1 (collagen 1alpha1) was expressed throughout the stromal response (panstromal expression), whereas the second (osteonectin) was specifically expressed within the juxtatumoral stromal cells, indicating a critical "regionality" of gene expression within the stromal response itself. A comparison of the gene expression profiles of the juxtatumoral stroma in breast and pancreatic carcinomas indicated important differences between the two, suggesting tumor-specific or organ-specific differences in the desmoplastic responses. Some of the genes presented are novel markers of the invasive process, imply communication at the host/tumor interface, and suggest potential therapeutic targets.

Breast Neoplasms↗

Analysis of fat body transcriptome from the adult tsetse fly, Glossina morsitans morsitans.

Tsetse flies (Diptera: Glossinidia) are vectors of pathogenic African trypanosomes. To develop a foundation for tsetse physiology, a normalized expressed sequence tag (EST) library was constructed from fat body tissue of immune-stimulated Glossina morsitans morsitans. Analysis of 20,257 high-quality ESTs yielded 6372 unique genes comprised of 3059 tentative consensus (TC) sequences and 3313 singletons (available at http://aksoylab.yale.edu). We analysed the putative fat body transcriptome based on homology to other gene products with known functions available in the public domain. In particular, we describe the immune-related products, reproductive function related yolk proteins and milk-gland protein, iron metabolism regulating ferritins and transferrin, and tsetse's major energy source proline biosynthesis. Expression analysis of the three yolk proteins indicates that all are detected in females, while only the yolk protein with similarity to lipases, is expressed in males. Milk gland protein, apparently important for larval nutrition, however, is primarily synthesized by accessory milk gland tissue.

Adipose Tissue↗

Zonal gene expression in mouse liver resembles expression patterns of Ha-ras and beta-catenin mutated hepatomas.

Hepatocytes of the periportal and perivenous zones of the liver lobule differ in their levels and activities of various enzymes and other proteins. We have recently suggested that beta-catenin- and Ras-dependent signaling pathways play an important role in the regulation of perivenous and periportal gene expression profiles. This hypothesis was primarily based on similarities in zonal differences in gene expression of hepatocytes from normal liver with gene expression patterns of liver tumors: several proteins and mRNAs preferentially expressed in periportal hepatocytes were often overexpressed in Ha-ras mutated mouse liver tumors, whereas perivenous markers were overexpressed in Ctnnb1 (encoding beta-catenin) mutated tumors. We have now extended this work by use of data from two previously conducted microarray analyses aimed to analyze 1) global gene expression patterns of Ha-ras and Ctnnb1 mutated mouse liver tumors and 2) transcriptome differences between periportal and perivenous mouse hepatocytes. By comparison of the datasets, 134 genes or expressed sequences were identified that were present in both datasets. Gene expression patterns in perivenous hepatocytes and Ctnnb1 mutated hepatoma cells were strongly correlated: 96.5% of the genes present in both datasets were regulated in the same direction. In analogy, expression of 74.1% of the genes deregulated in Ha-ras mutated tumors was correlated with the respective expression patterns in periportal hepatocytes. These findings favor the hypothesis that gene expression patterns in periportal and perivenous hepatocytes are regulated, at least in part, by Ras- and beta-catenin-dependent signaling pathways.

Animals↗

Iron-related transcriptomic variations in CaCo-2 cells, an in vitro model of intestinal absorptive cells.

Regulation of iron absorption by duodenal enterocytes is essential for the maintenance of homeostasis by preventing iron deficiency or overload. Despite the identification of a number of genes implicated in iron absorption and its regulation, it is likely that further factors remain to be identified. For that purpose, we used a global transcriptomic approach, using the CaCo-2 cell line as an in vitro model of intestinal absorptive cells. Pangenomic screening for variations in gene expression correlating with intracellular iron content allowed us to identify 171 genes. One hundred nine of these genes are clustered into five types of expression profile. This is the first time that most of these genes have been associated with iron metabolism. Functional annotation of these five clusters indicates potential links between the immune response, proteolysis processes, and iron depletion. In contrast, iron overload is associated with cellular metabolism, especially that of lipids and glutathione involving redox function and electron transfer.

Caco-2 Cells↗

Systematic identification of abundant A-to-I editing sites in the human transcriptome.

RNA editing by members of the ADAR (adenosine deaminases acting on RNA) family leads to site-specific conversion of adenosine to inosine (A-to-I) in precursor messenger RNAs. Editing by ADARs is believed to occur in all metazoa, and is essential for mammalian development. Currently, only a limited number of human ADAR substrates are known, whereas indirect evidence suggests a substantial fraction of all pre-mRNAs being affected. Here we describe a computational search for ADAR editing sites in the human transcriptome, using millions of available expressed sequences. We mapped 12,723 A-to-I editing sites in 1,637 different genes, with an estimated accuracy of 95%, raising the number of known editing sites by two orders of magnitude. We experimentally validated our method by verifying the occurrence of editing in 26 novel substrates. A-to-I editing in humans primarily occurs in noncoding regions of the RNA, typically in Alu repeats. Analysis of the large set of editing sites indicates the role of editing in controlling dsRNA stability.

Adenosine↗

Serial analysis of gene expression in methamphetamine- and phencyclidine-treated rodent cerebral cortices: are there common mechanisms?

Pharmacological actions of methamphetamine (METH) and phencyclidine (PCP) are different, but both of them can induce similar psychiatric disorders including abuse, intoxication, withdrawal, and psychotic symptoms like those of schizophrenia. These mental disorders are caused not only by their direct pharmacological effects, but also by secondary brain damage containing gene expression changes. In order to broadly grasp these alterations, we used serial analysis of gene expression (SAGE), a transcriptome analysis. We analyzed three cDNA libraries from cerebral cortices of saline (1 mL/kg)-, METH (4 mg/kg)-, or PCP (10 mg/kg)-treated Wistar rats (one hour after i.p. administration). The numbers of total tags were about 50,000 in each library, and approximately 18,000 kinds of tags were identified respectively. From the comparisons of three groups, we found both METH- and PCP-reactive genes. Upregulated genes contained calmodulin 2, stromal cell-derived factor receptor 1, brain-specific angiogenesis inhibitor 1-associated protein 2, ras homologue enriched in brain, basigin and thyrotropin-releasing hormone receptor. Downregulated genes contained lipocalin 2, aldolase A, importin 13, fatty acid binding protein 3, and glycine receptor alpha2 subunit. These data suggest important clues of common molecular basis in METH- and PCP-related psychiatric disorders.

Animals↗

Genomic inferences of the cis-regulatory nucleotide polymorphisms underlying gene expression differences between Drosophila melanogaster mating races.

Nucleotide sequence polymorphisms affecting gene expression occur in the regulatory region of genes (in cis) and elsewhere in the genome (in trans). Further study is required to weigh the relative importance of cis- and trans-acting mutations in mediating gene expression differences within and between species. Here, microarray hybridization experiments were used to isolate 363 gene expression differences between the female fly head transcriptomes of 2 Drosophila melanogaster strains. One strain (French) represented the cosmopolitan M mating race and the other strain (ZS30) represented the Z mating race derived from Zimbabwe, Africa. From chromosomal substitution strains engineered from the 2 strains, we inferred that the expression differences between M and Z alleles largely could be attributed to the genotype of the chromosomes where the differentially expressed genes were located, that is, cis-regulatory polymorphisms prominently influence gene expression differences between M and Z. The effects of trans-regulatory polymorphisms were apparent yet difficult to quantify. Results have implications for models of gene regulatory evolution as well as experimental studies trying to identify the nucleotide sequence polymorphisms underlying gene expression differences between Drosophila strains.

Animals↗

Gene expression profiles in young adult Ciona intestinalis.

Comparison of 12,230 expressed sequence tags (ESTs) of 3' ends of cDNA clones derived from young adults of Ciona intestinalis allowed us to categorize them into 976 independent clusters. When the 5'-end sequences of 10,400 ESTs of the 976 clusters were compared with the sequences in databases, 406 of the clusters showed significant matches ( P < E-15) with reported proteins with defined functions, while 117 showed matches with putative proteins for which there is not enough information to categorize their function, and 453 had no significant sequence similarities to known proteins. The 406 clusters with sequence similarity to proteins with defined functions consisted of 304 clusters related to proteins with functions common to many kinds of cells, 73 related to proteins associated with cell-cell communication and 29 related to transcription factors. Spatial expression of all of the 976 clusters was examined by a newly improved whole-mount in situ hybridization method. A total of 430 clusters did not show distinct in situ hybridization signals, while 122 clusters showed ubiquitous distribution of signals, and 253 clusters showed signals in multiple tissues. The remaining 171 clusters showed signals specific to a certain organ or tissue: 16 showed epidermis-specific expression, 3 were specific to the neural complex, 1 to heart, 6 to body-wall muscle, 94 to pharyngeal gill, 3 to esophagus, 26 to stomach, 1 to intestine and 21 to endostyle. Many of these organ-specific genes encode proteins with no sequence similarity to known proteins. The present analysis thus highlights characteristic gene expression profiles of Ciona young adults and provides not only molecular markers for organs and tissues but also transcriptomic information useful for further genomic analyses of this model organism.

Animals↗

A complete and near-perfect rhesus macaque reference genome: lessons from subtelomeric repeats and sequencing bias.

A truly complete, telomere-to-telomere (T2T), and error-free reference genome remains a foundational resource-and long-standing goal-for unbiased comparative and functional genomics. While recent T2T assemblies of humans and other primates have made substantial progress, most still contain thousands of base-level errors, particularly within highly repetitive regions. Here, we present T2T-MMU8v2.0, a near-perfect T2T assembly of the rhesus macaque (Macaca mulatta), representing the highest base-level accuracy reported in a primate genome to date. By employing an optimized ONT-only assembly strategy, we identify subtelomeric satellite-rich regions as the principal bottleneck to improving assembly quality, owing to technological biases in long-read platforms and limitations in current hybrid assembly frameworks. We discover 268 previously unannotated repeat families and resolve ~8 Mbp of SATR satellite arrays, with over 99-fold enrichment in historically misassembled subtelomeric regions. These satellites form four distinct genomic architectures, each with unique SATR satellite composition, segmental duplication organization, and epigenetic signatures, distinct from the subtelomeric architectures observed in hominid genomes. Notably, in contrast to the largely gene-poor subtelomeric regions in African hominids, the SATR architectures in macaques harbor 58 actively transcribed genes, supported by open chromatin and expression data, suggesting gene innovation within these repetitive regions. Functionally, T2T-MMU8v2.0 improves read mappability and accuracy across sequencing platforms, and results in a 19% improvement of transcription start site enrichment scores and 5,821 additional chromatin accessibility peaks on average, thereby enhancing variant detection, regulatory annotation, and transcriptomic resolution in population genetics or single-nucleus studies. Together, this work establishes a new benchmark for genomics, offers a roadmap for resolving complex repetitive regions, and reveals previously unrecognized features of subtelomeric genome structure and evolution.

Journal Article↗

Gene expression profiling of Escherichia coli expressing double Vitreoscilla haemoglobin.

In a recent investigation, expression of a double Vitreoscilla haemoglobin (two fused VHb molecules) in Escherichia coli grown in shake flasks resulted in higher final cell density and considerably higher levels of ribosomes and tRNA. In this study, we have investigated the E. coli transcriptome in cells expressing native VHb, double VHb and control cells lacking VHb by hybridising mRNA from the different constructs to high-density oligonucleotide arrays. Within the 95% confidence interval, 4 and 5% of all detected genes in native VHb cells were up- and down-regulated, respectively; in double VHb cells the corresponding numbers were 6 and 10%, respectively. Dividing the data into different functional groups revealed that genes involved in energy metabolism, central intermediary metabolism and cell processes were the most affected at the mRNA level. Particularly, the up-regulation of genes involved in translation and posttranslational modification observed in double VHb cells demonstrates a strong relationship between the regulation of ribosomal genes and the actual number of ribosomes.

Bacterial Proteins↗

In silico analysis of 2085 clones from a normalized rat vestibular periphery 3' cDNA library.

The inserts from 2400 cDNA clones isolated from a normalized Rattus norvegicus vestibular periphery cDNA library were sequenced and characterized. The Wackym-Soares vestibular 3' cDNA library was constructed from the saccular and utricular maculae, the ampullae of all three semicircular canals and Scarpa's ganglia containing the somata of the primary afferent neurons, microdissected from 104 male and female rats. The inserts from 2400 randomly selected clones were sequenced from the 5' end. Each sequence was analyzed using the BLAST algorithm compared to the Genbank nonredundant, rat genome, mouse genome and human genome databases to search for high homology alignments. Of the initial 2400 clones, 315 (13%) were found to be of poor quality and did not yield useful information, and therefore were eliminated from the analysis. Of the remaining 2085 sequences, 918 (44%) were found to represent 758 unique genes having useful annotations that were identified in databases within the public domain or in the published literature; these sequences were designated as known characterized sequences. 1141 sequences (55%) aligned with 1011 unique sequences had no useful annotations and were designated as known but uncharacterized sequences. Of the remaining 26 sequences (1%), 24 aligned with rat genomic sequences, but none matched previously described rat expressed sequence tags or mRNAs. No significant alignment to the rat or human genomic sequences could be found for the remaining 2 sequences. Of the 2085 sequences analyzed, 86% were singletons. The known, characterized sequences were analyzed with the FatiGO online data-mining tool (http://fatigo.bioinfo.cnio.es/) to identify level 5 biological process gene ontology (GO) terms for each alignment and to group alignments with similar or identical GO terms. Numerous genes were identified that have not been previously shown to be expressed in the vestibular system. Further characterization of the novel cDNA sequences may lead to the identification of genes with vestibular-specific functions. Continued analysis of the rat vestibular periphery transcriptome should provide new insights into vestibular function and generate new hypotheses. Physiological studies are necessary to further elucidate the roles of the identified genes and novel sequences in vestibular function.

Afferent Pathways↗

Transcriptome of axenic liver stages of Plasmodium yoelii.

Plasmodium liver stages or early exo-eythrocytic forms (EEFs) contain antigens that are essential for achieving sterile, protective immunity against malaria. Yet, attempts at identifying these antigens have been hampered by the challenge of obtaining large numbers of purified EEFs, uncontaminated with hepatocyte material. Using a recently described system for producing axenically cultured EEFs from Plasmodium yoelii, we have constructed a cDNA library and generated 1453 expressed sequence tags (ESTs) resulting in 652 unique transcripts. Analysis of the library provides insight into processes required for the initiation and development of Plasmodium liver stages, such as protein degradation, cell cycle progression and nutrient transport. Analysis of the gene expression profile of liver stages, as revealed by this library, suggests that liver stages represent a shift from "sporozoite-like" to "blood-stage-like". This is the first study of the transcriptional repertoire of Plasmodium liver stages.

Animals↗

Integrative multi-omics analysis proposes a metabolic classification of gliomas: distinct metabolic states, immune infiltration, and prognosis.

BACKGROUND: The tumor microenvironment (TME) of glioma harbors diverse cell types; however, cell metabolic heterogeneity remains to be explored. This study aims to characterize the metabolic features of different cell types in the TME by integrating multiple datasets, including genomics, bulk and single-cell transcriptomics, and metabolomics. METHODS: Unsupervised machine learning was used to construct an energy metabolic classifier based on the metabolic pathways identified from bulk RNA-seq of gliomas in the TCGA dataset. The classifier was externally validated using multiple datasets, including genomics, bulk RNA-seq, snRNA-seq, and the metabolomics data. Furthermore, metabolic heterogeneity associated with the classifier was further characterized at single-cell resolution. RESULTS: The energy metabolism-based classifier stratified patients into two prognostic clusters: patients in cluster 1 were characterized by high pathway activity of glycolysis, the pentose phosphate pathway (PPP), and fatty acid oxidation (FAO), whereas patients in cluster 2 exhibited higher activity in glutaminolysis. This metabolic classifier revealed both intratumoral and intertumoral metabolic heterogeneity, and the complexity was further validated by the metabolomics profiling and snRNA-seq data from the CPTAC dataset. Notably, OSMR, highly expressed in cluster 1, showed significant co-expression with key glycolytic enzyme genes. The OSM/OSMR/JAK1/STAT3 axis potently drives malignant progression of glioma cells, specially enhancing their invasive and migratory capabilities. Single-cell resolution analyses demonstrated that tumor metabolic heterogeneity is primarily driven by malignant cells rather than non-malignant components, while tumor microenvironment (TME) factors were also found to modulate malignant cell metabolism. Significantly, glycolytic activity in glioma cells increased during the phenotypic transition from PN (proneural) to MES (mesenchymal), with cluster 1 metabolic phenotypes predominating in the tumor core. Compared to cluster 2, cluster 1 patients exhibited higher mRNA expression of immunosuppressive checkpoint genes, which correlated with pronounced immunosuppression in the TME. Furthermore, various immune cells demonstrated distinct metabolic preferences at single-cell resolution. CONCLUSIONS: This study developed an energy metabolic-based classifier for gliomas with prognostic and therapeutic potential. Metabolic reprogramming was linked with the PN-to-MES transition of glioma cells and immunosuppression in the tumor microenvironment. Multi-omics data, especially snRNA-seq, offered insights into metabolism heterogeneity at single-cell resolution, enabling personalized treatment strategies.

Humans↗

Transcriptome analysis of the murine forelimb and hindlimb autopod.

To gain insight into the coordination of gene expression profiles during forelimb and hindlimb differentiation, a transcriptome analysis of mouse embryonic autopod tissues was performed using Affymetrix Murine Gene Chips (MOE-430). Forty-four transcripts with expression differences higher than 2-fold (T test, P < or = 0.05) were detected between forelimb and hindlimb tissues including 38 new transcripts such as Rdh10, Frzb, Tbx18, and Hip that exhibit differential limb expression. A comparison of gene expression profiles in the forelimb, hindlimb, and brain revealed 24 limb-signature genes whose expression was significantly enriched in limb autopod versus brain tissue (fold change >2, P < or = 0.05). Interestingly, the genes exhibiting enrichment in the developing autopod also segregated into significant fore- and hindlimb-specific clusters (P < or = 0.05) suggesting that by E 12.5, unique gene combinations are being used during the differentiation of each autopod type.

Animals↗

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified &#x223c;218,000 OGs supported by expression evidence, while &#x223c;330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans↗