PubMed HealthSearch

SEARCH · PubMed Health

Results for “Sequence Analysis, RNA”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genetic predisposition to systemic inflammatory proteins is causally associated with inflammatory bowel disease: Insights from multi-omics association study and single-cell RNA-sequencing analysis.

Systemic inflammatory proteins have been reported to be related to inflammatory bowel disease (IBD) in previous observational research. However, their causal links remain obscure. Herein, we performed a Mendelian randomization (MR) analysis to analyze the causality between systemic inflammatory proteins and IBD. Genetic variants related to systemic inflammatory proteins were extracted from a meta-analysis of genome-wide association study (GWAS) data of 8293 European participants. Summary statistics of IBD diverse subtypes were obtained from the international IBD genetic consortium (IIBDGC). We conducted multi-omics method and MR study to detect the causal links through integrating GWAS and protein quantity trait loci (pQTL) data. Inverse variance weighted (IVW) approach was utilized as the dominated analysis method. Moreover, complementary approaches such as MR-Egger intercept test, Cochran Q test and leave-one-out analysis were utilized to validate pleiotropy and heterogeneity. Finally, single-cell RNA-sequencing analysis was performed to detect the expression of significant genes. For IBD, IVW estimates suggested that genetically predicted IL-10 and IL-13 were suggestively associated with an elevated risk of IBD (IL-10: OR: 1.12, 95% CI: 1.00-1.24, P = .04; IL-13: OR: 1.09, 95% CI: 1.01-1.18, P = .023), while CXCL10 was suggestively linked to a lower risk of IBD (CXCL10: OR: 0.90, 95% CI: 0.82-0.99, P = .037). For Crohn disease (CD), the IVW approach provided evidence to sustain that genetically determined IL-13 and CCL3 had a suggestive association with a higher risk of CD (IL-13: OR: 1.13, 95% CI: 1.02-1.26, P = .023; CCL3: OR: 1.22, 95% CI: 1.03-1.45, P = .018). Sensitivity analysis did not explore any heterogeneity and pleiotropy. Our findings supported the causal relationships between 4 specific inflammatory proteins (IL-10, IL-13, CXCL10, and CCL3) and the risk of IBD and CD, thereby providing promising biomarkers of various subtypes stratification and new insights for the prevention and therapeutic target of IBD.

Humans

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis

Dual RNA isolation from blood: an optimized protocol for host and bacterial RNA purification for dual RNA-sequencing analysis in whole blood sepsis samples.

Dual RNA-sequencing (dual RNA-seq) holds significant promise for deciphering bacterial virulence mechanisms during systemic infections. However, its application in sepsis research is hindered by technical challenges, including a low bacterial burden in blood and limited sample volumes and RNA yield from vulnerable populations, such as neonates. We developed an optimized protocol [dual RNA isolation from blood (DRIB)] for simultaneous stabilization, isolation and purification of high-quality host leukocyte and bacterial RNA from low-volume whole blood samples (0.5 ml). This protocol is compatible with clinical sample collection workflows and high-throughput RNA sequencing. The feasibility of DRIB for dual RNA-seq was validated using a pilot cohort of clinical adult sepsis samples, enabling the investigation of host-bacterial gene expression during sepsis. The DRIB protocol yielded 2.10-6.91 µg of total RNA per clinical sample in our pilot cohort. Dual-species ribosomal RNA (rRNA) depletion and RNA-seq generated 16.6-24.8 million filtered reads per sample, with 63±7% of reads uniquely mapped to host or bacterial sequences. Host genes accounted for 51-68% (8.4-10.9 million) reads, while 0.5-6.7% (79,496-789,808 reads) mapped to bacterial genomes. Bioinformatic analysis revealed that both shared and individual transcriptional patterns were identified in host and bacterial responses, including pathways related to immune metabolism and metal-ion binding. Our optimized DRIB protocol and RNA-seq pipeline effectively captured both host and bacterial RNA transcription in clinical sepsis samples. Expanding this approach to larger cohorts and varying disease timepoints will provide crucial new insights into host-bacterial gene co-expression dynamics in sepsis progression and outcomes.

Humans

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis

CLN3 transcript complexity revealed by long-read RNA sequencing analysis.

BACKGROUND: Batten disease is a group of rare inherited neurodegenerative diseases. Juvenile CLN3 disease is the most prevalent type, and the most common pathogenic variant shared by most patients is the "1-kb" deletion which removes two internal coding exons (7 and 8) in CLN3. Previously, we identified two transcripts in patient fibroblasts homozygous for the 1-kb deletion: the 'major' and 'minor' transcripts. To understand the full variety of disease transcripts and their role in disease pathogenesis, it is necessary to first investigate CLN3 transcription in "healthy" samples without juvenile CLN3 disease. METHODS: We leveraged PacBio long-read RNA sequencing datasets from ENCODE to investigate the full range of CLN3 transcripts across various tissues and cell types in human control samples. Then we sought to validate their existence using data from different sources. RESULTS: We found that a readthrough gene affects the quantification and annotation of CLN3. After taking this into account, we detected over 100 novel CLN3 transcripts, with no dominantly expressed CLN3 transcript. The most abundant transcript has median usage of 42.9%. Surprisingly, the known disease-associated 'major' transcripts are detected. Together, they have median usage of 1.5% across 22 samples. Furthermore, we identified 48 CLN3 ORFs, of which 26 are novel. The predominant ORF that encodes the canonical CLN3 protein isoform has median usage of 66.7%, meaning around one-third of CLN3 transcripts encode protein isoforms with different stretches of amino acids. The same ORFs could be found with alternative UTRs. Moreover, we were able to validate the translational potential of certain transcripts using public mass spectrometry data. CONCLUSION: Overall, these findings provide valuable insights into the complexity of CLN3 transcription, highlighting the importance of studying both canonical and non-canonical CLN3 protein isoforms as well as the regulatory role of UTRs to fully comprehend the regulation and function(s) of CLN3. This knowledge is essential for investigating the impact of the 1-kb deletion and rare pathogenic variants on CLN3 transcription and disease pathogenesis.

Humans

Identification of GREM-1 and GAS6 as Specific Biomarkers for Cancer-Associated Fibroblasts Derived from Patients with Non-Small-Cell Lung Cancer.

Background/Objectives: Cancer-associated fibroblasts (CAFs) play a pivotal role in the tumor microenvironment. We conducted an analysis using RNA sequencing to identify specific markers for CAFs compared to normal fibroblasts (NFs) in non-small-cell carcinoma (NSCLC). Methods: CAFs and NFs were isolated and cultured from tumor tissues (primary tumor or metastatic lymph nodes) and matched non-tumor tissues, respectively. Bulk RNA sequencing was conducted on isolated CAFs and normal fibroblast NFs. Differential expressions, gene set enrichment, and CAF subpopulation prediction analyses were performed. Results: During the study period, 27 CAFs and 12 NFs were isolated and cultured from tumor and non-tumor tissues in patients with treatment-na&#xef;ve NSCLC. Among them, 22 CAFs and 11 NFs were included in the RNA sequencing analysis. The 22 CAF samples consisted of 12 adenocarcinomas and 10 squamous cell carcinomas (SqCC), with 16 samples from the lungs and 6 samples from the lymph nodes. Notably, COL11A1, GREM1, CD36, and GAS6 showed a higher expression in CAFs than in NFs, whereas TNC and CXCL2 were more abundantly expressed in NFs. CD36 levels were elevated in CAFs from lymph nodes (LN-CAFs) compared with those from lung specimens (Lung-CAFs) and NFs. COL11A1 levels in Lung-CAFs surpassed those in LN-CAFs and NFs. Both GREM1 and GAS6 showed a strong expression in Lung-CAFs and LN-CAFs relative to NFs. CAFs exhibited features of the myofibroblast CAF subpopulation, whereas NFs displayed traits of the antigen-presenting CAF subtype. In the co-culture model of CAFs and THP-1 cells, the knockdown of GREM1 or GAS6 in CAFs significantly decreased the M2 marker expression in macrophages. Conclusions: In NSCLC, GREM1 and GAS6 can be valuable diagnostic targets for CAFs from primary tumors and metastatic sites; they warrant further study.

cancer-associated fibroblast

From transcriptomic profiling to precision oncology: a bibliometric analysis of RNA sequencing in acute myeloid leukemia.

BACKGROUND: RNA sequencing (RNA-seq) has become an important tool for investigating the molecular heterogeneity of acute myeloid leukemia (AML); however, the global development and thematic evolution of this field remain inadequately characterized. OBJECTIVE: To map the global landscape of AML RNA-seq research and identify major knowledge domains, emerging themes, and temporal changes in research priorities. METHODS: Publications indexed in the Web of Science Core Collection and Scopus between January 1, 2007, and August 18, 2025, were retrieved. After database filtering, merging, and deduplication, 3,460 articles and reviews were included. CiteSpace, VOSviewer, the bibliometrix R package, and Microsoft Excel were used to analyze publication trends, collaboration networks, co-citation structures, keyword evolution, and citation bursts. RESULTS: Publication output increased steadily, accelerating after 2014. China contributed the largest number of publications (n&#x202f;=&#x202f;547, 15.8%), whereas the United States had the highest total citation count. Major publication outlets spanned hematology, oncology, genomics, and molecular biology. Co-citation analysis identified prominent themes involving next-generation sequencing, gene mutations, KMT2A rearrangements, epigenetic dysregulation, leukemia-initiating cells, drug resistance, biomarkers, T-cell biology, and single-cell sequencing. Earlier literature emphasized sequencing technologies, gene expression profiling, and molecular alterations, whereas recent publications show increasing representation of cellular heterogeneity, single-cell transcriptomics, drug resistance, biomarker applications, immune-related research, and computational interpretation. CONCLUSION: While molecular characterization remains foundational, AML RNA-seq research has broadened to encompass increasingly prominent cellular, functional, computational, and translational dimensions. This study provides a structured overview of the field; nevertheless, bibliometric prominence should not be interpreted as direct evidence of clinical utility.

RNA sequencing

Biallelic VPS41 Variants in Autosomal Recessive Spinocerebellar Ataxia 29 Resolved by Long-Read Sequencing and RNA Analysis.

BACKGROUND: Biallelic variants in VPS41, encoding a subunit of the HOPS complex, cause autosomal recessive spinocerebellar ataxia 29 (SCAR29), a rare neurodevelopmental disorder with an incompletely defined phenotypic and molecular spectrum. METHODS: We investigated a 24-year-old man with cerebellar ataxia, hypotonia, and intellectual disability. Exome sequencing identified four candidate VPS41 variants. Because maternal DNA was unavailable, long-read genome sequencing was performed to determine allelic configuration, followed by RNA and protein analyses. RESULTS: In addition to typical SCAR29 features, the patient showed previously unreported findings, including swan-neck deformities and pes cavus. Long-read genome sequencing demonstrated that two VPS41 variants were in trans. RNA analysis revealed distinct splicing consequences: one allele produced an out-of-frame transcript predicted to undergo nonsense-mediated decay, whereas the other generated an in-frame exon-skipped transcript. These complementary defects reduced VPS41 expression at both transcript and protein levels, supporting pathogenicity and variant reclassification. CONCLUSION: Our findings expand the phenotypic spectrum of VPS41-related disease and highlight the value of long-read allelic resolution in clarifying pathogenic mechanisms in rare genetic disorders.

Humans

TGIRT-seq to profile tRNA-derived RNAs and associated RNA modifications.

RNA modifications are key regulators for RNA processes. tRNA-derived RNAs are small RNAs with size between 15 and 50 bases long that are processed from mature or precursor tRNAs. Despite their more recent discovery, tRNA-derived RNAs have been found to play regulatory roles in many cellular processes including gene silencing, protein synthesis, stress response, and transgenerational inheritance. Furthermore, tRNA-derived RNAs are highly abundant in bodily fluids, posing as potential biomarkers. A unique feature of tRNA-derived RNAs is that they are rich in RNA modifications. Many of the RNA modifications on tRNA-derived RNAs disrupt Watson-Crick base pairing and will thus stall reverse transcriptase, such as N1-methyladenosine (m1A), N1-methylguanosine (m1G) and N2, N2-dimethylguanosine (m22G). These RNA modifications add another layer of regulation onto tRNA-derived RNAs' functions and are of interests for future research. However, these RNA modifications could also lead to lower detection of modification-containing RNAs in genome-wide small RNA sequencing analysis due to reverse transcriptase stall. To circumvent this bias, TGIRT (Thermostable Group II Intron Reverse Transcriptase) has been used to readthrough RNA modifications inserting mismatches. These mismatch signatures can then be used to precisely map the modification sites at base resolution. Here we describe the step-by-step experimental protocol to start with purified RNAs from cells or tissues and use TGIRT to make small RNA sequencing library for Illumina sequencing to profile the abundance of tRNA-derived RNAs and the associated RNA modifications.

RNA, Transfer

Analysis of ferroptosis-related genes in cerebral ischemic stroke via immune infiltration and single-cell RNA-sequencing.

Ischemic stroke (IS) represents a harmful neurological disorder with limited treatment options. Ferroptosis accounts for the iron-dependent, nonapoptotic cell death pattern, which shows the feature of fatal lipid ROS accumulation. Nonetheless, ferroptosis-related biomarkers for identifying IS early are currently lacking. The present study focused on investigating the possible ferroptosis-related biomarkers for IS and analyzing their effects on immune infiltration. Altogether five hub differentially expressed ferroptosis-related genes (DEFRGs) were identified from the relevant databases. Additionally, single-cell RNA-sequencing (seq) analysis was conducted for the comprehensive mapping of cell populations based on the IS database. These five hub DEFRGs were analyzed using gene set enrichment analysis, miRNA prediction, and single-cell RNA-seq analysis. A transient middle cerebral artery occlusion mouse model was constructed. We also adopted bioinformatics methods combined with western blot, changes to mitochondria, hematoxylin & eosin staining, Nissl staining, ROS fluorescence staining, immunohistochemistry, and quantitative real-time polymerase chain reaction (qRT-PCR) to show the involvement of ferroptosis in IS progression. The results revealed that nuclear factor erythroid-derived 2-like 2 (Nfe2l2) was the potential candidate biomarker for IS diagnosis, and ferroptosis may be suppressed via the Nfe2l2/HO-1 pathway. Thus, drug targeting Nfe2l2 can shed novel lights on IS treatment.

Ferroptosis

Tandem splice acceptor sites: Profiling their relevance to human disease.

PURPOSE: Interpretation of variation, particularly the creation or disruption of tandem splice acceptor sites (NAGNnAG variants), challenges genomic medicine practice. METHODS: We analyzed the creation and disruption of dinucleotide AG sites within &#xb1;30 bases of natural splice-acceptor sites in the GRCh37 human reference genome. These results were compared with variant data from the ClinVar and gnomAD databases, as well as with data from 779 National Institutes of Health Undiagnosed Diseases Program study participants. Using RNA sequencing, we assessed the splicing at NAGNnAG variants for 107 of the Undiagnosed Diseases Program participants and compared the empirical data with SpliceAI predictions. RESULTS: Creation or disruption of NAGNnAG sites within 30 bases of the natural splice acceptor are enriched in ClinVar compared with gnomAD; however, such variants in the 2 databases are rarely differentiated by SpliceAI scores. Empirical evaluation via RNA sequencing analysis supported novel acceptor site usage from -21 to +30; splice-altering variants did not predominate in a specific region or have SpliceAI scores invariantly, suggesting increased spliceogenicity. CONCLUSION: NAGNnAG variants within 30 bp of the natural splice acceptor have a high probability of clinical relevance and are poorly contextualized for clinical utility. Their interpretation benefits from empirical evaluation via RNA analysis.

Humans

Extraction, Purification, and Next-Generation Sequencing (NGS) Analysis of DNA and RNA from Formalin-Fixed and Paraffin-Embedded (FFPE) Tissue.

Formalin fixed paraffin embedded (FFPE) tissues have long been used for immunohistological analyses. FFPE tissues can be stored at room temperature for several years enabling analyses to be performed later. Ease of storage and transport makes these tissues an attractive source of biological material. However, formalin fixation results in chemical modifications of proteins and nucleic acids that poses a major challenge to any type of analysis. Recovery of nucleic acids for quantitative assays is rendered difficult due to degradation resulting from fixation and long-term storage, producing low usable yields. Extensive efforts in the last 20&#xa0;years have led to significant improvements in use of FFPE tissues for DNA and RNA analyses and resulted in development of sensitive assays for a wide range of applications, including next-generation sequencing. In this chapter, we describe the optimization of methods for sequential extraction of DNA and RNA from FFPE tissue and subsequent preparation of DNA-seq and RNA-seq libraries for use with the Illumina platform using commercially available reagents/kits.

Paraffin Embedding