PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “de novo transcriptome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Ulmus minor response to Dutch elm disease: de novo transcriptome assembly and annotation.

Dutch elm disease (DED), caused by Ophiostoma novo-ulmi (ONU), has devastated elm populations across Europe and North America since the 20th century. In this work, a de novo transcriptome assembly of Ulmus minor in response to ONU is presented. We used two DED-resistant genotypes, MDV2.3 and VAD2, and one DED-susceptible genotype, MDV1, to capture responses to ONU at four time points post-inoculation (6, 24, 72, and 144 hours). RNA from collected samples was isolated and sequenced producing 60.88 M 100 bp paired-end reads per sample. We performed a de novo transcriptome assembly combining data from the three genotypes. The assembly was functionally annotated and validated through differential gene expression analysis of the response. This dataset provides a valuable resource for studying molecular mechanisms of DED resistance in elms, contributing to broadening our understanding of tree immunity and facilitating potential applications in functional annotation of future genome assemblies.

Transcriptome↗

De novo transcriptome meta-analysis reveals candidate genes involved in life-stage transitions for RNAi-mediated management of the citrus root weevil (Diaprepes abbreviatus).

BACKGROUND: The citrus root weevil, Diaprepes abbreviatus, is a destructive agricultural pest for which molecular control options remain limited due to historically sparse genomic resources. Leveraging a comprehensive de novo transcriptome, we investigated developmental gene regulation across larval, pupal, and adult stages and identified essential targets for RNA interference (RNAi)-based intervention. RESULTS: Stage-resolved transcriptomic analyses revealed extensive transcriptional reprogramming associated with metabolism, detoxification, cuticle biosynthesis, endocrine signaling, and sensory perception. Among these, chitin synthase (DaCHS) emerged as a critical developmental gene, exhibiting pronounced up-regulation during late larval and pupal stages corresponding to intensive cuticle synthesis. Phylogenetic and structural analyses demonstrated that DaCHS is highly conserved among insects and retains canonical catalytic domains and transmembrane topology. Alpha Fold-based structural modeling and molecular docking confirmed stable interaction of DaCHS with its substrate, N-acetylglucosamine, supporting functional conservation of enzymatic activity. Oral delivery of DaCHS double-stranded RNA induced robust transcript suppression, leading to significant mortality and severe developmental defects, including larval and pupal abnormalities, and adults with disrupted wing and abdominal morphogenesis. CONCLUSION: These findings establish DaCHS as an indispensable gene for D. abbreviates development and validate transcriptome-guided RNAi as a powerful framework for target discovery. This work provides a strong molecular foundation for developing RNAi-based strategies that can be integrated into sustainable management programs for citrus root weevil control. © 2026 Society of Chemical Industry.

Animals↗

De novo transcriptome assembly and gene expression analysis of Cnidium officinale under high-temperature conditions.

BACKGROUND: The medicinal plant Cnidium officinale (CO) is widespread in Northeast Asia and vulnerable to heat stress. The naturally occurring composition of pharmacological ingredients of CO results in overall physiological consequences; therefore, it is crucial to have a comprehensive understanding of metabolic response to ambient heat in terms of acclimation to estimate how much CO is exposed to threatening environmental conditions. RESULTS: Transcriptome analysis is critical for understanding the consequences of long-term physiological adaptation of CO to abiotic stress. However, transcriptome analysis on this species, particularly under prolonged stress conditions, has remained limited. We employed a temperature gradient tunnel (TGT) to subject CO to high-temperature exposure for four months, enabling us to observe the cumulative effects of heat and assess its acclimation mechanisms. In the absence of genome sequencing data, we performed de novo transcriptome assembly and compared DEGs from temperature treatment plots of a TGT and a growth chamber (GC). Since interpreting transcriptomic data can be complex, we employed a sequential analytical approach, including DEG clustering, GO enrichment, KEGG pathway mapping, miRNA-target gene analysis, and multiple rounds of RNA sequencing validation. DEGs were classified into two categories: genes exhibiting significant fold changes and genes showing significant count changes rather than fold changes. Then, we analyzed the functional roles of DEGs to determine which pathways respond to ambient and stressful high temperatures and validated the findings through cross-comparison with GC. Additionally, we conducted miRNA analysis to investigate post-transcriptional regulation under high temperatures. CO grown under higher ambient temperatures exhibited slight upregulation of pathways related to protein stability and turnover, ABA biosynthesis, and energy production, such as photosynthesis and oxidative phosphorylation. However, under extreme heat stress, most metabolic pathways were downregulated except for those involved in transcription, translation, oxidative phosphorylation and the biosynthesis of cutin, suberin, and wax. CONCLUSION: This study demonstrated that proper clustering of genes based on expression levels and fold changes in two different experimental conditions, along with pathway mapping, may provide a comprehensive understanding of CO's response to heat stress. These insights could contribute to future research on heat tolerance and crop improvement.

Gene Expression Profiling↗

Transcriptomic responses of Porphyrophora sophorae larvae during licorice root colonization reveal coordinated remodeling of translation, mitochondrial energy metabolism and defense-related genes.

BACKGROUND: Porphyrophora sophorae is a subterranean piercing-sucking scale insect that damages licorice (Glycyrrhiza uralensis) roots, but the molecular responses associated with larval root colonization remain insufficiently defined. METHODS: We compared non-parasitic larvae (NP) and root-colonizing larvae (RC) using six RNA-seq libraries, de novo transcriptome assembly, DESeq2-based differential expression analysis, GO/KEGG enrichment, annotation-based candidate gene screening, and RT-qPCR validation of selected genes. RESULTS: Sequencing yielded 260.91 million clean reads, and de novo assembly produced 60,794 non-redundant transcripts. DESeq2 identified 703 FDR-significant DEGs, including 49 upregulated and 654 downregulated genes in RC larvae. Upregulated genes were mainly associated with translation- and ribosome-related processes, whereas downregulated genes were enriched in mitochondrial, oxidation-reduction, energy metabolism, and oxidative phosphorylation-related functions. Annotation-based screening identified 75 FDR-significant candidate genes associated with chemosensation, defense-related responses, and energy metabolism, with mitochondrial energy metabolism-related genes forming the largest module. RT-qPCR validation based on the raw Ct data showed concordant expression directions for ten selected transcript targets. CONCLUSIONS: Root colonization in P. sophorae larvae was associated with coordinated transcriptional remodeling involving selective activation of translation-related processes, adjustment of mitochondrial energy metabolism, and changes in defense-related gene expression. These results provide candidate molecular targets for future functional studies of host contact, feeding establishment, and physiological adjustment in this subterranean scale insect.

Animals↗

De novo assembly of transcriptomes of six Hua species (Semisulcospiridae, Cerithioidea, Gastropoda).

Species in Semisulcospiridae are important in freshwater ecology and have great research value, yet their genomic resources remain very limited. Here, we present de novo assembled transcriptomes from six species of Hua in Semisulcospiridae, including Hua textrix (Heude, 1888), H. yangi L.-N. Du, J.-X. Yang & Chen, 2023, H. wujiangensis L.-N. Du, J.-X. Yang & Chen, 2023, and three undescribed species. Assembly was performed using Trinity, resulting in average contig lengths ranging from 716.6 to 883.3 bp and transcript numbers ranging from 147,147 to 268,741. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis was used to assess the transcriptome completeness. The functional annotation of transcripts for each species had over 18,000 BLAST hits, 17,000 GO terms, 15,000 KEGG pathways, 8,000 Pfam accessions, and 140 COG functional categories. This study provides valuable transcriptomic resources for the six Hua species, which can be used for various research of Semisulcospiridae, including biodiversity, phylogeny, and comparative genomics.

Transcriptome↗

HERV Modulation in Colorectal Carcinoma Patients: A Snapshot of Endogenous Retroviral Transcriptome.

Human endogenous retroviruses (HERVs) are proviral relics of infections that affected primates' germ line. Many HERV elements retain a residual capacity to encode transcripts and proteins that have been occasionally domesticated for the host physiology. In addition, HERV transcriptional modulation is of great interest to clarify the etiology of complex disorders such as cancer, even if a few studies assessed the specific HERV loci modulated in tumor tissues. In the present work, we used a transcriptomic approach to investigate the specific expression of ~3300 HERV loci in paired tumor and normal tissues of 7 colorectal cancer (CRC) patients. A total of 102 HERVs were significantly modulated in CRC, with a general tendency towards downregulation. Of note, among the 42 upregulated HERVs 23 belonged to the HERV-H group, that is the most investigated in CRC. De novo transcriptome reconstruction and qPCR validation allowed to identify a transcript from a HERV-H locus on chromosome Xp22.3 with high specific expression in CRC samples, potentially encoding for a partial Pol protein. These results provide a detailed description of HERV transcriptional variations in CRC and its interindividual variability, identifying a HERV-H transcript that deserves further investigation for its possible impact on tumor progression.

Endogenous Retroviruses↗

Time-course transcriptome and proteomic dynamics during the de novo shoot organogenesis in Chinese fir (Cunninghamia lanceolata).

De novo shoot organogenesis (DNSO) enables plants to regenerate shoots from various explants, offering valuable opportunities for research and plant biotechnology applications. While significant progress has been made in understanding regeneration in angiosperms, the regulatory mechanisms in gymnosperms, particularly Chinese fir (Cunninghamia lanceolata), remain poorly understood, despite its importance as a key timber species in China. This study successfully established an efficient DNSO protocol for Chinese fir, identifying six distinct stages in the process through cellular-level analysis. Time-course transcriptome and proteomics analyses revealed dynamic changes in mRNA and protein levels during regeneration. Notably, proteins showed more significant alterations across a broad range of biological processes, often independent of corresponding mRNA changes. Key pathways associated with ethylene metabolism and abiotic stress responses were enriched, highlighting their critical roles in regeneration. Further experiments confirmed that moderate osmotic stress treatments (150 mm mannitol) and ethylene treatment (100 μm ACC and 5 μm AgNO3) substantially enhanced DNSO efficiency. In summary, this study uncovers the molecular mechanisms underlying Chinese fir DNSO, providing valuable insights into improving plant regeneration efficiency in this economically important species. These findings contribute to advancements in plant biotechnology and sustainable forestry practices.

Cunninghamia↗

Genomic and Transcriptomic Profiling of Radiation-Resistant, Locally Recurrent Prostate Cancer.

PURPOSE: The biology of locally radiorecurrent prostate cancer (LRR-PCa) is poorly understood. METHODS AND MATERIALS: We sought to explore the genomic and transcriptomic landscape of LRR-PCa with targeted DNA sequencing and RNA expression analysis from 41 biopsy-proven LRR-PCa tumors from 36 unique patients who had a recurrence at a median interval of 84 months (IQR, 70-124 months). Genomic alteration frequencies and transcriptomic data were compared between the LRR-PCa cohort and treatment-na&#xef;ve patients from the Cancer Genome Atlas (genomic; n = 496) and Gleason grade-at-recurrence-matched patients from the Decipher Genomics Resource for Intelligent Discovery (transcriptomic; n = 22,320). RESULTS: Twenty-five patients (69%) had pathologic upgrading at recurrence (17% vs 64% with Gleason grade 4-5 disease; P < .001). The LRR-PCa cohort demonstrated significantly greater single-nucleotide variations in 29 genes known to be associated with prostate cancer, including several associated with increased aggressiveness and DNA repair: FAT1 (58.5% vs 1.0%), RAD51B (36.6% vs 0.4%), POLQ (34.1% vs 1.4%), KMT2C (34.1% vs 4.9%), BRCA2 (29.3% vs 1.8%), ATRX (26.8% vs 0.8%), and BRCA1 (24.4% vs 0.4%) (Pvalues < .001 for all). The LRR-PCa cohort had a significantly higher Decipher score (median, 0.80 vs 0.66; P = .05) and demonstrated significantly greater basal subtype based on PAM50 (56% vs 20%; P < .001) and lower androgen receptor activity (61% for LRR vs 9%; P < .001). CONCLUSIONS: Overall, these results suggest that LRR-PCa has a distinct genomic and transcriptomic landscape from de novo prostate cancer. Specifically, LRR-PCa has an enrichment in SNVs in genes associated with tumor aggressiveness and/or DNA repair, has higher Decipher scores, a more basal subtype, and has transcriptomic evidence of lower androgen receptor activity and loss of tumor suppressor genes.

Humans↗

De novo clustering of large long-read transcriptome datasets with isONclust3.

MOTIVATION: Long-read sequencing techniques can sequence transcripts from end to end, greatly improving our ability to study the transcription process. Although there are several well-established tools for long-read transcriptome analysis, most are reference-based. This limits the analysis of organisms without high-quality reference genomes and samples or genes with high variability (e.g. cancer samples or some gene families). In such settings, analysis using a reference-free method is favorable. The computational problem of clustering long reads by region of common origin is well-established for reference-free transcriptome analysis pipelines. Such clustering enables large datasets to be split roughly by gene family and, therefore, an independent analysis of each cluster. There exist tools for this. However, none of those tools can efficiently process the large amount of reads that are now generated by long-read sequencing technologies. RESULTS: We present isONclust3, an improved algorithm over isONclust and isONclust2, to cluster massive long-read transcriptome datasets into gene families. Like isONclust, isONclust3 represents each cluster with a set of minimizers. However, unlike other approaches, isONclust3 dynamically updates the cluster representation during clustering by adding high-confidence minimizers from new reads assigned to the cluster and employs an iterative cluster-merging step. We show that isONclust3 yields results with higher or comparable quality to state-of-the-art algorithms but is 10-100 times faster on large datasets. Also, using a 256&#x2009;Gb computing node, isONclust3 was the only tool that could cluster 37 million PacBio reads, which is a typical throughput of the recent PacBio Revio sequencing machine. AVAILABILITY AND IMPLEMENTATION: https://github.com/aljpetri/isONclust3.

Algorithms↗

Divergent Lineage of Terpene Synthases Establishes Terpenoid Biosynthesis in Brown Macroalgae.

Brown algae of the order Dictyotales uniquely stand out among stramenopiles (heterokonts) as prolific producers of bioactive terpenoid molecules associated with chemical defense and antifouling. Although more than 200 sesquiterpenoids and diterpenoids have been reported, largely from the genera of Dictyota and Dictyopteris, their biosynthetic origin has remained unknown for decades. Leveraging de novo genome and transcriptome sequencing in the nonmodel alga Dictyota coriacea, we identified a brown algal-specific lineage of type I terpene synthases (TSs) that harbors novel catalytic motifs distinct from those characterized in plants, microbes, red algae, and metazoans. Across three brown algal species, we characterized 15 terpene synthases, including DcTS-2, which produces the diterpene alcohol dilophol, a proposed biosynthetic intermediate to the antifouling metabolite pachydictyol A. X-ray crystal structures of the monoterpene synthase DcTS-3 further revealed that the brown algal enzymes retain the canonical terpene synthase fold, and together with mutagenesis studies, suggest the catalytic role of the novel motifs defining this newly established evolutionary lineage. Brown algal terpene synthases separate into two subgroups, with mono- and diTSs containing putative chloroplast-targeting sequences while sesquiTSs lack them, suggesting convergent compartmentalization of terpene biosynthesis with land plants. Together, these findings establish the molecular basis of terpenoid biosynthesis in brown algae and highlight the challenges of adapting established biosynthetic logic to nonmodel marine algae.

Alkyl and Aryl Transferases↗

Towards a transcriptome definition of microglial cells.

This study provides an expression signature of interferon-gamma (IFN-gamma)-activated microglia. Microglia are macrophage precursor cells residing in the brain and spinal cord. The microglial phenotype is highly plastic and changes in response to numerous pathological stimuli. IFN-gamma has been established as a strong immunological activator of microglial cells both in vitro and in vivo. Affymetrix RG_U34A microarrays were used to determine the effect of IFN-gamma stimulation on migroglia cells isolated from newborn Lewis rat brains. More than 8,000 gene sequences were examined, i.e., 7,000 known genes and 1,000 expressed sequence tag (EST) clusters. Under baseline conditions, microglia expressed 326 of 8,000 genes examined (approximately 4% of all genes, 182 known and 144 ESTs). Transcription of only 34 of 7,000 known genes and 8 of 1,000 ESTs was induced by IFN-gamma stimulation. The majority of the newly expressed genes encode pro-inflammatory cytokines and components of the MHC-mediated antigen presentation pathway. The expression of 60 of 182 identified genes and of 9 of 144 ESTs was increased by IFN-gamma, whereas 29 of 182 known genes and 7 of 144 ESTs were down-regulated or undetectable in IFN-gamma-stimulated cultures. Overall, the activating effect of IFN-gamma on the microglial transcriptome showed restriction to pathways involved in antigen presentation, protein degradation, actin binding, cell adhesion, apoptosis, and cell signaling. In comparison, down-regulatory effects of IFN-gamma stimulation appeared to be confined to pathways of growth regulation, remodeling of the extracellular matrix, lipid metabolism, and lysosomal processing. In addition, transcriptomic profiling revealed previously unknown microglial genes that were de novo expressed, such as calponin 3, or indicated differential regulatory responses, such as down-regulation of cathepsins that are up-regulated in response to other microglia stimulators.

Animals↗

Antennal transcriptome analysis of chemosensory proteins in the raspberry weevil, Aegorhinus superciliosus (Coleoptera: Curculionidae).

Aegorhinus superciliosus (Coleoptera: Curculionidae) is a polyphagous pest of economic importance in southern Chile, the chemical ecology of which remains poorly characterized. Across insect species, chemosensory proteins, including odorant receptors (ORs), gustatory receptors (GRs), ionotropic receptors (IRs), odorant-binding proteins (OBPs), chemosensory proteins (CSPs), and sensory neuron membrane proteins (SNMPs), mediate the detection of chemical cues involved in host selection, reproduction, and other ecologically relevant behaviors. In this study, the antennal transcriptome of adult A. superciliosus was sequenced and analyzed using a de novo RNA-seq approach. Three independent biological replicates per sex were used for RNA-seq, and the same number of independent biological replicates was used for RT-qPCR validation; sequencing yielded 147,409,936 high-quality reads after quality filtering. A total of 112 candidate chemosensory genes were identified, comprising 43 ORs, 34 OBPs, 10 CSPs, 18 IRs, 5 GRs, and 2 SNMPs. Phylogenetic analyses assigned these candidate proteins to established clades, providing a comparative framework for functional inference for ORs and OBPs. Sex- and tissue-biased expression analyses revealed that several ORs, including AsupOR4, AsupOR19, and AsupOBP13, exhibit antennal enrichment and sex-specific expression patterns. Notably, AsupOR19 and AsupOBP13 displayed strong female-biased expression. In addition, transcripts of selected ORs and OBPs were detected in non-antennal tissues, such as the rostrum and legs, suggesting potential functional versatility beyond canonical olfaction. Together, these findings represent the first molecular identification of the chemosensory repertoire of A. superciliosus. This study establishes a foundation for reverse chemical ecology approaches aimed at identifying behaviorally active volatile organic compounds (VOCs) toward environmentally sustainable strategies for integrated pest management.

Animals↗

The UTRs of Leishmania donovani vary in length and are enriched in potential regulatory structures.

Leishmania spp. regulate gene expression largely post-transcriptionally, yet untranslated regions (UTRs) remain poorly delineated. We generated high-quality genome and transcriptome datasets for Leishmania donovani strain 1S2D (Ld1S) by combining PacBio HiFi de novo assembly with Oxford Nanopore direct RNA sequencing of promastigotes and axenic amastigotes. The genome assembly consists of 65 scaffolds totaling ~33.3 Mb. Structural comparisons to LdBPK282A1 revealed numerous rearrangements, including some reshuffling genes among polycistronic transcription units and validated by polycistronic reads from RNA sequencing. Promastigote and amastigote RNA sequencing produced 469,010 and 46,729 monocistronic reads containing a spliced-leader and a polyA tail sequences, defining 8,479 transcripts and supporting 7,415 of the 7,969 annotated protein coding genes, as well as 604 putative long non-coding RNAs. We annotated UTRs for 4,921 genes and observed that putative RNA G-quadruplexes were markedly enriched in UTRs. We also noted that 31.9% and 11.5% were expressed into multiple isoforms in promastigotes and amastigotes, respectively. Collectively, these data provide a comprehensive annotation of L. donovani genes and their UTRs and reveal widespread and stage-specific UTR length polymorphisms, and, overall, points to an important role of 3' UTR in post-transcriptional regulation in L. donovani.

Journal Article↗

A De Novo 16p13.3 Triplication Underlying Early-Onset Complex Neurodegeneration.

BACKGROUND: Neurodegenerative disorders are clinically and genetically heterogeneous, characterized by progressive neuronal loss and multidomain functional decline. Despite a presumed genetic etiology, a substantial proportion of cases remain molecularly undiagnosed. OBJECTIVE: The aim was to identify the genetic cause of an early-onset neurodegenerative disorder presenting with ataxia and cognitive impairment. METHODS: Rare copy-number variants were detected via short-read whole-genome sequencing (WGS), with candidate structural models inferred using long-read WGS. We performed transcriptomic profiling of peripheral blood leukocytes by RNA sequencing, with validation using reverse transcription-quantitative polymerase chain reaction (RT-qPCR). RESULTS: We identified a de novo copy-number gain at 16p13.3. Combined copy-number profiling and long-read WGS suggested a candidate model comprising a triplicated segment in tandem with a proximal duplication, joined to a distal duplication via an inverted junction. Transcriptomic analysis demonstrated significant upregulation of ATP6V0C, AMDHD2, and PDPK1. CONCLUSIONS: These findings support a role for structural variation in early-onset neurodegeneration and highlight the value of combining short-read copy-number profiling with long-read WGS to detect and characterize complex genomic rearrangements. &#xa9; 2026 International Parkinson and Movement Disorder Society.

16p13.3↗

De Novo TRIO Missense Variants Disrupt Ras-GEF Domains and Cause Congenital Ventriculomegaly and Hydrocephalus.

Congenital hydrocephalus (CH), characterized by congenital ventriculomegaly (CV), affects approximately 0.5-1 per 1000 live births and is a common cause of pediatric neurosurgical intervention, yet its genetic architecture remains incompletely defined. We report a child with syndromic CH requiring cerebrospinal fluid diversion who harbored a pathogenic de novo missense variant in TRIO (c.3232C > T; p.(Arg1078Trp)), a gene previously associated with autosomal dominant neurodevelopmental disorders featuring variable head circumference. This case prompted systematic evaluation of TRIO variation in our CV/CH cohort (2,697 patient-parent trios) using exome sequencing. We identified five additional unrelated probands with de novo TRIO variants, including two novel substitutions affecting the same residue within the Ras-GEF1 domain (p.(Glu1299Lys) and p.(Glu1299Gly)), yielding significant gene-level enrichment for protein-damaging de novo variants (adjusted&#x2009;p = 6.12 &#xd7; 10-5). All affected individuals exhibited CV, frequently accompanied by developmental delay and additional structural brain abnormalities. In silico structural modeling predicted that associated variants destabilize critical TRIO Ras-GEF domains required for Rho GTPase activation. Analysis of single-nucleus transcriptomic data from the developing human neocortex revealed enrichment of TRIO expression in multipotent progenitor populations. A systematic literature review identified six additional individuals with TRIO de novo variants and reported CV or CH, including an unrelated patient with the same p.(Arg1078Trp) substitution. Together, these findings expand the phenotypic spectrum associated with pathogenic TRIO variation to include CV/CH and support TRIO as a clinically relevant gene in the genetic evaluation of syndromic CV/CH patients.

Humans↗

Differential gene expression in egg cells and zygotes suggests that the transcriptome is restructed before the first zygotic division in tobacco.

We applied suppression subtractive hybridization and mirror orientation selection to compare gene expression profiles of isolated Nicotiana tabacum cv SR1 zygotes and egg cells. Our results revealed that many differentially expressed genes in zygotes were transcribed de novo after fertilization. Some of these genes are critical to zygote polarity and pattern formation during early embryogenesis. This suggests that the transcriptome is restructed in zygote and that the maternal-to-zygotic transition happens before the first zygotic division, which is much earlier in higher plants than in animals. The expressed sequence tags used in this study provide a valuable resource for future research on fertilization and early embryogenesis.

Body Patterning↗

Complex Genetics and Regulatory Drivers of Hypermobile Ehlers-Danlos Syndrome: Insights from Genome-Wide Association Study Meta-analysis.

BACKGROUND: Hypermobile Ehlers-Danlos syndrome (hEDS) is the most common subtype of EDS, a group of heritable connective tissue disorders. Clinically, hEDS is defined by generalized joint hypermobility and chronic musculoskeletal pain, but its impact extends beyond the musculoskeletal system. Affected individuals frequently experience autonomic, gastrointestinal, immune, and neuropsychiatric involvement, highlighting both the multisystemic nature of the condition and challenges of diagnosis. In contrast to other EDS subtypes with defined genetic causes, the molecular basis of hEDS has remained elusive. METHODS: We conducted a genome-wide association study (GWAS) of hEDS across three case controls studies, including 1,815 cases and 5,008 ancestry-matched controls. Fixed-effects meta-analysis of 6.2 million variants was complemented with LDAK gene-based association testing, transcriptome-wide association studies, and integrative annotation across multiple tissues and cell types including eQTLs, enhancer marks and open chromatin accessibility profiles, supported by luciferase assays on one candidate variant. LD-score genetic correlations were assessed between hEDS and 19 frequently reported comorbid conditions. RESULTS: Two loci reached genome-wide significance, including a regulatory region near the atypical chemokine receptor 3 gene (ACKR3) on chromosome 2. Functional annotation supports ACKR3 risk alleles colocalize with eQTLs in tibial nerve, alter enhancer activity, and generate a de novo AHR transcription factor regulatory site, implicating neuroimmune and pain signaling pathways. Gene-based and transcriptome-wide analyses identified common variants in a locus containing multiple candidates, including SLC39A13, a zinc transporter critical for connective tissue development previously implicated in a rare form of EDS, and PSMC3, a gene involved in central nervous system development. LD-score regression revealed significant genetic correlations between hEDS and joint hypermobility, myalgic encephalomyelitis/chronic fatigue syndrome, fibromyalgia, depression, anxiety, autism spectrum disorder, migraine, and gastrointestinal diseases. CONCLUSIONS: These results establish the first evidence of common variant contributions to hEDS, supporting a complex, multisystem model involving neuroimmune-stromal dysregulation. Our findings add novel indications to hEDS pathogenesis and provide solid foundations for future molecular definition and therapeutic discovery.

Genome-wide association study↗

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article↗