PubMed HealthSearch

SEARCH · PubMed Health

Results for “Massively parallel sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

27 records · Page 2Linked to original sources

Regulatory Evolution and the Genetic Basis of Human Brain Expansion.

The evolution of the human brain is characterized by profound changes in structure and function, despite relatively limited divergence in protein-coding genes compared to other primates. This paradox has led to increasing recognition of gene regulatory elements (GREs) as primary drivers of evolutionary innovation. In this review, we synthesize current knowledge on the role of conserved noncoding elements (CNEs), human accelerated regions (HARs), and transposable element (TE)-derived sequences in shaping gene regulatory networks (GRNs) underlying brain development. Comparative analyses across humans and closely related primates, including the chimpanzee, gorilla, and orangutan, reveal that while core regulatory architectures are highly conserved, subtle changes in regulatory elements drive species-specific gene expression patterns. We highlight how CNEs provide a stable regulatory framework, whereas HARs and TE-derived elements introduce lineage-specific modifications that fine-tune neurodevelopmental processes. Advances in functional genomics, including CRISPR-based perturbations, massively parallel reporter assays, and single-cell multi-omics, have enabled direct interrogation of regulatory function, linking sequence variation to cellular phenotypes. Furthermore, we discuss how regulatory evolution contributes to both cognitive innovation and susceptibility to neurological disorders. Despite significant progress, challenges remain in establishing causal relationships between regulatory variation and phenotypic outcomes. Future integration of multi-omics data and comparative models will be essential for resolving these complexities. Together, this review provides a comprehensive framework for understanding the molecular basis of primate brain evolution through the lens of gene regulation.

Brain evolution

Deciphering the prodrome of inflammatory bowel disease up to 10 years before disease onset by massively parallel serology.

BACKGROUND: Defining immune dysregulation during the asymptomatic prodrome of immune-mediated diseases offers opportunities for early disease detection and interception. In inflammatory bowel disease (IBD), prodromal immune changes remain poorly characterised. OBJECTIVE: To define preclinical immunological alterations by characterising longitudinal serum antibody repertoires using high-throughput phage-display immunoprecipitation sequencing (PhIP-Seq). DESIGN: We applied PhIP-Seq to profile antibody responses in 2000 longitudinal serum samples from 200 individuals who developed Crohn's disease (CD), 200 who developed ulcerative colitis (UC) and 100 matched healthy controls within the US military Proteomic Evaluation and Discovery in an IBD Cohort of Tri-service Subjects cohort, collected up to 10 years before diagnosis. Antibody repertoires were profiled against 357 000 microbial-associated, viral-associated, food-associated and immune-associated peptides. RESULTS: Antibody repertoire variability was increased up to ~4 years prediagnosis in pre-CD and pre-UC individuals. Differential analyses revealed elevated herpesvirus-directed responses (notably Epstein-Barr virus) and anti-flagellin antibodies up to 10 years prediagnosis in CD, particularly in individuals who later developed complicated or ileal disease. In contrast, responses to encapsulated bacteria (eg, Streptococcus pneumoniae, Haemophilus, Neisseria) progressively declined towards diagnosis. Pre-UC was characterised by combined antimicrobial, antiviral and autoantibody signatures, including antibodies against the MAP kinase-activating death domain protein. CONCLUSIONS: Large-scale serological profiling of archived prediagnostic samples identified disease-specific immune trajectories years before IBD onset, providing novel insights into disease pathogenesis in its prodromal phase.

ANTIGENS

Massively parallel characterization of adolescent idiopathic scoliosis risk variants.

Adolescent idiopathic scoliosis (AIS) is a common pediatric musculoskeletal disorder characterized by lateral spinal curvature, often leading to chronic pain and deformity. Although a significant genetic component to AIS is recognized, the functional impact of most associated genetic variants, particularly those in noncoding regions, remains largely unknown. Using massively parallel reporter assays, we characterize 1664 variant positions in linkage disequilibrium with 26 AIS lead variants identified by genome-wide association studies (GWASs) in chondrocytes, a major cell type implicated in AIS pathogenesis. Using a library of 7173 candidate regulatory sequences, we compare the 1664 reference alleles against 4708 alternate alleles in two human chondrocyte cell lines (TC28a2 and SW1353). Our analysis identifies 92 variants that exhibit significant differential regulatory activity between their reference and alternate alleles, 79 of which are predicted to disrupt transcription factor binding sites, often correlating with their observed regulatory effect. Notably, we validate rs9496392, a single-nucleotide variant near the ADGRG6 locus, which shows consistent differential regulatory activity in both cell lines. ADGRG6 is a key regulator of cartilage homeostasis, and its cartilage-specific knockout in mice results in a scoliosis-like phenotype. The AIS risk allele of rs9496392 (T) is predicted to strongly disrupt several TFBSs, including SP1. This study provides a foundational catalog of functional AIS-associated regulatory variants active in chondrocytes, offering crucial insights into the perturbed gene regulatory networks in AIS. These findings lay the groundwork for identifying biomarkers and potential therapeutic targets for this complex childhood disease.

Journal Article

A systematic strategy for identifying causal single nucleotide polymorphisms and their target genes on Juvenile arthritis risk haplotypes.

BACKGROUND: Although genome-wide association studies (GWAS) have identified multiple regions conferring genetic risk for juvenile idiopathic arthritis (JIA), we are still faced with the task of identifying the single nucleotide polymorphisms (SNPs) on the disease haplotypes that exert the biological effects that confer risk. Until we identify the risk-driving variants, identifying the genes influenced by these variants, and therefore translating genetic information to improved clinical care, will remain an insurmountable task. We used a function-based approach for identifying causal variant candidates and the target genes on JIA risk haplotypes. METHODS: We used a massively parallel reporter assay (MPRA) in myeloid K562 cells to query the effects of 5,226 SNPs in non-coding regions on JIA risk haplotypes for their ability to alter gene expression when compared to the common allele. The assay relies on 180 bp oligonucleotide reporters ("oligos") in which the allele of interest is flanked by its cognate genomic sequence. Barcodes were added randomly by PCR to each oligo to achieve > 20 barcodes per oligo to provide a quantitative read-out of gene expression for each allele. Assays were performed in both unstimulated K562 cells and cells stimulated overnight with interferon gamma (IFNg). As proof of concept, we then used CRISPRi to demonstrate the feasibility of identifying the genes regulated by enhancers harboring expression-altering SNPs. RESULTS: We identified 553 expression-altering SNPs in unstimulated K562 cells and an additional 490 in cells stimulated with IFNg. We further filtered the SNPs to identify those plausibly situated within functional chromatin, using open chromatin and H3K27ac ChIPseq peaks in unstimulated cells and open chromatin plus H3K4me1 in stimulated cells. These procedures yielded 42 unique SNPs (total = 84) for each set. Using CRISPRi, we demonstrated that enhancers harboring MPRA-screened variants in the TRAF1 and LNPEP/ERAP2 loci regulated multiple genes, suggesting complex influences of disease-driving variants. CONCLUSION: Using MPRA and CRISPRi, JIA risk haplotypes can be queried to identify plausible candidates for disease-driving variants. Once these candidate variants are identified, target genes can be identified using CRISPRi informed by the 3D chromatin structures that encompass the risk haplotypes.

Humans

Promoter identity shapes splicing outcomes and fidelity.

Gene expression is a complex process subject to regulation at multiple functionally interconnected levels. One prominent example is the crosstalk between transcription and splicing regulation. Past work has shown that transcription can influence splicing in multiple ways, but a systematic investigation of this complex interplay is lacking. Here we employ massively parallel reporter assays of large combinatorial promoter-splice site libraries to dissect how promoter identity and transcription dynamics affect alternative splicing in human cells. We find that promoter identity, rather than expression level, exerts strong and highly context-specific effects on cassette exon inclusion, exceeding the effect of pharmacological inhibitors of transcription initiation or elongation. Groups of exons display coordinated promoter-dependent splicing behavior, and we identified predictive sequence and structural features underlying this sensitivity. Promoter and gene architecture also shape isoform diversity by modulating cryptic splice site usage. These findings present promoters as central regulators of splicing outcomes and fidelity.

Humans

Massively parallel characterization and predictive modelling of neuronal regulatory variation.

Disease-associated variants reside frequently in noncoding cis-regulatory elements (CREs), yet their functional consequences remain poorly understood. We performed a large-scale lentiMPRA in human excitatory neurons, quantifying the impact of >46,000 naturally occurring variants across >27,000 candidate CREs near 524 disease-associated genes. These data improved regulatory variant effect predictions beyond state-of-the-art models. Significant allelic effects occurred at comparable rates across common, rare, and singleton variants, demonstrating that, within MPRA-measurable effects, population frequency carries limited information about per-variant regulatory impact. Variant effect detectability and magnitude were governed primarily by baseline activity of the enclosing regulatory element and local sequence context. Regulatory effects were distributed across numerous transcription factors rather than concentrated in master regulators, consistent with a combinatorial enhancer architecture. We establish a large-scale functional variant catalog and provide a complementary benchmark and resource for developing and evaluating models of noncoding regulatory variation.

Journal Article

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

Unequally Abundant Chromosomes and Unusual Collections of Transferred Sequences Characterize Mitochondrial Genomes of Gastrodia (Orchidaceae), One of the Largest Mycoheterotrophic Plant Genera.

The mystery of genomic alternations in heterotrophic plants is among the most intriguing in evolutionary biology. Compared to plastid genomes (plastomes) with parallel size reduction and gene loss, mitochondrial genome (mitogenome) variation in heterotrophic plants remains underexplored in many aspects. To further unravel the evolutionary outcomes of heterotrophy, we present a comparative mitogenomic study with 13 de novo assemblies of Gastrodia (Orchidaceae), one of the largest fully mycoheterotrophic plant genera, and its relatives. Analyzed Gastrodia mitogenomes range from 0.56 to 2.1 Mb, each consisting of numerous, unequally abundant chromosomes or contigs. Size variation might have evolved through chromosome rearrangements followed by stochastic loss of "dispensable" chromosomes, with deletion-biased mutations. The discovery of a hyper-abundant (∼15 times intragenomic average) chromosome in two assemblies represents the hitherto most extreme copy number variation in any mitogenomes, with similar architectures discovered in two metazoan lineages. Transferred sequence contents highlight asymmetric evolutionary consequences of heterotrophy: despite drastically reduced intracellular plastome transfers convergent across heterotrophic plants, their rarity of horizontally acquired sequences sharply contrasts parasitic plants, where massive transfers from their hosts prevail. Rates of sequence evolution are markedly elevated but not explained by copy number variation, extending prior findings of accelerated molecular evolution from parasitic to heterotrophic plants. Putative evolutionary scenarios for these mitogenomic convergence and divergence fit well with the common (e.g. plastome contraction) and specific (e.g. host identity) aspects of the two heterotrophic types. These idiosyncratic mycoheterotrophs expand known architectural variability of plant mitogenomes and provide mechanistic insights into their content and size variation.

Genome, Mitochondrial

Dosa: A method to covalently barcode proteins for high throughput biochemistry.

Deep mutational scanning couples a protein's activity to DNA sequencing for high throughput assessment of the effects of all single amino acid substitutions, but it largely uses indirect assays, like growth, as proxy for protein activity. Here, we covalently link variant proteins in vivo to an RNA barcode by fusing them to E. coli tRNA (m5U54) methyltransferase TrmA (E358Q), which forms a covalent bond with a tRNA stem-loop. Following cell lysis, variant proteins are separated in vitro according to their biochemical properties and identified by their barcodes. We use this method, Dosa, to analyze a large pool of FLAG epitope variants for binding to an anti-FLAG antibody, to profile the cleavage preferences of variants of enteropeptidase and human rhinovirus 3C protease, and to measure the solubility of several hundred Aβ(1-42) variants. This method should be amenable to numerous biochemical assays with proteins produced in E. coli or mammalian cells.

Protein display