PubMed HealthSearch

SEARCH · PubMed Health

Results for “Genome integrity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

AEGIS: an annotation extraction and genomic integration resource.

MOTIVATION: Genome annotation files (GFF3/GTF) are the standard for storing genomic feature data, yet their flexibility often results in formatting inconsistencies that create bottlenecks for downstream bioinformatics analyses. A robust, unified framework is required to parse, standardise, and validate these files to ensure interoperability and facilitate complex comparative genomic tasks. RESULTS: We present AEGIS (Annotation Extraction and Genomic Integration Suite), a comprehensive toolkit designed to parse, correct, and standardise genome annotations. Beyond quality control, AEGIS provides advanced modules for flexible feature extraction (e.g., coding sequences, promoters) and comparative genomic analysis. Uniquely, it integrates multiple lines of evidence, including sequence homology, synteny, and coordinate-based lift-overs, to assess gene model correspondence and infer orthology. We demonstrate the utility of AEGIS by quantifying complex structural changes between Arabidopsis annotation versions and identifying high-confidence orthologues across diverse plant genomes. AVAILABILITY: AEGIS is implemented in Python. Source code and documentation are freely available under the GPL-3 license at https://github.com/Tomsbiolab/aegis and as a Docker container at https://hub.docker.com/r/tomsbiolab/aegis. The package is also available on PyPI (pip install aegis-bio).

Software

Quantile Tensor Regression for Integrative Genomic Analysis of Oesophageal Carcinoma.

Recent integrative genomic studies have increasingly exploited the tensor structure of multi-omics data to develop statistical methods that jointly model the relationship between clinical outcomes and multiple genomes. However, genomic measurements and clinical outcomes are frequently contaminated by outliers or heavy-tailed noise, necessitating robust tensor-based inference approaches. In this paper, we investigate the quantile tensor regression with an emphasis on the region selection problem. We introduce a novel estimator that integrates quantile regression for robustness with a nonconvex penalty to encourage sparsity in the tensor coefficient, thereby enabling the identification of localized genomic regions that significantly influence the clinical response. To solve the resulting optimization problem, we devise an effective algorithm tailored to the nonconvex objective and tensor architecture. We establish the asymptotic properties of the proposed nonconvex penalized estimator. Extensive simulations demonstrate the excellent finite-sample performance of the proposed estimator. We further illustrate the practical utility of the proposed estimator through an application to esophageal carcinoma data, providing empirical validation.

Humans

The integrated genome of murine leukemia virus.

The Southern gel filter transfer technique has been used to characterize the integrated genome of Moloney murine leukemia virus (M-MuLV) and the genomes of the endogenous viruses of the mouse. Study of 10 clones of rat cell independently infected by M-MuLV indicates a minimum of 15 integration sites into which the M-MuLV provirus can be inserted. No common integration site is observed among these clones. Clones productively infected by M-MuLV acquire multiple proviruses, whereas infected cells unable to produce virus contain only one M-MuLV provirus. Once established, the integrated genomes are stable for at least two years after initial infection. The use of M-MuLV probe allows detection of a spectrum of Eco RI-cleaved mouse DNA fragments containing endogenous MuLV genomes. DNAs of different inbred laboratory mouse strains yield similar patterns of provirus with each strain showing minor characteristic differences. In some instances, mouse cells infected by M-MuLV reveal additional proviruses beyond those seen in the uninfected cell. DNAs from three different M-MuLV-induced thymomas indicate, as in rat cells, multiple possible integration sites.

Animals

Clinical Variant Interpretation with the Integrative Genomics Viewer (IGV) for Molecular Pathologists.

The integrative genomics viewer (IGV) is a pivotal tool in clinical genomics, enabling the visualization and interpretation of complex sequencing data. Bringing clinical knowledge to bear with visual evaluation of sequencing results is the primary means by which molecular pathologists and other professionals assess and finalize cases. A variety of software tools can assist, but their relationship to the underlying data must be understood and applied systematically. This study includes essential background on next-generation sequencing (NGS) data file types (e.g., FASTQ, BAM, VCF) with a discussion of their format and purpose. We then describe features of IGV that derive nuances from these files. We utilize a series of curated practical cases based on clinical vignettes through which the reader will interact with clinical NGS sequencing data using the IGV software to review various types of clinically relevant variants relative to the human reference genome. These clinical vignettes have been curated to describe examples of some of the complexities of interpretation of genomic data, and how utilizing IGV as part of a routine workflow can provide additional interpretive information for variants beyond routine bioinformatic software algorithm variant calls. The visual inspection of genomic variants utilizing the tools within IGV can unmask subtle contextual cues (i.e., variant allele frequency, strand bias, tissue-specific context) that can influence the interpretation of genomic variants. Although this study focuses on using IGV for the detection and interpretation of somatic variants, the provided applications can be extrapolated for use in the germline setting, including analysis of complex variants and detection of mosaicism.

Humans

Systematic modular engineering of genome-integrated Escherichia coli MG1655 for high-level 2'-fucosyllactose production.

2'-Fucosyllactose (2'-FL), the most abundant human milk oligosaccharide (HMO), has attracted considerable interest for its prebiotic and immunomodulatory functions, with broad applications in infant nutrition. In this study, we report the development of a high-yield, genome-integrated 2'-FL-producing strain based on Escherichia coli MG1655 through systematic modular optimization. Starting from a single-copy BKHT strain (MGC06), we first optimized the copy number of the α-1,2-fucosyltransferase (α-1,2-FT) gene BKHT. Subsequently, the GDP-L-fucose supply was enhanced through coordinated genomic integration of the gene clusters cpsG-cpsB and gmd-fcl, while the multidrug efflux transporter gene mdfA was integrated to improve product export and strain robustness. BKHT copy number was then re-evaluated in the optimized background, with four copies yielding the highest production. The final engineered strain, harboring all genetic modifications stably integrated into the chromosome, produced 17.18 g/L 2'-FL in shake-flask culture. In fed-batch fermentation using a 5-L bioreactor, this strain achieved a titer of 154.12 g/L after 60 h, with a productivity of 2.57 g/L/h. Notably, throughout the entire fermentation process, no antibiotics or inducers were supplemented, underscoring the genetic stability and regulatory compliance of this plasmid-free system. To our knowledge, this represents the highest 2'-FL titer reported to date, positioning our engineered strain as a promising candidate for commercial 2'-FL production.

Escherichia coli

BRCA1 safeguards genome integrity by activating chromosome asynapsis checkpoint to eliminate recombination-defective oocytes.

In the meiotic prophase, programmed DNA double-strand breaks are repaired by meiotic recombination. Recombination-defective meiocytes are eliminated to preserve genome integrity in gametes. BRCA1 is a critical protein in somatic homologous recombination, but studies have suggested that BRCA1 is dispensable for meiotic recombination. Here we show that BRCA1 is essential for meiotic recombination. Interestingly, BRCA1 also has a function in eliminating recombination-defective oocytes. Brca1 knockout (KO) rescues the survival of Dmc1 KO oocytes far more efficiently than removing CHK2, a vital component of the DNA damage checkpoint in oocytes. Mechanistically, BRCA1 activates chromosome asynapsis checkpoint by promoting ATR activity at unsynapsed chromosome axes in Dmc1 KO oocytes. Moreover, Brca1 KO also rescues the survival of asynaptic Spo11 KO oocytes. Collectively, our study not only unveils an unappreciated role of chromosome asynapsis in eliminating recombination-defective oocytes but also reveals the dual functions of BRCA1 in safeguarding oocyte genome integrity.

Oocytes

Integrating genomics, multi-omics, CRISPR and speed breeding for stress-resilient vegetable legume improvement.

Vegetable legumes are nutritionally and ecologically important crops. However, their genetic improvement has not kept pace with the increasing challenges posed by climate change due to the polygenic nature of stress tolerance, narrow genetic diversity, and the persistent gap between molecular discoveries and field-level cultivar development. Although recent reviews have examined individual genomic tools or specific stress responses, a comprehensive synthesis integrating genomics-assisted breeding, multi-omics technologies, genome editing, and speed breeding within a unified crop improvement framework has been lacking. This review addresses that gap by critically evaluating how these complementary approaches can accelerate the development of stress-resilient vegetable legumes, including pea, common bean, cowpea, faba bean, cluster bean, yard-long bean, and hyacinth bean. This review synthesizes advances in QTL mapping, genome-wide association studies, transcriptomics, metabolomics, and CRISPR-based functional genomics that have identified key regulators and pathways underlying resistance to major biotic and abiotic stresses. Rather than considering these technologies independently, the review emphasizes their convergence into a systems-level breeding framework integrating genomic discovery, functional validation, predictive breeding, and accelerated generation advancement to improve breeding efficiency. Speed breeding, enabling up to seven to eight generations annually under optimized controlled-environment experimental conditions in cowpea, is discussed as a complementary strategy with genomic selection and genome editing. The review further identifies major translational bottlenecks, including transformation recalcitrance, limited genomic resources for underutilized vegetable legumes, inadequate multi-environment validation, and fragmented omics integration, and presents an integrated systems-breeding framework to bridge the gap between gene discovery and cultivar development.

Fabaceae

The SARS-CoV-2 Integrated Genomic Epidemiology Database (IGED): Linking viral genomes with patient-level metadata to advance statewide genomic surveillance in California.

In July 2021, the California Code of Regulations Title 17 required all laboratories performing SARS‑CoV‑2 whole genome sequencing (WGS) to report their sequencing results to the California Department of Public Health (CDPH). These viral genomic data and patient metadata were compiled into the Integrated Genomic Epidemiology Database (IGED). Linking anonymized viral sequences with patient‑level information enabled monitoring of infectiousness, pathogenicity, transmission dynamics, evolution, and vaccine evasion among emerging SARS‑CoV‑2 lineages. Laboratories performing SARS-CoV-2 WGS transmitted sequencing results to CDPH through Electronic Laboratory Reporting (ELR) and non-ELR pathways. CDPH applied uniform reporting requirements but allowed flexibility in specific data formats to accommodate diverse data systems. To preserve data quality and interoperability across heterogeneous sources, CDPH implemented standardization, validation, and deduplication protocols. Snowflake, a cloud‑based data storage and analytics platform, and Posit Connect, a cloud deployment and automation platform, supported the management, processing, and integration of data within the IGED. The IGED established links between SARS‑CoV‑2 WGS data and epidemiologic metadata for 801,418 sequences, representing 81.7% of all sequences reported in California. Lineages reported to the IGED showed strong concordance with lineage proportions in GISAID. Sequences reported to the IGED had average turnaround times longer than one month, and the majority of sequencing was performed in Southern California and Los Angeles. The IGED enhanced genomic surveillance through predictive modeling and monitoring concerning evolutionary trends such as recombination and saltations in persistent infections. Development of the IGED highlighted the need for standardized data requirements, sustained funding for sequencing, incentives for data submission, and interdisciplinary collaboration to build an effective genomic surveillance system. This framework for linking genomic and epidemiologic data has not only generated critical insights for SARS‑CoV‑2 but also provided the foundation for CDPH and other public health organizations to develop similar IGED‑like systems for other priority pathogens as genomic surveillance expands.

Journal Article

Genomically integrated cassettes swapping: bringing modularity to the strain level in Saccharomyces cerevisiae.

A large variety of synthetic biology toolkits for the introduction of multiple expression cassettes is available for Saccharomyces cerevisiae. Unfortunately, none of these tools is designed to allow the modification - exchange or removal - of the cassettes already integrated into the genome in a standardized way. The application of the modularity principle therefore ends to the steps preceding the final host engineering, making microbial cell factories construction stiff and strictly sequential. In this work, we describe a system that easily allows CRISPR-mediated swapping or removal of previously integrated cassettes, thus bringing the modularity to the strain level, enhancing the possibility of modifying existing strains with a reduced number of steps. In the system, each cassette is tagged with specific barcodes, which can be used as targets for CRISPR nucleases (Cas9 and Cas12a), allowing the excision of the construct from the genome and its substitution with another expression cassette or the restoration of the wild type locus in one single standardized step. The system has been applied to the previously developed Easy-MISE toolkit and tested by swapping fluorescent protein expression cassettes with an efficiency of ∼90% quantified by PCR and flow cytometry.

Saccharomyces cerevisiae

Pragmatic Phenotype-Electrophysiology-Genomics Integration in Pediatric Congenital Myasthenic Syndromes: Insights From 36 Patients in a Single-Center Study in China.

AIMS: To characterize the clinical, electrophysiological, and genetic spectrum of pediatric CMS and evaluate genotype-informed outcomes using an integrated phenotype-electrophysiology-genomics approach. METHODS: We retrospectively reviewed 36 pediatric CMS patients evaluated at a single center between 2015 and 2025. Clinical features, RNS, targeted NGS/WES variants, ventilator use, treatments, ACMG/AMP classifications, and MG-ADL outcomes were analyzed. RESULTS: Of 36 patients, 28 (77.8%) developed symptoms in the neonatal period or infancy. Biallelic variants involved 17 CMS genes; postsynaptic CMS was most common (55.6%, 20/36). COLQ and CHRNE were the most frequent genes (13.9%, 5/36 each), followed by CHAT (11.1%, 4/36). VUS were detected in 19 patients (52.8%, 19/36), including 8 with biallelic VUS supported by phenotype, neuromuscular transmission findings, treatment response, and follow-up. RNS showed a ≥ 10% decrement in 16/21 tested patients (76.2%). CHAT-CMS was associated with higher ventilator use (3/4 vs. 6/32; p = 0.041) and early mortality (3/4 vs. 1/32; p = 0.002). Median MG-ADL improved from 5 to 3 after genotype-informed therapy. CONCLUSION: Pediatric CMS shows marked genetic heterogeneity and frequent VUS-related uncertainty. Integrating phenotype, electrophysiology, and genomics supports diagnosis and mechanism-guided therapy. CHAT-CMS is high risk for early respiratory failure and mortality.

Humans

Integrative genomic and transcriptomic analyses identify key regulators of skin pigmentation in Larimichthys crocea.

The yellow body coloration of large yellow croaker (Larimichthys crocea) constitutes a crucial economic trait, yet its underlying genetic regulatory mechanisms remain poorly understood. This study systematically elucidated the molecular basis of body color variation by integrating genome resequencing and skin transcriptome analyses, combined with the contextual analysis of key pigmentation-related genes and phenotypic histological validation. 200 phenotyped individuals (including yellow-selected lines, F1 progeny, and normal control groups, all derived from a well-characterized aquaculture stock) identified 39 significantly associated SNPs (-log₁₀(P) ≥ 6), mapping to multiple candidate genes. These genes were significantly enriched in pathways related to pigment deposition (GO:0033059), melanosome organization (GO:0032438), melanogenesis, and tyrosine metabolism. Cross-developmental stage transcriptome analysis revealed 2395 differentially expressed genes (DEGs). Multi-omics integration identified eight overlapping candidate genes, including tyrp1, slc45a2, oca2, and dgat2, among which tyrp1 was prioritized for in-depth validation based on its core regulatory role in eumelanin synthesis, significant SNP association signal, and consistent downregulation in transcriptomic data. Experimental validation demonstrated that the g.895C > T mutation in exon 2 of tyrp1b was strongly significantly associated with the yellow phenotype: the frequency of mutant genotypes (TT/CT) reached 92.86%in the yellow-selected group, whereas the control group exclusively exhibited the wild-type genotype (CC). qPCR confirmed significantly downregulated tyrp1b expression in the skin of yellow individuals, consistent with the transcriptome trend. Histological and stereomicroscopic observations of skin tissues further validated the physiological basis of the yellow phenotype, revealing a significant reduction in melanophore number and abnormal melanosome morphology in yellow-phenotype individuals, accompanied by increased xanthophore density. These results suggest that tyrp1b mutation is strongly associated with the yellow phenotype. However, the presence of a wild-type CC individual in the yellow group indicates that this mutation is not strictly required for yellow coloration, suggesting that other genetic or environmental factors may also contribute to the phenotype, Additionally, downregulation of the carotenoid metabolism gene bco2 coupled with upregulation of xdh, together with the functional changes of slc45a2 and oca2, may synergistically promote xanthophore pigment deposition, contributing to the yellow phenotype. As melanin synthesis in large yellow croaker relies on the conserved tyrosinase pathway and transporter proteins, mutations in associated genes (tyrp1b, slc45a2, oca2) represent a primary underlying cause for the loss of melanin-based coloration and transition to a yellow phenotype in L. crocea. These findings provide key molecular targets and a theoretical foundation for molecular breeding of body color in this species, and also enrich the understanding of xanthism regulatory mechanisms in teleosts.

Animals

Integrated Genome Mining and Bioactivity-Guided Isolation of Antimicrobial Peptides from Bacillus amyloliquefaciens BS4.

Bacterial resistance remains a critical global health challenge, driving the continuous search for novel antimicrobial agents. Bacillus amyloliquefaciens is a recognized repository of bioactive metabolites; however, its full biosynthetic potential requires integrated genomic and experimental validation. This study characterized the antimicrobial profile of B. amyloliquefaciens BS4 through a hybrid pipeline. Genome sequencing and de novo assembly revealed a 3.9 Mb chromosome with a G + C content of 46.14%. Functional annotation identified 3,887 coding sequences, including pathways for siderophore biosynthesis and a complete bacilysin biosynthetic cluster. BGC analysis using antiSMASH v7.1.0 and BAGEL4 identified 18 biosynthetic gene clusters, while similarity network analysis via BiG-SCAPE highlighted unique singleton BGCs, indicating untapped biosynthetic diversity. Although in silico screening via Macrel predicted two putative cationic antimicrobial peptides (AMPs), bioactivity-guided purification utilizing sequential RP-HPLC, and de novo sequencing revealed a distinct set of four active peptides. Notably, three of these sequences were identified as fragments derived from the BclA exosporium protein family, highlighting the structural proteome as a non-canonical source of antimicrobials. The purified fractions exhibited activity against M. luteus and E. coli, while displaying no significant hemolytic activity or cytotoxicity, even above the MIC values. Molecular docking further supported the interaction of these candidates with bacterial targets. Overall, this hybrid strategy effectively uncovers the antimicrobial complexity of BS4, revealing 'cryptic' peptide candidates with therapeutic potential.

Bacillus amyloliquefaciens BS4

SUMO: a regulator of gene expression and genome integrity.

Post-translational modification with the ubiquitin-like SUMO protein is involved in the regulation of many cellular key processes. The SUMO system modulates signal transduction pathways, including cytokine, Wnt, growth factor and steroid hormone signalling. SUMO frequently restrains the activity of downstream transcription factors in these pathways presumably by facilitating the recruitment of corepressors or mediating the assembly of repressor complexes. Additionally, evidence is accumulating that SUMO controls pathways important for the surveillance of genome integrity. SUMO regulates the PML/p53 tumour suppressor network, a key determinant in the cellular response to DNA damage. Moreover, proteins that maintain genomic stability by functioning at the interface between DNA replication, recombination and repair processes undergo SUMOylation. We will discuss some key findings that exemplify the role of SUMO in transcriptional regulation and genome surveillance.

Animals

Integrative Genomic, Transcriptomic and Epigenomic Analysis Reveals cis-regulatory Contributions to High-altitude Adaptation in Tibetan Pigs.

The Qinghai-Tibet Plateau, characterized by its extreme environmental conditions, presents significant challenges to life, making it an ideal region for studying adaptation and evolution. Tibetan pigs, known for their high genetic diversity and exceptional adaptability to high altitudes, serve as excellent models for investigating high-altitude adaptation. While previous studies have extensively identified genetic determinants associated with high-altitude adaptation, the molecular mechanisms, particularly cis-regulatory patterns, remain poorly understood. Here, we conducted a selective sweep analysis using 484 genomes from Chinese and Western pig breeds across various altitudes, revealing 38.56 Mb of genomic regions under selection in Tibetan pigs. Enrichment analysis identified the lung as the primary functional tissue involved in high-altitude adaptation, supported by tissue-specific transcriptional and regulatory patterns observed between Tibetan and Meishan pigs (low altitude). By integrating genomic, RNA-seq, ATAC-seq, and H3K27ac HiChIP data, we constructed comprehensive enhancer-promoter regulatory maps of candidate genes and pinpointed promising genetic determinants associated with high-altitude adaptation, including SNPs in EPAS1, KLF13, SPRED1, and CFD. These loci were predicted to influence chromatin accessibility and the interactions of regulatory elements, with altered binding strength of relevant transcription factors. Further in vitro experiments confirmed that these loci function as allele-specific enhancers, modulating the expression of target genes. Our findings elucidate the regulatory basis of high-altitude adaptation in Tibetan pigs and provide valuable insights for exploring hypoxia-related diseases in livestock and humans.

Animals

Integrated genomic, transcriptomic, and metabolomic analyses of Chrysanthemum aromaticum provide insights into the volatile terpene biosynthesis.

Chrysanthemum aromaticum is renowned for its uniformly emitted strong and attractive scent, primarily attributed to volatile terpenes. Despite its commercial and horticultural significance, the molecular mechanisms underlying volatile terpene production in C. aromaticum remain largely unexplored. Here, we present the haplotype-resolved genome assembly of C. aromaticum, with a total size of 3.10 Gb, comprising nine anchored chromosomes with a contig N50 of 30.66 Mb and a scaffold N50 of 350.58 Mb. Phylogenetic analyses revealed a distant relationship between C. aromaticum and C. indicum, suggesting that C. aromaticum likely represents a distinct species rather than a variety of C. indicum. Through integrated genomic, transcriptomic, metabolomic, and biochemical analyses, we identified seven TPS involved in monoterpene biosynthesis and six TPS for sesquiterpene biosynthesis. Notably, comparative genomic analysis revealed a gene cluster for α-bisabolol biosynthesis in C. aromaticum, which has specifically expanded in Chrysanthemum species through tandem gene duplications, contributing to the elevated accumulation of α-bisabolol in the leaves of C. aromaticum. Our study provides important insights into the biosynthesis of volatile terpenes, highlighting the genetic basis for C. aromaticum's unique aromatic profile.

Chrysanthemum

Integrative Genomic and Transcriptomic Insights into High-Altitude Adaptation in Changthangi Goats.

The Changthangi goat, native to the high-altitude Ladakh Plateau in northern India, thrives in oxygen-deficient environments above 4,000 m. This study investigated the genetic basis of high-altitude adaptation in Changthangi goats by integrating comparative genomics and transcriptomics, using the tropical lowland Jamunapari goat as a comparative model. Whole-genome sequence data from 15 individuals per breed were analyzed using complementary selection sweep metrics, including nucleotide diversity, Tajima's D, iHS, CLR, XP-EHH, and FST. These analyses identified candidate genomic regions under strong selective pressure, encompassing genes involved in hypoxia sensing (HIF-1α, HIF-2α/EPAS1, EGLN1), angiogenesis (VEGFA, AGGF1, ZEB1), cardiovascular regulation (PRKCB, ESR1, RYR2), mitochondrial and energy metabolism (ACADSB, ACSS3, ACSL1), cellular stress tolerance (BCL2, ATM), and thermogenesis (UCP1, FGF21). Unlike previous caprine studies that primarily infer hypoxia adaptation from genomic signals alone, our study integrates cardiac transcriptomics to demonstrate that genomic selection in Changthangi goats is accompanied by coordinated transcriptional remodeling across interconnected physiological systems in a physiologically relevant tissue. Comparative cardiac transcriptomic profiling revealed concordant expression divergence in genes associated with oxygen transport, vascular remodeling, mitochondrial function, substrate utilization, redox balance, and genome maintenance. This integrative multi-omics framework provides a mechanistic view of caprine high-altitude adaptation and highlights the value of combining genomic selection analyses with tissue-specific transcriptional profiling to resolve complex adaptive traits.

Animals

Integrating Genomic and Nongenomic Data to Stratify the Risk of Contralateral Breast Cancer After Radiation Therapy.

PURPOSE: Women treated with radiation therapy (RT) for breast cancer have an increased risk of developing radiation-associated contralateral breast cancer (CBC). Predicting CBC events is challenging because of the complex interplay of genomic, treatment, personal, and clinical factors. This study investigated computational methods that integrate genome-wide single-nucleotide polymorphisms and nongenomic data to develop a risk stratification model for developing CBC in women treated with RT for their first primary breast cancer. METHODS AND MATERIALS: This study used a subset of the population-based Women's Environmental Cancer and Radiation Epidemiology study that included 633 CBC cases and 1253 individually matched unilateral breast cancer controls who were treated with RT and had single-nucleotide polymorphism data available from a genome-wide association study. The study population was split into training, validation, and test sets for rigorous modeling and validation. Three data integration methods were compared in terms of their ability to stratify CBC risk: (1) naive integration; (2) sequential integration; and (3) sequential iterative integration. A biological analysis of the final model was performed using gene set enrichment analysis and protein-protein interaction analysis with gene annotation information informed by the model. RESULTS: The best-performing integration method was the sequential iterative integration equipped with the mixed-effect random forest algorithm. This approach achieved an area under the curve of 0.64 to stratify CBC risk in the test set, representing moderate predictive power. Calibration analysis showed good agreement between the lowest and highest risk bins stratified using sorted predicted values in the test set, resulting in an odds ratio of 3.27 for both predicted and observed CBC occurrence. Gene set enrichment analysis and protein-protein interaction analysis revealed that genes with high importance scores were associated with pathways relevant to lipid and fatty acid metabolism as well as breast cancer sensitivity to tamoxifen. CONCLUSIONS: The mixed-effect random forest approach demonstrated the potential for integrating high-dimensional genomic and low-dimensional nongenomic data to stratify CBC risk.

Humans

Targeted genomic integration and rearrangement using prime assembly.

Although therapeutic genome editing holds great potential to remedy diverse inherited and acquired disorders, targeted installation of medium-to-large genomic modifications in therapeutically relevant cells remains challenging1. Here we develop prime assembly, an approach that permits DNA sequence assembly and integration in human cells leveraging CRISPR-targeted dual flap synthesis. This method enables RNA-programmable site-specific integration of single or double-stranded DNA fragments. Unlike homology-directed repair, prime assembly is similarly active in dividing and non-dividing cells. We applied prime assembly to perform targeted exon recoding, transgene integration and megabase-scale rearrangements, including at therapeutically relevant loci in primary human cells. Prime assembly expands the capabilities of genome engineering by enabling the targeted integration of medium to large-sized DNA sequences without relying on double-stranded DNA donors, nuclease-driven double-strand breaks or cell cycle progression.

Journal Article