PubMed HealthSearch

Biomedical subjects

Benjamin M Neale

Publications and source records attributed to Benjamin M Neale.

10 recordsLinked to original sources

Contribution of copy number variants to schizophrenia in East Asian populations.

Studies on schizophrenia-associated rare copy number variants (CNVs) have predominantly focused on people of European (EUR) ancestry. Here we present a rare CNV study of schizophrenia in East Asian (EAS) populations, comprising 20,903 cases and 23,258 controls. We observed a significantly elevated genome-wide rare CNV burden in EAS cases compared with controls. Cross-population comparisons showed largely consistent rare CNV effects on schizophrenia risk. In the EAS sample, we identified nine genome-wide-significant schizophrenia-associated rare CNV loci. Meta-analysis with EUR data yielded 14 significant loci, including 8 that reached genome-wide significance for the first time. Genes within these 14 loci were significantly less tolerant to loss-of-function variants than genes in other CNV loci. The new rare CNVs associated with schizophrenia in EAS populations showed higher carrier frequencies in EAS than in EUR populations (0.38% versus 0.0017%). Overall, this study underscores the importance of increasing population diversity to fully capture the genetic underpinnings of schizophrenia.

Journal Article

Mechanism of age-related accumulation of mtDNA mutations in human blood.

Accumulation of mutant mitochondrial DNA (mtDNA) heteroplasmy is among the strongest signatures of ageing1. Here we investigated the underlying mechanism by calling mtDNA sequence, mtDNA abundance and mtDNA heteroplasmic variants in human blood using whole-genome sequences from approximately 750,000 individuals. We observed that mtDNA single-nucleotide variants (mtSNVs) accumulate sharply at age 60 years, occur at low levels of heteroplasmy, exhibit little evidence of positive selection and are likely to be predominantly neutral. The mutational spectrum of mtSNVs does not reflect oxidative lesions, as is commonly invoked, but is more consistent with mtDNA replication errors. To understand why mtSNVs become detectable with age, we performed a genome-wide association study for heteroplasmic mtSNV burden, identifying germline variants near TERT, TCL1A and SMC4, all of which have been linked to clonal haematopoiesis (CH)2. Rare-variant analysis also showed that high mtSNV burden is associated with mutations in numerous CH driver genes. These genetic associations persisted even after exclusion of individuals with known CH driver mutations. Our results support a model in which 'cryptic' mtDNA mutations initially arise randomly as replication errors but are undetectable in bulk. They then become apparent only through age-related expansion of cellular clones in blood. We propose that the high copy number and mutation rate of mtDNA make it a sensitive blood-based marker of somatic mosaicism due to CH. Our work mechanistically unifies three prominent signatures of ageing: common germline variants in TERT, CH and observed accrual of mtDNA mutations.

Humans

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article

Systematic common and rare variant association testing in 392,030 whole genomes in All of Us.

Large-scale genome-wide association studies (GWAS) and rare variant association studies (RVAS) from population biobanks provide valuable resources for gene discovery in complex human traits. We present an analysis of the All of Us Research Program v8 release, which includes whole genome sequencing data and harmonized phenotypic information of 392,030 participants after quality control, enabling a unified investigation of rare and common variants across a spectrum of human traits and diseases. We build an extensive phenome- and genome-wide ("All by All") computational framework to perform GWAS and RVAS on 3,602 phenotypes and identify 49,863 approximately independent, high-quality single-variant and gene-level associations. Meta-analyses of All of Us and UK Biobank, with sample sizes as large as 786,871 participants, further enhance statistical power and find 193 pLoF gene-phenotype associations that are not significant in either cohort alone, including 22 associations not highlighted by previous studies. We also present a public interactive browser that integrates association results for common and rare variants to facilitate interpretation and rapid querying of summary statistics, along with supporting documentation, and a Featured Workspace in the All of Us Researcher Workbench. Our framework will apply to iterative data releases as All of Us grows, empowering researchers worldwide to uncover insights into the functional effects of genetic components on complex traits and diseases.

Journal Article

Genomic analyses implicate hormonal and metabolic dysregulation in polycystic ovary syndrome.

Polycystic ovary syndrome (PCOS) and its underlying features remain poorly understood. In this genetic study (n = 544,513), we expand the number of genetic loci from 16 to 29, and additionally identify 31 associated plasma proteins. Many risk-increasing loci were associated with later age at menopause, underscoring the reproductive longevity related to an increased oocyte number and/or availability across the lifespan. Hormonal regulation in the etiology of this condition, through metabolic and reproductive features, was emphasized. The proteomic analysis highlighted metabolic biology known to be related to PCOS. A polygenic risk score (PRS) was associated with adverse cardiometabolic outcomes, with differing relevance of testosterone and body mass index in women and men. Finally, while oligo-anovulation and anovulatory infertility are features of PCOS, we observed no impact of PCOS susceptibility on childlessness. We suggest that PCOS susceptibility confers balanced pleiotropic influences on fertility in women, and life-long adverse metabolic consequences in both sexes.

Humans

Multipopulation GWAS for venous thromboembolism identifies novel loci followed by experimental validation in zebrafish.

Venous thromboembolisms (VTEs) are a leading cause of morbidity and mortality. Although many genetic risk factors have been identified, a substantial portion of the heritability remains unexplained. In this study, we employed a genome-wide association study (GWAS) for VTE across 9 international cohorts of the Global Biobank Meta-Analysis Initiative to address this question, along with in vivo functional validation. In this multipopulation GWAS (VTE cases, 27 987; controls, 1 035 290), 38 genome-wide significant loci were identified, 4 of which were potentially novel. For each autosomal locus, we performed gene prioritization using 7 independent, yet converging, lines of evidence. Through prioritization, we identified genes associated with VTE through GWAS and/or functional studies (eg, F5, F11, VWF, STAB2, PLCG2, TC2N), functionally validated those that did not have evidence other than GWAS (TC2N, TSPAN15), and discovered 1 not previously associated with coagulation (RASIP1). We evaluated the function of 6 prioritized genes with strong genetic evidence, including F7 as a positive control, using laser-mediated endothelial injury to induce thrombosis in zebrafish after CRISPR/Cas9 knockdown. From this assay, we have supportive evidence for the role of RASIP1 and TC2N in the modification of human VTE and suggestive evidence for STAB2 and TSPAN15. This study expands on the currently identified genomic architecture of VTE through biobank-based, multipopulation GWASs, in silico candidate gene predictions, and in vivo functional follow-up of candidate genes.

Zebrafish

Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank.

The phenotypic impact of compound heterozygous (CH) variation has not been investigated at the population scale. We phased rare variants (MAF &#x223c;0.001%) in the UK Biobank (UKBB) exome-sequencing data to characterize recessive effects in 175,587 individuals across 311 common diseases. A total of 6.5% of individuals carry putatively damaging CH variants, 90% of which are only identifiable upon phasing rare variants (MAF&#xa0;<&#xa0;0.38%). We identify six recessive gene-trait associations (p&#xa0;<&#xa0;1.68&#xa0;&#xd7;&#xa0;10-7) after accounting for relatedness, polygenicity, nearby common variants, and rare variant burden. Of these, just one is discovered when considering homozygosity alone. Using longitudinal health records, we additionally identify and replicate a novel association between bi-allelic variation in ATP2C2 and an earlier age at onset of chronic obstructive pulmonary disease (COPD) (p&#xa0;<&#xa0;3.58&#xa0;&#xd7;&#xa0;10-8). Genetic phase contributes to disease risk for gene-trait pairs: ATP2C2-COPD (p&#xa0;= 0.000238), FLG-asthma (p&#xa0;= 0.00205), and USH2A-visual impairment (p&#xa0;= 0.0084). We demonstrate the power of phasing large-scale genetic cohorts to discover phenome-wide consequences of compound heterozygosity.

Humans

Exome-wide evidence of compound heterozygous effects across common phenotypes in the UK Biobank.

Exome-sequencing association studies have successfully linked rare protein-coding variation to risk of thousands of diseases. However, the relationship between rare deleterious compound heterozygous (CH) variation and their phenotypic impact has not been fully investigated. Here, we leverage advances in statistical phasing to accurately phase rare variants (MAF ~ 0.001%) in exome sequencing data from 175,587 UK Biobank (UKBB) participants, which we then systematically annotate to identify putatively deleterious CH coding variation. We show that 6.5% of individuals carry such damaging variants in the CH state, with 90% of variants occurring at MAF < 0.34%. Using a logistic mixed model framework, systematically accounting for relatedness, polygenic risk, nearby common variants, and rare variant burden, we investigate recessive effects in common complex diseases. We find six exome-wide significant () and 17 nominally significant () gene-trait associations. Among these, only four would have been identified without accounting for CH variation in the gene. We further incorporate age-at-diagnosis information from primary care electronic health records, to show that genetic phase influences lifetime risk of disease across 20 gene-trait combinations (FDR < 5%). Using a permutation approach, we find evidence for genetic phase contributing to disease susceptibility for a collection of gene-trait pairs, including FLG-asthma () and USH2A-visual impairment (). Taken together, we demonstrate the utility of phasing large-scale genetic sequencing cohorts for robust identification of the phenome-wide consequences of compound heterozygosity.

Preprint

Genetic structure correlates with ethnolinguistic diversity in eastern and southern Africa.

African populations are the most diverse in the world yet are sorely underrepresented in medical genetics research. Here, we examine the structure of African populations using genetic and comprehensive multi-generational ethnolinguistic data from the Neuropsychiatric Genetics of African Populations-Psychosis study (NeuroGAP-Psychosis) consisting of 900 individuals from Ethiopia, Kenya, South Africa, and Uganda. We find that self-reported language classifications meaningfully tag underlying genetic variation that would be missed with consideration of geography alone, highlighting the importance of culture in shaping genetic diversity. Leveraging our uniquely rich multi-generational ethnolinguistic metadata, we track language transmission through the pedigree, observing the disappearance of several languages in our cohort as well as notable shifts in frequency over three generations. We find suggestive evidence for the rate of language transmission in matrilineal groups having been higher than that for patrilineal ones. We highlight both the diversity of variation within Africa as well as how within-Africa variation can be informative for broader variant interpretation; many variants that are rare elsewhere are common in parts of Africa. The work presented here improves the understanding of the spectrum of genetic variation in African populations and highlights the enormous and complex genetic and ethnolinguistic diversity across Africa.

Africa, Southern