PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “biobanks”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Integrating Biobanking Into Conservation Practice: The Development and Impact of the EAZA Biobank.

Zoological biobanks are becoming essential tools in conservation, offering a means to preserve genetic material and support in situ population management amid accelerating biodiversity loss. With rapid advances in genomics, cryopreservation, and assisted reproduction technologies, biobanks enable a proactive approach to providing insurance against genetic erosion and facilitating future research, supplementation, and genetic rescue. However, to be effective, zoological biobanks must be purposefully designed, strategically integrated into conservation frameworks such as the Convention on Biological Diversity (CBD) Kunming-Montreal Global Biodiversity Framework (KMGBF), and regularly evaluated for coverage and impact. Using the EAZA Biobank as an example, we outline the structure, development, and collaborative foundations that have enabled its rapid growth, built on community support and conservation impact. Leveraging EAZA's institutional network and data-sharing platforms such as ZIMS, the Biobank employs a decentralized, four-hub model of zoological institutions storing samples. A gap analysis, integrating threat status, breeding programs, genomic data repositories, and phylogenetic diversity, highlights current sampling strengths and deficiencies and guides future collection priorities. The integration of specimen-specific genomic data and the EAZA Biobank Cryonetwork of institutions with expertise in storing and generating gametes and cell lines will expand the Biobank's role in population management and conservation. Zoological biobanks must now evolve alongside advances in biotechnology and genomics. Sample collection strategies should serve conservation needs and anticipate future applications in genomics, cryobiology, and conservation medicine, linking biospecimens with the wealth of data generated from them. This approach should be scalable beyond EAZA, forming the foundation of a global standardized biobanking framework. Ultimately, zoological biobanks are not merely repositories of the past-they are essential infrastructures shaping the future potential of species conservation.

EAZA↗

Commercial biobanks and genetic research: ethical and legal issues.

Human biological material is recognized as an important tool in research, and the demand for collections that combine samples and data is increasing. For-profit companies have assumed a leading role in assembling and managing these collections. The emergence of commercial biobanks has raised significant ethical and legal issues. The growing awareness of the importance of human biological material in research has been accompanied by a growing awareness of the deficiencies of existing archives of tissue. Commercial biobanks are attempting to position themselves as a, if not the, solution to problems that include a lack of public trust in researchers and lack of financial resources to support the prospective creation of collections that meet the highest scientific and ethical standards in the non-profit sector. Broad social and policy questions surrounding the operation of commercial biobanks have been raised however. International documents, in particular, suggest discomfort with the idea of gain from the mere transfer or exchange of human genetic material and information. Commercial involvement in the development of useful products from tissue is generally not condemned, so long as there is attention to scientific and social norms. Views on the acceptability of commercial biobanks vary. Specific issues that arise when commercial biobanks are permitted--in the areas of consent, recruitment, confidentiality, and accountability--are also relevant to the operation of public and private, non-profit biobanks. Although many uncertainties remain, consensus seems to be forming on a number of issues. For example, there appears to be agreement that blanket consent to future unspecified research uses, with no conditions, is unacceptable. Indeed, many of the leading commercial biobanks have been attentive to concerns about consent, recruitment, and confidentiality. Unfortunately, the binding nature of assurances in these areas is unclear, especially given the risk of insolvency. Hence, accountability may be the most important area of concern in relation to commercial biobanks. A few countries have enacted general legislation providing for comprehensive regulation of biobanks, for example, through licensure. Efforts to achieve harmonization of standards at the international level, and cautions against an approach that focuses on biobanking for genetic research alone, are to be applauded.

Biological Specimen Banks↗

Benchmarking large language models for extracting biobank-derived insights into health and disease.

Biobank-scale datasets such as the UK Biobank have become foundational resources for advancing biomedical discovery. Yet the complexity and heterogeneity of these resources, spanning genomics, imaging, clinical records, and metadata, pose substantial barriers to access and interpretation. Large Language Models (LLMs) offer a promising avenue for making such datasets more navigable through natural language interfaces. However, the extent to which current general-purpose LLMs can retrieve and synthesize biobank-specific insights has not yet been systematically evaluated. In this study, we present a reproducible, multi-metric evaluation framework to benchmark the capabilities of leading LLMs. We evaluated six leading large language models: Gemini 3 Pro, Claude Opus 4.5, Claude Sonnet 4.5, GPT-5.2, Mistral Large 2, and DeepSeek V3, on four benchmark tasks designed to assess biobank-related knowledge retrieval. We evaluate model performance across six dimensions (semantic accuracy, factual correctness, domain knowledge, reasoning quality, response depth, and biobank specificity) and assessed output consistency using curated UK Biobank references and a robust random baseline. All models outperformed the baseline by 2&#xd7; to 3&#xd7;&#x2009;, with strong statistical separation (p&#x2009;<&#x2009;0.001), confirming meaningful biobank-specific knowledge retrieval. Gemini 3 Pro achieved the highest overall accuracy across tasks such as keyword synthesis, institution recognition, and topic inference, while Claude Sonnet 4.5 demonstrated the most uniform performance across evaluation dimensions. Our benchmark provides a rigorous framework for evaluating LLMs in biomedical settings. Using the UK Biobank as a real-world testbed, we highlight both the capabilities and limitations of current models, measuring their capacity to recall structured biomedical knowledge consistent with authoritative biobank metadata.

Large Language Models↗

Evaluation of biobank constitution and use: multicentre analysis in France and propositions for formalising the activities of research ethics committees.

Biobanks are collections of biological material and related files gathered and stored for clinical or research purposes. Here, we investigated the questions raised during the evaluation of biobanks by biomedical Research Ethics Committees (RECs), particularly in the context of genetic research. We sent a questionnaire to all RECs in France to survey their concerns and the ethical criteria used when evaluating research involving the storage of biological samples. Most of the RECs think that they should be consulted to evaluate the constitution of biobanks. The proportion of RECs of this opinion depended on whether the biobank is being constituted in the absence of an associated research project (initially created for clinical purposes or for undefined research) (14/28), whether the biobank is being constituted for research use (21/28) or whether an existing research biobank is being re-used (19/28). Views diverged concerning the way ethics principles are applied, showing that REC evaluations of biobanks might be formalised at each of the following steps: constitution, use and re-use. In this paper, we suggest concrete elements that could be integrated into the application of the new French law concerning the protection of the human beings participating in research as well as into international recommendations.

Biomedical Research↗

An empirical survey on biobanking of human genetic material and data in six EU countries.

Biobanks correspond to different situations: research and technological development, medical diagnosis or therapeutic activities. Their status is not clearly defined. We aimed to investigate human biobanking in Europe, particularly in relation to organisational, economic and ethical issues in various national contexts. Data from a survey in six EU countries (France, Germany, the Netherlands, Portugal, Spain and the UK) were collected as part of a European Research Project examining human and non-human biobanking (EUROGENBANK, coordinated by Professor JC Galloux). A total of 147 institutions concerned with biobanking of human samples and data were investigated by questionnaires and interviews. Most institutions surveyed belong to the public or private non-profit-making sectors, which have a key role in biobanking. This activity is increasing in all countries because few samples are discarded and genetic research is proliferating. Collections vary in size, many being small and only a few very large. Their purpose is often research, or research and healthcare, mostly in the context of disease studies. A specific budget is very rarely allocated to biobanking and costs are not often evaluated. Samples are usually provided free of charge and gifts and exchanges are the common rule. Good practice guidelines are generally followed and quality controls are performed but quality procedures are not always clearly explained. Associated data are usually computerised (identified or identifiable samples). Biobankers generally favour centralisation of data rather than of samples. Legal and ethical harmonisation within Europe is considered likely to facilitate international collaboration. We propose a series of recommendations and suggestions arising from the EUROGENBANK project.

Biological Specimen Banks↗

A biobank management model applicable to biomedical research.

BACKGROUND: The work of Research Ethics Boards (REBs), especially when involving genetics research and biobanks, has become more challenging with the growth of biotechnology and biomedical research. Some REBs have even rejected research projects where the use of a biobank with coded samples was an integral part of the study, the greatest fear being the lack of participant protection and uncontrolled use of biological samples or related genetic data. The risks of discrimination and stigmatization are a recurrent issue. In light of the increasing interest in biomedical research and the resulting benefits to the health of participants, it is imperative that practical solutions be found to the problems associated with the management of biobanks: namely, protecting the integrity of the research participants, as well as guaranteeing the security and confidentiality of the participant's information. METHODS: We aimed to devise a practical and efficient model for the management of biobanks in biomedical research where a medical archivist plays the pivotal role as a data-protection officer. The model had to reduce the burden placed on REBs responsible for the evaluation of genetics projects and, at the same time, maximize the protection of research participants. RESULTS: The proposed model includes the following: 1) a means of protecting the information in biobanks, 2) offers ways to provide follow-up information requested about the participants, 3) protects the participant's confidentiality and 4) adequately deals with the ethical issues at stake in biobanking. CONCLUSION: Until a governmental governance body is established in Quebec to guarantee the protection of research participants and establish harmonized guidelines for the management of biobanks in medical research, it is definitely up to REBs to find solutions that the present lack of guidelines poses. The model presented in this article offers a practical solution on a day-to-day basis for REBs, as well as researchers by promoting an archivist to a pivotal role in the process. It assures protection of all participants who altruistically donate their samples to generate and improve knowledge for better diagnosis and medical treatment.

Biological Specimen Banks↗

Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.

OBJECTIVE: To quantify and improve the performance of standard rheumatoid arthritis (RA) algorithms in a biobank setting. METHODS: This retrospective cohort study within the Mayo Clinic (MC) Biobank and MC Tapestry Study identified RA cases by presence of at least two RA codes OR positive anti-cyclic citrullinated peptide antibodies (CCP) plus disease-modifying anti-rheumatic drug (DMARD) prescription as of 7/18/2022. Rheumatology physicians manually verified all RA cases using RA criteria and/or rheumatology physician diagnosis plus DMARD use. All other biobank participants served as non-RA controls. We defined seropositivity as rheumatoid factor and/or anti-CCP positivity. We assessed rules-based and Electronic Medical Records and Genomics (eMERGE) RA algorithms using positive predictive value (PPV). Finally, we developed a novel RA algorithm using a LASSO-based machine learning approach with five-fold cross validation. RESULTS: We identified 1,316 confirmed RA cases (968 MC Biobank, 348 Tapestry, 70 % seropositive) and 82,123 non-RA controls (mean age 65, 61 % female). The PPV of 3 RA codes was 43 %, codes plus DMARD was 54 %, and codes plus DMARD plus seropositivity was 85 %. The PPV of eMERGE was 77 %. Available in the MC Biobank, self-reported RA (PPV 10 %) only minimally improved algorithm performance (PPV from 83 % to 85 %), whereas family history of RA (PPV 3 %) worsened performance. At 90 % PPV, the novel RA algorithm incorporating key variables such as anti-CCP and DMARD use increased sensitivity by 4-11 % compared to eMERGE. CONCLUSION: Rules-based and eMERGE RA algorithms had worse performance in biobank than administrative settings. Our novel RA algorithm outperformed these standard algorithms.

Humans↗

Using Large Genomic Biobanks to Generate Insights into Genetic Kidney Disease.

Chronic kidney disease (CKD) affects approximately 9% of the global population, leading to increased risks of end-stage kidney disease (ESKD), cardiovascular disease (CVD), and mortality. Patients with CKD are a huge burden on health care resources globally. CKD is a complex condition influenced by a combination of genetic, environmental, and traditional risk factors. Family studies have suggested heritability rates for CKD ranging from 30% to 75%, and large genomic biobank studies have proven essential in identifying genes with substantial effects on CKD risk and in capturing cumulative genetic risk through polygenic risk scores. These biobanks are crucial for discovering new genes associated with kidney health and disease, and their growing size enhances the power to detect novel genetic associations. Integrating multi-omics technologies such as transcriptomics, metabolomics, and proteomics further enriches our understanding of CKD, while advanced computational tools continue to expand our insights into genetic data. Polygenic risk scores, derived from hundreds of genetic variants with small effect sizes, can help identify individuals at high risk of CKD. Genomic biobanks offer valuable opportunities for early identification and personalized treatment of monogenic kidney disorders, such as autosomal dominant polycystic kidney disease and Alport syndrome. These biobanks help fill knowledge gaps, particularly in individuals with milder or asymptomatic presentations who are often underrepresented in traditional studies. Expanding genomic biobank efforts globally, especially in diverse populations, is vital to enhancing our understanding of the genetic underpinnings of kidney disease. This review highlights the significant contributions of genomic biobanks to advancing our comprehension of the genetics of CKD.

Humans↗

Underrepresented voices in a Colorado Biobank: Perspectives from focus groups on motivations, return of results, and data sharing.

Most participants in large cohorts, such as biobanks, are of European descent. This lack of representation has been an ongoing challenge in genomic research. Understanding the perspectives on genomics research and participation in biobanks of historically underrepresented populations could provide insight into ways to better engage with these groups. We conducted a series of virtual and in-person focus groups with individuals who self-identified as American Indian or Alaska Native (AI/AN), African American/Black (AA/B), or Hispanic/Latino (H/L) and who were enrolled in the Colorado Center for Personalized Medicine (CCPM) biobank. The focus group discussions were centered on participant experiences, including but not limited to their motivations, return of results, and data sharing. There was a total of 23 participants across the six focus groups. The majority of participants identified as AI/AN (60.9%), followed by H/L (39.1%), and AA/B (21.7%); many participants identified with multiple race/ethnicities. The motivations for participating in the biobank included the potential to advance science and health, the potential for return of results, to learn more about one's ancestry, and a few indicated that they were interested in helping the biobank be more representative of all populations. Notably, many expressed positive feedback of the focus groups and felt that their views were valued, illustrating the importance of community-centered work. Our findings can be used to guide recruitment and engagement of biobank participants, especially from diverse backgrounds, contributing to enhanced partnerships advancing knowledge and healthcare.

biobank↗

The Biobank Rare Variant consortium powers the discovery of rare genetic associations through global collaboration.

Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.

Journal Article↗

Securing our genetic health: engendering trust in UK Biobank.

The recent development of genetic databases, or 'biobanks', in a number of countries reflects scientists' and policy makers' beliefs in the future health benefits to be derived from genetics research. In Britain, however, a proposal for a genetic database, UK Biobank, has been the focus of a number of concerns. Establishing consent and legitimacy for any controversial biomedical research involving the participation of human subjects is difficult; it is however, acute for UK Biobank given the scale of the project and the criticisms levelled at it. Analysing recently published documents pertaining to UK Biobank, this article examines how consent for the project has been discursively framed and how this is reflected in its governance. It is argued that the problem of organising consent has been framed narrowly in terms of adherence to a well-established repertoire of institutional mechanisms which serves to limit debate on the substantive issues at stake. There is little evidence of reflection on the adequacy of such mechanisms for dealing with the unique challenges posed by UK Biobank, including achieving the confidence and participation of a population with diverse perspectives on genetic research. It is concluded that a restricted public discourse about UK Biobank may contribute to a decline in confidence in regulatory systems governing biotechnology and science more generally.

Aged↗

CLINICAL AND COGNITIVE PHENOTYPING OF COPY NUMBER VARIANTS ASSOCIATED WITH NEURODEVELOPMENTAL DISORDERS FROM A MULTI-ANCESTRY BIOBANK.

Clinical biobanks with electronic health records (EHRs) linked to genotype data continue to expand yielding an opportunity to further characterize disease-relevant genomic risk factors, yet few recall-by-genotype studies from biobanks have been published to date. For example, copy number variants (CNVs) that significantly increase risk for multiple neurodevelopmental disorders (NDDs) and negatively affect neurocognition, may present in up to 2% of population cohorts, with public health implications for ascertaining NDD CNV carriers. From BioMe, a multi-ancestry biobank derived from the Mount Sinai healthcare system (New York, NY), 892 adult participants were recontacted for deep phenotyping, including 335 NDD CNV carriers as well as comparators, 217 individuals with schizophrenia and 340 controls. Clinical and cognitive assessments were administered to each participant. There was no disclosure of genetic information. Eight percent of recontacted biobank participants completed the study (30 NDD CNV carriers across 15 unique loci, 20 schizophrenia and 23 controls). The study sample had a mean age of 48.8 (10.2) years, was 66% female and of diverse ancestry, 36% African, 34% Hispanic, and 26% European. Overall, 70% of 30 NDD-CNV carriers harbored at least one neuropsychiatric or developmental phenotype, including 40% with mood or anxiety disorders. Further, 22 NDD CNV carriers were significantly impaired compared to controls on digit span backwards (Beta=-1.76, FDR=0.04) and digit span sequencing (Beta=-2.01, FDR=0.04), but higher performing than schizophrenia on verbal learning (Beta=4.5, FDR=0.05). Thirty NDD CNV carriers were successfully recruited from a multi-ancestry biobank, as well as healthy controls and low-functioning individuals with schizophrenia. Deep phenotyping corroborated past reports, while also identifying discordance with EHRs. Future recall-by-genotype studies may further benchmark the study design and elucidate feasibility.

Biobank↗

The in vivo effects of the Pro12Ala PPARgamma2 polymorphism on adipose tissue NEFA metabolism: the first use of the Oxford Biobank.

AIMS/HYPOTHESIS: To investigate the phenotypic effects of common polymorphisms on adipose tissue metabolism and cardiovascular risk factors, we set out to establish a biobank with the unique feature of allowing a prospective recruit-by-genotype approach. The first use of this biobank investigates the effects of the peroxisome proliferator-activated receptor (PPAR) Pro12Ala polymorphism on integrative tissue-specific physiology. We hypothesised that Ala12 allele carriers demonstrate greater adipose tissue metabolic flexibility and insulin sensitivity. MATERIALS AND METHODS: From a comprehensive population register, subjects were recruited into a biobank, which was genotyped for the Pro12Ala polymorphism. Twelve healthy male Ala12 carriers and 12 matched Pro12 homozygotes underwent detailed physiological phenotyping using stable isotope techniques, and measurements of blood flow and arteriovenous differences in adipose tissue and muscle in response to a mixed meal containing [1,1,1-(13)C]tripalmitin. RESULTS: Of 6,148 invited subjects, 1,072 were suitable for inclusion in the biobank. Among Pro12 homozygotes, insulin sensitivity correlated with HDL-cholesterol concentrations, and inversely correlated with blood pressure, apolipoprotein B, triglyceride and total cholesterol concentrations. Ala12 carriers showed no such correlations. In the meal study, Ala12 carriers had lower plasma NEFA concentrations, higher adipose tissue and muscle blood flow, and greater insulin-mediated postprandial hormone-sensitive lipase suppression along with greater insulin sensitivity than Pro12 homozygotes. CONCLUSIONS/INTERPRETATION: This study shows that a recruit-by-genotype approach is feasible and describes the biobank's first application, providing tissue-specific physiological findings consistent with the epidemiological observation that the PPAR Ala12 allele protects against the development of type 2 diabetes.

Adipose Tissue↗

The consent problem within DNA biobanks.

Large prospective biobanks are being established containing DNA, lifestyle and health information in order to study the relationship between diseases, genes and environment. Informed consent is a central component of research ethics protection. Disclosure of information about the research is an essential element of seeking informed consent. Within biobanks, it is not possible at recruitment to describe in detail the information that will subsequently be collected because people will not know which disease they will develop. It will also be difficult to describe the specific research that will be performed using the biobank, other than to stipulate categories of research or diseases that are not included. Potential subjects can only be given information about the sorts of research that will be performed and by whom. Organisations responsible for biobanks usually argue that this disclosure of information is adequate when seeking informed consent, especially if coupled with a right to withdraw, as it would not be feasible or it would be too expensive to seek consent renewal on a regular basis. However, there are concerns about this 'blanket consent' approach'. Consent waivers have also been proposed in which research subjects entrust their consent with an independent third party to decide whether subsequent research using the biobank is consistent with the original consent provided by the subject.

Adult↗

The importance of family-based sampling for biobanks.

Biobanks aim to improve our understanding of health and disease by collecting and analysing diverse biological and phenotypic information in large samples. So far, biobanks have largely pursued a population-based sampling strategy, where the individual is the unit of sampling, and familial relatedness occurs sporadically and by chance. This strategy has been remarkably efficient and successful, leading to thousands of scientific discoveries across multiple research domains, and plans for the next wave of biobanks are underway. In this Perspective, we discuss the strengths and limitations of a complementary sampling strategy for future biobanks based on oversampling of close genetic relatives. Such family-based samples facilitate research that clarifies causal relationships between putative risk factors and outcomes, particularly in estimates of genetic effects, because they enable analyses that reduce or eliminate confounding due to familial and demographic factors. Family-based biobank samples would also shed new light on fundamental questions across multiple fields that are often difficult to explore in population-based samples. Despite the potential for higher costs and greater analytical complexity, the many advantages of family-based samples should often outweigh their potential challenges.

Humans↗

The IgA nephropathy Biobank. An important starting point for the genetic dissection of a complex trait.

BACKGROUND: IgA nephropathy (IgAN) or Berger's disease, is the most common glomerulonephritis in the world diagnosed in renal biopsied patients. The involvement of genetic factors in the pathogenesis of the IgAN is evidenced by ethnic and geographic variations in prevalence, familial clustering in isolated populations, familial aggregation and by the identification of a genetic linkage to locus IGAN1 mapped on 6q22-23. This study seems to imply a single major locus, but the hypothesis of multiple interacting loci or genetic heterogeneity cannot be ruled out. The organization of a multi-centre Biobank for the collection of biological samples and clinical data from IgAN patients and relatives is an important starting point for the identification of the disease susceptibility genes. DESCRIPTION: The IgAN Consortium organized a Biobank, recruiting IgAN patients and relatives following a common protocol. A website was constructed to allow scientific information to be shared between partners and to divulge obtained data (URL: http://www.igan.net). The electronic database, the core of the website includes data concerning the subjects enrolled. A search page gives open access to the database and allows groups of patients to be selected according to their clinical characteristics. DNA samples of IgAN patients and relatives belonging to 72 multiplex extended pedigrees were collected. Moreover, 159 trios (sons/daughters affected and healthy parents), 1068 patients with biopsy-proven IgAN and 1040 healthy subjects were included in the IgAN Consortium Biobank. Some valuable and statistically productive genetic studies have been launched within the 5th Framework Programme 1998-2002 of the European project No. QLG1-2000-00464 and preliminary data have been published in "Technology Marketplace" website: http://www.cordis.lu/marketplace. CONCLUSION: The first world IgAN Biobank with a readily accessible database has been constituted. The knowledge gained from the study of Mendelian diseases has shown that the genetic dissection of a complex trait is more powerful when combined linkage-based, association-based, and sequence-based approaches are performed. This Biobank continuously expanded contains a sample size of adequately matched IgAN patients and healthy subjects, extended multiplex pedigrees, parent-child trios, thus permitting the combined genetic approaches with collaborative studies.

Databases, Nucleic Acid↗

The phenotype-genotype reference map: Improving biobank data science through replication.

Population-scale biobanks linked to electronic health record data provide vast opportunities to extend our knowledge of human genetics and discover new phenotype-genotype associations. Given their dense phenotype data, biobanks can also facilitate replication studies on a phenome-wide scale. Here, we introduce the phenotype-genotype reference map (PGRM), a set of 5,879 genetic associations from 523 GWAS publications that can be used for high-throughput replication experiments. PGRM phenotypes are standardized as phecodes, ensuring interoperability between biobanks. We applied the PGRM to five ancestry-specific cohorts from four independent biobanks and found evidence of robust replications across a wide array of phenotypes. We show how the PGRM can be used to detect data corruption and to empirically assess parameters for phenome-wide studies. Finally, we use the PGRM to explore factors associated with replicability of GWAS results.

Humans↗