PubMed Health⌕ Search

Biomedical subjects

Frank Dudbridge

Publications and source records attributed to Frank Dudbridge.

17 recordsLinked to original sources

Proteome-wide Mendelian randomisation of lung function to identify potential therapeutic targets for respiratory disease.

BACKGROUND: Despite multiple clinical trials, disease-modifying treatments for COPD are currently limited. Since many drugs target proteins, identifying causality between proteins and lung function informs understanding of COPD pathophysiology and may suggest novel targets. We used Mendelian randomisation (MR) to prioritise proteins as potentially causal for imparied lung function. For prioritised proteins, we explored their potential suitability as drug targets by predicting their effects on a range of clinical outcomes. METHODS: We used genome-wide association study (GWAS) data on 2923 proteins (n=48&#x2009;195, UK Biobank) to identify single genetic variants (protein quantitative trait loci (cis-pQTLs)) associated with protein levels (p&#x2264;5&#xd7;10-9, variant &#x2264;100&#x2005;kb of a transcription start site). We performed cis-pQTL-MR analyses of four spirometric traits (n=149&#x2009;166, 36 independent cohorts). Sensitivity analyses included colocalisation and reverse direction MR. We report associations between cis-pQTLs for prioritised proteins and multiple clinical respiratory outcomes, and use phenome-wide analysis to explore potential adverse effects or drug repurposing opportunities. FINDINGS: 1841 proteins had a suitable cis-pQTL. We implicated 16 proteins as potentially causal for lung function (p<1.71&#xd7;10-5): seven proteins have not been implicated by previous lung function GWAS or MR (CCND2, DTD1, PILRA, PTPRK, TDRKH, GRHPR, NUDT5), and we provide corroborative evidence for 10 proteins. We add to the literature identifying surfactant protein D (SFTPD) as a candidate, yet predict that integrin subunit alpha V (ITGAV) inhibition could impair some lung function measures, mimicking adverse results from a recent trial. INTERPRETATION: Our approach identifies proteins (some novel) that are potentially therapeutic targets for respiratory disease, and which warrant follow-up for utility and safety.

Journal Article↗

Genome topology analysis and transcriptomics of human osteoclasts reveals enhancer-promoter interactions at loci for bone traits and diseases.

Genome-wide association studies (GWAS) relevant to osteoporosis have identified hundreds of loci; however, understanding how these variants influence the phenotype is complicated because most reside in non-coding DNA sequence that serves as transcriptional enhancers and repressors. To advance knowledge on these regulatory elements in osteoclasts (OCs), we performed Micro-C analysis, which informs on the genome topology of these cells and integrated the results with transcriptome and GWAS data to further define loci linked to BMD. Using blood cells isolated from 4 healthy participants aged 31-61&#xa0;yr, we cultured OC in vitro and generated a Micro-C chromatin conformation capture dataset. We characterized chromatin loops (CLs) in OC from among more than 69 million chromatin interactions identified in the genome. Of the CL identified in OC, >16&#x2009;000 were unique compared to precursor cells. When sentinel single nucleotide polymorphisms from osteoporosis and bone-related GWAS and those in linkage disequilibrium at r 2&#x2009;>&#x2009;0.6 were mapped to CL for OC, 12&#x2009;588 of these variants were observed within chromatin contact regions. Notable in differential gene ontology enrichment analyses of the topology data for OC and precursors were pathways regulating pluripotency of stem cells, Wnt signaling, nucleotide-binding oligomerization domain (NOD)-like receptor signaling and chemokine signaling. These data, in combination with other 3D genome architecture and epigenetic data (eg, histone modifications and chromatin accessibility), will be useful in modeling to predict genome-wide, which enhancers regulate which genes in OC. This data will therefore also be informative for resolving GWAS hits. In conclusion, we have generated a high-resolution genome topology dataset for human OC and have used this to identify CLs relevant to studies of the genetics of osteoporosis. This data will serve as a powerful resource to inform future functional studies of OC biology.

BMD↗

Epigenome-wide Association Study Shows Differential DNA Methylation of MDC1, KLF9, and CUTA in Autoimmune Thyroid Disease.

CONTEXT: Autoimmune thyroid disease (AITD) includes Graves disease (GD) and Hashimoto disease (HD), which often run in the same family. AITD etiology is incompletely understood: Genetic factors may account for up to 75% of phenotypic variance, whereas epigenetic effects (including DNA methylation [DNAm]) may contribute to the remaining variance (eg, why some individuals develop GD and others HD). OBJECTIVE: This work aimed to identify differentially methylated positions (DMPs) and differentially methylated regions (DMRs) comparing GD to HD. METHODS: Whole-blood DNAm was measured across the genome using the Infinium MethylationEPIC array in 32 Australian patients with GD and 30 with HD (discovery cohort) and 32 Danish patients with GD and 32 with HD (replication cohort). Linear mixed models were used to test for differences in quantile-normalized &#x3b2; values of DNAm between GD and HD and data were later meta-analyzed. Comb-p software was used to identify DMRs. RESULTS: We identified epigenome-wide significant differences (P < 9E-8) and replicated (P < .05) 2 DMPs between GD and HD (cg06315208 within MDC1 and cg00049440 within KLF9). We identified and replicated a DMR within CUTA (5 CpGs at 6p21.32). We also identified 64 DMPs and 137 DMRs in the meta-analysis. CONCLUSION: Our study reveals differences in DNAm between GD and HD, which may help explain why some people develop GD and others HD and provide a link to environmental risk factors. Additional research is needed to advance understanding of the role of DNAm in AITD and investigate its prognostic and therapeutic potential.

Humans↗

An investigation of the neurotrophic factor genes GDNF, NGF, and NT3 in susceptibility to ADHD.

Attention deficit hyperactivity disorder (ADHD) is a common, highly heritable, neurodevelopmental disorder with onset in early childhood. Genes involved in neuronal development and growth are, thus, important etiological candidates and neurotrophic factors have been hypothesized to play a role in the pathogenesis of ADHD. Glial derived neurotrophic factor (GDNF), nerve growth factor (NGF (beta subunit)), and neurotrophic factor 3 (NT3) are members of the neurotrophin family and are involved in the survival, differentiation, and maintenance of neuronal cells. We have examined 10 coding and intronic single nucleotide polymorphisms (SNPs) across GDNF, NGF, and NT3 in a family-based association sample of 120 DSM-IV ADHD probands and their biological parents, as well as a case-control analysis with 120 sex-matched controls. Borderline significant overtransmission of the C allele of a non-synonymous C/T SNP (rs6330) in NGF which codes an alanine/valine change was found in the family-based sample (Chi-square = 3.69, odds ratio (OR) = 1.65, P = 0.05). Although this SNP is located in the 5' pro-NGF sequence and not the mature NGF protein, it may affect intracellular processing and secretion of NGF.

Adolescent↗

Comparative gene expression profiling of in vitro differentiated megakaryocytes and erythroblasts identifies novel activatory and inhibitory platelet membrane proteins.

To identify previously unknown platelet receptors we compared the transcriptomes of in vitro differentiated megakaryocytes (MKs) and erythroblasts (EBs). RNA was obtained from purified, biologically paired MK and EB cultures and compared using cDNA microarrays. Bioinformatical analysis of MK-up-regulated genes identified 151 transcripts encoding transmembrane domain-containing proteins. Although many of these were known platelet genes, a number of previously unidentified or poorly characterized transcripts were also detected. Many of these transcripts, including G6b, G6f, LRRC32, LAT2, and the G protein-coupled receptor SUCNR1, encode proteins with structural features or functions that suggest they may be involved in the modulation of platelet function. Immunoblotting on platelets confirmed the presence of the encoded proteins, and flow cytometric analysis confirmed the expression of G6b, G6f, and LRRC32 on the surface of platelets. Through comparative analysis of expression in platelets and other blood cells we demonstrated that G6b, G6f, and LRRC32 are restricted to the platelet lineage, whereas LAT2 and SUCNR1 were also detected in other blood cells. The identification of the succinate receptor SUCNR1 in platelets is of particular interest, because physiologically relevant concentrations of succinate were shown to potentiate the effect of low doses of a variety of platelet agonists.

Cell Differentiation↗

Polymorphism in HSD17B6 is associated with key features of polycystic ovary syndrome.

OBJECTIVE: To investigate polymorphisms in androgen metabolism regulators that are implicated in the etiology of polycystic ovary syndrome (PCOS) in vitro; to investigate HSD17B6 and GATA6 to determine whether these genes are associated with susceptibility to PCOS or key phenotypic features of patients with PCOS. DESIGN: Case-control association study. SETTING: Participants with PCOS were recruited from a clinical-practice database, and controls, from the general community. PATIENT(S): One hundred seventy-three patients with PCOS and who were of Caucasian descent and conformed to the National Institutes of Health (NIH) diagnostic criteria; 107 normally ovulating women of Caucasian descent from the general community. INTERVENTION(S): Drawing of blood for DNA extraction. MAIN OUTCOME MEASURE(S): Frequency of HSD17B6 and GATA6 polymorphisms in cases and controls. Association of single-nucleotide polymorphisms from HSD17B6 in subjects with PCOS with key phenotypes of PCOS: androgen status, insulin resistance, and body mass index. RESULT(S): Allele distribution for the single-nucleotide polymorphism rs898611 in HSD17B6 was significantly different between PCOS and control subjects (P=.03). Presence of the polymorphic allele was associated with reduced fasting glucose-insulin ratio (P=.02) and increased homeostasis model assessment (P<.01) and body mass index (P<.001) as well as with reduced T (P=.03) in the PCOS group. No association was seen between GATA6 and any of the variables studied. CONCLUSION(S): These data suggest that polymorphisms in the HSD17B6 gene are associated with PCOS and key clinical phenotypes of the disorder.

Adult↗

Linkage and potential association of obesity-related phenotypes with two genes on chromosome 12q24 in a female dizygous twin cohort.

Obesity is a multifactorial disorder with a complex phenotype. It is a significant risk factor for diabetes and hypertension. We assessed obesity-related traits in a large cohort of twins and performed a genome-wide linkage scan and positional candidate analysis to identify genes that play a role in regulating fat mass and distribution in women. Dizygous female twin pairs from 1,094 pedigrees were studied (mean age 47.0+/-11.5 years (range 18-79 years)). Nonparametric multipoint linkage analyses showed linkage for central fat mass to 12q24 (141 cM) with LOD 2.2 and body mass index to 8q11 (67 cM) with LOD 1.3, supporting previously established linkage data. Novel areas of suggestive linkage were for total fat percentage at 6q12 (LOD 2.4) and for total lean mass at 2q37 (LOD 2.4). Data from follow-up fine mapping in an expanded cohort of 1243 twin pairs reinforced the linkage for central fat mass to 12q24 (LOD 2.6; 143 cM) and narrowed the -1 LOD support interval to 22 cM. In all, 45 single-nucleotide polymorphisms (SNPs) from 26 positional candidate genes within the 12q24 interval were then tested for association in a cohort of 1102 twins. Single-point Monks-Kaplan analysis provided evidence of association between central fat mass and SNPs in two genes - PLA2G1B (P = 0.0067) and P2RX4 (P = 0.017). These data provide replication and refinement of the 12q24 obesity locus and suggest that genes involved in phospholipase and purinoreceptor pathways may regulate fat accumulation and distribution.

Adolescent↗

Detecting multiple associations in genome-wide studies.

Recent developments in the statistical analysis of genome-wide studies are reviewed. Genome-wide analyses are becoming increasingly common in areas such as scans for disease-associated markers and gene expression profiling. The data generated by these studies present new problems for statistical analysis, owing to the large number of hypothesis tests, comparatively small sample size and modest number of true gene effects. In this review, strategies are described for optimising the genotyping cost by discarding promising genes at an earlier stage, saving resources for the genes that show a trend of association. In addition, there is a review of new methods of analysis that combine evidence across genes to increase sensitivity to multiple true associations in the presence of many non-associated genes. Some methods achieve this by including only the most significant results, whereas others model the overall distribution of results as a mixture of distributions from true and null effects. Because genes are correlated even when having no effect, permutation testing is often necessary to estimate the overall significance, but this can be very time consuming. Efficiency can be improved by fitting a parametric distribution to permutation replicates, which can be re-used in subsequent analyses. Methods are also available to generate random draws from the permutation distribution. The review also includes discussion of new error measures that give a more reasonable interpretation of genome-wide studies, together with improved sensitivity. The false discovery rate allows a controlled proportion of positive results to be false, while detecting more true positives; and the local false discovery rate and false-positive report probability give clarity on whether or not a statistically significant test represents a real discovery.

Genetic Techniques↗

Evaluation of Nyholt's procedure for multiple testing correction.

OBJECTIVE: A simple method for accounting efficiently for multiple testing of many SNPs in an association study was recently proposed by Nyholt, but its performance was not extensively evaluated. The method involves estimating an 'effective number' of independent tests and then adjusting the smallest observed p value using Sidák's formula based on this number of tests. We sought to carry out an empirical and theoretical evaluation of Nyholt's method. METHODS: Nyholt's method was applied to a sample of 31 genes typed at a total of 291 SNPs and permutation used to determine the type-I error rate for each gene. Based on our empirical results, we algebraically investigated the effective number of independent tests for a simple model of haplotype block structure. RESULTS: The nominal 5% type I error rate varied from under 3% to over 7%, and was dependent on linkage disequilibrium. Theoretical considerations show further that the method can be very conservative in the presence of haplotype block structure. CONCLUSION: Although Nyholt's approach may be useful as an exploratory tool, it is not an adequate substitute for permutation tests.

Case-Control Studies↗

The use of edge-betweenness clustering to investigate biological function in protein interaction networks.

BACKGROUND: This paper describes an automated method for finding clusters of interconnected proteins in protein interaction networks and retrieving protein annotations associated with these clusters. RESULTS: Protein interaction graphs were separated into subgraphs of interconnected proteins, using the JUNG implementation of Girvan and Newman's Edge-Betweenness algorithm. Functions were sought for these subgraphs by detecting significant correlations with the distribution of Gene Ontology terms which had been used to annotate the proteins within each cluster. The method was implemented using freely available software (JUNG and the R statistical package). Protein clusters with significant correlations to functional annotations could be identified and included groups of proteins know to cooperate in cell metabolism. The method appears to be resilient against the presence of false positive interactions. CONCLUSION: This method provides a useful tool for rapid screening of small to medium size protein interaction datasets.

Algorithms↗

Efficient computation of significance levels for multiple associations in large studies of correlated data, including genomewide association studies.

Large exploratory studies, including candidate-gene-association testing, genomewide linkage-disequilibrium scans, and array-expression experiments, are becoming increasingly common. A serious problem for such studies is that statistical power is compromised by the need to control the false-positive rate for a large family of tests. Because multiple true associations are anticipated, methods have been proposed that combine evidence from the most significant tests, as a more powerful alternative to individually adjusted tests. The practical application of these methods is currently limited by a reliance on permutation testing to account for the correlated nature of single-nucleotide polymorphism (SNP)-association data. On a genomewide scale, this is both very time-consuming and impractical for repeated explorations with standard marker panels. Here, we alleviate these problems by fitting analytic distributions to the empirical distribution of combined evidence. We fit extreme-value distributions for fixed lengths of combined evidence and a beta distribution for the most significant length. An initial phase of permutation sampling is required to fit these distributions, but it can be completed more quickly than a simple permutation test and need be done only once for each panel of tests, after which the fitted parameters give a reusable calibration of the panel. Our approach is also a more efficient alternative to a standard permutation test. We demonstrate the accuracy of our approach and compare its efficiency with that of permutation tests on genomewide SNP data released by the International HapMap Consortium. The estimation of analytic distributions for combined evidence will allow these powerful methods to be applied more widely in large exploratory studies.

Association↗

Pelican: pedigree editor for linkage computer analysis.

SUMMARY: Linkage analysis software requires an input text file that describes the structure of the pedigrees to be analysed. Manual creation of these files is tedious and error-prone, and a graphical input tool is desirable. This is currently only available in commercial packages that include much greater functionality. We have therefore developed Pelican, a lightweight graphical pedigree editor for rapid construction of linkage pedigree files and diagrams. AVAILABILITY: The software runs on any Java-enabled machine (version 1.2 or higher). A Java Web Start launch, class files, a demonstration applet, source code and documentation are freely available at http://www.rfcgr.mrc.ac.uk/Software/PELICAN/

Chromosome Mapping↗

Linkage and association mapping of the LRP5 locus on chromosome 11q13 in type 1 diabetes.

Linkage of chromosome 11q13 to type 1 diabetes (T1D) was first reported from genome scans (Davies et al. 1994; Hashimoto et al. 1994) resulting in P <2.2 x 10(-5) (Luo et al. 1996) and designated IDDM4 ( insulin dependent diabetes mellitus 4). Association mapping under the linkage peak using 12 polymorphic microsatellite markers suggested some evidence of association with a two-marker haplotype, D11S1917*03-H0570POLYA*02, which was under-transmitted to affected siblings and over-transmitted to unaffected siblings ( P=1.5 x 10(-6)) (Nakagawa et al. 1998). Others have reported evidence for T1D association of the microsatellite marker D11S987, which is approximately 100 kb proximal to D11S1917 (Eckenrode et al. 2000). We have sequenced a 400-kb interval surrounding these loci and identified four genes, including the low-density lipoprotein receptor related protein (LRP5) gene, which has been considered as a functional candidate gene for T1D (Hey et al. 1998; Twells et al. 2001). Consequently, we have developed a comprehensive SNP map of the LRP5 gene region, and identified 95 SNPs encompassing 269 kb of genomic DNA, characterised the LD in the region and haplotypes (Twells et al. 2003). Here, we present our refined linkage curve of the IDDM4 region, comprising 32 microsatellite markers and 12 SNPs, providing a peak MLS=2.58, P=5 x 10(-4), at LRP5 g.17646G>T. The disease association data, largely focused in the LRP5 region with 1,106 T1D families, provided no further evidence for disease association at LRP5 or at D11S987. A second dataset, comprising 1,569 families from Finland, failed to replicate our previous findings at LRP5. The continued search for the variants of the putative IDDM4 locus will greatly benefit from the future development of a haplotype map of the genome.

Chromosome Mapping↗

Pedigree disequilibrium tests for multilocus haplotypes.

Association tests of multilocus haplotypes are of interest both in linkage disequilibrium mapping and in candidate gene studies. For case-parent trios, I discuss the extension of existing multilocus methods to include ambiguous haplotypes in tests of models which distinguish between the cis and trans phase. A likelihood-ratio test is proposed, using the expectation-maximization (E-M) algorithm to account for haplotype ambiguities. Assumptions about the population structure are required, but realistic situations, including population stratification, which violate the assumptions lead to conservative tests. I describe a permutation procedure for the null hypothesis of interest, which controls for violation of the assumptions. For general pedigrees, I describe extensions of the pedigree disequilibrium test to include uncertain haplotypes. The summary statistics are replaced by their expected values over prior distributions of haplotype frequencies. If prior distributions are not available, a valid test is possible by using the E-M algorithm to estimate the null distribution of haplotype frequencies. Similar methods are available for quantitative traits. Exact permutation tests are difficult to construct in small samples, but an approximate procedure is appropriate in large samples, and can be used to account for dependencies between tests of multiple haplotypes and loci.

Algorithms↗

Rank truncated product of P-values, with application to genomewide association scans.

Large exploratory studies are often characterized by a preponderance of true null hypotheses, with a small though multiple number of false hypotheses. Traditional multiple-test adjustments consider either each hypothesis separately, or all hypotheses simultaneously, but it may be more desirable to consider the combined evidence for subsets of hypotheses, in order to reduce the number of hypotheses to a manageable size. Previously, Zaykin et al. ([2002] Genet. Epidemiol. 22:170-185) proposed forming the product of all P-values at less than a preset threshold, in order to combine evidence from all significant tests. Here we consider a complementary strategy: form the product of the K most significant P-values. This has certain advantages for genomewide association scans: K can be chosen on the basis of a hypothesised disease model, and is independent of sample size. Furthermore, the alternative hypothesis corresponds more closely to the experimental situation where all loci have fixed effects. We give the distribution of the rank truncated product and suggest some methods to account for correlated tests in genomewide scans. We show that, under realistic scenarios, it provides increased power to detect genomewide association, while identifying a candidate set of good quality and fixed size for follow-up studies.

Algorithms↗

A survey of current software for linkage analysis.

There is now a wide choice of software available for linkage analysis. The most well known packages are briefly reviewed here. The package with the most extensive range of analyses is GENEHUNTER, but for many of its functions there are other programs with better performance. These include FASTLINK and VITESSE for parametric analysis ALLEGRO and MERLIN for non-parametric analysis and SOLAR for variance components analysis. The computational limits of current approaches can be improved with SIMWALK2 and the promising new SUPERLINK program. Directions for future work include improved user interfaces and consensus formats for data input and exchange.

Algorithms↗