PubMed Health⌕ Search

Biomedical subjects

Lars Feuk

Publications and source records attributed to Lars Feuk.

At least 19 recordsLinked to original sources

Nallo: a Nextflow pipeline for comprehensive human long-read genome analysis.

MOTIVATION: Long-read sequencing (LRS) is increasingly used for human medical research and clinical diagnostics due to its capacity to generate complete genome information. However, there is a lack of robust and easy-to-use pipelines for comprehensive LRS data analysis. RESULTS: Here we present Nallo, a Nextflow pipeline for analysis of PacBio and Oxford Nanopore data, with additional support for rare disease research projects. The pipeline detects a wide range of genetic variants, performs genome assembly, and reports CpG methylation. It also enables annotation and ranking of variants based on their predicted functional consequences. AVAILABILITY AND IMPLEMENTATION: Nallo is available from GitHub: https://github.com/genomic-medicine-sweden/nallo.

Humans↗

Global variation in copy number in the human genome.

Copy number variation (CNV) of DNA sequences is functionally significant but has yet to be fully ascertained. We have constructed a first-generation CNV map of the human genome through the study of 270 individuals from four populations with ancestry in Europe, Africa or Asia (the HapMap collection). DNA from these individuals was screened for CNV using two complementary technologies: single-nucleotide polymorphism (SNP) genotyping arrays, and clone-based comparative genomic hybridization. A total of 1,447 copy number variable regions (CNVRs), which can encompass overlapping or adjacent gains or losses, covering 360 megabases (12% of the genome) were identified in these populations. These CNVRs contained hundreds of genes, disease loci, functional elements and segmental duplications. Notably, the CNVRs encompassed more nucleotide content per genome than SNPs, underscoring the importance of CNV in genetic diversity and evolution. The data obtained delineate linkage disequilibrium patterns for many CNVs, and reveal marked variation in copy number among populations. We also demonstrate the utility of this resource for genetic disease studies.

Chromosome Mapping↗

Genome assembly comparison identifies structural variants in the human genome.

Numerous types of DNA variation exist, ranging from SNPs to larger structural alterations such as copy number variants (CNVs) and inversions. Alignment of DNA sequence from different sources has been used to identify SNPs and intermediate-sized variants (ISVs). However, only a small proportion of total heterogeneity is characterized, and little is known of the characteristics of most smaller-sized (<50 kb) variants. Here we show that genome assembly comparison is a robust approach for identification of all classes of genetic variation. Through comparison of two human assemblies (Celera's R27c compilation and the Build 35 reference sequence), we identified megabases of sequence (in the form of 13,534 putative non-SNP events) that were absent, inverted or polymorphic in one assembly. Database comparison and laboratory experimentation further demonstrated overlap or validation for 240 variable regions and confirmed >1.5 million SNPs. Some differences were simple insertions and deletions, but in regions containing CNVs, segmental duplication and repetitive DNA, they were more complex. Our results uncover substantial undescribed variation in humans, highlighting the need for comprehensive annotation strategies to fully interpret genome scanning and personalized sequencing projects.

Base Sequence↗

Accurate and reliable high-throughput detection of copy number variation in the human genome.

This study describes a new tool for accurate and reliable high-throughput detection of copy number variation in the human genome. We have constructed a large-insert clone DNA microarray covering the entire human genome in tiling path resolution that we have used to identify copy number variation in human populations. Crucial to this study has been the development of a robust array platform and analytic process for the automated identification of copy number variants (CNVs). The array consists of 26,574 clones covering 93.7% of euchromatic regions. Clones were selected primarily from the published "Golden Path," and mapping was confirmed by fingerprinting and BAC-end sequencing. Array performance was extensively tested by a series of validation assays. These included determining the hybridization characteristics of each individual clone on the array by chromosome-specific add-in experiments. Estimation of data reproducibility and false-positive/negative rates was carried out using self-self hybridizations, replicate experiments, and independent validations of CNVs. Based on these studies, we developed a variance-based automatic copy number detection analysis process (CNVfinder) and have demonstrated its robustness by comparison with the SW-ARRAY method.

Algorithms↗

Absence of a paternally inherited FOXP2 gene in developmental verbal dyspraxia.

Mutations in FOXP2 cause developmental verbal dyspraxia (DVD), but only a few cases have been described. We characterize 13 patients with DVD--5 with hemizygous paternal deletions spanning the FOXP2 gene, 1 with a translocation interrupting FOXP2, and the remaining 7 with maternal uniparental disomy of chromosome 7 (UPD7), who were also given a diagnosis of Silver-Russell Syndrome (SRS). Of these individuals with DVD, all 12 for whom parental DNA was available showed absence of a paternal copy of FOXP2. Five other individuals with deletions of paternally inherited FOXP2 but with incomplete clinical information or phenotypes too complex to properly assess are also described. Four of the patients with DVD also meet criteria for autism spectrum disorder. Individuals with paternal UPD7 or with partial maternal UPD7 or deletion starting downstream of FOXP2 do not have DVD. Using quantitative real-time polymerase chain reaction, we show the maternally inherited FOXP2 to be comparatively underexpressed. Our results indicate that absence of paternal FOXP2 is the cause of DVD in patients with SRS with maternal UPD7. The data also point to a role for differential parent-of-origin expression of FOXP2 in human speech development.

Apraxias↗

Frequent appearance of novel protein-coding sequences by frameshift translation.

Genomic duplication, followed by divergence, contributes to organismal evolution. Several mechanisms, such as exon shuffling and alternative splicing, are responsible for novel gene functions, but they generate homologous domains and do not usually lead to drastic innovation. Major novelties can potentially be introduced by frameshift mutations and this idea can explain the creation of novel proteins. Here, we employ a strategy using simulated protein sequences and identify 470 human and 108 mouse frameshift events that originate new gene segments. No obvious interspecies overlap was observed, suggesting high rates of acquisition of evolutionary events. This inference is supported by a deficiency of TpA dinucleotides in the protein-coding sequences, which decreases the occurrence of translational termination, even on the complementary strand. Increased usage of the TGA codon as the termination signal in newer genes also supports our inference. This suggests that tolerated frameshift changes are a prevalent mechanism for the rapid emergence of new genes and that protein-coding sequences can be derived from existing or ancestral exons rather than from events that result in noncoding sequences becoming exons.

Amino Acid Sequence↗

Copy number variation: new insights in genome diversity.

DNA copy number variation has long been associated with specific chromosomal rearrangements and genomic disorders, but its ubiquity in mammalian genomes was not fully realized until recently. Although our understanding of the extent of this variation is still developing, it seems likely that, at least in humans, copy number variants (CNVs) account for a substantial amount of genetic variation. Since many CNVs include genes that result in differential levels of gene expression, CNVs may account for a significant proportion of normal phenotypic variation. Current efforts are directed toward a more comprehensive cataloging and characterization of CNVs that will provide the basis for determining how genomic diversity impacts biological function, evolution, and common human diseases.

Animals↗

Structural variants: changing the landscape of chromosomes and design of disease studies.

The near completeness of human chromosome sequences is facilitating accurate characterization and assessment of all classes of genomic variation. Particularly, using the DNA reference sequence as a guide, genome scanning technologies, such as microarray-based comparative genomic hybridization (array CGH) and genome-wide single nucleotide polymorphism (SNP) platforms, have now enabled the detection of a previously unrecognized degree of larger-sized (non-SNP) variability in all genomes. This heterogeneity can include copy number variations (CNVs), inversions, insertions, deletions and other complex rearrangements, most of which are not detected by standard cytogenetics or DNA sequencing. Although these genomic alterations (collectively termed structural variants or polymorphisms) have been described previously, mainly through locus-specific studies, they are now known to be more global in occurrence. Moreover, as just one example, CNVs can contain entire genes and their number can correlate with the level of gene expression. It is also plausible that structural variants may commonly influence nearby genes through chromosomal positional or domain effects. Here, we discuss what is known of the prevalence of structural variants in the human genome and how they might influence phenotype, including the continuum of etiologic events underlying monogenic to complex diseases. Particularly, we highlight the newest studies and some classic examples of how structural variants might have adverse genetic consequences. We also discuss why analysis of structural variants should become a vital step in any genetic study going forward. All these progresses have set the stage for a golden era of combined microscopic and sub-microscopic (cytogenomic)-based research of chromosomes leading to a more complete understanding of the human genome.

Chromosomes↗

Speech and language impairment and oromotor dyspraxia due to deletion of 7q31 that involves FOXP2.

We report detailed clinical, cytogenetic, and molecular findings in a girl with a deletion of chromosome 7q31-q32. This child has a severe communication disorder with evidence of oromotor dyspraxia, dysmorphic features, and mild developmental delay. She is unable to cough, sneeze, or laugh spontaneously. Her deletion is on the paternally inherited chromosome and includes the FOXP2 gene, which has recently been associated with speech and language impairment and a similar form of oromotor dyspraxia in at least three other published cases. We hypothesize that our patient's communication disorder and oromotor deficiency are due to haploinsufficiency for FOXP2 and that her dysmorphism and developmental delay are a consequence of the absence of the other genes involved in the microdeletion. We propose that this patient, together with others reported in the literature, may define a new contiguous gene deletion syndrome encompassing the 7q31-FOXP2 region. Cytogenetic and molecular analysis of this region should be considered for other individuals displaying similar characteristics.

Abnormalities, Multiple↗

Longitudinal memory performance during normal aging: twin association models of APOE and other Alzheimer candidate genes.

The APOE gene (apolipoprotein E) is a major risk factor for Alzheimer's Disease (AD) but has been inconsistently associated with memory in nondemented adults. Two other genes with mixed support as genetic risk factors for AD, A2M (alpha-2-macroglobulin) and LRP (low-density lipoprotein receptor-related protein), have not been studied in relation to memory among nondemented adults. The present study examined these three genes and latent growth parameters estimated from memory performance spanning 13 years in 478 twins from the Swedish Adoption/Twin Study of Aging (SATSA). APOE was associated with working and recall memory ability levels and working memory rate of change, with e4 homozygotes exhibiting the worst performance at all ages. Homozygotes for the rare A2M insertion/deletion variant exhibited accelerating decline on delayed figural recognition. There were no significant findings for LRP. Dominance, often untested in previous studies, was important in the current study's findings.

Age Factors↗

Structural variation in the human genome.

The first wave of information from the analysis of the human genome revealed SNPs to be the main source of genetic and phenotypic human variation. However, the advent of genome-scanning technologies has now uncovered an unexpectedly large extent of what we term 'structural variation' in the human genome. This comprises microscopic and, more commonly, submicroscopic variants, which include deletions, duplications and large-scale copy-number variants - collectively termed copy-number variants or copy-number polymorphisms - as well as insertions, inversions and translocations. Rapidly accumulating evidence indicates that structural variants can comprise millions of nucleotides of heterogeneity within every genome, and are likely to make an important contribution to human diversity and disease susceptibility.

Genetic Variation↗

Strategies for the detection of copy number and other structural variants in the human genome.

Advances in genome scanning technologies are revealing that copy number variants (CNVs) and polymorphisms, ranging from a few kilobases to several megabases in size, are present in genomes at frequencies much greater than previously known. Discoveries of additional forms of genomic variation, including inversions, insertions, deletions and complex rearrangements, are also occurring at an increased rate. Along with CNVs, these sequence alterations are collectively known as structural variants, and their discovery has had an immediate impact on the interpretation of basic research and clinical diagnostic data. This paper discusses different methods, experimental strategies and technologies that are currently available to study copy number variation and other structural variants in the human genome.

Gene Dosage↗

Towards compendia of negative genetic association studies: an example for Alzheimer disease.

Most genetic sequence variants that contribute to variability in complex human traits will have small effects that are not readily detectable with population samples typically used in genetic association studies. A potentially valuable tool in the gene discovery process is meta-analysis of the accumulated published data, but in order to be valid these require a sample of studies representative of the true genetic effect and thus hypothetically should include some positive and an abundance of negative reports. A survey of the literature on association studies for Alzheimer disease (AD) from January 2004-April 2005, identified 138 studies, 86 of which reported positive findings other than for apolipoprotein E (APOE), strongly indicative of publication bias. We report here an analysis of 62 genetic markers, tested for association with AD risk as well as for possible effects upon quantitative indices of AD severity (mini-mental state examination scores, age-at-onset, and cerebrospinal fluid (CSF) beta-amyloid (Abeta) and CSF tau proteins). Within this set, only modest signals were present that, with the exception of APOE are easily lost when corrections for multiple hypotheses are applied. In isolation, results are thus broadly negative. Genes studied encompass both novel candidates as well as several recently claimed to be associated with AD (e.g. urokinase plasminogen activator (PLAU) and acetyl-coenzyme A acetyltransferase 1 (ACAT1)). By reporting these data we hope to encourage the publication of gene compendia to guide further studies and aid future meta-analyses aimed at resolving the involvement of genes in complex human traits.

Aged↗

Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.

With a draft genome-sequence assembly for the chimpanzee available, it is now possible to perform genome-wide analyses to identify, at a submicroscopic level, structural rearrangements that have occurred between chimpanzees and humans. The goal of this study was to investigate chromosomal regions that are inverted between the chimpanzee and human genomes. Using the net alignments for the builds of the human and chimpanzee genome assemblies, we identified a total of 1,576 putative regions of inverted orientation, covering more than 154 mega-bases of DNA. The DNA segments are distributed throughout the genome and range from 23 base pairs to 62 mega-bases in length. For the 66 inversions more than 25 kilobases (kb) in length, 75% were flanked on one or both sides by (often unrelated) segmental duplications. Using PCR and fluorescence in situ hybridization we experimentally validated 23 of 27 (85%) semi-randomly chosen regions; the largest novel inversion confirmed was 4.3 mega-bases at human Chromosome 7p14. Gorilla was used as an out-group to assign ancestral status to the variants. All experimentally validated inversion regions were then assayed against a panel of human samples and three of the 23 (13%) regions were found to be polymorphic in the human genome. These polymorphic inversions include 730 kb (at 7p22), 13 kb (at 7q11), and 1 kb (at 16q24) fragments with a 5%, 30%, and 48% minor allele frequency, respectively. Our results suggest that inversions are an important source of variation in primate genome evolution. The finding of at least three novel inversion polymorphisms in humans indicates this type of structural variation may be a more common feature of our genome than previously realized.

Animals↗

Mutation screening of a haplotype block around the insulin degrading enzyme gene and association with Alzheimer's disease.

Genetic and biological studies point to a role for insulin-degrading enzyme (IDE) in Alzheimer's disease (AD). Two SNP-based studies recently reported evidence for association with AD using markers in a approximately 270 kb haplotype block on chromosome 10q. This haplotype block region harbors three known genes; insulin-degrading enzyme (IDE), kinesin family member 11 (KIF11), and hematopoietically expressed homeobox (HHEX). In an attempt to search for susceptibility variants we have sequenced all coding exons, 2 kb of 5' and 3'-flanking sequence, and all regions showing a high degree of human-mouse conservation in these three genes in 30 individuals. We found a total of 40 single nucleotide polymorphisms and 8 insertion/deletion polymorphisms. No coding variants were identified in any of the three genes. Nine polymorphisms in IDE and four polymorphisms in KIF11 situated in conserved regions or near coding exons were subsequently genotyped in a set of AD cases and controls. Two markers in KIF11 yielded borderline significant results in the ApoE4 non-carrier subgroup, but the results were otherwise not significant in this small set of samples. This study of multiple new markers in the region will facilitate further association studies in this important AD region.

Alzheimer Disease↗

Sequence variants of IDE are associated with the extent of beta-amyloid deposition in the Alzheimer's disease brain.

Insulin degrading enzyme, encoded by IDE, plays a primary role in the degradation of amyloid beta-protein (A beta), the deposition of which in senile plaques is one of the defining hallmarks of Alzheimer's disease (AD). We recently identified haplotypes in a broad linkage disequilibrium (LD) block encompassing IDE that associate with several AD-related quantitative traits. Here, by examining 32 polymorphic markers extending across IDE and testing quantitative measures of plaque density and cognitive function in three independent Swedish AD samples, we have refined the probable position of pathogenic sequences to a 3' region of IDE, with local maximum effects in the proximity of marker rs1887922. To replicate these findings, a subset of variants were examined against measures of brain A beta load in an independent English AD sample, whereby maximum effects were again observed for rs1887922. For both Swedish and English autopsy materials, variation at rs1887922 explained approximately 10% of the total variance in the respective histopathology traits. However, across all clinical materials studied to date, this variant site does not appear to associate directly with disease, suggesting that IDE may affect AD severity rather than risk. Results indicate that alleles of IDE contribute to variability in A beta deposition in the AD brain and suggest that this relationship may have relevance for the degree of cognitive dysfunction in AD patients.

Alzheimer Disease↗

Linkage disequilibrium patterns vary substantially among populations.

A major initiative to create a global human haplotype map has recently been launched as a tool to improve the efficiency of disease gene mapping. The 'HapMap' project will study common variants in depth in four (and to a lesser degree in up to 12) populations to catalogue haplotypes that are expected to be common to all populations. A hope of the 'HapMap' project is that much of the genome occurs in regions of limited diversity such that only a few of the SNPs in each region will capture the diversity and be relevant around the world. In order to explore the implications of studying only a limited number of populations, we have analyzed linkage disequilibrium (LD) patterns of three 175-320 kb genomic regions in 16 diverse populations with an emphasis on African and European populations. Analyses of these three genomic regions provide empiric demonstration of marked differences in frequencies of the same few haplotypes, resulting in differences in the amount of LD and very different sets of haplotype frequencies. These results highlight the distinction between the statistical concept of LD and the biological reality of haplotypes and their frequencies. The significant quantitative and qualitative variation in LD among populations, even for populations within a geographic region, emphasizes the importance of studying diverse populations in the HapMap project to assure broad applicability of the results.

Asian People↗

Elevated amyloid beta protein (Abeta42) and late onset Alzheimer's disease are associated with single nucleotide polymorphisms in the urokinase-type plasminogen activator gene.

Plasma amyloid beta protein (Abeta42) levels and late onset Alzheimer's disease (LOAD) have been linked to the same region on chromosome 10q. The PLAU gene within this region encodes urokinase-type plasminogen activator, which converts plasminogen to plasmin. Abeta aggregates induce PLAU expression thereby increasing plasmin, which degrades both aggregated and non-aggregated forms of Abeta. We evaluated single nucleotide polymorphisms (SNPs) in PLAU for association with Abeta42 and LOAD. PLAU SNP compound genotypes composed of haplotype pairs showed significant association with AD in three independent case-control series. PLAU SNP haplotypes associated significantly with plasma Abeta42 in 10 extended LOAD families. One of the SNPs analyzed was a missense C/T polymorphism in exon 6 of PLAU (PLAU_1=rs2227564), which causes a proline to leucine change (P141L). We analyzed PLAU_1 for association with AD in six case-control series and 24 extended LOAD families. The CT and TT PLAU_1 genotypes showed association (P=0.05) with an overall estimated odds ratio of 1.2 (1.0-1.5). The CT and TT genotypes of PLAU_1 were also associated with significant age-dependent elevation of plasma Abeta42 in 24 extended LOAD families (P=0.0006). In knockout mice lacking the PLAU gene, plasma--but not brain--Abeta42 as well as Abeta40 was significantly elevated, also in an age-dependent manner. The PLAU_1 associations were independent of the associations we found among plasma Abeta42, LOAD and variants in the IDE or VR22 region. These results provide strong evidence that PLAU or a nearby gene is involved in the development of LOAD. PLAU_1 is a plausible pathogenic mutation that could act by increasing Abeta42, but additional biological experiments are required to show this definitively.

Age Factors↗