PubMed Health⌕ Search

Biomedical subjects

E E Eichler

Publications and source records attributed to E E Eichler.

At least 19 recordsLinked to original sources

Molecular evolution of the human chromosome 15 pericentromeric region.

We present a detailed molecular evolutionary analysis of 1.2 Mb from the pericentromeric region of human 15q11. Sequence analysis indicates the region has been subject to extensive interchromosomal and intrachromosomal duplications during primate evolution. Comparative FISH analyses among non-human primates show remarkable quantitative and qualitative differences in the organization and duplication history of this region - including lineage-specific deletions and duplication expansions. Phylogenetic and comparative analyses reveal that the region is composed of at least 24 distinct segmental duplications or duplicons that have populated the pericentromeric regions of the human genome over the last 40 million years of human evolution. The value of combining both cytogenetic and experimental data in understanding the complex forces which have shaped these regions is discussed.

Animals↗

BAC microarray analysis of 15q11-q13 rearrangements and the impact of segmental duplications.

Chromosome 15q11-q13 is one of the most variable regions of the human genome, with numerous clinical rearrangements involving a dosage imbalance. Multiple clusters of segmental duplications are found in the pericentromeric region of 15q and at the breakpoints of proximal 15q rearrangements. Using sequence maps and previous global analyses of segmental duplications in the human genome, a targeted microarray was developed to detect a wide range of dosage imbalances in clinical samples. Clones were also chosen to assess the effect of paralogous sequences in the array format. In 19 patients analysed, the array data correlated with microsatellite and FISH characterisation. The data showed a linear response with respect to dosage, ranging from one to six copies of the region. Paralogous sequences in arrayed clones appear to respond to the total genomic copy number, and results with such clones may seem aberrant unless the sequence context of the arrayed sequence is well understood. The array CGH method offers exquisite resolution and sensitivity for detecting large scale dosage imbalances. These results indicate that the duplication composition of BAC substrates may affect the sensitivity for detecting dosage variation. They have important implications for effective microarray design, as well as for the detection of segmental aneusomy within the human population.

Chromosome Aberrations↗

Identification of four highly conserved genes between breakpoint hotspots BP1 and BP2 of the Prader-Willi/Angelman syndromes deletion region that have undergone evolutionary transposition mediated by flanking duplicons.

Prader-Willi and Angelman syndromes (PWS and AS) typically result from an approximately 4-Mb deletion of human chromosome 15q11-q13, with clustered breakpoints (BP) at either of two proximal sites (BP1 and BP2) and one distal site (BP3). HERC2 and other duplicons map to these BP regions, with the 2-Mb PWS/AS imprinted domain just distal of BP2. Previously, the presence of genes and their imprinted status have not been examined between BP1 and BP2. Here, we identify two known (CYFIP1 and GCP5) and two novel (NIPA1 and NIPA2) genes in this region in human and their orthologs in mouse chromosome 7C. These genes are expressed from a broad range of tissues and are nonimprinted, as they are expressed in cells derived from normal individuals, patients with PWS or AS, and the corresponding mouse models. However, replication-timing studies in the mouse reveal that they are located in a genomic domain showing asynchronous replication, a feature typically ascribed to monoallelically expressed loci. The novel genes NIPA1 and NIPA2 each encode putative polypeptides with nine transmembrane domains, suggesting function as receptors or as transporters. Phylogenetic analyses show that NIPA1 and NIPA2 are highly conserved in vertebrate species, with ancestral members in invertebrates and plants. Intriguingly, evolutionary studies show conservation of the four-gene cassette between BP1 and BP2 in human, including NIPA1/2, CYFIP1, and GCP5, and proximity to the Herc2 gene in both mouse and Fugu. These observations support a model in which duplications of the HERC2 gene at BP3 in primates first flanked the four-gene cassette, with subsequent transposition of these four unique genes by a HERC2 duplicon-mediated process to form the BP1-BP2 region. Duplicons therefore appear to mediate genomic fluidity in both disease and evolutionary processes.

Adaptor Proteins, Signal Transducing↗

Using a pericentromeric interspersed repeat to recapitulate the phylogeny and expansion of human centromeric segmental duplications.

Despite considerable advances in sequencing of the human genome over the past few years, the organization and evolution of human pericentromeric regions have been difficult to resolve. This is due, in part, to the presence of large, complex blocks of duplicated genomic sequence at the boundary between centromeric satellite and unique euchromatic DNA. Here, we report the identification and characterization of an approximately 49-kb repeat sequence that exists in more than 40 copies within the human genome. This repeat is specific to highly duplicated pericentromeric regions with multiple copies distributed in an interspersed fashion among a subset of human chromosomes. Using this interspersed repeat (termed PIR4) as a marker of pericentromeric DNA, we recovered and sequence-tagged 3 Mb of pericentromeric DNA from a variety of human chromosomes as well as nonhuman primate genomes. A global evolutionary reconstruction of the dispersal of PIR4 sequence and analysis of flanking sequence supports a model in which pericentromeric duplications initiated before the separation of the great ape species (>12 MYA). Further, analyses of this duplication and associated flanking duplications narrow the major burst of pericentromeric duplication activity to a time just before the divergence of the African great ape and human species (5 to 7 MYA). These recent duplication exchange events substantially restructured the pericentromeric regions of hominoid chromosomes and created an architecture where large blocks of sequence are shared among nonhomologous chromosomes. This report provides the first global view of the series of historical events that have reshaped human pericentromeric regions over recent evolutionary time.

Animals↗

Positive selection of a gene family during the emergence of humans and African apes.

Gene duplication followed by adaptive evolution is one of the primary forces for the emergence of new gene function. Here we describe the recent proliferation, transposition and selection of a 20-kilobase (kb) duplicated segment throughout 15 Mb of the short arm of human chromosome 16. The dispersal of this segment was accompanied by considerable variation in chromosomal-map location and copy number among hominoid species. In humans, we identified a gene family (morpheus) within the duplicated segment. Comparison of putative protein-encoding exons revealed the most extreme case of positive selection among hominoids. The major episode of enhanced amino-acid replacement occurred after the separation of human and great-ape lineages from the orangutan. Positive selection continued to alter amino-acid composition after the divergence of human and chimpanzee lineages. The rapidity and bias for amino-acid-altering nucleotide changes suggest adaptive evolution of the morpheus gene family during the emergence of humans and African apes. Moreover, some genes emerge and evolve very rapidly, generating copies that bear little similarity to their ancestral precursors. Consequently, a small fraction of human genes may not possess discernible orthologues within the genomes of model organisms.

Animals↗

Lessons from the human genome: transitions between euchromatin and heterochromatin.

The publication of the human genome draft sequence provides, for the first time, a global view of the structural properties of the human genome. Initial sequence analysis, in combination with previous published reports, reveals that more than half of the transition regions between euchromatin and centromeric heterochromatin contain duplicated segments. The individual duplications originate from diverse euchromatic regions of the human genome, often containing intron-exon structure of known genes. Multiple duplicons are concatenated together to form larger blocks of wall-to-wall duplications. For a single chromosome, these paralogous segments can span >1 Mb of sequence and define a buffer zone between unique sequence and tandemly repeated satellite sequences. Unusual pericentromeric interspersed repeat elements have been identified at the junctions of many of these duplications. Phylogenetic and comparative studies of pericentromeric sequences suggest that this peculiar genome organization has emerged within the last 30 million years of human evolution and is a source of considerable genomic variation between closely related primate species. Interestingly, not all human pericentromeric regions show this proclivity to duplicate and transpose genomic sequence, suggesting at least two different models for the organization of these regions.

Euchromatin↗

Integration of cytogenetic landmarks into the draft sequence of the human genome.

We have placed 7,600 cytogenetically defined landmarks on the draft sequence of the human genome to help with the characterization of genes altered by gross chromosomal aberrations that cause human disease. The landmarks are large-insert clones mapped to chromosome bands by fluorescence in situ hybridization. Each clone contains a sequence tag that is positioned on the genomic sequence. This genome-wide set of sequence-anchored clones allows structural and functional analyses of the genome. This resource represents the first comprehensive integration of cytogenetic, radiation hybrid, linkage and sequence maps of the human genome; provides an independent validation of the sequence map and framework for contig order and orientation; surveys the genome for large-scale duplications, which are likely to require special attention during sequence assembly; and allows a stringent assessment of sequence differences between the dark and light bands of chromosomes. It also provides insight into large-scale chromatin structure and the evolution of chromosomes and gene families and will accelerate our understanding of the molecular bases of human disease and cancer.

Chromosome Aberrations↗

Recent duplication, domain accretion and the dynamic mutation of the human genome.

An estimated 5% of the human genome consists of interspersed duplications that have arisen over the past 35 million years of evolution. Two categories of such recently duplicated segments can be distinguished: segmental duplications between nonhomologous chromosomes (transchromosomal duplications) and duplications mainly restricted to a particular chromosome (chromosome-specific duplications). Many of these duplications exhibit an extraordinarily high degree of sequence identity at the nucleotide level (>95%) and span large genomic distances (1-100 kb). Preliminary analyses indicate that these same regions are targets for rapid evolutionary turnover among the genomes of closely related primates. The dynamic nature of these regions because of recurrent chromosomal rearrangement, and their ability to create fusion genes from juxtaposed cassettes suggest that duplicative transposition was an important force in the evolution of our genome.

Biological Evolution↗

Sequence variation within the fragile X locus.

The human genome provides a reference sequence, which is a template for resequencing studies that aim to discover and interpret the record of common ancestry that exists in extant genomes. To understand the nature and pattern of variation and linkage disequilibrium comprising this history, we present a study of approximately 31 kb spanning an approximately 70 kb region of FMR1, sequenced in a sample of 20 humans (worldwide sample) and four great apes (chimp, bonobo, and gorilla). Twenty-five polymorphic sites and two insertion/deletions, distributed in 11 unique haplotypes, were identified among humans. Africans are the only geographic group that do not share any haplotypes with other groups. Parsimony analysis reveals two main clades and suggests that the four major human geographic groups are distributed throughout the phylogenetic tree and within each major clade. An African sample appears to be most closely related to the common ancestor shared with the three other geographic groups. Nucleotide diversity, pi, for this sample is 2.63 +/- 6.28 x 10(-4). The mutation rate, mu is 6.48 x 10(-10) per base pair per year, giving an ancestral population size of approximately 6200 and a time to the most recent common ancestor of approximately 320,000 +/- 72,000 per base pair per year. Linkage disequilibrium (LD) at the FMR1 locus, evaluated by conventional LD analysis and by the length of segment shared between any two chromosomes, is extensive across the region.

Animals↗

High-throughput variation detection and genotyping using microarrays.

The genetic dissection of complex traits may ultimately require a large number of SNPs to be genotyped in multiple individuals who exhibit phenotypic variation in a trait of interest. Microarray technology can enable rapid genotyping of variation specific to study samples. To facilitate their use, we have developed an automated statistical method (ABACUS) to analyze microarray hybridization data and applied this method to Affymetrix Variation Detection Arrays (VDAs). ABACUS provides a quality score to individual genotypes, allowing investigators to focus their attention on sites that give accurate information. We have applied ABACUS to an experiment encompassing 32 autosomal and eight X-linked genomic regions, each consisting of approximately 50 kb of unique sequence spanning a 100-kb region, in 40 humans. At sufficiently high-quality scores, we are able to read approximately 80% of all sites. To assess the accuracy of SNP detection, 108 of 108 SNPs have been experimentally confirmed; an additional 371 SNPs have been confirmed electronically. To access the accuracy of diploid genotypes at segregating autosomal sites, we confirmed 1515 of 1515 homozygous calls, and 420 of 423 (99.29%) heterozygotes. In replicate experiments, consisting of independent amplification of identical samples followed by hybridization to distinct microarrays of the same design, genotyping is highly repeatable. In an autosomal replicate experiment, 813,295 of 813,295 genotypes are called identically (including 351 heterozygotes); at an X-linked locus in males (haploid), 841,236 of 841,236 sites are called identically.

Algorithms↗

Segmental duplications: organization and impact within the current human genome project assembly.

Segmental duplications play fundamental roles in both genomic disease and gene evolution. To understand their organization within the human genome, we have developed the computational tools and methods necessary to detect identity between long stretches of genomic sequence despite the presence of high copy repeats and large insertion-deletions. Here we present our analysis of the most recent genome assembly (January 2001) in which we focus on the global organization of these segments and the role they play in the whole-genome assembly process. Initially, we considered only large recent duplication events that fell well-below levels of draft sequencing error (alignments 90%-98% similar and > or =1 kb in length). Duplications (90%-98%; > or =1 kb) comprise 3.6% of all human sequence. These duplications show clustering and up to 10-fold enrichment within pericentromeric and subtelomeric regions. In terms of assembly, duplicated sequences were found to be over-represented in unordered and unassigned contigs indicating that duplicated sequences are difficult to assign to their proper position. To assess coverage of these regions within the genome, we selected BACs containing interchromosomal duplications and characterized their duplication pattern by FISH. Only 47% (106/224) of chromosomes positive by FISH had a corresponding chromosomal position by comparison. We present data that indicate that this is attributable to misassembly, misassignment, and/or decreased sequencing coverage within duplicated regions. Surprisingly, if we consider putative duplications >98% identity, we identify 10.6% (286 Mb) of the current assembly as paralogous. The majority of these alignments, we believe, represent unmerged overlaps within unique regions. Taken together the above data indicate that segmental duplications represent a significant impediment to accurate human genome assembly, requiring the development of specialized techniques to finish these exceptional regions of the genome. The identification and characterization of these highly duplicated regions represents an important step in the complete sequencing of a human reference genome.

Base Sequence↗

Molecular evidence for a relationship between LINE-1 elements and X chromosome inactivation: the Lyon repeat hypothesis.

X inactivation is a chromosome-specific form of genetic regulation in which thousands of genes on one homologue become silenced early in female embryogenesis. Although many aspects of X inactivation are now understood, the spread of the X inactivation signal along the entire length of the chromosome remains enigmatic. Extending the Gartler-Riggs model [Gartler, S. M. & Riggs, A. D. (1983) Annu. Rev. Genet. 17, 155-190], Lyon recently proposed [Lyon, M. F. (1998) Cytogenet. Cell Genet. 80, 133-137] that a nonrandom organization of long interspersed element (LINE) repetitive sequences on the X chromosome might be responsible for its facultative heterochromatization. In this paper, we present data indicating that the LINE-1 (L1) composition of the human X chromosome is fundamentally distinct from that of human autosomes. The X chromosome is enriched 2-fold for L1 repetitive elements, with the greatest enrichment observed for a restricted subset of LINE-1 elements that were active <100 million years ago. Regional analysis of the X chromosome reveals that the most significant clustering of these elements is in Xq13-Xq21 (the center of X inactivation). Genomic segments harboring genes that escape inactivation are significantly reduced in L1 content compared with X chromosome segments containing genes subject to X inactivation, providing further support for the association between X inactivation and L1 content. These nonrandom properties of L1 distribution on the X chromosome provide strong evidence that L1 elements may serve as DNA signals to propagate X inactivation along the chromosome.

Base Pairing↗

Molecular structure and evolution of an alpha satellite/non-alpha satellite junction at 16p11.

We have determined the detailed molecular structure and evolution of an alpha satellite junction from human chromosome 16p11. The analysis reveals that the alpha satellite sequence bordering the transition lacks higher-order structure and that the non-alpha satellite portion consists of a mosaic of duplicated segments of complex evolutionary origin. The 16p11 junction was formed recently (5-10 million years ago) by the duplication and transposition of genomic segments from Xq28 and 4q24. Once this mosaic structure was formed, a larger complex was spread among multiple pericentromeric regions. This resulted in the formation of large (>62 kb) paralogous segments that share a high degree ( approximately 97%) of sequence similarity. Both phylogenetic and comparative analyses indicate that these pericentromeric-directed duplications occurred around the time of the divergence of the human, gorilla and chimpanzee lineages, resulting in the subtle restructuring of the primate genome among these species. The available data suggest that such chimeric structures are a general property of several different human chromosomes near their alpha satellite junctions.

Animals↗

Structure of chromosomal duplicons and their role in mediating human genomic disorders.

Chromosome-specific low-copy repeats, or duplicons, occur in multiple regions of the human genome. Homologous recombination between different duplicon copies leads to chromosomal rearrangements, such as deletions, duplications, inversions, and inverted duplications, depending on the orientation of the recombining duplicons. When such rearrangements cause dosage imbalance of a developmentally important gene(s), genetic diseases now termed genomic disorders result, at a frequency of 0.7-1/1000 births. Duplicons can have simple or very complex structures, with variation in copy number from 2 to >10 repeats, and each varying in size from a few kilobases in length to hundreds of kilobases. Analysis of the different duplicons involved in human genomic disorders identifies features that may predispose to recombination, including large size and high sequence identity between the recombining copies, putative recombination promoting features, and the presence of multiple genes/pseudogenes that may include genes expressed in germ cells. Most of the chromosome rearrangements involve duplicons near pericentromeric regions, which may relate to the propensity of such regions to accumulate duplicons. Detailed analyses of the structure, polymorphic variation, and mechanisms of recombination in genomic disorders, as well as the evolutionary origin of various duplicons will further our understanding of the structure, function, and fluidity of the human genome.

Animals↗

The mosaic structure of human pericentromeric DNA: a strategy for characterizing complex regions of the human genome.

The pericentromeric regions of human chromosomes pose particular problems for both mapping and sequencing. These difficulties are due, in large part, to the presence of duplicated genomic segments that are distributed among multiple human chromosomes. To ensure contiguity of genomic sequence in these regions, we designed a sequence-based strategy to characterize different pericentromeric regions using a single (162 kb) 2p11 seed sequence as a point of reference. Molecular and cytogenetic techniques were first used to construct a paralogy map that delineated the interchromosomal distribution of duplicated segments throughout the human genome. Monochromosomal hybrid DNAs were PCR amplified by primer pairs designed to the 2p11 reference sequence. The PCR products were directly sequenced and used to develop a catalog of sequence tags for each duplicon for each chromosome. A total of 685 paralogous sequence variants were generated by sequencing 34.7 kb of paralogous pericentromeric sequence. Using PCR products as hybridization probes, we were able to identify 702 human BAC clones, of which a subset, 107 clones, were analyzed at the sequence level. We used diagnostic paralogous sequence variants to assign 65 of these BACs to at least 9 chromosomal pericentromeric regions: 1q12, 2p11, 9p11/q12, 10p11, 14q11, 15q11, 16p11, 17p11, and 22q11. Comparisons with existing sequence and physical maps for the human genome suggest that many of these BACs map to regions of the genome with sequence gaps. Our analysis indicates that large portions of pericentromeric DNA are virtually devoid of unique sequences. Instead, they consist of a mosaic of different genomic segments that have had different propensities for duplication. These biologic properties may be exploited for the rapid characterization of, not only pericentromeric DNA, but also other complex paralogous regions of the human genome.

Centromere↗

Genome duplications and other features in 12 Mb of DNA sequence from human chromosome 16p and 16q.

Several publicly funded large-scale sequencing efforts have been initiated with the goal of completing the first reference human genome sequence by the year 2005. Here we present the results of analysis of 11.8 Mb of genomic sequence from chromosome 16. The apparent gene density varies throughout the region, but the number of genes predicted (84) suggests that this is a gene-poor region. This result may also suggest that the total number of human genes is likely to be at the lower end of published estimates. One of the most interesting aspects of this region of the genome is the presence of highly homologous, recently duplicated tracts of sequence distributed throughout the p-arm. Such duplications have implications for mapping and gene analysis as well as the predisposition to recurrent chromosomal structural rearrangements associated with genetic disease.

Animals↗