PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “DNA Copy Number Variations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 865 records · Page 48Linked to original sources

Large scale copy number variation (CNV) at 14q12 is associated with the presence of genomic abnormalities in neoplasia.

BACKGROUND: Advances made in the area of microarray comparative genomic hybridization (aCGH) have enabled the interrogation of the entire genome at a previously unattainable resolution. This has lead to the discovery of a novel class of alternative entities called large-scale copy number variations (CNVs). These CNVs are often found in regions of closely linked sequence homology called duplicons that are thought to facilitate genomic rearrangements in some classes of neoplasia. Recently, it was proposed that duplicons located near the recurrent translocation break points on chromosomes 9 and 22 in chronic myeloid leukemia (CML) may facilitate this tumor-specific translocation. Furthermore, approximately 15-20% of CML patients also carry a microdeletion on the derivative 9 chromosome (der(9)) and these patients have a poor prognosis. It has been hypothesised that der(9) deletion patients have increased levels of chromosomal instability. RESULTS: In this study aCGH was performed and identified a CNV (RP11-125A5, hereafter called CNV14q12) that was present as a genomic gain or loss in 10% of control DNA samples derived from cytogenetically normal individuals. CNV14q12 was the same clone identified by Iafrate et al. as a CNV. Real-time polymerase chain reaction (Q-PCR) was used to determine the relative frequency of this CNV in DNA from a series of 16 CML patients (both with and without a der(9) deletion) together with DNA derived from 36 paediatric solid tumors in comparison to the incidence of CNV in control DNA. CNV14q12 was present in approximately 50% of both tumor and CML DNA, but was found in 72% of CML bearing a der(9) microdeletion. Chi square analysis found a statistically significant difference (p <or= 0.001) between the incidence of this CNV in cancer and normal DNA and a slightly increased incidence in CML with deletions in comparison to those CML without a detectable deletion. CONCLUSION: The increased incidence of CNV14q12 in tumor samples suggests that either acquired or inherited genomic variation of this new class of variation may be associated with onset or progression of neoplasia.

Child↗

Identifying uniformly mutated segments within repeats.

Given a long string of characters from a constant size alphabet we present an algorithm to determine whether its characters have been generated by a single i.i.d. random source. More specifically, consider all possible n-coin models for generating a binary string S, where each bit of S is generated via an independent toss of one of the n coins in the model. The choice of which coin to toss is decided by a random walk on the set of coins where the probability of a coin change is much lower than the probability of using the same coin repeatedly. We present a procedure to evaluate the likelihood of a n-coin model for given S, subject a uniform prior distribution over the parameters of the model (that represent mutation rates and probabilities of copying events). In the absence of detailed prior knowledge of these parameters, the algorithm can be used to determine whether the a posteriori probability for n=1 is higher than for any other n>1. Our algorithm runs in time O(l4logl), where l is the length of S, through a dynamic programming approach which exploits the assumed convexity of the a posteriori probability for n. Our test can be used in the analysis of long alignments between pairs of genomic sequences in a number of ways. For example, functional regions in genome sequences exhibit much lower mutation rates than non-functional regions. Because our test provides means for determining variations in the mutation rate, it may be used to distinguish functional regions from non-functional ones. Another application is in determining whether two highly similar, thus evolutionarily related, genome segments are the result of a single copy event or of a complex series of copy events. This is particularly an issue in evolutionary studies of genome regions rich with repeat segments (especially tandemly repeated segments).

Algorithms↗

A survey of homozygous deletions in human cancer genomes.

Homozygous deletions of recessive cancer genes and fragile sites are known to occur in human cancers. We identified 281 homozygous deletions in 636 cancer cell lines. Of these deletions, 86 were homozygous deletions of known recessive cancer genes, 17 were of sequenced common fragile sites, and 178 were in genomic regions that do not overlap known recessive oncogenes or fragile sites ("unexplained" homozygous deletions). Some cancer cell lines have multiple homozygous deletions whereas others have none, suggesting intrinsic variation in the tendency to develop this type of genetic abnormality (P < 0.001). The 178 unexplained homozygous deletions clustered into 131 genomic regions, 27 of which exhibit homozygous deletions in more than one cancer cell line. This degree of clustering indicates that the genomic positions of the unexplained homozygous deletions are not randomly determined (P < 0.001). Many homozygous deletions, including those that are in multiple clusters, do not overlap known genes and appear to be in intergenic DNA. Therefore, to elucidate further the pathogenesis of homozygous deletions in cancer, we investigated the genome landscape within unexplained homozygous deletions. The gene count within homozygous deletions is low compared with the rest of the genome. There are also fewer short interspersed nuclear elements (SINEs), long interspersed nuclear elements (LINEs), and low-copy-number repeats (LCRs). However, DNA within homozygous deletions has higher flexibility. These features may signal the presence of currently unrecognized zones of susceptibility to DNA rearrangement. They may also reflect a tendency to reduce the adverse effects of homozygous deletions by minimizing the number of genes removed.

Cell Line, Tumor↗

Ribosomal DNA in the grasshopper Podisma pedestris: escape from concerted evolution.

Eukaryote nuclear ribosomal DNA (rDNA) typically exhibits strong concerted evolution: a pattern in which several hundred rDNA sequences within any one species show little or no genetic diversity, whereas the sequences of different species diverge. We report a markedly different pattern in the genome of the grasshopper Podisma pedestris. Single individuals contain several highly divergent ribosomal DNA groups. Analysis of the magnitude of divergence indicates that these groups have coexisted in the Podisma lineage for at least 11 million years. There are two putatively functional groups, each estimated to be at least 4 million years old, and several pseudogene groups, many of which are transcribed. Southern hybridization and real-time PCR experiments show that only one of the putatively functional types occurs at high copy number. However, this group is scarcely amplified under standard PCR conditions, which means that phylogenetic inference on the basis of standard PCR would be severely distorted. The analysis suggests that concerted evolution has been remarkably ineffective in P. pedestris. We propose that this outcome may be related to the species' exceptionally large genome and the associated low rate of deletion per base pair, which may allow pseudogenes to persist.

Animals↗

Variant mapping of the Apo(B) AT rich minisatellite. Dependence on nucleotide sequence of the copy number variations. Instability of the non-canonical alleles.

Because of its variations in length, the AT rich Hyper-Variable Region (HVR) of the 3' end of the Apolipoprotein B gene is used as a polymorphic maker in genetic studies. It contains a SspI site in its repeated motif and we used this feature to precisely analyse the internal structure of the different alleles found at this locus in a Caucasian population. We performed total digestion on 194 alleles as well as Minisatellite Variant Repeat mapping (MVR mapping: partial digestion) on 54. The results show that the level of length variability (in copy number) of the 5' end of this locus is at least two times higher than that of the 3' end. This could be correlated with the difference in nucleotide sequence between the two parts of the HVR and suggests the dependence on the primary structure of the mechanism that produces length variability. A molecular model is proposed to explain this result. Moreover, the sharp analysis of the minisatellite structure by the distribution of SspI sites reveals differences between long and short alleles, indicating that in most cases, no recombination occurs between alleles of different sizes. Finally the rare alleles exhibit a non-canonical structure. These important points could explain the bimodal distribution of the frequencies of the alleles in the population.

Alleles↗

Reading between the LINEs: human genomic variation induced by LINE-1 retrotransposition.

The insertion of mobile elements into the genome represents a new class of genetic markers for the study of human evolution. Long interspersed elements (LINEs) have amplified to a copy number of about 100,000 over the last 100 million years of mammalian evolution and comprise approximately 15% of the human genome. The majority of LINE-1 (L1) elements within the human genome are 5' truncated copies of a few active L1 elements that are capable of retrotransposition. Some of the young L1 elements have inserted into the human genome so recently that populations are polymorphic for the presence of an L1 element at a particular chromosomal location. L1 insertion polymorphisms offer several advantages over other types of polymorphisms for human evolution studies. First, they are typed by rapid, simple, polymerase chain reaction (PCR)-based assays. Second, they are stable polymorphisms that rarely undergo deletion. Third, the presence of an L1 element represents identity by descent, because the probability is negligible that two different young L1 repeats would integrate independently between the exact same two nucleotides. Fourth, the ancestral state of L1 insertion polymorphisms is known to be the absence of the L1 element, which can be used to root plots/trees of population relationships. Here we report the development of a PCR-based display for the direct identification of dimorphic L1 elements from the human genome. We have also developed PCR-based assays for the characterization of six polymorphic L1 elements within the human genome. PCR analysis of human/rodent hybrid cell line DNA samples showed that the polymorphic L1 elements were located on several different chromosomes. Phylogenetic analysis of nonhuman primate DNA samples showed that all of the recently integrated "young" L1 elements were restricted to the human genome and absent from the genomes of nonhuman primates. Analysis of a diverse array of human populations showed that the allele frequencies and level of heterozygosity for each of the L1 elements was variable. Polymorphic L1 elements represent a new source of identical-by-descent variation for the study of human evolution. [The sequence data described in this paper have been submitted to the GenBank data library under accession nos. AF242435-AF242451.]

Animals↗

A multicopy plasmid of the extremely thermophilic archaeon Sulfolobus effects its transfer to recipients by mating.

A plasmid of 45 kb, designated pNOB8, was found in high copy number in a new heterotrophic Sulfolobus isolate, NOB8H2, from Japan. Dissemination of the plasmid occurred in six cultures of nine different Sulfolobus strains when small amounts of the donor were added. These mixed cultures exhibited a high average copy number of the plasmid, between 20 and 40 per chromosome, and showed a marked growth retardation. Horizontal transfer of pNOB8 was proved by isolating transcipients from mating mixtures via single colonies. In these isolates, the copy number of the plasmid appeared to be subject to a control mechanism. Cell-free filtrates of donor cultures did not transmit the plasmid, and plating of the donor on lawns of recipients did not result in plaque formation, suggesting that the transfer was not mediated by a virus. Rapid formation of cell-to-cell contacts between differently stained donor and recipient partners was demonstrated after the two strains were mixed. Electron microscopic analysis of mating mixtures revealed many cell aggregates made up of 2 to 30 cells and intercellular cytoplasmic bridges connecting two or more cells. Cells that had been transformed with purified plasmid DNA as well as transcipients isolated from mating mixtures were shown to serve as donors for further transmission of pNOB8. The plasmid undergoes extensive genetic variations, since deletions and insertions were frequently observed in plasmid preparations from the donor strain and from mating mixtures.

Cell Aggregation↗

Organisation and molecular analysis of repeated DNA sequences in the rice blast fungus Magnaporthe grisea.

The distribution of a previously described repeated DNA sequence present as a 1.3-kb PstI fragment in the genome of the rice blast fungus Magnaporthe grisea was analysed by carrying out DNA fingerprint analysis of 36 isolates including rice, non-rice and laboratory strains. The analysis of various higher-molecular-weight PstI fragments with homology to the 1.3-kb repeat revealed that these may arise predominantly from transposon insertions or point mutations. Analysis of a 5.1-kb derivative revealed both a point mutation at a PstI site and an insertion of a putative transposable element which caused an increase in molecular weight from 1.3 to 5.1 kb. Another repeat element of 1.4 kb was identified and found to exist in association with the 1.3-kb repeat. Both 1.3- and 1.4-kb elements were found to be parts of MGR583 (Hamer et al. 1989), a LINE-like element. These elements were present in a high copy number in all the rice and a majority of non-rice pathogens indicating that MGR583 is not a host-specific sequence as reported earlier. Our results suggest that repeated DNA elements in M. grisea have amplified independently of one another and further indicate that different isolates of M. grisea may have evolved from several distinct lines of origin.

Ascomycota↗

Molecular genetic studies of tumor suppressor gene regions on chromosomes 13 and 17 in colorectal tumors.

BACKGROUND: In the majority of colorectal carcinomas, both copies of the tumor suppressor gene TP53 (tumor protein 53) are known to be inactivated. In contrast to a loss of tumor suppressor function, it has been suggested that an increased copy number of the RB1 gene is involved in the progression of these tumors. PURPOSE: To determine genetic alterations at chromosomes 13 and 17 in colorectal tumors, we have studied several loci on these chromosomes, with special focus on the RB1 and TP53 genes at both the level of DNA sequence and the level of gene expression. METHODS: Restriction fragment length polymorphism analysis was performed after alkaline Southern blotting of the DNA fragments and hybridization (in 7% sodium dodecyl sulfate and 0.5 M NaPO4) of the nylon membranes with multiprimed, radioactively labeled probes. Total RNA was extracted from tissue biopsy specimens by homogenization of the samples in guanidinium thiocyanate followed by separation in a CsCl gradient. By use of an image-processing system, x-ray film signals were measured densitometrically. Point mutations within the TP53 gene were detected by use of polymerase chain reaction (PCR) in combination with constant denaturant gel electrophoresis. Direct sequencing of PCR products revealed the exact nature of the mutations. Protein expression of TP53 was seen by immunostaining of sections from paraffin-embedded material using a mouse monoclonal antibody. The two-sided Fisher's Exact Test was used for statistical analysis. RESULTS: An increase in allelic copy number at 13q loci was seen in 10 (32%) of 31 tumors. In the majority of the cases, this increase probably reflected a change in the diploid status of chromosome 13; in some cases, however, only part of the 13q seemed to be involved. The RB1 gene showed an elevated level of RNA compared with the beta-actin signal. Fourteen (48%) of 29 tumors showed loss of heterozygosity at loci on 17p, and base mutations within the TP53 gene were seen in 14 (42%) of 33 tumors. RNA and protein analyses of TP53 revealed an increased level of expression in the tumors compared with normal mucosa. Allelic variations seen at 13q and 17p were not associated (P = .7). CONCLUSIONS: Our results suggest that, in addition to aneuploidy, gain of specific chromosome 13 sequences is involved in the tumorigenesis of the colon and rectum. In addition, they confirm the importance of TP53 mutations for the progression of such tumors and support the view that accumulation of events is more important than the order of events. The genetic changes observed at chromosome arms 13q and 17p seem to be independent of each other.

Adolescent↗

Analysis of the Piv recombinase-related gene family of Neisseria gonorrhoeae.

Neisseria gonorrhoeae (the gonococcus) is an obligate human pathogen and the causative agent of the disease gonorrhea. The gonococcal pilus undergoes antigenic variation through high-frequency recombination events between unexpressed pilS silent copies and the pilin expression locus pilE. The machinery involved in pilin antigenic variation identified to date is composed primarily of genes involved in homologous recombination. However, a number of characteristics of antigenic variation suggest that one or more recombinases, in addition to the homologous recombination machinery, may be involved in mediating sequence changes at pilE. Previous work has identified several genes in the gonococcus with significant identity to the pilin inversion gene (piv) from Moraxella species and transposases of the IS110 family of insertion elements. These genes were candidates for a recombinase system involved in pilin antigenic variation. We have named these genes irg for invertase-related gene family. In this work, we characterize these genes and demonstrate that the irg genes do not complement for Moraxella lacunata Piv invertase or IS492 MooV transposase activities. Moreover, by inactivation of all eight gene copies and overexpression of one gene copy, we conclusively show that these recombinases are not involved in gonococcal pilin variation, DNA transformation, or DNA repair. We propose that the irg genes encode transposases for two different IS110-related elements given the names ISNgo2 and ISNgo3. ISNgo2 is located at multiple loci on the chromosome of N. gonorrhoeae, and ISNgo3 is found in single and duplicate copies in the N. gonorrhoeae and Neisseria meningitidis genomes, respectively.

Amino Acid Sequence↗

Structure and evolution of the mitochondrial control region and flanking sequences in the European cave salamander Proteus anguinus.

The European cave salamander Proteus anguinus Laurenti 1768 is one of the best-known subterranean animals, yet its evolutionary history and systematic relationships remain enigmatic. This is the first comprehensive study on molecular evolution within the taxon, using an mtDNA segment containing the control region (CR) and adjacent sequences. Two to seven tandem repeats of 24-32 bp were found in the intergenic spacer region (VNTR1), and three, four or six repeats, 59-77 bp each, in the 3' end of the CR (VNTR2). Different molecular mechanisms account for VNTR2 formation in different lineages of Proteus. The overall CR variation was lower than that of the spacer region, the 3' end of the cytb gene, or the tRNA genes. Individual genes and the concatenated non-repetitive sequences produced similar, well resolved maximum likelihood, Bayesian inference and parsimony trees. The numbers of repeat elements as well as the genealogy of the VNTR2 repeat units were mostly inconsistent with the groupings of the non-repetitive sequences. Different degrees of repeat array homogenization were detected in all major groups. Orthology was established for the first and the second VNTR2 elements of some populations. These two copies may therefore be used for analyses at the population level. The pattern of CR sequence variation points to strong genetic isolation of hydrographically separated populations. Genetic separation of the major groups of populations is incongruent with the current division into subspecies.

Animals↗

Integrating molecular subtypes, genomics and functional dependencies to identify context-specific therapeutic vulnerabilities in small cell lung cancer.

Small cell lung cancer is one of the most aggressive malignancies, characterized by rapid tumor growth, early metastatic spread and extremely poor survival. Although most patients initially respond to platinum-based chemotherapy, relapse is almost inevitable and treatment options at recurrence remain limited. The recent introduction of immune checkpoint inhibitors has provided only modest clinical benefit, largely due to the fact that these tumors are immunologically cold. These limitations highlight the urgent need to better understand the molecular features of small cell lung cancer in order to identify more effective therapeutic strategies. In this review, we summarize current knowledge of the molecular landscape of small cell lung cancer, with particular emphasis on transcriptome-based classifications that have identified four major molecular subtypes defined by distinct transcriptional regulators and gene expression programs. We discuss how these classifications have improved the biological understanding of the disease and stimulated efforts to develop subtype-specific therapeutic strategies. At the same time, we highlight important limitations of this framework, including the remarkable transcriptional plasticity of tumor cells, which allows dynamic transitions between subtypes and may contribute to therapeutic resistance. To address these challenges, we examine additional molecular features that may represent more stable vulnerabilities, including recurrent genomic alterations, such as the widespread loss of tumor suppressor genes or oncogene amplifications through extrachromosomal DNA. We also discuss emerging approaches aimed at identifying novel context-specific cancer dependencies, including genome-scale functional screens in vitro and in vivo and genetic restraint analyses. Finally, we consider the growing potential of liquid biopsy strategies, which exploit the high level of circulating tumor DNA in patients with this disease to detect clinically relevant genomic alterations and monitor tumor evolution. Overall, this review highlights both the opportunities and challenges associated with molecular stratification in small cell lung cancer. The integration of transcriptional classifications with genomic and functional approaches may help identify more robust therapeutic vulnerabilities and guide the development of more effective treatments for this highly aggressive disease.

Cancer vulnerabilities↗

A statistical method to detect chromosomal regions with DNA copy number alterations using SNP-array-based CGH data.

Single nucleotide polymorphism (SNP) arrays were used to detect chromosomal regions with DNA copy number alterations. Current statistical methods for microarray-based comparative genomic hybridization (array-CGH) analysis generally assume certain relationships among adjacent markers on the same chromosome, and these assumptions may be questionable. For an SNP-array-based CGH study, multiple normal reference SNP arrays were collected. In order to utilize these normal reference SNP arrays, we derived an empirical distribution of signal ratios for each SNP marker. With an assumed threshold value for the overall error rate control and the defined signal ratio ranges for chromosomal amplification and deletion, we proposed a procedure to identify chromosomal alteration regions based on several bootstrapped one-sample t-tests and the false discovery rate control. When we have multiple arrays for different individuals with the same disease, our method can also be used to detect SNP markers for chromosomal alteration regions that are common among these individuals. We applied our method to a published SNP array data set for breast carcinoma cell lines. For an individual with breast cancer, numerous chromosomal alteration regions were identified. Compared to results of previous studies, our method identified more chromosomal alteration regions, with some being implicated in the literature to harbor genes associated with breast cancer. For multiple cancer arrays, our results suggested the existence of common chromosomal alteration regions. However, a high proportion of false positives also indicated that genetic variations among different individuals with breast cancer can be present.

Breast Neoplasms↗

A cDNA-based comparison of dehydration-induced proteins (dehydrins) in barley and corn.

Several cDNAs related to an ABA-induced cDNA from barley aleurone were isolated from barley and corn seedlings that were undergoing dehydration. Four different barley polypeptides with sizes of 22.6, 16.2, 14.4 and 14.2 kDa and a single corn polypeptide with a size of 17.0 kDa were predicted from the nucleotide sequences of the cDNAs. These dehydration-induced proteins (dehydrins) are very similar to each other and to a previously identified rice protein induced by ABA and salt, and have at least some similarity to a previously identified cotton embryo protein. Each dehydrin is extremely hydrophilic, glycine-rich, cysteine- and tryptophan-free and contains repeated units in a conserved linear order. A lysine-rich repeating unit occurs twice in each protein, once at the carboxy terminus and once partway through the polypeptide, adjacent to a succession of serines. This repeating unit and the adjacent flanking run of serines are conserved with minimal variation among all dehydrins. Another repeating unit is flanked by the two copies of the lysine-rich unit, and varies in number from one to five copies. This latter repeating unit is less conserved than the former, varying even within a singly dehydrin. The messenger RNAs corresponding to each cDNA are abundant in dehydrating, but not in well-watered seedlings. The amino acid sequence of tryptic peptides from purified dehydration-induced proteins of corn established that the corn cDNAs correspond to a protein that is produced in abundance during the response of corn seedlings to dehydration.

Amino Acid Sequence↗

An investigation of the effect of antisense RNA gene on bovine leukaemia virus reproduction in cell culture.

A possible approach to control of bovine lymphoproliferative disease caused by bovine leukaemia virus (BLV) may be the development of an "antiviral information immunity" based on the effect of anti-sense RNA (asRNA). A numbers of constructs were obtained, under control of various promotors (herpesvirus thymidine kinase, T-antigen SV40 promoter), carrying as DNA against gene X, the expression product of which is a transactivator of viral transcription from the BLV LTR promotor. As a model system for the analysis of antiviral activity of constructs developed, cloned continuous cell lines of BLV-producing FLK cells were used. The level of BLV expression in cells transfected with the constructs was determined by various parameters. Differences were detected in different clones obtained from non-transfected cells, as well as variation between transfected clones, as measured by reverse transcriptase, competitive radio-immunoassay for BLV p24, the viral particle count on agar membrane, and the tumorigenicity for nude mice. The differences in inhibition of expression of BLV genes and their products may be explained in terms of the site of integration of asDNA and the number of integrated copies.

Animals↗

Cell-specific ribosomal DNA spacer variability in human urothelial carcinoma cultures.

Length variation of a ribosomal DNA "spacer" region in four chromosomally characterized transitional cell carcinoma cultures was analyzed by restriction endonuclease cleavage and Southern blotting. Cell lines with relative karyotypic conservation, such as UM-UC-2 (modal chromosome number 48, four marker chromosomes) demonstrate little change in the genetically regulated pattern of rDNA spacer length polymorphisms (7.6, 6.7 and 6.0 kilobases) which may be found in normal cells. Cell lines with more aberrant karyotypes, such as UM-UC-3 (modal chromosome number 86, 12 marker chromosomes) and UM-UC-4 (modal number 51, ten marker chromosomes) show fewer ribosomal DNA length variants (7.6, 6.7 kilobases for the former, 7.6 kilobases for the latter), consistent with relaxed constraints on the drive for ribosomal gene homogeneity through inter and intrachromosomal exchange. Uncharacterized rDNA length variants of low copy number were observed in cell lines with many marker chromosomes. Analysis of repetitive DNA structure provides an additional criterion for tumor diagnosis and staging, and a characterized series of tumor cell lines may provide a useful system for understanding repetitive DNA evolution.

Carcinoma, Transitional Cell↗

Error-prone replication for better or worse.

Precise genome duplication requires accurate copying by DNA polymerases and the elimination of occasional mistakes by proofreading exonucleases and mismatch repair enzymes. The commonly held belief that 'if something is worth doing, then it's worth doing well' normally applies to DNA replication and repair, however, there are exceptions. This review describes elements that are crucial to cell fitness, evolution and survival in the recently discovered error-prone DNA polymerases. Large numbers of errant DNA polymerases, spanning microorganisms to humans, are used to rescue stalled replication forks by copying damaged DNA and even undamaged DNA to generate 'purposeful' mutations that generate genetic diversity in times of stress. Here we focus on low-fidelity polymerases from bacteria, comparing Escherichia coli, archeabacteria and those most recently discovered in Gram-positive Bacilli, Streptococcus, pathogenic Mycobacterium and intein-containing cyanobacteria.

Adaptation, Biological↗

Long contiguous stretches of homozygosity in the human genome.

Single nucleotide polymorphisms (SNPs) are the most common sequence variation in the human genome; they have been successfully used in mapping disease genes and more recently in studying population genetics and cancer genetics. In a population-based association study using high-density oligonucleotide arrays for whole-genome SNP genotyping, we discovered that in the genomes of unrelated Han Chinese, 34 out of 515 (6.6%) individuals contained long contiguous stretches of homozygosity (LCSHs), ranging in the size from 2.94 to 26.27 Mbp (10.22+/-5.95 Mbp). Four out of four (100%) Taiwan aborigines also demonstrated this genetic characteristic. The number of LCSH regions increased markedly in the offspring of consanguineous marriages. LCSH was also detected in Caucasian samples (11/42; 26.2%) and African American samples (2/42; 4.76%). A total of 26 LCSH regions were recurrently detected among Han Chinese, Taiwan aborigines, and Caucasians. DNA copy number determination by hybridization intensity analysis and real-time quantitative PCR (qPCR) excluded deletion as the cause of LCSH. Our results suggest that LCSHs are common in the human genome of the outbred population and this genetic characteristic could have a significant impact on population genetics and disease gene studies.

Black or African American↗