PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Copy Number Variations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Mapping of DNA instability at the fragile X to a trinucleotide repeat sequence p(CCG)n.

The sequence of a Pst I restriction fragment was determined that demonstrate instability in fragile X syndrome pedigrees. The region of instability was localized to a trinucleotide repeat p(CCG)n. The sequence flanking this repeat were identical in normal and affected individuals. The breakpoints in two somatic cell hybrids constructed to break at the fragile site also mapped to this repeat sequence. The repeat exhibits instability both when cloned in a nonhomologous host and after amplification by the polymerase chain reaction. These results suggest variation in the trinucleotide repeat copy number as the molecular basis for the instability and possibly the fragile site. This would account for the observed properties of this region in vivo and in vitro.

Base Sequence↗

Organization and evolution of highly repeated satellite DNA sequences in plant chromosomes.

A major component of the plant nuclear genome is constituted by different classes of repetitive DNA sequences. The structural, functional and evolutionary aspects of the satellite repetitive DNA families, and their organization in the chromosomes is reviewed. The tandem satellite DNA sequences exhibit characteristic chromosomal locations, usually at subtelomeric and centromeric regions. The repetitive DNA family(ies) may be widely distributed in a taxonomic family or a genus, or may be specific for a species, genome or even a chromosome. They may acquire large-scale variations in their sequence and copy number over an evolutionary time-scale. These features have formed the basis of extensive utilization of repetitive sequences for taxonomic and phylogenetic studies. Hybrid polyploids have especially proven to be excellent models for studying the evolution of repetitive DNA sequences. Recent studies explicitly show that some repetitive DNA families localized at the telomeres and centromeres have acquired important structural and functional significance. The repetitive elements are under different evolutionary constraints as compared to the genes. Satellite DNA families are thought to arise de novo as a consequence of molecular mechanisms such as unequal crossing over, rolling circle amplification, replication slippage and mutation that constitute "molecular drive".

Centromere↗

Phenotypic and genotypic characteristics of recently adapted isolates of Plasmodium falciparum from Thailand.

The drug sensitivity characteristics and Plasmodium falciparum pfmdr1 status of five isolates of P. falciparum recently isolated from patients presenting for treatment from the Thailand/Myanmar border have been investigated. The aim of the study was to avoid the criticisms of some earlier studies by focusing on newly collected isolates from a specific geographic location. Three of the isolates studied exhibited clear resistance to chloroquine similar to that observed in the K1 Thai standard isolate obtained in the 1970s, and the other two isolates were of intermediate sensitivity to chloroquine with concentrations of drug that inhibit parasite growth by 50% of 50 and 43 nmol. The sensitivity of all isolates was enhanced by verapamil but we found no clear association between chloroquine sensitivity and gene copy number or intra-allelic variation of pfmdr1. In contrast, clear cross-resistance was seen between mefloquine and halofantrine, with the most sensitive isolates carrying the K1 mutation in pfmdr1.

Animals↗

Antithrombin cambridge II (Ala384Ser): clinical, functional and haplotype analysis of 18 families.

Thirty-one individuals from 18 unrelated families with antithrombin deficiency have been identified as having a single point mutation within codon 384 (13268 GCA-->TCA) resulting in an alanine to serine substitution. Six families (11 individuals) were identified by the screening of individuals with thromboembolic disease or with a family history of thromboembolic disease, whilst the remaining 12 families (20 individuals) were identified by screening of asymptomatic blood donors. Four individuals had a history of venous thrombotic disease, a further 2 gave a history of superficial thrombophlebitis but the remaining 25 individuals were asymptomatic. Affected individuals demonstrated normal immunological levels of antithrombin but a decrease in anti-IIa activity in the presence of heparin. Haplotype analysis was used to examine the possibility of a founder effect to explain the high frequency of this non-CpG mutation. 29/31 individuals showed a single common "core" haplotype, the only variation existing in the number of copies of an (ATT)n repeat polymorphism--13, 14, 15 or 17. The results suggest that at most there are four independent origins for this mutation.

Alanine↗

Organisation of the pericentromeric region of chromosome 15: at least four partial gene copies are amplified in patients with a proximal duplication of 15q.

Clinical cytogenetic laboratories frequently identify an apparent duplication of proximal 15q that does not involve probes within the PWS/AS critical region and is not associated with any consistent phenotype. Previous mapping data placed several pseudogenes, NF1, IgH D/V, and GABRA5 in the pericentromeric region of proximal 15q. Recent studies have shown that these pseudogene sequences have increased copy numbers in subjects with apparent duplications of proximal 15q. To determine the extent of variation in a control population, we analysed NF1 and IgH D pseudogene copy number in interphase nuclei from 20 cytogenetically normal subjects by FISH. Both loci are polymorphic in controls, ranging from 1-4 signals for NF1 and 1-3 signals for IgH D. Eight subjects with apparent duplications, examined by the same method, showed significantly increased NF1 copy number (5-10 signals). IgH D copy number was also increased in 6/8 of these patients (4-9 signals). We identified a fourth pseudogene, BCL8A, which maps to the pericentromeric region and is coamplified along with the NF1 sequences. Interphase FISH ordering experiments show that IgH D lies closest to the centromere, while BCL8A is the most distal locus in this pseudogene array; the total size of the amplicon is estimated at approximately 1 Mb. The duplicated chromosome was inherited from either sex parent, indicating no parent of origin effect, and no consistent phenotype was present. FISH analysis with one or more of these probes is therefore useful in discriminating polymorphic amplification of proximal pseudogene sequences from clinically significant duplications of 15q.

Adult↗

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans↗

Restriction fragment length polymorphisms among uropathogenic Escherichia coli isolates: pap-related sequences compared with rrn operons.

Among the adhesin-encoding virulence operons associated with uropathogenic Escherichia coli, only pap (pyelonephritis-associated pilus)-related gene clusters typically exhibit variation in their structure and chromosomal copy number. To access further such variability, we compared pap restriction fragment length polymorphisms (RFLPs) with those detected among rRNA (rrn) operons, which encode an essential host function unrelated to virulence. To place such findings in a phylogenetic perspective, the E. coli isolates were also characterized by using multilocus enzyme electrophoresis. Variation in the rrn RFLP profiles correlated with evolutionary divergence resolved by multilocus enzyme electrophoresis; isolates with identical rrn profiles represented the same or closely related electrophoretic types. In contrast, such isolates frequently had different pap-related RFLPs, indicating that these genetic variations have developed recently relative to the changes associated with essential rrn operons or metabolic enzymes. Despite such fluctuations, two lines of evidence indicate conditions under which the pap-related RFLPs can be stably maintained. First, for each of 20 patients with urosepsis, both the primary urinary tract isolate and the concurrent blood isolate were identical. Second, although obtained from different patients, some isolates representing the same electrophoretic type also had identical pap-related RFLPs. Thus, the genotypic diversity of this virulence adhesin operon was not generated during the course of acute infection or during laboratory manipulations. Since fecal E. coli isolates frequently carry chromosomally encoded pap-related gene clusters, these findings suggest that the intra- and interchromosomal recombination events generating the polymorphisms associated with the pap-related sequences likely occur among the E. coli of the commensal reservoir.

Base Sequence↗

Chromosome-specific molecular organization of maize (Zea mays L.) centromeric regions.

A set of oat-maize chromosome addition lines with individual maize (Zea mays L.) chromosomes present in plants with a complete oat (Avena sativa L.) chromosome complement provides a unique opportunity to analyze the organization of centromeric regions of each maize chromosome. A DNA sequence, MCS1a, described previously as a maize centromere-associated sequence, was used as a probe to isolate cosmid clones from a genomic library made of DNA purified from a maize chromosome 9 addition line. Analysis of six cosmid clones containing centromeric DNA segments revealed a complex organization. The MCS1a sequence was found to comprise a portion of the long terminal repeats of a retrotransposon-like repeated element, termed CentA. Two of the six cosmid clones contained regions composed of a newly identified family of tandem repeats, termed CentC. Copies of CentA and tandem arrays of CentC are interspersed with other repetitive elements, including the previously identified maize retroelements Huck and Prem2. Fluorescence in situ hybridization revealed that CentC and CentA elements are limited to the centromeric region of each maize chromosome. The retroelements Huck and Prem2 are dispersed along all maize chromosomes, although Huck elements are present in an increased concentration around centromeric regions. Significant variation in the size of the blocks of CentC and in the copy number of CentA elements, as well as restriction fragment length variations were detected within the centromeric region of each maize chromosome studied. The different proportions and arrangements of these elements and likely others provide each centromeric region with a unique overall structure.

Base Sequence↗

Identification of microsatellite markers from Cicer reticulatum: molecular variation and phylogenetic analysis.

Microsatellite sequences were cloned and sequenced from Cicer reticulatum, the wild annual progenitor of chickpea (C. arietinum L.). Based on the flanking sequences of the microsatellite motifs, 11 sequence-tagged microsatellite site (STMS) markers were developed. These markers were used for phylogenetic analysis of 29 accessions representing all the nine annual Cicer species. The 11 primer pairs amplified distinct fragments in all the annual species demonstrating high levels of sequence conservation at these loci. Efficient marker transferability (97%) of the C. reticulatum STMS markers across other species of the genus was observed as compared to microsatellite markers from the cultivated species. Variability in the size and number of alleles was obtained with an average of 5.8 alleles per locus. Sequence analysis at three homologous microsatellite loci revealed that the microsatellite allele variation was mainly due to differences in the copy number of the tandem repeats. However, other factors such as (1) point mutations, (2) insertion/deletion events in the flanking region, (3) expansion of closely spaced microsatellites and (4) repeat conversion in the amplified microsatellite loci were also responsible for allelic variation. An unweighted pairgroup method with arithmetic averages (UPGMA)-based dendrogram was obtained, which clearly distinguished all the accessions (except two C. judaicum accessions) from one another and revealed intra- as well as inter-species variability in the genus. An annual Cicer phylogeny was depicted which established the higher similarity between C. arietinum and C. reticulatum. The placement of C. pinnatifidum in the second crossability group and its closeness to C. bijugum was supported. Two species, C. yamashitae and C. chorassanicum, were grouped distinctly and seemed to be genetically diverse from members of the first crossability group. Our data support the distinct placement of C. cuneatum as well as a revised classification regarding its placement.

Alleles↗

Ecological and evolutionary physiology of heat shock proteins and the stress response in Drosophila: complementary insights from genetic engineering and natural variation.

Classical adaptational and genetic engineering approaches offer complementary insights to understanding biological variation: the former elucidates the origins, magnitude and ecological context of natural variation, while the latter establishes which genes can underlie natural variation. Studies of the stress or heat shock response in Drosophila illustrate this point. At the cellular level, heat shock proteins (Hsps) function as molecular chaperones, minimizing aggregation of peptides in non-native conformations. To understand the adaptive significance of Hsps, we have characterized thermal stress that Drosophila experience in nature, which can be substantial. We used these findings to design ecologically relevant experiments with engineered Drosophila strains generated by unequal site-specific homologous recombination; these strains differ in hsp70 copy number but share sites of transgene integration. hsp70 copy number markedly affects Hsp70 levels in intact Drosophila, and strains with extra hsp70 copies exhibit corresponding differences in inducible thermotolerance and reactivation of a key enzyme after thermal stress. Elevated Hsp70 levels, however, are not without penalty; these levels retard growth and increase mortality. Transgenic variation in hsp70 copy number has counterparts in nature: isofemale lines from nature vary significantly in Hsp70 expression, and this variation is also correlated with both inducible thermotolerance and mortality in the absence of stress.

Animals↗

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans↗

Expression of vascular endothelial growth factor in uveal melanoma is independent of 6p21-region copy number.

PURPOSE: Overexpression of vascular endothelial growth factor (VEGF) and overrepresentation of the 6p region have been reported with a wide variation in uveal melanoma. The aim of the current study is to identify the frequency of copy number alteration in the 6p21 region and its correlation with the expression of VEGF in uveal melanoma. EXPERIMENTAL DESIGN: We studied 88 uveal melanomas for copy number change in the 6p region by comparative genomic hybridization and/or chromogenic in situ hybridization. Expression of VEGF protein was estimated by immunohistochemistry. In 15 tumors, VEGF mRNA expression was also studied by quantitative reverse transcription-PCR (RT-PCR) and VEGF splice variants were detected by RT-PCR. RESULTS: Copy number of the 6p21 region was successfully estimated in 37 tumors. In 10 (27%) of those, overrepresentation of the 6p21 region was detected. There was no statistically significant difference in VEGF expression between tumors with and without gain of 6p21 (P = 0.82). VEGF expression was not confined to the tumors and was also detected in the surrounding normal tissue. Expression of VEGF, detected by quantitative RT-PCR, was concordant with expression of VEGF protein. Different VEGF isoforms were expressed in different tumors with no obvious correlation with disease status. CONCLUSION: VEGF is overexpressed in a significant number of uveal melanomas. It should be noted that VEGF is not a candidate oncogene in uveal melanoma with 6p gain/amplification. VEGF overexpression other than structural amplification is probably significant in the pathogenesis of uveal melanomas, and its mechanism must be sought.

Alternative Splicing↗

A pseudolikelihood approach for simultaneous analysis of array comparative genomic hybridizations.

DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based comparative genomic hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure and with random effects to allow for intertumoral variation, as well as intratumoral clonal variation. For ease of computation, we base estimation on a pseudolikelihood function. The method produces quantitative assessments of the likelihood of genetic alterations at each clone, along with a graphical display for simple visual interpretation. We assess the characteristics of the method through simulation studies and analysis of a brain tumor aCGH data set. We show that the pseudolikelihood approach is superior to existing methods both in detecting small regions of copy number alteration and in accurately classifying regions of change when intratumoral clonal variation is present. Software for this approach is available at http://www.biostat.harvard.edu/ approximately betensky/papers.html.

Algorithms↗

Jointly analyzing gene expression and copy number data in breast cancer using data reduction models.

With the growing surge of biological measurements, the problem of integrating and analyzing different types of genomic measurements has become an immediate challenge for elucidating events at the molecular level. In order to address the problem of integrating different data types, we present a framework that locates variation patterns in two biological inputs based on the generalized singular value decomposition (GSVD). In this work, we jointly examine gene expression and copy number data and iteratively project the data on different decomposition directions defined by the projection angle theta in the GSVD. With the proper choice of theta, we locate similar and dissimilar patterns of variation between both data types. We discuss the properties of our algorithm using simulated data and conduct a case study with biologically verified results. Ultimately, we demonstrate the efficacy of our method on two genome-wide breast cancer studies to identify genes with large variation in expression and copy number across numerous cell line and tumor samples. Our method identifies genes that are statistically significant in both input measurements. The proposed method is useful for a wide variety of joint copy number and expression-based studies. Supplementary information is available online, including software implementations and experimental data.

Biomarkers, Tumor↗

Degeneracy in human multicopy RBM (YRRM), a candidate spermatogenesis gene.

In order to search for mutations in the multicopy RBM genes that might be associated with male infertility, we have used sequence data from the reported cDNA clone to determine the intron exon boundaries of the YRRM 1 gene. This gene has 12 exons, three of which encode the putative RNA binding domain of the protein. Different copies of the gene contain sequence variations and, additionally, give rise to transcripts with different numbers of copies of the repeated SRGY motif. Since mutations in the RNA binding domain would seem likely to have an effect on the activity of the protein, we have scanned these exons for mutations by SSCP on DNA from normal and infertile men. Sequence differences in the exon encoding the N-terminal part of the RNA binding domain account for at least four different classes of the gene and give rise to different SSCP conformers. Sequence analysis shows that one of these classes is a pseudogene and that the members of another class are nonfunctional. RT-PCR shows that all classes are transcribed and that the A class is most abundant. We have found a point mutation that alters the highly conserved RNP2 motif in one infertile patient. This mutation is also found in his father. We have used PCR followed by SSCP analysis to map RBM on a Y Chromosome (Chr) YAC contig and have demonstrated a distribution that spans a major part of this chromosome's euchromatin.

Amino Acid Sequence↗

Interphasic analysis of aneuploidy in cancer cell lines using primed in situ labeling.

The primed in situ (PRINS) labeling technique has been adapted to chromosomal screening of interphasic tumoral cells. A panel of ten chromosome-specific alpha-satellite DNA primers was used to evaluate numerical chromosome abnormalities in two colon cancer cell lines (Caco-2 and HT-29) and in three of their subpopulations (PF11, TC7, and HT29-MTX). In each cell line, the copy number distribution for different chromosomes showed different patterns. The observation of significant variations in the chromosome constitutions between subpopulations derived from the same original tumor suggests the common occurrence of chromosome copy number heterogeneity in tumoral cell lines. This study demonstrates that the PRINS procedure offers a simple and reliable method for in situ chromosomal screening, which could be efficiently used for karyotypic analysis of tumoral cells.

Aneuploidy↗

The cloning of FRAXF: trinucleotide repeat expansion and methylation at a third fragile site in distal Xqter.

Three fragile sites, FRAXA, FRAXE and FRAXF lie in the Xq27-28 region of the human X chromosome. The expression of FRAXA is associated with the fragile X syndrome, the most prevalent form of inherited mental retardation whilst the expression of FRAXE is associated with a rarer and comparatively milder form of mental handicap. Both the FRAXA and FRAXE sites have been cloned and the fragile site expression found to be due to the expansion of analogous CGG/GCC trinucleotide repeat arrays. We describe here the cloning of the third fragile site, FRAXF, and demonstrate that it involves the expansion of a (GCCGTC)n(GCC)n compound array. PCR analyses across the repeat of normal individuals show that the number of triplets in the array ranges from 12-26 and the most common allele consists of 14 triplet units. Sequencing analyses show that 95% of normal individuals have three copies of the GCCGTC motif and in these individuals, the size variation observed by PCR is due to copy number alterations in the GCC array. In a cytogenetically positive male with developmental delay, the array is expanded by > 900 triplets and the adjacent CpG-rich region is methylated. The array is also expanded in cytogenetically positive carrier females from the family originally used to define the FRAXF site. We conclude that the expanded array corresponds to the FRAXF fragile site.

Base Sequence↗

Analysis of methicillin-resistant Staphylococcus aureus by IS1181 profiling.

Variation in the genomic location and copy number of the insertion element IS1181 in methicillin-resistant Staphylococcus aureus (MRSA) was investigated. Sixty-three isolates representing the Jevons type strain (NCTC 10442), phage-propagating strains, and epidemic strains were examined. A PCR amplicon of the insertion element was used to probe genomic restriction endonuclease digests. HindIII genomic digests gave 25 distinct IS1181 patterns, while EcoRI digests gave 20 patterns. EMRSA-01, -02, -04, -06, -07, -09, -10, -11, -13 and -14 contained the element but could not be subtyped by profiling it. EMRSA-16 did not contain IS1181, consistent with a unique evolutionary origin for this major UK epidemic strain. Marked heterogeneity was observed among isolates of EMRSA-03. Each EMRSA-03 strain examined gave a unique pattern, thereby allowing subtyping of an important epidemic phage type for the purposes of hospital cross-infection control.

Bacteriophage Typing↗