PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Copy Number Variations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Transposon-induced rearrangements in the duplicated locus ph of Drosophila melanogaster can create new chimeric genes functionally identical to the wild type.

Variation in the number of gene copies can play a major role in changing the coding capacities of eukaryotic genomes. Different mechanisms, such as unequal recombination or transposon-induced chromosome rearrangements, are believed to be responsible for these events. We have used the direct tandem duplication at the complex locus polyhomeotic (ph) of Drosophila melanogaster as a model system to study functional redundancy associated with chromosomal rearrangements, such as duplications or deletions. The locus covers 28.6 kb and comprises two independent units, ph proximal and ph distal, which are not only similar on the molecular level, but appear to be functionally redundant [Dura et al., Cell 51 (1987) 829-839; Deatrick et al., Gene 105 (1991) 185-195]. We present a molecular and phenotypic analysis of two hypomorphic ph mutants, ph2 and ph4, induced during hybrid dysgenesis. Each corresponds to an internal deletion in the ph locus that overlaps both transcription units. We show that the deletions are likely due to a P/M hybrid dysgenesis-induced rearrangement between proximal and distal ph, that created a single new chimerical ph gene. At least one of the breakpoints must be located in a 1247-bp region that is rich in single sequence, and 100% identical between proximal and distal ph. Junction points between units are in the protein-coding regions, but could not be exactly localized on the genomic sequence of either mutant, because of the precise molecular mechanism that caused the deletions. Protein products of the hybrid genes contain the same functional domains as either wild-type (wt) product.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Size-polymorphism of mini-exon gene-bearing chromosomes among natural populations of Leishmania, subgenus Viannia.

In order to explore genomic plasticity at the level of the mini-exon gene-bearing chromosome in natural populations of Leishmania, the molecular karyotype of 84 Leishmania stocks belonging to subgenus Viannia, originating mostly from Peru and Bolivia, and differing according to eco-geographical and clinical parameters, was resolved and hybridised with a mini-exon probe. The results suggest that size variation of the mini-exon gene-bearing chromosome is frequent and important (up to 245-kb size-difference), and partially involves variation (up to 50%) in copy number of mini-exon genes. There is no significant size-difference between mini-exon-bearing chromosomes of Peruvian and Bolivian populations of cutaneous and mucosal isolates of Leishmania (Viannia) braziliensis, but there is between eco-geographical populations of Leishmania (Viannia) peruviana. Leishmania (V.) peruviana presented a significantly smaller mini-exon-bearing chromosome than the other species of subgenus Viannia. The contrast between the general chromosome size heterogeneity and the homogeneity observed in some Peruvian Andean areas is discussed in terms of selective pressure.

Animals↗

Network analysis provides insights into evolution of 5S rDNA arrays in Triticum and Aegilops.

We have used network analysis to study gene sequences of the Triticum and Aegilops 5S rDNA arrays, as well as the spacers of the 5S-DNA-A1 and 5S-DNA-2 loci. Network analysis describes relationships between 5S rDNA sequences in a more realistic fashion than conventional tree building because it makes fewer assumptions about the direction of evolution, the extent of sexual isolation, and the pattern of ancestry and descent. The networks show that the 5S rDNA sequences of Triticum and Aegilops species are related in a reticulate manner around principal nodal sequences. The spacer networks have multiple principal nodes of considerable antiquity but the gene network has just one principal node, corresponding to the correct gene sequence. The networks enable orthologous groups of spacer sequences to be identified. When orthologs are compared it is seen that the patterns of intra- and interspecific diversity are similar for both genes and spacers. We propose that 5S rDNA arrays combine sequence conservation with a large store of mutant variations, the number of correct gene copies within an array being the result of neutral processes that act on gene and spacer regions together.

Algorithms↗

NUMTs in sequenced eukaryotic genomes.

Mitochondrial DNA sequences are frequently transferred to the nucleus giving rise to the so-called nuclear mitochondrial DNA (NUMT). Analysis of 13 eukaryotic species with sequenced mitochondrial and nuclear genomes reveals a large interspecific variation of NUMT number and size. Copy number ranges from none or few copies in Anopheles, Caenorhabditis, Plasmodium, Drosophila, and Fugu to more than 500 in human, rice, and Arabidopsis. The average size is between 62 (baker's yeast) and 647 bps (Neurospora), respectively. A correlation between the abundance of NUMTs and the size of the nuclear or the mitochondrial genomes, or of the nuclear gene density, is not evident. Other factors, such as the number and/or stability of mitochondria in the germline, or species-specific mechanisms controlling accumulation/loss of nuclear DNA, might be responsible for the interspecific diversity in NUMT accumulation.

Cell Nucleus↗

Mapping of DNA instability at the fragile X to a trinucleotide repeat sequence p(CCG)n.

The sequence of a Pst I restriction fragment was determined that demonstrate instability in fragile X syndrome pedigrees. The region of instability was localized to a trinucleotide repeat p(CCG)n. The sequence flanking this repeat were identical in normal and affected individuals. The breakpoints in two somatic cell hybrids constructed to break at the fragile site also mapped to this repeat sequence. The repeat exhibits instability both when cloned in a nonhomologous host and after amplification by the polymerase chain reaction. These results suggest variation in the trinucleotide repeat copy number as the molecular basis for the instability and possibly the fragile site. This would account for the observed properties of this region in vivo and in vitro.

Base Sequence↗

Phenotypic and genotypic characteristics of recently adapted isolates of Plasmodium falciparum from Thailand.

The drug sensitivity characteristics and Plasmodium falciparum pfmdr1 status of five isolates of P. falciparum recently isolated from patients presenting for treatment from the Thailand/Myanmar border have been investigated. The aim of the study was to avoid the criticisms of some earlier studies by focusing on newly collected isolates from a specific geographic location. Three of the isolates studied exhibited clear resistance to chloroquine similar to that observed in the K1 Thai standard isolate obtained in the 1970s, and the other two isolates were of intermediate sensitivity to chloroquine with concentrations of drug that inhibit parasite growth by 50% of 50 and 43 nmol. The sensitivity of all isolates was enhanced by verapamil but we found no clear association between chloroquine sensitivity and gene copy number or intra-allelic variation of pfmdr1. In contrast, clear cross-resistance was seen between mefloquine and halofantrine, with the most sensitive isolates carrying the K1 mutation in pfmdr1.

Animals↗

Antithrombin cambridge II (Ala384Ser): clinical, functional and haplotype analysis of 18 families.

Thirty-one individuals from 18 unrelated families with antithrombin deficiency have been identified as having a single point mutation within codon 384 (13268 GCA-->TCA) resulting in an alanine to serine substitution. Six families (11 individuals) were identified by the screening of individuals with thromboembolic disease or with a family history of thromboembolic disease, whilst the remaining 12 families (20 individuals) were identified by screening of asymptomatic blood donors. Four individuals had a history of venous thrombotic disease, a further 2 gave a history of superficial thrombophlebitis but the remaining 25 individuals were asymptomatic. Affected individuals demonstrated normal immunological levels of antithrombin but a decrease in anti-IIa activity in the presence of heparin. Haplotype analysis was used to examine the possibility of a founder effect to explain the high frequency of this non-CpG mutation. 29/31 individuals showed a single common "core" haplotype, the only variation existing in the number of copies of an (ATT)n repeat polymorphism--13, 14, 15 or 17. The results suggest that at most there are four independent origins for this mutation.

Alanine↗

Organisation of the pericentromeric region of chromosome 15: at least four partial gene copies are amplified in patients with a proximal duplication of 15q.

Clinical cytogenetic laboratories frequently identify an apparent duplication of proximal 15q that does not involve probes within the PWS/AS critical region and is not associated with any consistent phenotype. Previous mapping data placed several pseudogenes, NF1, IgH D/V, and GABRA5 in the pericentromeric region of proximal 15q. Recent studies have shown that these pseudogene sequences have increased copy numbers in subjects with apparent duplications of proximal 15q. To determine the extent of variation in a control population, we analysed NF1 and IgH D pseudogene copy number in interphase nuclei from 20 cytogenetically normal subjects by FISH. Both loci are polymorphic in controls, ranging from 1-4 signals for NF1 and 1-3 signals for IgH D. Eight subjects with apparent duplications, examined by the same method, showed significantly increased NF1 copy number (5-10 signals). IgH D copy number was also increased in 6/8 of these patients (4-9 signals). We identified a fourth pseudogene, BCL8A, which maps to the pericentromeric region and is coamplified along with the NF1 sequences. Interphase FISH ordering experiments show that IgH D lies closest to the centromere, while BCL8A is the most distal locus in this pseudogene array; the total size of the amplicon is estimated at approximately 1 Mb. The duplicated chromosome was inherited from either sex parent, indicating no parent of origin effect, and no consistent phenotype was present. FISH analysis with one or more of these probes is therefore useful in discriminating polymorphic amplification of proximal pseudogene sequences from clinically significant duplications of 15q.

Adult↗

Federated learning for the pathogenicity annotation of genetic variants in multi-site clinical settings.

MOTIVATION: Rare diseases collectively affect 5% of the population. However, fewer than 50% of rare disease patients receive a molecular diagnosis after whole genome sequencing. Supervised machine learning is a valuable approach for the pathogenicity scoring of human genetic variants. However, existing methods are often trained on curated but limited central repositories, resulting in poor accuracy when tested on external cohorts. Yet, large collections of variants generated at hospitals and research institutions remain inaccessible to machine-learning purposes because of privacy and legal constraints. Federated learning (FL) algorithms have been recently developed enabling institutions to collaboratively train models without sharing their local datasets. RESULTS: Here, we present a proof-of-concept study evaluating the effectiveness of FL for the clinical classification of genetic variants. A comprehensive array of diverse FL strategies was assessed for coding and non-coding Single Nucleotide Variants as well as Copy Number Variants. Our results showed that federated models generally achieved comparable or superior performance to traditional centralized learning. In addition, federated models reached a robust generalization to independent sets with smaller data fractions as compared to their centralized model counterparts. Our findings support the adoption of FL to establish secure multi-institutional collaborations in human variant interpretation. AVAILABILITY AND IMPLEMENTATION: All source code required to reproduce the results presented in this article, implemented in Python, is available under the GNU General Public License v3 at https://github.com/RausellLab/FedLearnVar.

Humans↗

Restriction fragment length polymorphisms among uropathogenic Escherichia coli isolates: pap-related sequences compared with rrn operons.

Among the adhesin-encoding virulence operons associated with uropathogenic Escherichia coli, only pap (pyelonephritis-associated pilus)-related gene clusters typically exhibit variation in their structure and chromosomal copy number. To access further such variability, we compared pap restriction fragment length polymorphisms (RFLPs) with those detected among rRNA (rrn) operons, which encode an essential host function unrelated to virulence. To place such findings in a phylogenetic perspective, the E. coli isolates were also characterized by using multilocus enzyme electrophoresis. Variation in the rrn RFLP profiles correlated with evolutionary divergence resolved by multilocus enzyme electrophoresis; isolates with identical rrn profiles represented the same or closely related electrophoretic types. In contrast, such isolates frequently had different pap-related RFLPs, indicating that these genetic variations have developed recently relative to the changes associated with essential rrn operons or metabolic enzymes. Despite such fluctuations, two lines of evidence indicate conditions under which the pap-related RFLPs can be stably maintained. First, for each of 20 patients with urosepsis, both the primary urinary tract isolate and the concurrent blood isolate were identical. Second, although obtained from different patients, some isolates representing the same electrophoretic type also had identical pap-related RFLPs. Thus, the genotypic diversity of this virulence adhesin operon was not generated during the course of acute infection or during laboratory manipulations. Since fecal E. coli isolates frequently carry chromosomally encoded pap-related gene clusters, these findings suggest that the intra- and interchromosomal recombination events generating the polymorphisms associated with the pap-related sequences likely occur among the E. coli of the commensal reservoir.

Base Sequence↗

Chromosome-specific molecular organization of maize (Zea mays L.) centromeric regions.

A set of oat-maize chromosome addition lines with individual maize (Zea mays L.) chromosomes present in plants with a complete oat (Avena sativa L.) chromosome complement provides a unique opportunity to analyze the organization of centromeric regions of each maize chromosome. A DNA sequence, MCS1a, described previously as a maize centromere-associated sequence, was used as a probe to isolate cosmid clones from a genomic library made of DNA purified from a maize chromosome 9 addition line. Analysis of six cosmid clones containing centromeric DNA segments revealed a complex organization. The MCS1a sequence was found to comprise a portion of the long terminal repeats of a retrotransposon-like repeated element, termed CentA. Two of the six cosmid clones contained regions composed of a newly identified family of tandem repeats, termed CentC. Copies of CentA and tandem arrays of CentC are interspersed with other repetitive elements, including the previously identified maize retroelements Huck and Prem2. Fluorescence in situ hybridization revealed that CentC and CentA elements are limited to the centromeric region of each maize chromosome. The retroelements Huck and Prem2 are dispersed along all maize chromosomes, although Huck elements are present in an increased concentration around centromeric regions. Significant variation in the size of the blocks of CentC and in the copy number of CentA elements, as well as restriction fragment length variations were detected within the centromeric region of each maize chromosome studied. The different proportions and arrangements of these elements and likely others provide each centromeric region with a unique overall structure.

Base Sequence↗

Ecological and evolutionary physiology of heat shock proteins and the stress response in Drosophila: complementary insights from genetic engineering and natural variation.

Classical adaptational and genetic engineering approaches offer complementary insights to understanding biological variation: the former elucidates the origins, magnitude and ecological context of natural variation, while the latter establishes which genes can underlie natural variation. Studies of the stress or heat shock response in Drosophila illustrate this point. At the cellular level, heat shock proteins (Hsps) function as molecular chaperones, minimizing aggregation of peptides in non-native conformations. To understand the adaptive significance of Hsps, we have characterized thermal stress that Drosophila experience in nature, which can be substantial. We used these findings to design ecologically relevant experiments with engineered Drosophila strains generated by unequal site-specific homologous recombination; these strains differ in hsp70 copy number but share sites of transgene integration. hsp70 copy number markedly affects Hsp70 levels in intact Drosophila, and strains with extra hsp70 copies exhibit corresponding differences in inducible thermotolerance and reactivation of a key enzyme after thermal stress. Elevated Hsp70 levels, however, are not without penalty; these levels retard growth and increase mortality. Transgenic variation in hsp70 copy number has counterparts in nature: isofemale lines from nature vary significantly in Hsp70 expression, and this variation is also correlated with both inducible thermotolerance and mortality in the absence of stress.

Animals↗

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans↗

Degeneracy in human multicopy RBM (YRRM), a candidate spermatogenesis gene.

In order to search for mutations in the multicopy RBM genes that might be associated with male infertility, we have used sequence data from the reported cDNA clone to determine the intron exon boundaries of the YRRM 1 gene. This gene has 12 exons, three of which encode the putative RNA binding domain of the protein. Different copies of the gene contain sequence variations and, additionally, give rise to transcripts with different numbers of copies of the repeated SRGY motif. Since mutations in the RNA binding domain would seem likely to have an effect on the activity of the protein, we have scanned these exons for mutations by SSCP on DNA from normal and infertile men. Sequence differences in the exon encoding the N-terminal part of the RNA binding domain account for at least four different classes of the gene and give rise to different SSCP conformers. Sequence analysis shows that one of these classes is a pseudogene and that the members of another class are nonfunctional. RT-PCR shows that all classes are transcribed and that the A class is most abundant. We have found a point mutation that alters the highly conserved RNP2 motif in one infertile patient. This mutation is also found in his father. We have used PCR followed by SSCP analysis to map RBM on a Y Chromosome (Chr) YAC contig and have demonstrated a distribution that spans a major part of this chromosome's euchromatin.

Amino Acid Sequence↗

Interphasic analysis of aneuploidy in cancer cell lines using primed in situ labeling.

The primed in situ (PRINS) labeling technique has been adapted to chromosomal screening of interphasic tumoral cells. A panel of ten chromosome-specific alpha-satellite DNA primers was used to evaluate numerical chromosome abnormalities in two colon cancer cell lines (Caco-2 and HT-29) and in three of their subpopulations (PF11, TC7, and HT29-MTX). In each cell line, the copy number distribution for different chromosomes showed different patterns. The observation of significant variations in the chromosome constitutions between subpopulations derived from the same original tumor suggests the common occurrence of chromosome copy number heterogeneity in tumoral cell lines. This study demonstrates that the PRINS procedure offers a simple and reliable method for in situ chromosomal screening, which could be efficiently used for karyotypic analysis of tumoral cells.

Aneuploidy↗

The cloning of FRAXF: trinucleotide repeat expansion and methylation at a third fragile site in distal Xqter.

Three fragile sites, FRAXA, FRAXE and FRAXF lie in the Xq27-28 region of the human X chromosome. The expression of FRAXA is associated with the fragile X syndrome, the most prevalent form of inherited mental retardation whilst the expression of FRAXE is associated with a rarer and comparatively milder form of mental handicap. Both the FRAXA and FRAXE sites have been cloned and the fragile site expression found to be due to the expansion of analogous CGG/GCC trinucleotide repeat arrays. We describe here the cloning of the third fragile site, FRAXF, and demonstrate that it involves the expansion of a (GCCGTC)n(GCC)n compound array. PCR analyses across the repeat of normal individuals show that the number of triplets in the array ranges from 12-26 and the most common allele consists of 14 triplet units. Sequencing analyses show that 95% of normal individuals have three copies of the GCCGTC motif and in these individuals, the size variation observed by PCR is due to copy number alterations in the GCC array. In a cytogenetically positive male with developmental delay, the array is expanded by > 900 triplets and the adjacent CpG-rich region is methylated. The array is also expanded in cytogenetically positive carrier females from the family originally used to define the FRAXF site. We conclude that the expanded array corresponds to the FRAXF fragile site.

Base Sequence↗

Analysis of methicillin-resistant Staphylococcus aureus by IS1181 profiling.

Variation in the genomic location and copy number of the insertion element IS1181 in methicillin-resistant Staphylococcus aureus (MRSA) was investigated. Sixty-three isolates representing the Jevons type strain (NCTC 10442), phage-propagating strains, and epidemic strains were examined. A PCR amplicon of the insertion element was used to probe genomic restriction endonuclease digests. HindIII genomic digests gave 25 distinct IS1181 patterns, while EcoRI digests gave 20 patterns. EMRSA-01, -02, -04, -06, -07, -09, -10, -11, -13 and -14 contained the element but could not be subtyped by profiling it. EMRSA-16 did not contain IS1181, consistent with a unique evolutionary origin for this major UK epidemic strain. Marked heterogeneity was observed among isolates of EMRSA-03. Each EMRSA-03 strain examined gave a unique pattern, thereby allowing subtyping of an important epidemic phage type for the purposes of hospital cross-infection control.

Bacteriophage Typing↗

Chromosomal localization of six repeated DNA sequences among species of Microtus (Rodentia).

C-banding and fluorescence in situ hybridization (FISH) document the distribution of constitutive heterochromatin and six highly repeated DNA families (MSAT2570, MSAT21, MSAT160, MS2, MS4 and STR47) in the chromosomes of nine species of Microtus (M. chrotorrhinus, M. rossiaemeridionalis, M. arvalis, M. ilaeus, M. transcaspicus, M. cabrerae, M. pennsylvanicus, M. miurus and M. ochrogaster). Autosomal heterochromatin is largely centromeric and contains different repeated families in different species. Similarly, large C-band positive blocks on the sex chromosomes of four species contain different repeated DNAs. This interspecific variation in the chromosomal distribution and copy number of the repeats suggests that a common ancestor to modern species contained most of the repetitive families, and that descendant species selectively amplified or deleted different repeats on different chromosomes.

Animals↗