PubMed HealthSearch

SEARCH · PubMed Health

Results for “segmental duplications”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Strategic targeting of Cas9 nickase induces large segmental duplications.

Gene/segmental duplications play crucial roles in genome evolution and variation. Here, we introduce paired nicking-induced amplification (PNAmp) for their experimental induction. PNAmp strategically places two Cas9 nickases upstream and downstream of a replication origin on opposite strands. This configuration directs the sister replication forks initiated from the origin to break at the nicks, generating a pair of one-ended double-strand breaks. If homologous sequences flank the two break sites, then end resection converts them to single-stranded DNAs that readily anneal to drive duplication of the region bounded by the homologous sequences. PNAmp induces duplication of segments as large as ∼1 Mb with efficiencies exceeding 10% in the budding yeast Saccharomyces cerevisiae. Furthermore, appropriate splint DNAs allow PNAmp to duplicate/multiplicate even segments not bounded by homologous sequences. We also provide evidence for PNAmp in mammalian cells. Therefore, PNAmp provides a prototype method to induce structural variations by manipulating replication fork progression.

Saccharomyces cerevisiae

GenomeDecoder: inferring segmental duplications in highly repetitive genomic regions.

MOTIVATION: The emergence of the 'telomere-to-telomere' genomics brought the challenge of identifying segmental duplications (SDs) in complete genomes. It further opened a possibility for identifying the differences in SDs across individual human genomes and studying the SD evolution. These newly emerged challenges require algorithms for reconstructing SDs in the most complex genomic regions that evaded all previous attempts to analyze their architecture, such as rapidly evolving immunoglobulin loci. RESULTS: We describe the GenomeDecoder algorithm for inferring SDs and apply it to analyzing genomic architectures of various loci in primate genomes. Our analysis revealed that multiple duplications/deletions led to a rapid birth/death of immunoglobulin genes within the human population and large changes in genomic architecture of immunoglobulin loci across primate genomes. Comparison of immunoglobulin loci across primate genomes suggests that they are subjected to diversifying selection. AVAILABILITY AND IMPLEMENTATION: GenomeDecoder is available at https://github.com/ZhangZhenmiao/GenomeDecoder. The software version and test data used in this paper are uploaded to https://doi.org/10.5281/zenodo.14753844.

Humans

The Role of Small Segmental Duplications in Generating Identical Isoforms Through Alternative Splicing Sites.

Alternative splicing plays a crucial role in expanding proteomic diversity but can also generate identical isoforms under certain conditions. While mutually exclusive splicing of tandem exons has occasionally been reported to produce identical isoforms, the extent to which other splicing events contribute to this phenomenon remains unclear. In this study, we demonstrate that alternative 5' and 3' splice site selection can also lead to the formation of identical isoforms, providing an additional type of splicing event for functional redundancy in transcriptomes. To address this, we analyzed reference genome annotations from 15 plant species, including Arabidopsis thaliana and wheat (Triticum aestivum), obtained from the RefSeq database. Identical isoforms were computationally defined as transcripts with distinct exon-intron structures but identical coding sequences. Our analysis reveals that the majority of alternative 5' and 3' fragments originate from small segmental duplications, suggesting that sequence repetition within gene regions facilitates the emergence of such splicing patterns. We also observed differences in the annotated 5' UTRs of some identical isoforms. However, since the alternative splicing sites themselves were not located within UTRs, these differences may reflect annotation uncertainty rather than genuine AS-derived variation. Given that UTR predictions in reference databases are not always precise, such observations should be interpreted cautiously. Expression analysis using an isoform-specific k-mer approach confirmed that identical isoforms can be differentially regulated. These findings suggest that, beyond expanding protein diversity, alternative splicing can also generate redundant isoforms that are differentially expressed at the RNA level, indicating potential regulatory roles. By elucidating the structural and regulatory factors contributing to the formation and retention of identical isoforms, our study provides new insights into the evolutionary and functional significance of alternative splicing in plants.

Alternative Splicing

[Congenital segmental duplication of the lumbar ureter (author's transl)].

In a 23-year-old man attacks of nephritic colic led to the discovery of an obstruction on the left lumbar ureter. Segmental resection of the ureter was performed, removing 10 mm of malformed, obstructed ureter. This was an incomplete duplication, the two ureteral segments lying side-by-side, each with its own musculature, for a distance of 7mm. Above and below the anomaly, the ureter was normal. This exceptional malformation is compared with other internal obstructions of the ureter.

Adult

SegMantX: A Novel Tool for Detecting DNA Duplications Uncovers Prevalent Duplications in Plasmids.

Segmental duplications play an important role in genome evolution via their contribution to copy-number variation, gene-family diversification, and the emergence of novel functions. The detection of segmental duplications is challenging due to heterogeneous amelioration of sequence similarity among duplicates, which hinders the reconstruction of continuous sequence alignment. Here we introduce SegMantX, a novel approach for the identification of diverged segmental duplications in prokaryote genomes using local alignment chaining. In this approach, local alignments resulting from a preliminary sequence similarity search (e.g. BLASTn) are chained into continuous segments. Evaluating the performance of SegMantX using simulated sequences shows that the tool can detect diverged duplications beyond the sensitivity limits of standard alignment-based methods. Applying SegMantX to 6,784 enterobacterial plasmids, we find that 65% plasmids contain duplicated regions and gene duplications, most of which correspond either to dispersed, noncoding regions or duplicated mobile genetic elements (MGEs; e.g. transposons and insertion sequences). Furthermore, we demonstrate the applicability of SegMantX for the identification of diverged gene transfers between replicons and plasmid hybridization events. Our findings highlight MGEs as drivers of segmental duplications in plasmid evolution, leading to the amplification of their cargo genes, including antibiotic resistance genes. SegMantX provides a powerful framework for reconstructing diverged segmental duplications and other alignment problems.

Plasmids

The effects of temperature on genetic instability in Aspergillus nidulans.

Previous work has shown that strains of Aspergillus nidulans with a chromosome segment in duplicate (one in normal position, one translocated to another chromosome) are unstable. Deletions occur from either duplicate segment. The present work has shown that most deletions occur from the translocated duplicate segment. Furthermore, it has been found that the overall frequency of deletions from a duplication is dependent upon the temperature of growth. The overall frequency of deletions from a chromosome III duplication is greatly enhanced by low temperatures, while the overall frequency of deletions from a chromosome I duplication is markedly enhanced by high temperatures. A temperature of 39.5 degrees C appears to enhance to overall frequency of deletions from the I duplication to the greatest extent. With regard to the non-translocated duplicate I segment, an increase in temperature progressively enhances the frequency of those deletions to which it is subject to far more deletions during a particular period of growth than during any other period, and at 42 degrees C, a section of the III duplication is subject to far more deletions during a given period of growth than during any other period. Comparisons with other cases of genetic instability are made and common underlying connections are proposed.

Aneuploidy

Genome-wide identification and expression profiling of the MADS-box gene family in Lavandula angustifolia.

BACKGROUND: MADS-box genes encode transcription factors critical for plant development, particularly floral organogenesis, flowering time regulation, and adaptation to environmental stresses. Among these, the MIKCC-type genes are pivotal regulators in floral developmental processes. Although the evolutionary diversification and functional dynamics of MADS-box genes have been extensively characterized in model plants such as Arabidopsis thaliana and Oryza sativa, their evolutionary relationships and functional profiles in Lavandula angustifolia, an economically significant aromatic plant, remain poorly understood. RESULTS: Genome-wide analysis identified 173 MADS-box genes in L. angustifolia, categorized into type I (Mα: 26; Mβ: 0; Mγ: 10) and type II (MIKCC: 125; MIKC*: 12) based on phylogenetic comparisons with A. thaliana. The MIKCC subgroup was further subdivided into 12 subclasses, including genes central to the ABCDE model of floral organ specification. Structural analyses revealed distinct conserved motifs and exon-intron configurations specific to each subgroup, indicative of functional divergence. Synteny analysis demonstrated Whole Genome Duplication (WGD) and segmental duplications as major contributors to MIKCC gene family expansion, notably among genes linked to floral organ development. Expression profiling via RNA-seq and quantitative real-time PCR (qPCR) showed type II MADS-box genes exhibited higher expression levels with pronounced tissue-specific and developmental stage-specific expression patterns compared to type I genes. Many type II genes displayed significant associations with floral organogenesis, floral transition, and abiotic stress responses, underscoring their essential roles in reproductive development and environmental adaptability in L. angustifolia. CONCLUSIONS: The identification and comprehensive characterization of 173 MADS-box genes in L. angustifolia highlight the significant expansion of the MIKCC subgroup driven primarily by WGD and segmental duplications. The distinct structural features and specific expression patterns observed provide insights into the functional divergence and complexity of these genes, particularly regarding floral organogenesis and adaptation to environmental stress. This study establishes a robust molecular basis for further functional analysis and genetic improvement of aromatic plants.

MADS Domain Proteins

Genetic control of mitochondrial malate dehydrogenases: evidence for duplicated chromosome segments.

The genetic control of the major mitochondrial isoenzymes of malate dehydrogenase (L-malate:NAD+ oxidoreductase; EC 1.1.1.37) has been investigated in Zea mays. The mitochondrial isozymes are coded at four nuclear gene loci. Two of the loci (mdh1 and mdh2) are diallelic and tightly linked. The other two loci (mdh3 and mdh4) appear to have arisen by duplication of the chromosome segment carrying mdh1 and mdh2, but are not linked to them. The segregation of such a duplicate segment can explain anomalous backcross and F2 segregation ratios.

Alleles

Where Did the Y Chromosome in the Spiny Rat Go, and How Did It Get There?

The XX/XY sex chromosome system is highly conserved across mammals, with rare exceptions where males lack a Y chromosome. Among these is the genus Tokudaia, a group of spiny rats comprising three species with unique sex chromosome systems deviating from the typical XX/XY pattern. While Tokudaia osimensis and Tokudaia tokunoshimensis have completely lost the Y chromosome, they retain some Y-linked genes on the X chromosome. In contrast, Tokudaia muenninki retains large sex chromosomes where both the X and Y chromosomes have fused with an autosome pair, carrying multi-copied Y-linked genes, including Sry. In this study, we generated chromosome-level genome assemblies for male individuals of all three Tokudaia species. By investigating loci typically associated with rodent Y-linked genes, we characterized sequences derived from the Tokudaia Y-chromosomal most recent common ancestor (Tokudaia Y-MRCA) and traced their evolutionary trajectories. Our analyses revealed that an initial X-to-Y translocation of a sequence containing the boundary-associated segmental duplication in a common ancestor of Tokudaia marked the beginning of their unique sex chromosome evolution. The boundary-associated segmental duplication, uniquely multi-copied in Tokudaia, facilitated further rearrangements through nonallelic homologous recombination and duplications. These processes culminated in subsequent Y-to-X translocations and duplications, leading to the complete loss of the Y chromosome as a distinct entity while preserving Y-linked genes in a multicopy state on the X chromosome. These findings highlight Tokudaia's rapid sex chromosome evolution within 3 million years and provide insights into the mechanisms underlying Y chromosome loss, contributing to a broader understanding of sex chromosome evolution in rodents.

Animals

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24 h of exposure to 34 °C, whereas DNAJC3 was significantly induced only at 12 h. All six genes reached their highest transcript abundance at 12 h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

The SMN locus in the T2T era: Structure, gene conversion, and clinical implications.

Long-read sequencing, paralog-aware variant calling, and telomere-to-telomere (T2T) human genome assemblies now enable the resolution of copy-, haplotype-, and nucleotide-level complexities in segmentally duplicated loci, which were previously inaccessible with short-read sequencing. In this review, we highlight how current technologies and analysis methods reveal extensive diversity in copy number (CN), structure, and gene conversion within the spinal muscular atrophy-associated survival motor neuron (SMN) locus. We summarize how understanding population-level structural variation could be translated into clinical practice, where a nucleotide-level view of the SMN locus may refine prognostic accuracy beyond SMN2 CN and explain variable treatment responses. Finally, we discuss how the approaches and methodologies required to study the SMN locus may be applied elsewhere, providing a scaffold to characterize other complex human genetic regions.

Humans

Genome-wide identification, characterization, evolutionary analysis, and expression profiling of the FCS-like zinc finger (FLZ) gene family in soybean (Glycine max L.) under abiotic stresses.

Drought and salinity limit soybean yield. Despite their role in the SnRK1 energy-sensing complex, a systematic study of FCS-Like Zinc Finger (FLZ) proteins in soybean has not been reported. We performed a genome-wide identification of the GmFLZ gene family, identifying 40 members distributed across 18 of the 20 soybean chromosomes. Phylogenetic analysis of 87 FLZ proteins from Glycine max, Arabidopsis thaliana, and Oryza sativa revealed four major evolutionary clades, suggesting that diversification predates the separation of monocots and dicots. Structural analysis identified ten conserved motifs, with Motifs 1 and 2 present in all family members. Gene duplication analysis identified 304 paralogous pairs, most arising from segmental duplication. Ka/Ks analysis indicated localized positive selection in six gene pairs and purifying selection in 97.9% of pairs. Tissue-specific expression profiling across nine tissues showed that GmFLZ5, GmFLZ15, GmFLZ25, and GmFLZ34 had the highest expression levels detected across the GmFLZ family, with GmFLZ5 the most highly expressed member in leaves, nodules, and stem and showing moderate expression in pod, root, and root hairs, whereas GmFLZ18, GmFLZ23, and GmFLZ37 showed root-preferential expression. RT-qPCR validation under drought (20% PEG-6000) and salt (200 mM NaCl) treatments in the Giza 5 cultivar showed that 36 and 34 of the 40 GmFLZ genes, respectively, exhibited at least a two-fold change in expression, with GmFLZ21 and GmFLZ35 among the most strongly induced under salt stress. These findings provide an evolutionary and functional framework for the GmFLZ family and identify candidate genes for future functional studies in soybean stress tolerance.

Glycine max

Diverse evolutionary rates and gene duplication patterns among families of functional olfactory receptor genes in humans.

In humans, odors are detected by ~400 functional olfactory receptor (OR) genes. The superfamily of functional OR genes can be further divided into tens of families. In large part, the OR genes have experienced extensive tandem duplications, which have led to gene gains and losses. However, whether different OR gene families have experienced distinct modes of gene duplication has yet to be reported. We conducted comparative genomic and evolutionary analyses for human functional OR genes. Based on analysis of human-mouse 1-1 orthologs, we found that human functional OR genes show higher-than-average evolutionary rates, and there are significant differences among families of functional OR genes. Via comparison with seven vertebrate outgroups, families of human functional OR genes show different extents of gene synteny conservation. Although the superfamily of human functional OR genes is enriched in tandem and proximal duplications, there are particular families which are enriched in segmental duplications. These findings suggest that human functional OR genes may be governed by different evolutionary mechanisms and that large-scale gene duplications have contributed to the early evolution of human functional OR genes.

Humans

The nucleotide sequence of oocyte 5S DNA in Xenopus laevis. II. The GC-rich region.

The primary sequence of the GC-rich half of the repeating unit in X. laevis 5S DNA has been determined in both a single plasmid-cloned repeating unit and in the total population of repeatig units. The GC-rich half of the repeating unit contains a single long duplication of 174 nucleotides. The duplicated segment commences 73 nucleotides preceding the 5' end of the gene and terminates at nucleotide 101 of the gene. The duplicated portion of the gene, termed the pseudogene, differs by 10 nucleotides from the corresponding portion of the gene, and the remaining duplicated sequence of 73 nucleotides differs by 13 nucleotides. The plasmid-cloned repeating unit differs from the dominant sequence in the total population repeating units by 6 nucleotides in the GC-rich region. Evidence is provided that most of the CpG dinucleotides in 5S DNA are at least partially methylated.

Animals

Genetic and segregation analysis of Escherichia coli strains containing a tandem duplication of the trpD-purB region of the chromosome.

Genetic and segregation analysis of Escherichia coli strains containing a partial duplication of the trp operon reveal that the 2.5-min-long region trpD-purB is duplicated in tandem in the chromosome. The adjacent loci cysB and fabD are not duplicated. Although one copy of the duplicated region is longer than the maximum size of bacteriophage P1kc transducing fragments, the frequency at which the duplicated segment trpDCBA is transferred by transduction to tonB-trp deletion strains is equal to that observed for transfer of the normal trp operon. This suggests that three-point recombination events believed to account for transduction of long duplications occur as frequently as two-point recombination events believed to account for normal transduction. Cotransduction frequencies of trpDCBA with the duplicated loci tonB, galU, tyrT, and hemA are very similar to those for the trp operon with the same loci. This indicates that normal genetic linkage is maintained during the three-point recombination event. However, purB, which is normally unlinked to trp by transduction, is closely linked to trpDCBA and thus must be near the repeat point of the duplication. Transduction tests with point mutations in the trp operon indicated that the repeat point occurs near the normal boundary between trpE and trpD. Segregation analysis of heterogenotes constructed from tonB-trp deletion strains shows that the frequency at which a marker is lost is approximately proportional to its distance from the repeat point. This finding is consistent with a random, singlesite crossover event during segregation. Several observations indicate that non-reciprocal genetic exchange also occurs between copies of the duplication. Analysis of heterogenotes containing dadR1 and dadR(+) demonstrate that the mutant allele is transdominant.

Chromosome Aberrations

Genetic analysis of partial duplication of the long arm of chromosome 16.

BACKGROUND: Pure partial trisomy 16q12.1q22.1 is a rare chromosome copy number variant (CNV). The primary clinical phenotypes associated with this syndrome include abnormal facial morphology, global developmental delay (GDD), short stature, and reported predisposing factors for atypical behavior, autism, the development of learning disabilities, and neuropsychiatric disorders. The dosage-sensitive genes associated with partial trisomy are not disclosed preventing to establish a genotype-phenotype correlation. METHODS: We report a case of a Chinese patient diagnosed with GDD and an abnormal facial shape, who was found to have partial trisomy 16 through karyotyping and high-throughput sequencing analysis. Karyotype and CNV tracing analyses were also conducted on the biological parents of the patient to assess for any chromosomal structural abnormalities. Additionally, we included 29 patients with pure partial trisomy 16q, reported in the DECIPHER database and the literature. We and performed a genotype-phenotype correlation analysis. RESULTS: The proband, a 2-year-old female, was found to have a de novo 21.96 Mb duplication located between 16q12.1q22.1, with no other deletions observed on other chromosomes, indicating a pure partial trisomy of 16q. Through genotype and phenotype analysis of 29 individuals, we found that patients with the duplicated region located at the distal region of 16q may exhibit more severe symptoms than those with duplication at the proximal region; however, no relationship was identified between phenotype and the size of the duplicated segment. CONCLUSION: We report, for the first time, a patient with partial trisomy 16q validated by multiple genetic tests, including CNV-seq, whole exome sequencing (WES), and karyotyping. It is speculated that partial trisomy of 16q may be associated with continuous gene duplication. However, functional studies are necessary to identify the causative gene or critical region linked to duplication syndrome of chromosome 16q.

Child, Preschool

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article

Partial trisomy 13q21toqter de novo due to a recombinant chromosome rec(13)dup q.

A female is described who has a karyotype with an additional distal half of 13q in a recombinant rec(13)dup q chromosome. Since her parents have normal karyotypes, the origin of her karyotype is assumed to be a premeiotic pericentric inversion de novo with crossing-over within the inversion loop at meiosis. By means of various banding techniques, the breaks preceding the rearrangement could be located exactly. The joint between the duplicated segment and the satellites of the receptor chromosome is of special note. The phenotype of the patient stated at the age of 9 months and at the age of 7 1/2 years was found to be related to the segments involved in the partial trisomy. The clinical features were largely in accordance with previous case reports having an identical extent of the triplicated 13q segment.

Child