PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomic Structural Variation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

A worldwide survey of haplotype variation and linkage disequilibrium in the human genome.

Recent genomic surveys have produced high-resolution haplotype information, but only in a small number of human populations. We report haplotype structure across 12 Mb of DNA sequence in 927 individuals representing 52 populations. The geographic distribution of haplotypes reflects human history, with a loss of haplotype diversity as distance increases from Africa. Although the extent of linkage disequilibrium (LD) varies markedly across populations, considerable sharing of haplotype structure exists, and inferred recombination hotspot locations generally match across groups. The four samples in the International HapMap Project contain the majority of common haplotypes found in most populations: averaging across populations, 83% of common 20-kb haplotypes in a population are also common in the most similar HapMap sample. Consequently, although the portability of tag SNPs based on the HapMap is reduced in low-LD Africans, the HapMap will be helpful for the design of genome-wide association mapping studies in nearly all human populations.

Chromosome Mapping↗

Pseudo-periodic partitions of biological sequences.

MOTIVATION: Algorithm development for finding typical patterns in sequences, especially multiple pseudo-repeats (pseudo-periodic regions), is at the core of many problems arising in biological sequence and structure analysis. In fact, one of the most significant features of biological sequences is their high quasi-repetitiveness. Variation in the quasi-repetitiveness of genomic and proteomic texts demonstrates the presence and density of different biologically important information. It is very important to develop sensitive automatic computational methods for the identification of pseudo-periodic regions of sequences through which we can infer, describe and understand biological properties, and seek precise molecular details of biological structures, dynamics, interactions and evolution. RESULTS: We develop a novel, powerful computational tool for partitioning a sequence to pseudo-periodic regions. The pseudo-periodic partition is defined as a partition, which intuitively has the minimal bias to some perfect-periodic partition of the sequence based on the evolutionary distance. We devise a quadratic time and space algorithm for detecting a pseudo-periodic partition for a given sequence, which actually corresponds to the shortest path in the main diagonal of the directed (acyclic) weighted graph constructed by the Smith-Waterman self-alignment of the sequence. We use several typical examples to demonstrate the utilization of our algorithm and software system in detecting functional or structural domains and regions of proteins. A big advantage of our software program is that there is a parameter, the granularity factor, associated with it and we can freely choose a biological sequence family as a training set to determine the best parameter. In general, we choose all repeats (including many pseudo-repeats) in the SWISS-PROT amino acid sequence database as a typical training set. We show that the granularity factor is 0.52 and the average agreement accuracy of pseudo-periodic partitions, detected by our software for all pseudo-repeats in the SWISS-PROT database, is as high as 97.6%.

Algorithms↗

Targeted population genomics uncovers demographic history and genetic divergence in north American wild cranberry.

Wild populations of North American cranberry (Vaccinium macrocarpon Aiton) are reservoirs of genetic variation that may contribute to the improvement of breeding-relevant traits. However, the extent to which wild genetic variation is geographically structured and represented in elite germplasm remains unclear. We analysed 179 wild cranberry accessions from the upper Midwest and Eastern North America to estimate nucleotide diversity (π), population structure, and loci associated with genetic differentiation and environmental variables using a genome-informed targeted genotyping panel. Additionally, 14 demographic scenarios were evaluated using site-frequency-spectrum-based inference to identify historical events that could explain current genetic diversity. We observed extremely low nucleotide diversity within the targeted panel (π = 5 × 10-6). Rare allele distributions strongly influenced π and Tajima's D values, suggesting constrained diversity in the genomic regions assayed that is not captured by heterozygosity-based estimates alone. However, we interpreted these results as conservative lower bounds on genome-wide neutral diversity because the targeted panel is enriched for genic and conserved regions. A clear separation between the Midwest and East populations was observed, with inbreeding coefficients ranging from -0.13 to 0.15. Furthermore, site frequency spectrum inference from the targeted panel supported a demographic scenario consistent with a significant population reduction ≈15-14 thousand years ago (kya), followed by a divergence between the two regions ≈12 kya, and an asymmetric gene flow ≈1.3 kya. We detected 254 candidate loci showing regional allele-frequency differentiation. Several of these loci colocalized with candidate genes linked to stress response, development, and metabolic processes. To evaluate the representation of geographically differentiated wild alleles in a breeding context, we analysed Rutgers breeding materials (n = 484) and found that this panel is enriched for common alleles in Eastern wild populations. These findings indicate regionally structured allele-frequency variation in wild cranberry, with potential relevance to environmental response and breeding. This study extends prior wild cranberry population-genetic research by providing targeted-panel estimates of diversity, comparisons of demographic models, and breeding insights on geographically differentiated alleles, while highlighting the importance of conserving wild cranberry germplasm for use in modern breeding programs.

Journal Article↗

Mitogenomic Insights Into the Population Structure and Demographic History of Tree Shrews (Tupaia belangeri) in China.

The northern tree shrew (Tupaia belangeri) exhibits significant morphological and geographical variations, but its evolutionary history and subspecies boundaries remain controversial. Here, we analyzed the complete mitochondrial genomes of 63 individuals, representing 12 populations in China to study phylogenetic relationships, genetic diversity, and population history. Phylogenetic analysis consistently restored four mitochondrial branches with strong geographic structures and significant differences. The three lineages correspond to geographically restricted subspecies (T. b. tonquinia, T. b. modesta, and T. b. gaoligongensis), while individuals assigned to several traditional subspecies cluster in a broad mainland lineage (T. b. chinensis, T. b. yunalis, and T. b. yaoshanensis). The divergence time estimate places the origin of the main lineage in the Miocene, consistent with major tectonic and geomorphological events. Demographic analysis revealed different population histories, including varying degrees of expansion in recent continental and island lineages, as well as the long-term stability of T. b. gaoligongensis. Genetic diversity varied markedly among lineages, with the highest diversity observed in the T. b. gaoligongensis and the lowest diversity observed in the T. b. modesta. These findings demonstrate that landscape complexity and demographic history are key drivers of evolutionary diversification in T. belangeri, challenging classical morphology-based subspecies classifications and underscoring the need for comprehensive sampling across both domestic and international ranges.

Tupaia belangeri↗

Identification of a novel non-coding deletion in Allan-Herndon-Dudley syndrome by long-read HiFi genome sequencing.

BACKGROUND: Allan-Herndon-Dudley syndrome (AHDS) is an X-linked disorder caused by pathogenic variants in the SLC16A2 gene. Although most reported variants are found in protein-coding regions or adjacent junctions, structural variations (SVs) within non-coding regions have not been previously reported. METHODS: We investigated two male siblings with severe neurodevelopmental disorders and spasticity, who had remained undiagnosed for over a decade and were negative from exome sequencing, utilizing long-read HiFi genome sequencing. We conducted a comprehensive analysis including short-tandem repeats (STRs) and SVs to identify the genetic cause in this familial case. RESULTS: While coding variant and STR analyses yielded negative results, SV analysis revealed a novel hemizygous deletion in intron 1 of the SLC16A2 gene (chrX:74,460,691 - 74,463,566; 2,876 bp), inherited from their carrier mother and shared by the siblings. Determination of the breakpoints indicates that the deletion probably resulted from Alu/Alu-mediated rearrangements between homologous AluY pairs. The deleted region is predicted to include multiple transcription factor binding sites, such as Stat2, Zic1, Zic2, and FOXD3, which are crucial for the neurodevelopmental process, as well as a regulatory element including an eQTL (rs1263181) that is implicated in the tissue-specific regulation of SLC16A2 expression, notably in skeletal muscle and thyroid tissues. CONCLUSIONS: This report, to our knowledge, is the first to describe a non-coding deletion associated with AHDS, demonstrating the potential utility of long-read sequencing for undiagnosed patients. Although interpreting variants in non-coding regions remains challenging, our study highlights this region as a high priority for future investigation and functional studies.

Humans↗

Genomic characterization and mutation rate of hepatitis C virus isolated from a patient who contracted hepatitis during an epidemic of non-A, non-B hepatitis in Japan.

To investigate the genomic characterization of hepatitis C virus (HCV) isolated from patient who contracted hepatitis during an epidemic of non-A, non-B (NANB) hepatitis in Shimizu city, Japan, we have cloned the nucleotide sequence of the viral genome (HCV-KF) spanning the structural domain. When compared to other previously reported HCV isolates, HCV-KF showed an overall identity at the amino acid level of 90.0 to 92.1% with Japanese isolates and 80.9 to 82.1% with American-like isolates. The HCV-KF genome displays an insertion of three nucleotides in-frame (corresponding to one amino acid) found at the junction between the E1 and E2/NS1 region. The mutation rate of the HCV-KF genome was assessed by comparing the nucleotide and deduced amino acid sequences of the viral RNA obtained from the serum of the original patient with viral sequences derived from the serum of a chimpanzee inoculated with the same serum 9 years previously. The substitution rate of the viral genome was estimated at 0.9 x 10(-3) nucleotides per site per year for the HCV structural region. The highest mutation rate was found in the hypervariable region within the E2/NS1 domain. It is suggested that the outbreak in Shimizu city was caused by a strain of HCV closely related to the Japanese-like subgroup of isolates.

Americas↗

Developmental variation in epidermal growth factor receptor size and localization in the malaria mosquito, Anopheles gambiae.

The AGER gene encoding the epidermal growth factor receptor (EGFR) of the malaria mosquito Anopheles gambiae was cloned and sequenced. It represents a canonical member of this family of tyrosine kinase proteins exhibiting many similarities to orthologues from other species, both on the level of genomic organization and protein structure. The mRNA can be detected throughout development. Western analysis with an antibody raised against the extracellular domain of the mosquito protein suggests developmental variation in protein size and location that may be involved in the function of EGFR in the mosquito.

Animals↗

Growth phase-dependent variation in protein composition of the Escherichia coli nucleoid.

The genome DNA of Escherichia coli is associated with about 10 DNA-binding structural proteins, altogether forming the nucleoid. The nucleoid proteins play some functional roles, besides their structural roles, in the global regulation of such essential DNA functions as replication, recombination, and transcription. Using a quantitative Western blot method, we have performed for the first time a systematic determination of the intracellular concentrations of 12 species of the nucleoid protein in E. coli W3110, including CbpA (curved DNA-binding protein A), CbpB (curved DNA-binding protein B, also known as Rob [right origin binding protein]), DnaA (DNA-binding protein A), Dps (DNA-binding protein from starved cells), Fis (factor for inversion stimulation), Hfq (host factor for phage Q(beta)), H-NS (histone-like nucleoid structuring protein), HU (heat-unstable nucleoid protein), IciA (inhibitor of chromosome initiation A), IHF (integration host factor), Lrp (leucine-responsive regulatory protein), and StpA (suppressor of td mutant phenotype A). Intracellular protein levels reach a maximum at the growing phase for nine proteins, CbpB (Rob), DnaA, Fis, Hfq, H-NS, HU, IciA, Lrp, and StpA, which may play regulatory roles in DNA replication and/or transcription of the growth-related genes. In descending order, the level of accumulation, calculated in monomers, in growing E. coli cells is Fis, Hfq, HU, StpA, H-NS, IHF*, CbpB (Rob), Dps*, Lrp, DnaA, IciA, and CbpA* (stars represent the stationary-phase proteins). The order of abundance, in descending order, in the early stationary phase is Dps*, IHF*, HU, Hfq, H-NS, StpA, CbpB (Rob), DnaA, Lrp, IciA, CbpA, and Fis, while that in the late stationary phase is Dps*, IHF*, Hfq, HU, CbpA*, StpA, H-NS, CbpB (Rob), DnaA, Lrp, IciA, and Fis. Thus, the major protein components of the nucleoid change from Fis and HU in the growing phase to Dps in the stationary phase. The curved DNA-binding protein, CbpA, appears only in the late stationary phase. These changes in the composition of nucleoid-associated proteins in the stationary phase are accompanied by compaction of the genome DNA and silencing of the genome functions.

Bacterial Proteins↗

Fibroblast growth factor receptor 2 (FGFR2): genomic sequence and variations.

Fibroblast growth factor receptors (FGFRs) play an important role in development and tumorigenesis. Mutations in FGFR2 cause more than five craniosynostosis syndromes. The FGFR2 genomic structure is the largest of the FGFR family. We have refined and extended the genomic organization of the FGFR2 gene by sequencing more than 119 kb of PACs, cosmids, and PCR products and assembling a region of approximately 175 kb. Although the gene structure has been reported to include only 20 exons, we have verified the presence of at least 22 exons, some of which are alternatively spliced. The sizes of six exons differed from those reported previously. Comparison of our sequence and those in the NCBI database detected more than 300 potential single nucleotide polymorphisms (SNPs). However, sequencing regions containing 52 of these potential SNPs verified only 14 in PCR products generated from 16 CEPH alleles. In contrast, direct sequencing of the CEPH DNAs revealed 21 other polymorphisms. Only one SNP was found in the 2,926 bp of coding sequence. Twenty-seven SNPs, two insertion polymorphisms and five microsatellite polymorphisms are contained in approximately 16.6 kb of non-coding sequence. These data yield an average of one polymorphism for approximately 488 bp of non-coding sequence examined. This collection of SNP, insertion, and repeat polymorphisms will aid future association studies between the FGFR2 gene and human disease and will enhance mutation detection.

Alleles↗

Evolution of the terminal regions of the Streptomyces linear chromosome.

Comparative analysis of the Streptomyces chromosome sequences, between Streptomyces coelicolor, Streptomyces avermitilis, and Streptomyces ambofaciens ATCC23877 (whose partial sequence is released in this study), revealed a highly compartmentalized genetic organization of their genome. Indeed, despite the presence of specific genomic islands, the central part of the chromosome appears highly syntenic. In contrast, the chromosome of each species exhibits large species-specific terminal regions (from 753 to 1,393 kb), even when considering closely related species (S. ambofaciens and S. coelicolor). Interestingly, the size of the central conserved region between species decreases as the phylogenetic distance between them increases, whereas the specific terminal fraction reciprocally increases in size. Between highly syntenic central regions and species-specific chromosomal parts, there is a notable degeneration of synteny due to frequent insertions/deletions. This reveals a massive and constant genomic flux (from lateral gene transfer and DNA rearrangements) affecting the terminal contingency regions. We speculate that a gradient of recombination rate (i.e., insertion/deletion events) toward the extremities is the force driving the exclusion of essential genes from the terminal regions (i.e., chromosome compartmentalization) and generating a fast gene turnover for strong adaptation capabilities.

Chromosome Structures↗

Genotyping and sequence analysis of apolipoprotein E isoforms.

Apolipoprotein E (apoE), a polymorphic plasma protein, is essential for catabolism of lipoproteins by receptor-mediated endocytosis. One of the apoE isoforms (E2) differs in its binding affinity to specific receptors and contributes to variations in lipoprotein metabolism. Diagnosis of apoE isoforms is done by isoelectric focusing, but it is hindered by various degrees of post-translational sialylation of the apoE protein. Electrophoretically silent structural variations may also escape detection by this technique. We describe a method for genotyping apoE based on hybridization of allele-specific oligonucleotides with enzymatically amplified genomic DNA, which permits unambiguous diagnosis of six common apoE phenotypes within 24 h. Among 100 E2 alleles present in 81 unrelated individuals genotyped by this technique, we found two rare structural mutants of apoE in addition to the common E2 form, E2(158Arg----Cys). Automated sequencing of amplified DNA identified the rare mutants as E2(136Arg----Ser) and E2(145Arg----Cys). The genotypic method may complement or even replace isoelectric focusing for routine determination of apoE phenotypes and for identification of rare structural variants.

Alleles↗

Microarrays in ecology and evolution: a preview.

Microarray technology provides a new tool with which molecular ecologists and evolutionary biologists can survey genome-wide patterns of gene expression within and among species. New analytical approaches based on analysis of variance will allow quantification of the contributions of among individual variation, genotype, sex, microenvironment, population structure, and geography to variation in gene expression. Applications of this methodology are reviewed in relation to studies of mechanisms of adaptation and divergence; delineation of developmental and physiological pathways and networks; characterization of quantitative genetic parameters at the level of transcription ('quantitative genomics'); molecular dissection of parasitism and symbiosis; and studies of the diversification of gene content. Establishment of microarray resources is neither prohibitively expensive nor technologically demanding, and a commitment to development of gene expression profiling methods for nonmodel organisms could have a tremendous impact on molecular and genetic research at the interface of organismal and population biology.

Animals↗

Low genetic variation in muskoxen (Ovibos moschatus) from western Greenland using microsatellites.

Muskoxen are large herbivores living in Arctic environments. Lack of genetic variation in allozymes has made it difficult to study the social and genetic structure of this species. In this study, we have tried to find polymorphic microsatellite loci using both cattle-derived microsatellite primers and primers developed from a genomic plasmid library of muskoxen. Only limited variation was found for both sets of microsatellite loci. We conclude that this consistent low genetic variation is probably due to demographic features of the muskoxen populations rather than to methodological constraints caused by the transfer of microsatellites between species.

Animals↗

Target SNP selection in complex disease association studies.

BACKGROUND: The massive amount of SNP data stored at public internet sites provides unprecedented access to human genetic variation. Selecting target SNP for disease-gene association studies is currently done more or less randomly as decision rules for the selection of functional relevant SNPs are not available. RESULTS: We implemented a computational pipeline that retrieves the genomic sequence of target genes, collects information about sequence variation and selects functional motifs containing SNPs. Motifs being considered are gene promoter, exon-intron structure, AU-rich mRNA elements, transcription factor binding motifs, cryptic and enhancer splice sites together with expression in target tissue. As a case study, 396 genes on chromosome 6p21 in the extended HLA region were selected that contributed nearly 20,000 SNPs. By computer annotation ~2,500 SNPs in functional motifs could be identified. Most of these SNPs are disrupting transcription factor binding sites but only those introducing new sites had a significant depressing effect on SNP allele frequency. Other decision rules concern position within motifs, the validity of SNP database entries, the unique occurrence in the genome and conserved sequence context in other mammalian genomes. CONCLUSION: Only 10% of all gene-based SNPs have sequence-predicted functional relevance making them a primary target for genotyping in association studies.

Amino Acid Substitution↗

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals↗

Generation of protein lineages with new sequence spaces by functional salvage screen.

A variety of different methods to generate diverse proteins, including random mutagenesis and recombination, are currently available and most of them accumulate the mutations on the target gene of a protein, whose sequence space remains unchanged. On the other hand, a pool of diverse genes, which is generated by random insertions, deletions and exchange of the homologous domains with different lengths in the target gene, would present the protein lineages resulting in new fitness landscapes. Here we report a method to generate a pool of protein variants with different sequence spaces by employing green fluorescent protein (GFP) as a model protein. This process, designated functional salvage screen (FSS), comprises the following procedures: a defective GFP template expressing no fluorescence is first constructed by genetically disrupting a predetermined region(s) of the protein and a library of GFP variants is generated from the defective template by incorporating the randomly fragmented genomic DNA from Escherichia coli into the defined region(s) of the target gene, followed by screening of the functionally salvaged, fluorescence-emitting GFPs. Two approaches, sequence-directed and PCR-coupled methods, were attempted to generate the library of GFP variants with new sequences derived from the genomic segments of E.coli. The functionally salvaged GFPs were selected and analyzed in terms of the sequence space and functional properties. The results demonstrate that the functional salvage process not only can be a simple and effective method to create protein lineages with new sequence spaces, but also can be useful in elucidating the involvement of a specific region(s) or domain(s) in the structure and function of protein.

Amino Acid Sequence↗

Osteoarthritis phenotypes: advancing precision medicine through clinical, structural, and molecular stratification.

PURPOSE: Osteoarthritis (OA) is now understood as a heterogeneous syndrome driven by diverse biological, biomechanical, metabolic, genetic, and molecular mechanisms. This variability explains differences in disease progression and treatment response, challenging the traditional "one-size-fits-all" approach. This review highlights OA phenotyping as a key step toward precision medicine, focusing on clinical, structural, and molecular classifications that inform individualized care. METHODS: A narrative review was conducted using a non-systematic search of major databases and Osteoarthritis Research Society International sources (2010-2026). Evidence was thematically synthesized across clinical, imaging, and molecular domains to characterize OA phenotypes and their potential relevance to precision medicine. RESULTS: Multiple OA phenotypes were identified: inflammatory, metabolic, biomechanical, cartilage-subchondral, pain-sensitization, and aging/senescence. These exhibit distinct clinical features, risk factors, and therapeutic responses. Imaging-based phenotypes (e.g., inflammatory, meniscus-cartilage, subchondral bone, atrophic, hypertrophic) and molecular endotypes (low turnover, structural damage, systemic inflammation) further refine stratification. Pain-structure discordance is notable in sensitization phenotypes and may predict poorer surgical outcomes. Joint-specific variations and emerging genomic and epigenetic insights underscore disease complexity. Advances in imaging, biomarkers, and machine learning may enable earlier detection and patient clustering, though clinical application remains limited. CONCLUSION: Phenotype- and endotype-based classification represents a critical advancement toward precision OA management. Tailored interventions based on stratification hold promise for improving outcomes; however, clinical translation remains limited by overlapping phenotypes, lack of validated biomarkers, and inconsistent results from phenotype-driven trials. Wider clinical adoption requires standardized definitions, validation across joints, and integration of multimodal diagnostic tools into routine practice.

Humans↗