PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Low-coverage whole-genome sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

15 recordsLinked to original sources

Large-scale low-coverage whole-genome sequencing reveals the genetic architecture of wool and growth traits in fine-wool sheep.

Breeding sheep with superior growth performance and wool quality is essential for the sustainability of the fine-wool sheep industry. In this study, we perform low-coverage whole-genome sequencing (lcWGS) on 3842 individuals from 5 sheep breeds (4 fine-wool and 1 semi-fine wool) and generate a large genomic dataset. By comparing these breeds with coarse-wool sheep, we characterize the genomic landscape and selection signatures of fine-wool sheep. We identify several known functional genes associated with hair follicle development and skin morphology, including EGFR, KRT74, EDAR, EREG, and GLI2. Furthermore, GWAS of 19 traits identifies 156 candidate genes significantly associated with growth and wool characteristics, including LCORL for body size, EGFR for clean wool yield, and PRDM1 for fiber diameter. Notably, EGFR is detected in both GWAS and selection signature analyses, indicating its important role in phenotype formation and historical selection. Overall, our findings reveal the genetic basis of growth and wool traits in fine-wool and semi-fine wool sheep, highlight EGFR, LCORL, and PRDM1 as candidate genes, and provide valuable genomic resources and candidate markers for future functional validation and molecular breeding.

Body size↗

Chromosomal instability by low-coverage whole-genome sequencing assay predicts prognosis in bladder cancer patients underwent radical cystectomy.

PURPOSE: To investigate chromosomal instability (CIN) in tumor tissue from radical bladder resection and to evaluate whether it can be used as a biomarker for the molecular typing of (BC). METHODS: DNA was extracted from formalin-fixed paraffin-embedded samples of 50 BC patients who were followed up to March 23 2023 using the Qiagen nucleic acid kits. We analyzed CIN in tumor of bladder by low-coverage whole genome sequencing (LC-WGS). Kaplan-Meier log-rank test was used to perform survival analysis. The association between variables and overall and progression-free survival was analyzed using the Cox proportional hazards model. RESULTS: There were 44 genome segments with statistically significant changes in copy number. CIN was significantly correlated with tumor stage, lymph node metastasis, relapse and survival status. Patients with high CIN were found to have a worse survival, with a median overall survival (OS) of 15 months. In addition, patients with high CIN were more likely to relapse, with a median progression-free survival (PFS) of 7 months. Patients with low CIN showed better OS and PFS. However, there was no significant difference in OS and PFS between T2 and T3-T4 patients. Multivariate cox regression analysis showed that high CIN was an independent predictor of OS, and high CIN and muscle invasion were independent predictors of PFS. Furthermore, patients with abnormal copy number of a single chromosome also had a poor prognosis, with a median survival of 14-30 months for OS and 5-10 months for PFS, while negative patients had a better prognosis. CONCLUSION: CIN was significantly correlated with tumor stage, lymph node metastasis, relapse and survival status of BC. Patients with high CIN or abnormal copy numbers of a single chromosome have a poor prognosis. CIN might be better than T stage in predicting the prognosis of patients with BC. Molecular typing of CIN can be used as an independent prognostic factor for BC.

Humans↗

Genetic rescue stabilizes diversity in small isolated populations of Bonneville cutthroat trout.

Genetic diversity loss due to anthropogenic factors is occurring rapidly on a global scale, putting many species at risk of extirpation and extinction. Different management strategies have been developed to slow this loss; however, it is often unknown whether these strategies reach their intended goals. In this study, we evaluate population structure and changes in nucleotide diversity (π) in isolated populations of Bonneville cutthroat trout (Oncorhynchus clarkii utah) from the Snake Range (Nevada, USA). Starting in the 1990s, three of these populations were used to reestablish populations in the Snake Range because many of the historic populations were extirpated. Some populations were stocked using a single-source and others were stocked using multiple-sources. Using low-coverage whole-genome sequencing coupled with historic samples (2003-2010) and contemporary samples (2019-2022), we find that single-source populations lost nucleotide diversity while mixed-source populations maintained nucleotide diversity. Further, source populations used to restore populations throughout the Snake Range lost the most nucleotide diversity over the time span evaluated. Our findings provide insight into how small, isolated populations can be managed to maintain genetic diversity.

Animals↗

Genetic analysis of three patients from two unrelated Chinese families with autosomal recessive spastic ataxia of Charlevoix-Saguenay.

Autosomal recessive spastic ataxia of Charlevoix-Saguenay (ARSACS) is a rare early-onset neurodegenerative disorder characterized by progressive cerebellar ataxia, spasticity, and sensorimotor peripheral neuropathy. This disorder is caused by homozygous or compound heterozygous variants in the sacsin (SACS) gene on chromosome 13q12.12. Three patients with ARSACS from two unrelated Chinese families were recruited for this study. Patient #1 was an 18-year-old male who had been walking unstably for 12 years. Patient #2, the younger sister of Patient #1, was a 5-year-old girl who had been walking unstably for 2 years. Patient #3 was a 19-year-old female who had been walking unstably and a tendency to fall for 17 years. For Patient #1, whole-exome sequencing (WES) identified a hemizygous variant c.8310_8313delAGAT (p.Asp2771fs4*) in SACS (NM_014363.6), with the father being heterozygous, the mother wild-type, and Patient #2 hemizygous, as verified by Sanger sequencing. Additional copy number variant analysis of the WES data indicated that Patient #1 had a heterozygous gross deletion of chr13q12.12 (chr13:23,808,732 - 24,890,322). Low-coverage whole-genome sequencing results revealed that Patient #2 carried a chr13q12.12 deletion (chr13:23,520,000-24,940,000). Together with Sanger sequencing results, this gross deletion was speculated to have been inherited from the mother, further explaining the hemizygous state of c.8310_8313delAGAT (p.Asp2771fs4*) in Patients #1 and #2. Through WES, Patient #3 was identified as having suspected compound heterozygous variants of c.2881 C > T (p.Arg961*) and c.6409 C > T (p.Gln2137*), inherited from the father and mother, respectively, as confirmed by Sanger sequencing. This study identified three variants in SACS. The c.8310_8313delAGAT (p.Asp2771fs4*) is novel, whereas c.2881 C > T (p.Arg961*) and c.6409 C > T (p.Gln2137*) have been reported previously. Moreover, this study highlights the growing trend that ARSACS has become increasingly prevalent worldwide rather than being localized to a specific region or race. As an increasing number of patients with ARSACS are diagnosed, the genetic spectrum of ARSACS will gradually broaden, providing an accurate genetic basis for prenatal diagnosis of mothers in the years ahead, if possible.

Adolescent↗

Genetic Adaptation to Brackish Water and Spawning Season in European Cisco.

How species adapt to diverse environmental conditions is essential for understanding evolution and the maintenance of biodiversity. The European cisco (Coregonus albula) is a salmonid that occurs in both fresh and brackish water, and this together with the presence of sympatric spring- and autumn-spawning lacustrine populations provides an opportunity for studying the genetics of adaptation in relation to salinity and timing of reproduction. Here, we present a high-quality reference genome of the European cisco based on PacBio HiFi long read sequencing and HiC-directed scaffolding. We generated low-coverage whole-genome sequencing data from 336 individuals across 12 population samples to explore population structure and genetics of ecological adaptation. We found a major subdivision between two groups of populations most likely reflecting colonisation from different glacial refugia. Within the two major groups, we detected further genetic differentiation between spring- and autumn-spawning populations and between populations from freshwater lakes, rivers and brackish water (Bothnian Bay). A genome-wide screen for genetic differentiation among populations identified a set of outlier SNPs strongly correlated with spawning timing and salinity. Several of the genes associated with spawning time, including BHLHE40, TIMELESS and CPT1A, have previously been shown to have a role in circadian rhythm biology. As many as 17 loci were associated with genetic differentiation between populations reproducing in fresh and brackish water. This study provides insights into the genomic basis of ecological adaptation in European cisco with implications for sustainable fishery management.

Animals↗

Genomics Detects Japanese and Pacific Sardine (Sardinops spp.) Hybrids in the Northeast Pacific Ocean.

Sardine (Sardinops spp.) are ecologically important forage fishes distributed globally across temperate, coastal upwelling zones and, when abundant, they support major fisheries. Previous genomic analyses of Pacific Sardine (S. sagax) in the Northeast Pacific detected the presence of Japanese Sardine (S. melanosticta), a species typically found in the Northwest Pacific, along the west coast of North America starting in 2022 over multiple years. To facilitate continued monitoring, we developed a highly accurate species identification Genotyping-in-Thousands-by-sequencing (GT-seq) panel consisting of 88 single nucleotide polymorphisms (SNPs). This panel was constructed by utilizing low-coverage, whole-genome sequence data to identify highly divergent, genome-wide loci between Pacific and Japanese Sardine, enabling unambiguous identification of these morphologically indistinguishable species as well as the detection of hybrids. Using the novel panel, we genotyped 1821 sardine samples from the Northeast Pacific and identified 35 hybrid individuals, which were collected from 2023 to 2025 and are the first known observation of such hybrids. Most were F1 hybrids (33); however, two individuals collected in 2025 appeared to be backcrosses with S. sagax. Our newly developed panel is a powerful resource for the continued monitoring of Japanese Sardine in the Eastern Pacific and is critical for investigating the potential fitness consequences of hybridization with Pacific Sardine.

Journal Article↗

Genome sequence analysis provides evidence that a boreal crustacean colonised Svalbard well before the ongoing Atlantification of the Arctic.

The study of present-day species distributions often raises questions about historical demography. A particularly interesting phenomenon to put in historical context is contemporary human-induced atlantification and its role in reshaping Arctic ecosystems. Despite this, the colonisation history of the Arctic remains generally understudied. In this study, we investigated the demographic history of the northern acorn barnacle, Semibalanus balanoides, a typically boreal species on the Svalbard Archipelago. Our focus was to determine the source and timing of its colonisation of this Arctic archipelago. Using low-coverage whole-genome sequence data, we evaluated two competing hypotheses: whether S. balanoides populations colonised Svalbard through ancient natural processes before the Anthropocene, or if their appearance is more recent, either natural or a consequence of growing anthropogenic influences, such as increased connectivity and global warming. Our results suggest that this boreal species expanded into the Arctic during the later phase of the Holocene Thermal Optimum, well before human-induced climate change.

Animals↗

Genomic consequences of admixture in an experimentally founded sand lizard population.

Conservation interventions are increasingly required for species threatened by population declines and isolation due to anthropogenic pressures. Small, isolated populations are particularly vulnerable to the loss of genetic diversity, increased inbreeding, and the accumulation of deleterious mutations. Translocations or supplementation of allopatric individuals for genetic rescue may be the only way to increase genetic diversity and increase population persistence via increased adaptive potential. Here, we use an experimentally admixed population of sand lizards on a small island in Sweden as a valuable model of genetic rescue. This population was established approximately 20 years ago (5-6 generations), resulting in increased fecundity and hatchling viability. This population was founded from crossings between individuals from an inbred population from the nearby mainland and individuals sourced from populations in southern Sweden. Low-coverage whole-genome sequencing revealed elevated genetic diversity and reduced realized genetic load in this admixed population relative to the source populations. Ancestry analyses indicated a greater contribution of southern Swedish genetic variation, potentially reflecting the contribution of beneficial adaptive variation from this region that may underlie the positive population effects. This system provides valuable empirical insights into the long-term genomic consequences of genetic rescue in this model vertebrate population.

Journal Article↗

Non-invasive strategy for gastric cancer detection: Integration of cell-free DNA fragmentomics and protein biomarkers.

Gastric cancer (GC) ranks as the fifth most common cancer worldwide, however, accurate and non-invasive diagnostic modalities for GC remain limited. Cell-free DNA (cfDNA) fragmentomics has emerged as a promising tool for cancer cell detection. Here we develop a gastric cancer detection model, named GaFraD model. The GaFraD model uses four cfDNA fragmentomics features, including fragment size ratio (FSR), copy number variation (CNV), 9-bp end motif (Motif), and fragment size at transcription start sites (TF). This model achieves an area under the receiver-operating characteristic curve (AUC) of 0.970 (95% CI: 0.944 - 0.990), a sensitivity of 95.0% and a specificity of 80.9%. By combining the GaFraD model and conventional protein biomarkers CA19-9 and PG-I/PG-II, the CONFIRM model was generated. The CONFIRM model attained an AUC of 0.986 (95% CI: 0.966 - 1.000), a sensitivity of 95.0% and a specificity of 95.6% in detecting GC. Moreover, the CONFIRM model achieved remarkable performance (AUC = 0.983, sensitivity 95.6%, specificity 94.2%) in distinguishing patients with early-stage GC from controls. Our work showed the high discriminatory power in distinguishing GC patients from controls, indicating the clinical potential of using cfDNA fragmentomics combined with protein biomarkers for non-invasive GC detection. The results of the study provide a new avenue for early, accurate, and non-invasive clinical diagnosis of GC.

Cell-free DNA↗

Whole Genome Development of Specific Alien-Chromosome Oligo (SAO) Markers for Wild Peanut Chromosomes Based on Chorus2.

The cultivated peanut (Arachis hypogaea L.) is a globally important oilseed and economic crop, but its narrow genetic base limits breeding progress. Wild Arachis species represent valuable genetic resources for enhancing the resilience of the peanut cultigen. While wild species from section Arachis are widely used in breeding programs, the detection of alien chromosomes in hybrids remains challenging due to limited molecular tools. In this study, a cost-effective and efficient system was established for generating species-specific molecular markers using low-coverage next-generation sequencing data, bypassing the need for whole-genome assembly. Utilizing the Chorus2 software, specific alien-chromosome oligo (SAO) markers were developed for four wild species, A. duranensis (accession A19), A. pusilla (A10), A. appresipilla (A33), and A. glabrata (G2 and G3). A total of 1166 primer pairs were designed, resulting in 220 SAO markers specific to A. duranensis, 77 to A. pusilla, 112 to A. appresipilla, 69 to A. glabrata G2, and 59 to A. glabrata G3, with the highest development efficiency observed in A. duranensis (55.0%). These markers span all chromosomes of the five wild accessions. Genome-wide, chromosome-specific SAO markers enable the efficient detection of introgressed alien chromosomes and provide insight into syntenic relationships among homoeologous chromosomes. These markers offer an effective tool for identifying favorable genes and facilitating targeted introgression for the genetic improvement of the cultivated peanut.

Chorus2↗

Genomic peculiarity of coding sequences and metabolic potential of probiotic Escherichia coli strain Nissle 1917 inferred from raw genome data.

Probiotic Escherichia coli strain Nissle 1917 (O6:K5:H1) is a commensal E. coli isolate that has a long tradition in medicine for the treatment of various intestinal disorders in humans. To elucidate the molecular basis of its probiotic nature, we started sequencing the genome of this organism with a whole-genome shotgun approach. A 7.8-fold coverage of the genomic sequence has been generated and is now in the finishing stage. To exploit the genome data as early as possible and to generate hypotheses for functional studies, the unfinished sequencing data were analyzed in this work using a new method [Sun, J., Zeng, A.P., 2004. IdentiCS--identification of coding sequence and in silico reconstruction of the metabolic network directly from unannotated low-coverage bacterial genome sequence. BMC Bioinformatics 5, 112] which is particularly suitable for the prediction of coding sequences (CDSs) from unannotated genome sequence. The CDSs predicted for E. coli Nissle 1917 were compared with those of all five other sequenced E. coli strains (E. coli K-12 MG1655, E. coli K-12 W3110, E. coli CFT073, EHEC O157:H7 EDL933 and EHEC O157:H7 Sakai) published to date. Five thousand one hundred and ninety-two CDSs were predicted for E. coli Nissle 1917, of which 1065 were assigned with enzyme EC numbers. The comparison of all predicted CDSs of E. coli Nissle 1917 to the other E. coli strains revealed 108 CDSs specific for this isolate. They are organized as four big genome islands and many other smaller gene clusters. Based on CDSs with EC numbers for enzymes, the potential metabolic network of Nissle 1917 was reconstructed and compared to those of the other five E. coli strains. Overall, the comparative genomic analysis sheds light on the genomic peculiarity of the probiotic E. coli strain Nissle 1917 and is helpful for designing further functional studies long before the sequencing project is completely finished.

Chromosome Mapping↗

THOR: targeted high-throughput ortholog reconstructor.

Low-coverage genomes (LCGs) are becoming an increasingly important source of data for phylogenetic studies. However, assembly of these genomes is time consuming, difficult and lags behind sequence generation. THOR is a fast, stringent application for targeted reconstruction of sequence orthologs in unassembled LCGs. Using a 4x coverage set of mouse whole-genome sequence reads, THOR could partially or completely reconstruct 416/1000 human promoter ortholog regions in approximately 7.3 min/promoter. THOR's reconstruction rate improves markedly with both higher-coverage, and less divergent target species.

Algorithms↗

Benchmarking Assembly-Free K-mer Methods for Species Identification in Complex Plant Groups: A Case Study in Populus.

Species identification in taxonomically complex plant groups is frequently limited by the inadequacy of organellar markers, whose phylogenetic signal is disrupted by cytonuclear discordance and chloroplast capture. Using the taxonomically complex genus Populus as a model, we evaluated an assembly-free k-mer workflow against a curated SNP reference benchmark. Whole-genome resequencing data from 235 Populus individuals were curated to a 202-individual, 34-species reference dataset in which all retained species are strictly monophyletic in a genome-wide SNP analysis. Independent maximum likelihood analyses further confirmed that the 31 non-hybrid backbone species each maintained high-support monophyly, while taxa of documented reticulate origin showed placement patterns consistent with their reticulate histories. ABBA-BABA D-statistics detected widespread residual allele sharing within the backbone, though the strongest signals did not correspond to the species pairs responsible for the few k-mer identification failures. Against this benchmark, complete plastomes showed limited resolution, recovering only 3.0% species monophyly and 71.1% nearest-neighbor assignment. The optimized k-mer workflow, operating directly on raw reads without assembly or alignment, recovered 91.2% species monophyly, 99.0% nearest-neighbor assignment, and 98.0% group-average assignment. K-mer length was the primary accuracy-controlling parameter, with k = 31 falling within a stable accuracy plateau. Distance-based metrics reached near-saturation at 0.2× sequencing depth, indicating that low-coverage genome skimming can support scalable nuclear genome-based identification with standard computational resources. K-mer distance heatmaps also flagged unusual genomic affinities in hybrid-origin and outlier samples, providing a rapid screen for subsequent population genomic analyses. These results support assembly-free k-mer distances as an efficient tool for reference-based species identification and sample screening in complex plant groups, with residual limitations concentrated near recently diverged species boundaries. Model-based phylogenomic, coalescent, and network analyses remain necessary for resolving deeper species relationships and detailed introgression histories.

Populus↗

Hierarchical scaffolding with Bambus.

The output of a genome assembler generally comprises a collection of contiguous DNA sequences (contigs) whose relative placement along the genome is not defined. A procedure called scaffolding is commonly used to order and orient these contigs using paired read information. This ordering of contigs is an essential step when finishing and analyzing the data from a whole-genome shotgun project. Most recent assemblers include a scaffolding module; however, users have little control over the scaffolding algorithm or the information produced. We thus developed a general-purpose scaffolder, called Bambus, which affords users significant flexibility in controlling the scaffolding parameters. Bambus was used recently to scaffold the low-coverage draft dog genome data. Most significantly, Bambus enables the use of linking data other than that inferred from mate-pair information. For example, the sequence of a completed genome can be used to guide the scaffolding of a related organism. We present several applications of Bambus: support for finishing, comparative genomics, analysis of the haplotype structure of genomes, and scaffolding of a mammalian genome at low coverage. Bambus is available as an open-source package from our Web site.

Algorithms↗

Development of a low-coverage whole genome sequencing screen for apomixis using a diverse set of Malus germplasm.

In the past decade, plant biologists have made several major discoveries pertaining to the genetic basis of apomixis (clonal propagation by seed) that have shown promise in preserving high-value hybrid rice and sorghum genotypes. This progress was made possible by foundational gene discovery efforts in model species and natural apomicts, but pleiotropic obstacles still limit its broad agricultural adoption, especially in eudicots. Thus, it follows that investigations of novel apomicts should lead to the development of new molecular tools for plant breeding. The two most common ways to identify clonal seed production are flow-cytometry seed screens and genome sequencing to compare the DNA sequences of the maternal parent and progeny, traditionally using low-throughput markers. While flow-cytometry has been the dominant method for more than two decades, it provides indirect information on the genetics of a resulting embryo and can be ineffective in certain species. Here we developed a method using short-read whole-genome sequencing at moderately low coverage (averaging 3X and 6X) to screen diverse Malus genotypes maintained in a USDA germplasm collection for clonal seed production. In total, we sequenced 55 genotypes, 1,216 of their embryos, and identified 17 previously undescribed apomictic genotypes. Several more were detected with the flow cytometry seed screen, which helped resolve certain types of reproduction and sources of noise in low-coverage datasets. This low-pass screening-by-sequencing method is a relatively low-cost, rapid method for detecting apomictic genotypes in diverse plant germplasm and when used thoughtfully in conjunction with flow cytometry, provides a new way to visualize the genetic outcomes of sexual and asexual reproduction in plants.

Apomixis↗