PubMed HealthSearch

SEARCH · PubMed Health

Results for “genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri

Low-pass whole-genome sequencing reveals genomic diversity and ecotype-specific adaptation in indigenous Tigrayan chickens.

Indigenous chickens play a critical role in food security and climate resilience in smallholder systems, yet their genomic diversity and adaptive potential remain insufficiently characterised. This study employed low-pass whole-genome sequencing (LP-WGS; 0.2-1.99×) to investigate genomic diversity, population structure, inbreeding and candidate environment-associated genomic variation in 33 chickens from highland, midland, and lowland agroecologies in the Tigray region of northern Ethiopia. After imputation and stringent filtering, 23.4 million high-confidence SNPs were retained, including ~ 17% novel variants, indicating substantial uncharacterised genetic diversity in these populations. SNP density (13.8 ± 8.6 SNPs/kb) was comparable to values reported from high-coverage Ethiopian chicken datasets, demonstrating the suitability of LP-WGS for population genomics in resource-limited settings. Marked differences in genomic diversity were observed among ecotypes: midland chickens showed the highest nucleotide diversity (π = 0.00267), followed by lowland (π = 0.00233), whereas highland chickens showed the lowest diversity (π = 0.00203) and elevated genomic inbreeding (FROH and FHOM ≈ 0.18). Population structure analyses revealed clear genetic separation among ecotypes. PCA (13.91% variation explained) distinguished lowland chickens along PC1 and separated highland from midland along PC2, while ADMIXTURE and FST patterns supported three major ancestral genomic backgrounds. Functional annotation of private missense variants uncovered distinct adaptive signatures reflecting the contrasting agroecological conditions. Highland chickens showed enrichment of candidate genes potentially involved in physiological processes relevant to high-altitude environments, including cold response, angiogenesis, cardiovascular regulation and metabolic homeostasis (eg., PARP1, ACOX2, ITGB3, EDNRB, SOX8, and SOX10). Midland chickens exhibited candidate signals of selection in genes with known roles in innate antiviral immunity, bacterial defence and inflammatory regulation (eg., BAK1, CLSTN1, CYSLTR1, CYSLTR2, CXCR7, GIPR, DSCAM, GDAP1, TLR3, TLR4, TLR7, IFIH1, ADORA1, EPHB1, and TMPRSS2). Lowland chickens displayed candidate variants associated with heat-stress response, DNA damage repair, oxidative balance and cardiovascular support under extreme temperatures (e.g., MLH1, BDKRB1, GPR19, FLT1, CCL18, TGM2, and RAMP3). Overall, the results indicate substantial genomic differentiation among ecotypes and suggest candidate environment-associated genetic divergence across Tigray's diverse agroecological zones. These populations may represent important reservoirs of adaptive genetic variation for climate-resilient poultry breeding, warranting further functional validation and conservation-oriented management.

Animals

Genome sequencing and population genomics provide insights into the demographic history, genetic load, and local adaptation of an endangered Tertiary relict.

Endangered Tertiary relict trees represent an exceptional evolutionary heritage with small and isolated populations, yet little is known about how demographic history, local adaptation, and genetic load have affected their long-term survival and extinction risk. We performed whole-genome sequencing and population genomic analyses on Ulmus elongata L. K. Fu & C. S. Ding, an endangered Tertiary relict tree endemic to East Asia. By integrating genomes from U. elongata and seven other endangered trees from public databases, we identified rate-decelerated genes across endangered trees and genes under positive selection of U. elongata associated with tissue development, detoxification, and immune response, and signal transduction and regulation mechanisms potentially leading to endangered status. Demographic analyses revealed continuous population decline from the late Miocene to present, especially during the last glacial maximum (LGM) and last 10&#x2009;000&#x2009;years. Spearman correlation indicated a strong negative relationship between effective population size and human population density (rpopulation density&#x2009;=&#x2009;-0.90, P&#x2009;<&#x2009;0.001) as well as cropland use (rcropland use&#x2009;=&#x2009;-0.89, P&#x2009;<&#x2009;0.001). Genotype-environment association (GEA) analyses identified a set of candidate genes associated with temperature and precipitation, supporting a polygenic adaptation model in U. elongata. Overall, our findings underscore the severe population bottlenecks that have led to the fixation of strongly deleterious mutations and inbreeding, further compromising the adaptive potential and long-term viability of U. elongata. Furthermore, assessments of genomic vulnerability under future climate scenarios revealed higher genetic offsets in northern region of Fujian and Jiangxi populations, suggesting these regions require prioritized conservation efforts due to reduced adaptive capacity.

Endangered Species

META-DIFF: a k-mer-based pipeline that detects differentially abundant sequences in metagenomics whole genome sequencing.

Traditional case-control metagenomic studies are constrained by their dependence on taxonomic and functional databases. Because annotation occurs before differential analysis, they are limited to known elements and keep function and taxonomy separate. Although binning strategies have emerged to reconstruct genomes and mitigate this issue, they still require an assembly step, preventing the use of all available sequencing data. Here, we introduce META-DIFF, a pipeline based on differentially abundant k-mers independently of any prior annotation. From those k-mers, it reconstructs longer sequences and provides biological context, as well as the best set of unitigs to discriminate between conditions. Across both taxonomy-centric and functionally-centric benchmarks, it showed robust performance and displayed great reproducibility. It also behaved more conservatively than did other univariate methodologies, i.e. it maintained a high precision at the expense of recall, particularly in conditions of low fold-change and limited sequencing depth. The efficacy of META-DIFF was further validated through its application to a real-world colorectal cancer dataset, which produced both confirmatory and novel results compared with those of previous publications. The pipeline is able to exploit all reads and identify differentially abundant elements, including unknown DNA, prior to annotation. With the guidelines provided, META-DIFF provides users with great exploratory power to unravel microbiome changes.

Metagenomics

Comparing the performance of exome and genome sequencing for rare disease diagnostics: A randomized implementation effectiveness trial.

PURPOSE: Exome sequencing (ES) and genome sequencing (GS) can improve rare disease diagnosis but are not routinely available in many jurisdictions. To inform implementation, we report on a randomized implementation effectiveness trial comparing ES and GS. METHODS: Eligible trios were randomized to receive ES or GS in the same clinically accredited laboratory. Patient-level data on diagnostic utility and turnaround times were collected. Outcomes were compared statistically between clinically important subgroups. RESULTS: Of 1048 patients, 68.5% had syndromic intellectual disability/developmental delay (ID/DD) and 20.5% had multisystem disorders without ID/DD. Most had prior genetic test(s) that were nondiagnostic (95.5%), and of these, 91.6% included chromosome microarray. Diagnostic yields were 33.8% and 33.6%, for ES (n = 526) and GS (n = 522), respectively. Within sequencing groups, diagnostic results were more frequent among those with ID/DD than those without (P < .005). For routine (ie, nonexpedited) patients (n = 1020), 87.0% were reported in <12 weeks, and the mean turnaround time was 55.5 days (SD: 24.0). Turnaround time for ES and GS did not differ; however, result type (P < .001) and age of onset (P < .005) significantly affected turnaround time. CONCLUSION: Findings provide robust evidence of diagnostic utility and timeliness of ES and GS and will inform policy related to the organization, delivery, and reimbursement of clinical-grade genome diagnostics for rare diseases.

Adolescent

Costs and cost-effectiveness of returning secondary findings from genomic sequencing based on the return of additional findings in the 100,000 Genomes Project.

PURPOSE: To assess costs and cost-effectiveness of returning additional findings from genome sequencing using data from the 100,000 Genomes Project (100kGP). METHODS: A model-based cost-utility analysis combining yield, consent rates, and cost data from the 100kGP with published estimates of downstream costs and quality-adjusted life years expected to accrue over a lifetime, after the identification of a pathogenic variant. RESULTS: The cost of returning additional findings to participants in the 100kGP was &#xa3;7.1m or &#xa3;81 per participant, with a yield of 0.85% for consented participants. The estimated lifetime incremental cost per participant was &#xa3;125 and quality-adjusted life years 0.004, giving an incremental cost-effectiveness ratio of &#xa3;28,830. Implementing a policy of returning additional findings is unlikely to be cost-effective (ie, 13%) at a willingness-to-pay threshold of &#xa3;20,000. A short-term cost of returning findings of &#xa3;43 per participant or lower (compared with the base case of &#xa3;81) would result in an incremental cost-effectiveness ratio of less than &#xa3;20,000. Alternatively, cost-effectiveness may be improved by returning additional findings to younger patient populations. CONCLUSION: Return of additional findings following genome sequencing for this group of conditions may not be a cost-effective use of health care system resources. Our cost-effectiveness outcomes rely on published estimates and should be validated through long-term follow-up data.

Humans

A tiled amplicon protocol for culture-free whole-genome sequencing of M. tuberculosis from clinical specimens.

Whole-genome sequencing of Mycobacterium tuberculosis can be a valuable tool for TB surveillance and treatment, providing insights into transmission patterns and comprehensive drug susceptibility testing. However, the slow growth of M. tuberculosis means traditional culture-based sequencing methods can take weeks to return results, which has limited the widespread adoption of these techniques and limited their use in clinical decision-making. Tiled amplicon sequencing is a fast, reliable, and cost-effective method of whole-genome sequencing that can be done directly on clinical specimens and has been implemented at scale in academic and public health laboratories across the world; it was the cornerstone of SARS-CoV-2 sequencing and has been adapted for a wide range of viral pathogens. However, similar methods are not yet available for far larger bacterial genomes. Extending this approach to M. tuberculosis would significantly reduce the cost, labor, and turnaround time for whole-genome sequencing. We designed a tiled amplicon panel consisting of 5,128 primers that covers the entire M. tuberculosis genome, the largest tiled amplicon sequencing panel we are aware of to date. Applying our amplicon panels to clinical samples of sputum, we show the ability to recover whole-genome bacterial sequences without the need for culture. The resulting sequence data can be used to determine M. tuberculosis lineage and reliably identify markers of drug resistance. Using this approach in clinical settings could reduce the time needed for comprehensive drug susceptibility testing from weeks to days and enable genomic epidemiology to be performed at scale, even in resource-limited settings.IMPORTANCEWe have developed and tested an amplicon panel, TB-seq, for the priority pathogen Mycobacterium tuberculosis, demonstrating recovery of near-full genomes directly from patient sputum, including mixed and low-concentration samples. This approach significantly reduces the turnaround time for this slow-growing bacterium while maintaining high accuracy in detecting clinically relevant mutations, including those associated with drug resistance. Given the global burden of tuberculosis and the critical need for faster diagnostic solutions, we believe our method has the potential to improve clinical decision-making and public health strategies.

Mycobacterium tuberculosis

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral

ZIPcnv: accurate and efficient inference of copy number variations from shallow whole-genome sequencing.

MOTIVATION: Shallow whole-genome sequencing (sWGS), a rapid and cost-effective sequencing technology, has gradually been widely adopted for CNV analyses. However, with genome&#x2011;wide coverage of only 0.1-5&#xd7;, sWGS data display a pronounced zero&#x2011;inflation phenomenon-a large fraction of loci has zero sequencing reads. Zero inflation causes read counts to fluctuate by several&#x2011;fold between adjacent windows. As a result, random upward blips in coverage can be misinterpreted as copy&#x2011;number gains (false positives), and true deletions often become indistinguishable from pervasive zero&#x2011;coverage noise. In addition, existing CNV detection tools developed for sWGS data often struggle to adapt across different CNV sizes. These combined effects severely constrain the accuracy of CNV inference. RESULTS: To address above challenges, we propose ZIPcnv, a novel CNV detection tool specifically designed for sWGS data. First, we apply a segment sliding window to smooth the raw read depth signal, which transforms the original zero-inflated statistical characteristics into approximately normal distribution characteristics. We then design a statistical process model that robustly detects persistent shifts under high background noise using a cumulative sum strategy, classifying genomic regions into candidate and non-candidate CNV regions. Finally, dynamic sliding windows are used for one-pass detection of CNVs of varying lengths, with window size adapting to the CNV region size. We evaluated the performance of ZIPcnv on simulated data and 190 real whole-genome sequencing samples. Experimental results show that ZIPcnv consistently outperforms currently popular CNV detection tools. AVAILABILITY AND IMPLEMENTATION: The ZIPcnv source code is freely available at https://github.com/Nevermore233/ZIPcnv.

DNA Copy Number Variations

The chromosomal genome sequence of the spiny sea fan, Muricea muricata (Pallas, 1766) (Malacalcyonacea: Plexauridae) and its associated microbial metagenome sequences.

We present a genome assembly from a Muricea muricata specimen (spiny sea fan; Cnidaria; Anthozoa; Malacalcyonacea; Plexauridae). The genome sequence has a total length of 453.40 megabases. Most of the assembly (98.45%) is scaffolded into 16 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 19.29 kilobases. Gene annotation of this assembly by Ensembl identified 52 164 protein-coding genes. From the metagenome data, we recovered five bins, of which three were high-quality MAGs.

Malacalcyonacea

The chromosomal genome sequence of the lesser starlet coral, Siderastrea radians (Pallas, 1766) (Scleractinia: Rhizangiidae) and its associated microbial metagenome sequences.

We present a genome assembly from a specimen of Siderastrea radians (lesser starlet coral; Cnidaria; Anthozoa; Scleractinia; Rhizangiidae). The genome sequence has a total length of 807.19 megabases. Most of the assembly (94.17%) is scaffolded into 14 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 19.38 kilobases. Gene annotation of this assembly by Ensembl identified 47 051 protein-coding genes. From the metagenome data, we recovered two binned metagenomes assigned to the bacterial phylum Bacteroidota and class Bacteroidia.

Scleractinia

The chromosomal genome sequence of the maze coral, Meandrina meandrites (Linnaeus, 1758) (Scleractinia: Meandrinidae) and its associated microbial metagenome sequences.

We present a genome assembly from a specimen of Meandrina meandrites (maze coral; Cnidaria; Anthozoa; Scleractinia; Meandrinidae). The genome sequence has a total length of 551.16 megabases. Most of the assembly (99.25%) is scaffolded into 14 chromosomal pseudomolecules. The mitochondrial genome has also been assembled, with a length of 17.2 kilobases. Gene annotation of this assembly by Ensembl identified 30 464 protein-coding genes. We recovered two bins from the metagenome data.

Meandrina meandrites

Genomic Sequencing in Neonatal Encephalopathy and Suspected Hypoxic-Ischaemic Encephalopathy: A Systematic Review.

BACKGROUND: Neonatal encephalopathy (NE) is a major cause of neonatal mortality and long-term neurological disability. Although hypoxic-ischaemic encephalopathy (HIE) is the most common cause, several genetic disorders may mimic or coexist with hypoxic-ischaemic injury. Next-generation sequencing has emerged as a promising diagnostic tool in this setting. This systematic review evaluated the current evidence on genomic sequencing in NE. MATERIAL AND METHODS: A systematic review was conducted according to PRISMA 2020 guidelines and prospectively registered in PROSPERO. PubMed/MEDLINE, Embase, and Scopus were searched from inception to June 2026. Eligible studies included neonates (&#x2264;28 days) with NE, suspected or confirmed HIE, HIE mimics, or unexplained NE who underwent genomic sequencing. Whole-exome sequencing (WES), whole-genome sequencing (WGS), clinical exome sequencing (CES), rapid genomic sequencing, and targeted next-generation sequencing panels were considered. Study quality was assessed using the Newcastle-Ottawa Scale. RESULTS: Seven studies met the inclusion criteria. Considerable heterogeneity was observed regarding patient selection, sequencing strategies, and reported outcomes. Among diagnostic sequencing studies, diagnostic yield ranged from 23.5% to 53.1%. Pathogenic and likely pathogenic variants were identified in genes associated with developmental and epileptic encephalopathies, metabolic disorders, mitochondrial diseases, and neurodevelopmental syndromes, including SCN2A, KCNQ2, CACNA1A, STXBP1, PTPN11, BCOR, MMUT, COQ2, and GBE1. Genomic sequencing frequently refined or changed the initial diagnosis, improved prognostic assessment and genetic counselling, and, in selected cases, guided disease-specific treatment. One study investigated genetic susceptibility to hypoxic-ischaemic injury rather than diagnostic sequencing. CONCLUSIONS: Genomic sequencing provides clinically meaningful diagnoses in a substantial proportion of neonates with unexplained NE or atypical HIE presentations. Current evidence supports integrating genomic sequencing into the diagnostic evaluation of selected infants, although larger prospective studies are needed to define its optimal timing, clinical utility, and cost-effectiveness.

Humans

Targeted reflex RNA sequencing for enhanced variant classification on exome and genome sequencing improves patient outcomes.

RNA sequencing (RNA-seq) has been utilized to provide functional evidence regarding the impact of splicing variants. This study explores the utility of targeted reflex RNA-seq to inform classification of predicted splicing variants identified through clinical exome sequencing (ES) and genome sequencing (GS). A retrospective analysis was conducted on consecutive ES/GS cases completed at a single center in which targeted reflex RNA-seq was performed following identification of eligible variants. There were 131 cases (4.1%) that had at least one RNA-seq eligible variant reported, with eight of these cases having two unique eligible variants. Of the 139 eligible variants, 125 were classified as variants of uncertain significance (VUS). Sixty-four cases had targeted reflex RNA-seq completed with 27 cases having at least one variant reclassified (42.2%). After reclassification, 23 cases had positive results, and two cases had a likely diagnosis of an autosomal recessive condition. Clinical outcomes data regarding positive RNA-seq cases showed that 71% (10/14) had clinical management changes and 43% (6/14) had treatment changes. Incorporation of targeted reflex RNA-seq analysis into the diagnostic pipeline of rare diseases enhances variant classification and resolves uncertainty regarding predicted splice variants, leading to an estimated 1.6% increase in diagnostic yield of clinical ES/GS.

Journal Article

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans