PubMed HealthSearch

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 73 records · Page 4Linked to original sources

Genomic Footprints of Historical Introgression Between Ancient Lineages of Wild Oryza AA-Genome Species With Widely Separated Contemporary Distributions.

Phylogenetic incongruence is increasingly recognized as pervasive, yet the extent to which reticulate evolution occurs between groups separated by substantial geographical distances and deep phylogenetic divergence remains poorly characterized. In the Oryza AA-genome group-a model for plant speciation and domestication-the traditional bifurcation model posits that Australian Oryza meridionalis and African Oryza longistaminata occupy basal branches, distinct from the more recently diversified monophyletic clade comprising Asian and other African lineages, including major cultivars. However, recent evidence from endogenous viral sequences has hinted at unexpected genetic relatedness between African O. longistaminata and Asian Oryza sativa, which are geographically and phylogenetically distant. Here, we conducted a genome-wide survey across 11 Oryza species to systematically identify genomic regions exhibiting phylogenetic incongruence. Widespread phylogenetic discordance was observed, notably involving genomic segments in which O. longistaminata showed phylogenetic proximity to Asian species, contradicting their established deep divergence. To distinguish between introgression and incomplete lineage sorting, we performed four-taxon ABBA-BABA tests, which provided statistical support for introgression. Furthermore, divergence time estimates for these incongruent regions were younger than the species divergence times, suggesting historical introgression between the ancestors of lineages that are currently separated by vast geographical distances. Systematic assessments indicated that potential analytical artifacts, such as compositional bias and substitution saturation, were unlikely to explain the observations. These convergent lines of evidence suggest that ancient introgression had occurred between currently geographically separated and evolutionarily divergent Oryza lineages, leaving detectable footprints across their modern genomes.

Oryza

Molecular cloning and physical mapping of the genome of simian herpes B virus and comparison of genome organization with that of herpes simplex virus type 1.

The molecular structure of the genome of simian herpes B virus (SHBV) was determined by restriction endonuclease mapping studies. Genomic DNA was cleaved with restriction endonucleases BamHI and SalI into 41 and 58 fragments, respectively. Most of these fragments were cloned into the plasmid vector pACYC184; uncloned fragments were identified following isolation from agarose gels. Terminal fragments were identified by exonuclease digestion and radioactive end-labelling, and linkage of fragments was deduced by a combination of single and double digest experiments and cross-blot hybridizations. The genome is larger than that of herpes simplex virus type 1 (HSV-1), being approximately 165 kilobase pairs. Like that of HSV-1, the SHBV genome is composed of a long and a short unique region each flanked by inverted repeat sequences, which allow the unique regions to invert relative to one another, resulting in four possible isomeric arrangements of the molecule. Genome locations of several SHBV genes were compared with their HSV-1 homologues.

Animals

Population-level genomic surveillance of human norovirus using wastewater-based whole-genome sequencing.

Wastewater-based surveillance has garnered increasing attention as a valuable approach for capturing community-level infection dynamics that are often difficult to detect through clinical reporting systems alone. In this study, we analyzed human norovirus genotype distributions and whole-genome-level variations in wastewater samples collected in Gwangju, Korea. These results were interpreted in conjunction with a documented foodborne outbreak to evaluate the epidemiological relevance of wastewater-based monitoring. Human norovirus concentrations were quantified using TaqMan Array Card-based RT-qPCR, and whole-genome next-generation sequencing (NGS) was performed to obtain viral read counts and reads per kilobase per million filtered reads values. Overall, strong correlations were observed between RT-qPCR-based concentrations and NGS-derived metrics. Genotype dynamics varied among wastewater treatment plants, reflecting differences in catchment size and local population characteristics. In particular, the relative abundance of GII.17[P17] increased during epidemiological week 50, temporally coinciding with a documented local foodborne outbreak. Variant analysis revealed that wastewater samples exhibited mixed nucleotide patterns, with multiple alleles coexisting at varying relative frequencies rather than fixed substitutions. Notably, some nonsynonymous variants detected in clinical samples were also observed in wastewater samples collected surrounding the outbreak period. Together, these findings demonstrate that wastewater-based whole-genome surveillance can capture both genotype-level shifts and nucleotide-level dynamics at the population scale, highlighting its potential as a complementary tool for monitoring community-level norovirus circulation and outbreak-associated genotype dynamics.IMPORTANCEWastewater-based surveillance is increasingly recognized as a promising approach for capturing community-level infection dynamics that are often missed by clinical surveillance. In this study, we applied whole-genome sequencing to wastewater samples collected in Gwangju, South Korea, to comprehensively characterize human norovirus genotype distributions and genetic variation. Distinct genotype patterns were observed across wastewater treatment plants, reflecting differences in catchment population size and local characteristics. Notably, an increase in the GII.17[P17] genotype detected in wastewater coincided with a foodborne outbreak investigated in Gwangju, demonstrating the potential of wastewater surveillance to reflect ongoing community transmission and emerging outbreak-associated genotypes. In addition, wastewater samples contained diverse and coexisting genetic variants, capturing population-level viral diversity and evolutionary dynamics that are not readily detected through clinical surveillance alone. These findings highlight the value of wastewater-based whole-genome surveillance for monitoring community-level viral circulation and support its integration as a complementary strategy to existing clinical surveillance systems.

genotype dynamics

Construction of a Helicobacter pylori genome map and demonstration of diversity at the genome level.

Genomic DNA from 30 strains of Helicobacter pylori was subjected to pulsed-field gel electrophoresis (PFGE) after digestion with NotI and NruI. The genome sizes of the strains ranged from 1.6 to 1.73 Mb, with an average size of 1.67 Mb. By using NotI and NruI, a circular map of H. pylori UA802 (1.7 Mb) which contained three copies of 16S and 23S rRNA genes was constructed. An unusual feature of the H. pylori genome was the separate location of at least two copies of 16S and 23S rRNA genes. Almost all strains had different PFGE patterns after NotI and NruI digestion, suggesting that the H. pylori genome possesses a considerable degree of genetic variability. However, three strains from different sites (the fundus, antrum, and body of the stomach) within the same patient gave identical PFGE patterns. The genomic pattern of individual isolates remained constant during multiple subcultures in vitro. The reason for the genetic diversity observed among H. pylori strains remains to be explained.

Base Sequence

Culture-free genomics: a shift toward genome-wide applications in Chagas disease and leishmaniasis.

INTRODUCTION: Chagas disease and leishmaniasis remain major neglected tropical diseases, with diagnosis and surveillance constrained by low parasite burden, multiclonal infections, and complex parasite biology. Traditional culture-dependent and targeted molecular approaches fail to capture the full genomic diversity of Trypanosoma cruzi and Leishmania spp. limiting clinical and epidemiological utility. AREAS COVERED: We review the evolution from early sequencing to second- and third-generation platforms, highlighting culture-free detection and genomic surveillance. We discuss enrichment strategies (selective whole-genome amplification (SWGA) and capture-enrichment sequencing (CES)) addressing low parasite DNA abundance in complex samples, alongside metagenomics and portable sequencing for field-based surveillance and diagnostics. We further explore how direct-from-host data can improve diagnostics, enhance transmission surveillance, support treatment monitoring, and guide control strategies. EXPERT OPINION: Culture-free genomic approaches represent a transformative advance in kinetoplastid research, providing resolution that culture-dependent methods cannot deliver. Their diagnostic contribution is at present largely indirect, operating through the identification of improved molecular and serological targets rather than through sequencing as the assay itself. Persistent barriers of cost, infrastructure, standardization, and bioinformatics capacity, together with the absence of formal clinical validation, currently confine these methods to research and surveillance settings.

Capture-enrichment sequencing

Doubled Genomes, Divergent Fates: Genomic Insights Into Diversification in an Allotetraploid Cavefish.

Cave environments impose unique challenges that drive remarkable genetic and phenotypic changes in cave-dwelling organisms. In this study, we investigated the genomic basis of adaptation in the small eye golden-line fish (Sinocyclocheilus microphthalmus), an allotetraploid cavefish endemic to Guangxi, China. Using whole-genome resequencing data from 47 individuals across six cave locations, we examined how neutral and selective forces influence diversification. Our analyses uncovered significant population structure indicative of allopatric divergence, along with evidence of locus-specific selection contributing to genomic differentiation. We identified seven single outlier clusters (SOCs), each tied to the divergence of specific populations, underscoring the role of local processes in driving diversity. Genes associated with vision showed relaxed selection, likely reflecting adaptation to darkness, while positive selection on other loci revealed additional functional shifts. Notably, allopolyploidy was found to fuel divergence through subgenome-specific patterns and asymmetric evolution within SOCs and among homoeologs. Taken together, these findings provide valuable insights into mechanisms of cave evolution and illustrate how allotetraploid genomes can facilitate diversification, potentially contributing to speciation in extreme environments.

Animals

A practical guide to studying genome function using single-molecule genomics.

Single-molecule genomics (SMG) has transformed our ability to study the mechanisms that regulate the genome by enabling profiling of the activity of regulatory factors on individual DNA molecules genome-wide. SMG is able to quantify molecular heterogeneity and the co-occurrence of regulatory events, including epigenetic modifications, transcription factor binding and chromatin organization on single DNA molecules. SMG reveals dynamics of chromatin interactions that cannot be measured by conventional genomics assays. Therefore, SMG offers a unique platform to study how regulatory events combine to control genome activity. In this Expert Recommendation article, we provide a practical guide for adopting SMG and outline best practices.

Journal Article

Structure of Herpesvirus saimiri genomes: arrangement of heavy and light sequences in the M genome.

Herpesvirus saimiri contains two species of DNA molecules. (i) The M genome is composed of 70% light (L) DNA (36% cytosine plus guanine; density in CsCl, 1.695 g/ml), which consists of unique sequences, and 30% heavy (H) DNA (71% cytosine plus guanine; density, 1.729 g/ml). (ii) The H genome contains heavy sequences exclusively. H sequences in M and H genomes cross-hybridize completely and are cleaved identically by restriction endonuclease R-Sma I into four classes of fragments with molecular weights of about 360,000, 300,000, 130,000 and 40,000, respectively. H sequences are chains of identical repeat units in tandem arrangement. The molecular weight of each repeat unit is about 830,000. L sequences have no cleavage site for endo R-Sma I H sequences are terminally arranged at both ends of the M genome, as seen by electron microscopy after partial denaturation. The length of the individual heavy ends varies between 21 mum and less than 1 mum, whereas the light region is uniform in size (35.3+/-0.35 mum). As a rule, molecules with a long heavy end at one side have a short heavy end at the other side, thus giving rise to a limited size heterogeneity. Orientation of M DNA molecules by the denaturation map of the light region shows that the longer heavy end may be located at the left or at the right side of the M genome.

Base Sequence

Complete genome of multiply antibiotic resistant ST10 Acinetobacter baumannii isolate NL6 from Vietnam and relationship to available ST10 genomes.

The genome of NL6, a multiply antibiotic-resistant Acinetobacter baumannii ST10:KL49:OCL2 carriage isolate from Vietnam, was sequenced using Nanopore technology, and complete chromosome and plasmid sequences were assembled from the long reads and available short reads. Resistance genes and their locations were identified, and transfer of a conjugative plasmid carrying several resistance genes into a new host was tested. The acquired resistance genes in NL6 were distributed between the chromosome and two of three plasmids present. The chromosome carries multiple copies of several insertion sequences, an incomplete copy of the ISAba1-bounded Tn6250 that includes the sul2 and strAB genes, and an integrative element carrying copper resistance genes designated IECuR. Plasmid pNL6-2 (r3-T5; 15 Kbp) is a Rep_3/OrfX plasmid that includes a tet39 dif module, and pNL6-3 (r3-T20; 66.9 Kbp) carries aacC2d, aphA6, and blaCARB-16 and a second ampC gene preceded by an ISAba1. Conjugation of pNL6-3 into derivatives of ATCC17978 was demonstrated, confirming that the ampC gene confers resistance to third-generation cephalosporins. NL6 was compared to other complete ST10 genomes. Several acquired elements in the chromosome were shared with the ST10 isolate LAC-4 (USA), indicating shared ancestry, but the plasmid content differed. The KL and plasmid content were variable in 17 further complete ST10 genomes downloaded from GenBank. Tn6250 and IECuR were only found together in the chromosome of two further KL49 isolates. Antibiotic resistance in ST10 A. baumannii was acquired mainly via plasmid acquisition, but resistance genes varied, and a variety of plasmids was involved.IMPORTANCEMembers of the CC10 clonal complex of Acinetobacter baumannii comprising ST10 plus single and double locus variants are known to be particularly virulent. However, antibiotic resistance in members of this group has rarely been examined. Here, determination of the complete genome (chromosome and plasmids) of a representative ST10 isolate from Vietnam allowed the context and location of acquired antibiotic resistance genes and of other mobile genetic elements to be determined. Mobile genetic element locations in completed chromosomes facilitate comparisons of potentially related genomes, revealing those with recent shared ancestry. Differences in plasmid content can also be examined.

Acinetobacter baumannii

Genomic Insights Into Multidrug-Resistant Foodborne Serratia liquefaciens Strains Carrying mcr-9 and Comparative Genomic Analysis of Novel Biosynthetic Gene Clusters.

Serratia liquefaciens is an opportunistic nosocomial pathogen with a wide range of antibiotic resistance patterns. This study reports the characterization of the first mcr-9-positive S. liquefaciens strains, 35E-19E1 and CST-066, isolated from meat products in Japan. The strains were screened for the presence of β-lactamases, plasmid-mediated mobile colistin resistance (mcr) genes, and carbapenemase-encoding genes using PCR. Antimicrobial susceptibility was tested using the broth microdilution method. The strains exhibited multidrug resistance (MDR) phenotypes to third-generation cephalosporins, cephamycin, fosfomycin, and other clinically important antimicrobials. Genomic DNA sequencing showed that the genome sizes of CST-066 and 35E-19E1 are 5,529,704 and 5,261,506 bps, respectively. mcr-9 was identified on a chromosome within a genetic environment that included the two-component system qseBC, which plays a key role in the signaling network that triggers colistin resistance in Enterobacterales. Downstream genome analysis revealed a 1695-bp eptB-like kdo2-lipid phosphoethanolamine transferase, which is involved in intrinsic polymyxin resistance mechanisms in Serratia spp. The strain 35E-19E1 carries five CRISPR-Cas enzymes that are essential for adaptive immunity in bacteria, allowing defense against invading elements. Functional analysis using subsystem technology revealed that both strains possess subsystem features responsible for invasion and adhesion within the host biomes. Genome mining using antiSMASH and BAGL4 revealed various biosynthetic gene clusters, responsible for secondary metabolite synthesis. Notably, we identified novel gene clusters, mainly nonribosomal peptide synthetases, in both the strains, indicating their potential to produce bioactive compounds. Although the presence of mcr-9 in Serratia may not be of clinical significance because of natural resistance of the strain to polymyxins, we shed light on the genomic characteristics of this MDR pathogen and the potential spread of mcr-9 among other bacterial species. The emergence of mcr-9 in drug-resistant S. liquefaciens provides significant insights, underscoring the need for increased surveillance of this pathogen.

biosynthetic gene cluster

Pinpointing genomic regions conferring herbicide tolerance in cassava via genome-wide association mapping.

Cassava (Manihot esculenta Crantz) is a tropical crop of major socioeconomic importance, whose productivity can be limited by sensitivity to herbicides used for weed management. This study aimed to perform a genome-wide association study (GWAS) in 194 cassava genotypes to identify genomic regions associated with tolerance to the herbicides mesotrione, S-metolachlor, and chloransulam-methyl. The evaluations performed at 3, 6, 9, 15, and 30 days after application (DAA) were used to characterize the temporal progression of phytotoxicity. Based on this analysis, the phenotype obtained at 9 days after application (PhytoX9DAA) was selected for genome-wide association analyses because it represented the period of greatest symptom expression and the highest discrimination among genotypes. GWAS analyses were performed using de-regressed BLUPs and the MLM, MLMM, and BLINK models, incorporating kinship (K) and population structure (Q) matrices. Significant markers were detected across multiple chromosomes, and the corresponding genomic windows contained candidate genes with functional annotations related to herbicide response. The predominant functional categories included membrane transport, channel activity, signal peptide processing, protein phosphorylation, cellular signaling, and metabolic regulation. Key candidate genes included Manes.02G151900 and Manes.02G152700 (chromosome 2), associated with transmembrane transport and signal peptide processing; Manes.09G060900 (chromosome 9), associated with protein kinase activity, ATP binding, and protein phosphorylation; and Manes.15G083800 and Manes.15G084000 (chromosome 15), associated with S-adenosylmethionine-dependent methyltransferase activity, membrane-related functions, and protein phosphorylation. These genes participate in biochemical pathways involved in cellular signaling, membrane transport, and metabolic regulation that may contribute to herbicide tolerance. Overall, the results demonstrate that herbicide tolerance in cassava is a quantitative and polygenic trait governed by numerous small-effect loci. The integration of cellular signaling, metabolic regulation, and membrane transport supports the physiological resilience of the species under chemical exposure, providing valuable insights for breeding strategies and marker-assisted selection.

Genome-Wide Association Study

The "genetic test request": A genomic stewardship intervention for inpatient exome and genome orders at a tertiary pediatric hospital.

PURPOSE: Exome sequencing (ES) and genome sequencing (GS) are useful tests to diagnose rare diseases in pediatric patients in critical care settings. Genomic test stewardship can increase the appropriate use of these tests leading to improved diagnostics and cost savings. METHODS: A mandatory review of ES and GS orders for admitted patients was implemented in March 2023. Outcomes of the reviews, cost analysis, and subsequent test results through February 2024 were analyzed with descriptive statistics. RESULTS: There were 444 genetic test request orders placed for 412 unique patients. Of these, 81 (18.2%) were redirected and 57 (12.8%) required modification after approval, leading to an overall cost savings of $345,821.00 or $778.88 per order. The combined diagnostic rate was 28.2% in this patient population. CONCLUSION: Stewardship of ES/GS orders for pediatric inpatients is an effective tool to improve the appropriate usage of these genomic tests. Additional collaboration with stakeholders and expansion of genomic stewardship initiatives may shorten the diagnostic odyssey for critically ill pediatric patients and result in cost savings.

Humans

Robust and accurate Bayesian inference of genome-wide genealogies for hundreds of genomes.

The Ancestral Recombination Graph (ARG), which describes the genealogical history of a sample of genomes, is a vital tool in population genomics and biomedical research. Recent advancements have substantially increased ARG reconstruction scalability, but they rely on approximations that can reduce accuracy, especially under model misspecification. Moreover, they reconstruct only a single ARG topology and cannot quantify the considerable uncertainty associated with ARG inferences. Here, to address these challenges, we introduce SINGER (sampling and inferring of genealogies with recombination), a method that accelerates ARG sampling from the posterior distribution by two orders of magnitude, enabling accurate inference and uncertainty quantification for hundreds of whole-genome sequences. Through extensive simulations, we demonstrate SINGER's enhanced accuracy and robustness to model misspecification compared to existing methods. We demonstrate the utility of SINGER by applying it to individuals of British and African descent within the 1000 Genomes Project, identifying signals of population differentiation, archaic introgression and strong support for ancient polymorphism in the human leukocyte antigen region shared across primates.

Humans

Streamlining large-scale genomic data management: Insights from the UK Biobank whole-genome sequencing data.

Biobank-scale whole-genome sequencing (WGS) studies are increasingly pivotal in unraveling the genetic bases of diverse health outcomes. However, managing and analyzing these datasets' sheer volume and complexity presents significant challenges. We highlight the annotated genomic data structure (aGDS) format, substantially reducing the WGS data file size while enabling seamless integration of genomic and functional information for comprehensive WGS analyses. The aGDS format yielded 23 chromosome-specific files for the UK Biobank 500k WGS dataset, occupying only 1.10 tebibytes of storage. We develop the vcf2agds toolkit that streamlines the conversion of WGS data from VCF to aGDS format. Additionally, the STAARpipeline equipped with the aGDS files enabled scalable, comprehensive, and functionally informed WGS analysis, facilitating the detection of common and rare coding and noncoding phenotype-genotype associations. Overall, the vcf2agds toolkit and STAARpipeline provide a streamlined solution that facilitates efficient data management and analysis of biobank-scale WGS data across hundreds of thousands of samples.

Humans

Delivering effective genome sequencing in pediatric care: From research in the 100,000 Genomes Project to routine clinical practice.

PURPOSE: Genome sequencing (GS) is increasingly used to investigate rare conditions, primarily in children. The 100,000 Genomes Project (100KG) evaluated GS ahead of implementation in the English National Health Service. In 2020, the National Health Service Genomic Medicine Service (GMS) became the first public health care system to offer GS in routine clinical care. We investigate how learning from 100KG informed GMS service delivery. METHODS: We compare GS outcomes in children tested at a large pediatric hospital via GMS (n = 501) and 100KG research (n = 1759). RESULTS: GMS diagnostic yield (29%) was higher than that in 100KG (22%) (P < .0016). Median age at testing was 8 years in 100KG and 6 in the GMS (P < .05). In 100KG, the diagnostic yield was <10% for 15 indications, none of which are included in GMS testing. 100KG data showed little benefit to application of >3 panels. Use of fewer but larger GMS panels resulted in a significantly higher number of genes tested per patient: median 2801 vs 1373 in 100KG (P < .001). In 100KG, diagnostic yield was not significantly increased by testing more than 3 family members (n = 34/142, 24%). CONCLUSION: Learning from 100KG has informed GS clinical service delivery, resulting in higher diagnostic yields and earlier age at testing. Lessons are broadly applicable to all services providing GS, enabling earlier access to tailored management with fewer investigations.

Humans

The Taiwanese hepatitis C virus genome: sequence determination and mapping the 5' termini of viral genomic and antigenomic RNA.

The complete nucleotide sequence of hepatitis C virus (HCV) cloned from the liver tissue of a Taiwanese patient with post-transfusion type C hepatitis was determined. The 5' end of HCV genomic RNA was located 341 nucleotides upstream from the initiation codon for the viral polyprotein open reading frame. The 5' end of the viral antigenomic RNA was shown to have 13 consecutive As. Thus the 3' terminus of the viral genome is a stretch of U which ends about 50 nucleotides downstream from the stop codon of the large open reading frame. The nucleotide sequence homology between this HCV strain and two Japanese isolates was 90.5 and 90.7%, respectively. Homology with the United States strain, however, was only 77.8%. Accordingly, the indigenous Taiwanese HCV strain is of the same subtype as the Japanese isolates. Novel features of the viral genome termini are possibly relevant to HCV genome replication.

Amino Acid Sequence

Folding a broken genome: the versatile roles of cohesin in genome maintenance.

Cohesin is a protein complex that shapes 3D genome organization through two distinct mechanisms. First, cohesin tethers replicated chromatids from DNA replication until mitosis. This process, known as sister chromatid cohesion, ensures accurate chromosome segregation and enables high-fidelity DNA repair through homologous recombination between the sister chromatids. Second, cohesin organizes the genome during interphase by dynamically extruding chromatin loops, structures that have key roles in gene regulation. Recent work has shown that, in addition to the well-established repair functions of sister chromatid cohesion, cohesin-mediated chromatin looping is closely linked to the repair of DNA double-strand breaks - one of the most toxic DNA lesions. In this Review, we discuss the central roles of cohesin in maintaining genome stability, with emphasis on the cellular response to DNA double-strand breaks. We review how dynamic loop structures facilitate signalling of repair events and promote long-range chromatin motions that underpin the repair process. Overall, its dual mode of action - cohesion and loop extrusion - positions cohesin as a central regulator of chromatin architecture and genome maintenance.

Cohesins

Hepatitis C virus (HCV) circulates as a population of different but closely related genomes: quasispecies nature of HCV genome distribution.

Sequencing of multiple recombinant clones generated from polymerase chain reaction-amplified products demonstrated that the degree of heterogeneity of two well-conserved regions of the hepatitis C virus (HCV) genome within individual plasma samples from a single patient was consistent with a quasispecies structure of HCV genomic RNA. About half of circulating RNA molecules were identical, while the remaining consisted of a spectrum of mutants differing from each other in one to four nucleotides. Mutant sequence diversity ranged from silent mutations to appearance of in-frame stop codons and included both conservative and nonconservative amino acid substitutions. From the relative proportion of essentially defective sequences, we estimated that most circulating particles should contain defective genomes. These observations might have important implications in the physiopathology of HCV infection and underline the need for a population-based approach when one is analyzing HCV genomes.

Amino Acid Sequence