PubMed HealthSearch

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 289 records · Page 16Linked to original sources

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning

The Genome Sequence DataBase: towards an integrated functional genomics resource.

During 1998 the primary focus of the Genome Sequence DataBase (GSDB; http://www.ncgr.org/gsdb ) located at the National Center for Genome Resources (NCGR) has been to improve data quality, improve data collections, and provide new methods and tools to access and analyze data. Data quality has been improved by extensive curation of certain data fields necessary for maintaining data collections and for using certain tools. Data quality has also been increased by improvements to the suite of programs that import data from the International Nucleotide Sequence Database Collaboration (IC). The Sequence Tag Alignment and Consensus Knowledgebase (STACK), a database of human expressed gene sequences developed by the South African National Bioinformatics Institute (SANBI), became available within the last year, allowing public access to this valuable resource of expressed sequences. Data access was improved by the addition of the Sequence Viewer, a platform-independent graphical viewer for GSDB sequence data. This tool has also been integrated with other searching and data retrieval tools. A BLAST homology search service was also made available, allowing researchers to search all of the data, including the unique data, that are available from GSDB. These improvements are designed to make GSDB more accessible to users, extend the rich searching capability already present in GSDB, and to facilitate the transition to an integrated system containing many different types of biological data.

Animals

Characterization of the genome of molluscum contagiosum virus type 1 between the genome coordinates 0.045 and 0.075 by DNA nucleotide sequence analysis of a 5.6-kb HindIII/MluI DNA fragment.

The complete DNA nucleotide sequence of a HindIII/MluI genomic DNA fragment (0.045-0.075 viral map units) from molluscum contagiosum virus type 1 (MCV-1) was determined. The HindIII/MluI DNA fragment comprises 5,646 bp with a base composition of 64.4% G + C and 35.6% A + T. The DNA sequence contains many perfect direct repeats. A cluster of three repetitive DNA elements R1, R2 and R3, with a complex structural arrangement was detected between nucleotide positions 1802 and 2107. The unit length (box) of the repetitive DNA sequences was found to be 6 bp (15 boxes) and 9 bp (24 boxes) for R1 and R2, respectively. The repetitive DNA element R3 is organized in fifteen boxes (15 bp) in which a unit length of R1 is combined with a unit length of R2. The arrangement of the repetition R3 within the DNA sequences of this particular region of the MCV-1 genome was found to be (5 x R3) + (2 x R2) + (1 x R3) + (6 x R2) + (1 x R3) + (1 x R2) + (8 x R3). Twenty-three open reading frames (ORFs) of 60-1,175 amino acid (AA) residues were detected. The largest ORF (number 17) comprises 1,175 AA with a predicted molecular weight of 126 kD. This ORF harbors a promoter signal which is located 21 nucleotides upstream from the start codon and is very similar to the early promoter signals known for vaccinia virus. This putative protein contains glutamine-enriched regions between AA residues 427 and 682 which show homologies to the corresponding glutamine-enriched regions of a variety of cellular genes like human transcriptional initiation factor (TFIID: TATA box factor).

Amino Acid Sequence

Global genomic and antimicrobial resistance profiling of Neisseria gonorrhoeae: Insights from whole genome sequencing and minimum inhibitory concentration analysis.

BACKGROUND: The rising antimicrobial resistance (AMR) of Neisseria gonorrhoeae is a major global health concern that limits treatment options and complicates disease management. Efflux pump systems and resistance genes are key to bacteria's ability to evade antibiotics. This study examined the genetic and phenotypic resistance landscape using a large dataset of whole-genome sequences to identify key resistance mechanisms, assess efflux pump gene prevalence, and analyze regional variations in Minimum Inhibitory Concentration (MIC) values to inform treatment strategies and public health interventions. METHODS: A total of 38,585 whole-genome sequences of N. gonorrhoeae were analyzed to identify AMR determinants. This study focused on the presence and distribution of efflux pump genes (mtrC, farB, norM, and mtrA) and specific resistance genes, including tet(C) (tetracycline resistance) and aph(3')-Ia (aminoglycoside resistance). The MIC values were assessed for multiple antibiotics to evaluate resistance trends and regional variations, including penicillin, spectinomycin, zoliflodacin, gentamicin, and fluoroquinolones. RESULTS: This analysis revealed widespread resistance to multiple antibiotics. Efflux pump genes (mtrC, farB, norM, and mtrA) were found in nearly all isolates, highlighting their essential roles in resistance and adaptation. The presence of tet(C) and aph (3')-Ia varied across different Gene Presence Patterns, suggesting that regional or therapeutic factors may influence tetracycline and aminoglycoside resistance. High MIC values for penicillin were observed, likely because of blaTEM, a beta-lactamase gene responsible for beta-lactam resistance. Resistance to spectinomycin is also widespread, raising concerns about the diminishing efficacy of this antibiotic. In contrast, zoliflodacin, gentamicin, and fluoroquinolones exhibited relatively low MIC values, indicating their sustained effectiveness against N. gonorrhoeae. DISCUSSION: Efflux pump systems are key to N. gonorrhoeae resistance and adaptability. Regional MIC variations indicate that local antibiotic use shapes resistance patterns. The high resistance to penicillin and spectinomycin highlights the need for alternative treatments, whereas zoliflodacin and fluoroquinolones remain effective but require monitoring. This study emphasizes global AMR surveillance, novel therapies, and targeted antimicrobial stewardship to address multidrug-resistant infections.

Neisseria gonorrhoeae

De Novo Whole Genome Assemblies of Unusual Case-Making Caddisflies (Trichoptera) Highlight Genomic Convergence in the Composition of the Major Silk Gene (h-fibroin).

Trichoptera (caddisflies) is one of the most species-rich orders of aquatic insects. Species of caddisflies cover a broad ecological diversity as exemplified by various uses of underwater silk secretions. Diversity of silk use generally aligns with the evolution of major caddisfly lineages, specifically at the subordinal level: Annulipalpia (retreat makers) and Integripalpia (cocoon and tube-case makers). However, silk use within suborders differs for a few exceptional species in these clades. In this study, we provide the first whole genome assemblies and annotations for two unusual Integripalpia species: Limnocentropus insolitus, whose hard tube-case is anchored to boulders by a rigid, elongated silken stalk, and Phryganopsyche brunnea which builds a "floppy" cylindrical case that lacks the typical robustness of tube-cases. Its texture rather resembles that of the flexible retreats built by Annulipalpia. Using the two high-quality genome assemblies, we identified and annotated the major silk gene, h-fibroin, and compared its amino acid composition across various groups, including retreat, cocoon, and tube-case makers. Our phylogenetic analysis confirmed the phylogenetic position of the two species in the tube-case-making clade. The major silk gene of L. insolitus shows a similar amino acid composition to other tube-case-making species. In contrast, the amino acid composition of P. brunnea resembles that of retreat-making species, in particular with regard to the high content of proline. This is consistent with the hypothesis that proline could be linked to enhanced extensibility of silk fibers. Taken together, our results underscore the role of silk genes in shaping the evolutionary ecology of retreat- and tube-case-making in caddisflies.

Animals

Genomic insights into end-use grain quality and nutritional traits of an ancient Indian dwarf wheat ( Triticum sphaerococcum Percival) population using a multi-locus genome-wide association study.

BACKGROUND: Triticum sphaerococcum, an ancient hexaploid wheat species, is renowned for its stress resilience and superior nutritional quality. A panel of 116 T. sphaerococcum accessions (the largest known collection at a single site globally), with six bread wheat released varieties, was evaluated for its potential for genetic quality improvement. Field experiments were conducted under standard, heat and moisture-deficit conditions across two cropping seasons for ten grain end-use quality and nutritional traits. RESULTS: Genotypes showed highly significant differences (P ≤ 0.001) for measured traits, with high broad-sense heritability resulting from substantial genotypic variance contributions. Triticum sphaerococcum consistently outperformed T. aestivum across environments, with moisture-deficit stress proving more detrimental to quality parameters than heat stress, while micronutrient content increased under stressed conditions. Trait correlations revealed that the gluten index (GI) correlated negatively with the grain hardness index (GHI), wet gluten (WG), and water-binding capacity (WB), while positively correlating with dry gluten (DG) and protein content (PRO), whereas grain iron (GFE), zinc (GZN), and protein showed consistent positive interrelationships. Two superior accessions, PAUTS10 (WG 35.13%, DG 13.71%, PRO 16.42%, GZN 50.89 ppm) and Sonamoti (WG 33.33%, DG 12.92%, PRO 16.27%, GZN 56.03 ppm), were identified, surpassing the best check variety HD3226 for quality and nutritional parameters. Multi-locus genome-wide association studies identified 30 stable quantitative trait nucleotides across environments, with candidate gene analysis revealing genes involved in transcription regulation, biosynthetic processes, metal ion homeostasis, and transport. CONCLUSIONS: Triticum sphaerococcum demonstrated superior grain quality and micronutrient potential compared with modern wheat, highlighting its value as a genetic resource for biofortification. The identification of elite accessions and stable quantitative trait nucleotides (QTNs) provides useful targets for breeding programs aimed at improving protein and micronutrient content. Integrating ancient germplasm with modern genomic tools can accelerate the development of nutritionally enhanced wheat varieties. © 2026 Society of Chemical Industry.

Triticum

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence

The detection of nucleotide sequences with strong similarity to hormone responsive elements in the genome of eubacteria and archaebacteria and their possible relation to similar sequences present in the mitochondrial genome.

To account for the presence of nucleotide sequences in mitochondria with similarity to the Hormone Response Elements (HREs) of the nuclear genomes of man, rat and mouse, the genomes of several procaryotes have been screened for the presence of the sequences AGAACA NNN TGTTCT and GGTACA NNN TGTTCT, which represent perfect palindromic and consensus class I HREs, respectively, and for the sequence AGGTCA NNN TGACCT, which represents class II HRE. In many of the examined procaryotes, eubacteria and archaebacteria, almost perfect palindromic class I HREs and perfect or almost perfect class II half palindromic HREs have been detected in various genes, some of which encode proteins involved in energy metabolism, in replication and in transcription control. These findings support the hypothesis that the similar sequences found in mitochondria, potentially involved in hormonal regulation of respiratory enzyme biosynthesis, were introduced into eucaryotic cell by the procaryotic endosymbionts.

Animals

Analysis of six DNA components of the faba bean necrotic yellows virus genome and their structural affinity to related plant virus genomes.

Faba bean necrotic yellows virus (FBNYV) has a multicomponent circular ssDNA genome. In addition to a previously described genome component (C1) coding for a replicase-associated protein (Rep), five further components (C2 to C6) have now been identified. Each of the six components is about 1 kb in size, contains one major open reading frame (ORF) in the virion sense with a TATA box and polyadenylation signal, and has a noncoding region containing a highly conserved sequence possibly forming a stem-loop structure. Similar to C1, C2 encodes another putative Rep of 33.1 kDa, which is closely related to the Rep of banana bunchy top virus (BBTV). Based on bacterial expression and immunoblot analysis, the ORF of C5 encodes the capsid protein (CP) with a deduced molecular mass of 19 kDa. The FBNYV CP shares the highest amino acid (aa) identity (56.2%) with that of subterranean clover stunt virus (SCSV). The ORF of C4 potentially codes for a hydrophobic protein which appears to be structurally and functionally similar to the BBTV-C4 and SCSV-C1 proteins. No protein sequence similarities were found in databases for the C3 and C6 ORFs of FBNYV. FBNYV is clearly distinct from any known virus but is taxonomically related to BBTV and SCSV.

Amino Acid Sequence

Phylogenetic relationship of the complete Rauscher murine leukemia virus genome with other murine leukemia virus genomes.

We report the complete nucleotide sequence of the genome of Rauscher murine leukemia virus (R-MuLV), the replication-competent helper virus present in the Rauscher virus complex, and its phylogenetic relationship with other murine leukemia virus genomes. An overall sequence identity of 97.6% was found between R-MuLV and the Friend helper virus (F-MuLV), and the two viruses were closely related on the phylogenetic trees constructed from either gag, pol, or env sequences. Moloney murine leukemia virus (Mo-MuLV) was the next closest relative to R-MuLV and F-MuLV on all trees, followed by Akv and radiation leukemia virus (RadLV). The most distantly related helper virus was Hortulanus murine leukemia virus (Ho-MuLV). Interestingly, Cas-Br-E branched with Mo-MuLV on the gag and pol trees, whereas on the env tree, it revealed the highest degree of relatedness to Ho-MuLV, possibly due to an ancient recombination with an Ho-MuLV ancestor. In summary, a phylogenetic analysis involving various MuLVs has been performed, in which the postulated close relationship between R-MuLV and F-MuLV has been confirmed, consistent with the pathobiology of the two viruses.

Algorithms

The genome nucleotide sequence of a contemporary wild strain of measles virus and its comparison with the classical Edmonston strain genome.

The only complete genome nucleotide sequences of measles virus (MeV) reported to date have been for the Edmonston (Ed) strain and derivatives, which were isolated decades ago, passaged extensively under laboratory conditions, and appeared to be nonpathogenic. Partial sequencing of many other strains has identified >/=15 genotypes. Most recent isolates, including those typically pathogenic, belong to genotypes distinct from the Edmonston type. Therefore, the sequence of Ed and related strains may not be representative of those of pathological measles circulating at that or any time in human populations. Taking into account these issues as well as the fact that so many studies have been based upon Ed-related strains, we have sequenced the entire genome of a recently isolated pathogenic strain, 9301B. Between this recent isolate and the classical Ed strain, there were 465 nucleotide differences (2.93%) and 114 amino acid differences (2.19%). Computation of nonsynonymous and synonymous substitutions in open reading frames as well as direct comparisons of noncoding regions of each gene and extracistronic regulatory regions clearly revealed the regions where changes have been permissible and nonpermissible. Notably, considerable nonsynonymous substitutions appeared to be permissible for the P frame to maintain a high degree of sequence conservation for the overlapping C frame. However, the cause and the effect were largely unclear for any substitution, indicating that there is a considerable gap between the two strains that cannot be filled. The sequence reported here would be useful as a reference of contemporary wild-type MeV.

3' Untranslated Regions

Characteristics of nucleotide substitution in the hepatitis C virus genome: constraints on sequence change in coding regions at both ends of the genome.

Comparison of complete genome sequences for different variants of hepatitis C virus (HCV) reveals several different constraints on sequence change. Synonymous changes are suppressed in coding regions at both 5' and 3' ends of the genome. No evidence was found for the existence of alternative reading frames or for a lower mutation frequency in these regions. Instead, suppression may be due to constraints imposed by RNA secondary structures identified within the core and NS5b genes. Nonsynonymous substitutions are less frequent than synonymous ones except in the hypervariable region of E2 and, to a lesser extent, in E1, NS2, and NS5b. Transitions are more frequent than transversions, particularly at the third position of codons where the bias is 16:1. In addition, nucleotide substitutions may not occur symmetrically since there is a bias toward G or C at the third position of codons, while T left and right arrow C transitions were twice as frequent as A left and right arrow G transitions. These different biases do not affect the phylogenetic analysis of HCV variants but need to be taken into account in interpreting sequence change in longitudinal studies.

Base Sequence

Complete nucleotide sequence and genome organization of sweet potato feathery mottle virus (S strain) genomic RNA: the large coding region of the P1 gene.

The complete nucleotide sequence of a sweet potato feathery mottle virus severe strain (SPFMV-S) genomic RNA was determined from overlapping cDNA clones and by directly sequencing viral RNA. The viral RNA genome is 10,820 nucleotides long, excluding the poly(A) tail and contains one open reading frame (ORF) starting at nucleotide 118 and ending at 10,599, potentially encoding a polyprotein of 3,493 amino acids (Mr 393,800). The ORF was followed by a 3' untranslated region of 221 nucleotides. The deduced polyprotein includes P1 (74K), HC-Pro (52K), P3 (46K), 6K1, CI (72K), 6K2, NIa-VPg (22K), NIa-Pro (28K), NIb (60K) and coat (35K) proteins, after an analysis of protein cleavage sites analogous to other potyvirus polyproteins. The polyprotein had a high level of amino acid identity with those of other potyviruses, except in the regions of P1 and P3. The P1 of SPFMV-S RNA has 664 amino acid residues, and is the largest and least similar to those of other potyviruses. HC-Pro and CI show high identity with those of other potyviruses. P3 has relatively low identity, however, the length of P3 was within the range of variability among other potyviruses. The 6K1 protein between P3 and C1 is also highly similar to those of other potyviruses. This is the first report on the complete nucleotide sequence of the sweet potato-infecting virus.

Genome, Viral

Identification of four genomic loci highly related to casein-kinase-2-alpha cDNA and characterization of a casein kinase-2-alpha pseudogene within the mouse genome.

Using the coding region of the human CK-2 alpha cDNA as a probe for screening a genomic mouse library, positive clones representing four different genomic loci were isolated. Partial DNA sequences of these loci encompassing the first 120 nucleotides of the putative coding region are reported. One positive clone was further analyzed by sequencing a 3.1 kb XbaI fragment. This clone displays the characteristics of a pseudogene, i.e. lack of introns and several nucleotide insertions and deletions. In its 3' region it contains a 91 bp large CT-rich stretch which consists of (CCTT) and (CT) repeats; in the 5' region three (CCCCCT) repeats.

Animals

Cereal genome evolution: pastoral pursuits with 'Lego' genomes.

The rapid progress in comparative analysis of cereal genomes reveals that they are composed of similar genomic building blocks. It seems that by simply rearranging these blocks and amplifying some of the repetitive sequences contained within them, it is possible to reconstitute the 56 different chromosomes found in wheat, rice, maize, sorghum, millet and sugarcane. Comparison of the orders of blocks in these reconstituted chromosomes reveals that the cleavage of a single chromosome formed from the blocks could give rise to all the combinations found in the chromosomes of the above species. A framework is now in place for collating all the information which has been generated from studying the individual cereals.

Biological Evolution

Genome sequences: genome sequence of a model prokaryote.

The complete Escherichia coli genome sequence is now known; it should greatly facilitate the analysis of other genomes, but a lot remains to be learnt about E. coli itself. About half the genes were previously uncharacterized, but expanding databases and improving analysis methods will help predict their functions.

Bacterial Proteins

Whole genome analysis: experimental access to all genome sequenced segments through larger-scale efficient oligonucleotide synthesis and PCR.

The recent ability to sequence whole genomes allows ready access to all genetic material. The approaches outlined here allow automated analysis of sequence for the synthesis of optimal primers in an automated multiplex oligonucleotide synthesizer (AMOS). The efficiency is such that all ORFs for an organism can be amplified by PCR. The resulting amplicons can be used directly in the construction of DNA arrays or can be cloned for a large variety of functional analyses. These tools allow a replacement of single-gene analysis with a highly efficient whole-genome analysis.

Animals

Chromosomal losses and gains in meningiomas: comparative genomic hybridization (CGH) study of the whole genome.

We investigated chromosomal aberrations in meningiomas using newly developed comparative genomic hybridization (CGH) technique and compared the results with the proliferating potential of the tumors. This technique permits the entire genome to be surveyed in one session of experiments. Our results revealed chromosomal aberrations in 5 out of 10 (50%) of the tumor samples studied. Losses of the distal parts of chromosome 1p (5 out of 10) and 22q (3 out of 10) were the two most frequent chromosomal aberrations. Losses and/or gains in other regions were only sporadic. The MIB-1 staining indices (MIB-SI, %) were 1.9 +/- 0.9% (mean +/- SD) in benign (n = 8), 4.5% in atypical (n = 1), and 11.7% in anaplastic (n = 1) meningiomas. The comparison of MIB-SI between the tumors with (2.3 +/- 0.6%) and without (1.6 +/- 0.3%) chromosomal aberrations demonstrated a trend towards an increased MIB-SI in meningiomas with chromosomal aberrations (p < 0.07) by unpaired Student's t-test. This study suggests that alterations in chromosomes 1p and 22q could be a primary focus of further detailed assessment of tumorigenesis and in understanding the biological behavior of meningiomas.

Adolescent