PubMed HealthSearch

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Whole blood transcriptome profile identifies motor neurone disease RNA biomarker signatures.

Blood-based biomarkers for motor neuron disease are needed for better diagnosis, progression prediction, and clinical trial monitoring. We used whole blood-derived total RNA and performed whole transcriptome analysis to compare the gene expression profiles in (motor neurone disease) MND patients to the control subjects. We compared 42 MND patients to 42 aged and sex-matched healthy controls and described the whole transcriptome profile characteristic for MND. In addition to the formal differential analysis, we performed functional annotation of the genomics data and identified the molecular pathways that are differentially regulated in MND patients. We identified 12,972 genes differentially expressed in the blood of MND patients compared to age and sex-matched controls. Functional genomic annotation identified activation of the pathways related to neurodegeneration, RNA transcription, RNA splicing and extracellular matrix reorganisation. Blood-based whole transcriptomic analysis can reliably differentiate MND patients from controls and can provide useful information for the clinical management of the disease and clinical trials.

Humans

Optimization of protoplast based DNA isolation and genome analysis in a gamma-irradiated Aspergillus niger mutant strain.

Aspergillus niger is an important industrial fungus widely used for citric acid production and a range of biotechnological applications. In this study, a protoplast-based DNA isolation protocol was optimized for a gamma-irradiated A. niger AN-L103_M1 mutant strain, followed by whole-genome sequencing and functional genome analysis. Protoplast yield was strongly influenced by enzyme concentration and the molarity of the osmotic stabilizer. The highest yield was achieved at an enzyme concentration of 50&#xa0;mg/mL (2.487&#x2009;&#xb1;&#x2009;0.04&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL) and 0.8&#xa0;M KCl (2.550&#x2009;&#xb1;&#x2009;0.06&#x2009;&#xd7;&#x2009;10&#x2078; cells/mL), with both factors showing significant effects (p&#x2009;<&#x2009;0.0001) in GraphPad Prism 11.0.0. Whole-genome sequencing performed using an Illumina NovaSeq 6000 platform yielded a 37.06&#xa0;Mb draft genome assembled into 537 contigs, with an N50 of 363,084&#xa0;bp and a GC content of 48.2%. BUSCO 14 analysis showed high completeness (97.95% complete BUSCOs). Functional annotation and KEGG pathway mapping identified genes involved in glycolysis, the tricarboxylic acid cycle, and citrate biosynthesis, while biosynthetic gene cluster analysis revealed diverse potential for secondary metabolite production. These findings provide an optimized workflow for protoplast-based DNA isolation and genome-scale functional analysis in A. niger, proposing a basis for future comparative genomics, transformation studies, and experimentally validated metabolic engineering.

Aspergillus niger

Pinpointing genomic regions conferring herbicide tolerance in cassava via genome-wide association mapping.

Cassava (Manihot esculenta Crantz) is a tropical crop of major socioeconomic importance, whose productivity can be limited by sensitivity to herbicides used for weed management. This study aimed to perform a genome-wide association study (GWAS) in 194 cassava genotypes to identify genomic regions associated with tolerance to the herbicides mesotrione, S-metolachlor, and chloransulam-methyl. The evaluations performed at 3, 6, 9, 15, and 30 days after application (DAA) were used to characterize the temporal progression of phytotoxicity. Based on this analysis, the phenotype obtained at 9 days after application (PhytoX9DAA) was selected for genome-wide association analyses because it represented the period of greatest symptom expression and the highest discrimination among genotypes. GWAS analyses were performed using de-regressed BLUPs and the MLM, MLMM, and BLINK models, incorporating kinship (K) and population structure (Q) matrices. Significant markers were detected across multiple chromosomes, and the corresponding genomic windows contained candidate genes with functional annotations related to herbicide response. The predominant functional categories included membrane transport, channel activity, signal peptide processing, protein phosphorylation, cellular signaling, and metabolic regulation. Key candidate genes included Manes.02G151900 and Manes.02G152700 (chromosome 2), associated with transmembrane transport and signal peptide processing; Manes.09G060900 (chromosome 9), associated with protein kinase activity, ATP binding, and protein phosphorylation; and Manes.15G083800 and Manes.15G084000 (chromosome 15), associated with S-adenosylmethionine-dependent methyltransferase activity, membrane-related functions, and protein phosphorylation. These genes participate in biochemical pathways involved in cellular signaling, membrane transport, and metabolic regulation that may contribute to herbicide tolerance. Overall, the results demonstrate that herbicide tolerance in cassava is a quantitative and polygenic trait governed by numerous small-effect loci. The integration of cellular signaling, metabolic regulation, and membrane transport supports the physiological resilience of the species under chemical exposure, providing valuable insights for breeding strategies and marker-assisted selection.

Genome-Wide Association Study

Exploring the ecological drivers of bacteriophage diversity and functional viral potential in the skin of the axolotl Ambystoma altamirani.

Bacteriophages play important roles in shaping microbial community dynamics across diverse environments. In the amphibian skin, most microbiome studies have focused on bacteria and their interactions with the fungus Batrachochytrium dendrobatidis (Bd), leaving other microbial components, including viruses, largely unexplored. Here, we present the first characterization of the viral community in the amphibian skin microbiome, focusing on ecological drivers of bacteriophage diversity and functional potential in the axolotl Ambystoma altamirani. Using public shotgun metagenomes, we found that the viral fraction was dominated by bacteriophages of the class Caudoviricetes. Bacteriophage diversity was significantly associated with local physicochemical parameters at the time of sampling, and showed a strong positive correlation with bacterial diversity, whereas no significant associations were detected with the presence of Bd. In addition, seasonality influenced the composition and properties of bacteria-bacteriophage co-abundance networks. Functional annotation of assembled bacteriophage sequences revealed a diverse functional potential, including putative auxiliary metabolic genes, superinfection exclusion, toxin-antitoxin, and virulence factors. Overall, these findings highlight the ecological relevance of bacteriophages in amphibian skin microbiomes and underscore the need for further studies on their role in the amphibian host's health.

Animals

Functional mapping and annotation of genetic associations with FUMA.

A main challenge in genome-wide association studies (GWAS) is to pinpoint possible causal variants. Results from GWAS typically do not directly translate into causal variants because the majority of hits are in non-coding or intergenic regions, and the presence of linkage disequilibrium leads to effects being statistically spread out across multiple variants. Post-GWAS annotation facilitates the selection of most likely causal variant(s). Multiple resources are available for post-GWAS annotation, yet these can be time consuming and do not provide integrated visual aids for data interpretation. We, therefore, develop FUMA: an integrative web-based platform using information from multiple biological resources to facilitate functional annotation of GWAS results, gene prioritization and interactive visualization. FUMA accommodates positional, expression quantitative trait loci (eQTL) and chromatin interaction mappings, and provides gene-based, pathway and tissue enrichment results. FUMA results directly aid in generating hypotheses that are testable in functional experiments aimed at proving causal relations.

Chromatin

Systematic discovery of pathogen effector functions across human pathogens and pathways.

Pathogens deploy effector proteins to exploit host cell biology, and most effector open reading frames (ORFs) are rapidly evolving and lack functional annotation. We developed the effector ORFeome (eORFeome), a scalable functional genomics platform encompassing 3,835 effector ORFs from diverse viruses, bacteria, and parasites. High-throughput barcoded screens across nuclear factor &#x3ba;B (NF-&#x3ba;B), apoptosis, p53, cGAS-STING, and major histocompatibility complex class I (MHC class I) pathways revealed novel pathway-modulating functions for hundreds of uncharacterized eORFs, unexpected activities of known effectors, and distinct pathway-specific functions encoded by single ORFs. Illustrating the power of this approach, we identified HHV6A U14 as a p53 antagonist, HHV7 U21 as a dual-function STING antagonist and MHC-I antigen display inhibitor, and adenoviral 13.6K/i-leader protein as a de novo-evolved TAP inhibitor that suppresses MHC-I display. These results establish a general framework for systematic effector annotation, uncover new mechanisms of host-pathogen interaction across kingdoms, and highlight pathogen effectors as a versatile toolkit for rewiring and probing human cellular pathways.

Humans

Improving the Annotations of JCVI-Syn3a Proteins.

The JCVI-Syn3 organism is a minimal organism derived from Mycoplasma mycoides capri, which is capable of self-replication. While the ancestor has 863 genes, the synthetic progeny has only 473, with 434 of these coding for proteins. Despite initial efforts to understand all functions of the organism, a significant number of these protein-coding genes still have unknown functions, and subsequent studies have been only partially successful in elucidating their roles. In this study, we employ our innovative method PROST to identify homologs and better understand these previously unidentified genes. PROST employs protein language embeddings and enables the identification of remote homologs with as low as 16% sequence identity. PROST successfully finds functionally annotated homologs for 93% of the minimal genome with a high level of accuracy, both confirming previously identified functions, as well as proposing new functions for others. The results of our study can be accessed at https://bit.ly/prost-syn3a .

Molecular Sequence Annotation

Whole-genome sequencing and analysis of the endophytic fungus Alternaria alternata Y-2 from Leymus chinensis.

To explore the genetic basis and functional potential of beneficial symbiosis between the endophytic fungus Alternaria alternata Y-2 and its host Leymus chinensis, we performed Illumina-based draft whole-genome sequencing and systematic bioinformatic analysis. Although this assembly does not reach telomere-to-telomere completeness, it provides high-quality gene-level information for gene prediction, functional annotation, carbohydrate-active enzyme (CAZyme) identification, and secondary metabolite biosynthetic gene cluster analysis. The final genome size of A. alternata Y-2 was 34,383,676&#xa0;bp with a GC content of 51.0%, containing 12,724 predicted protein-coding genes, 90 tRNAs, and 12 rRNAs. BUSCO assessment showed 98.9% completeness, supporting the high quality of this draft genome. A total of 12,627 genes were successfully annotated in the NCBI NR database, and 17,183 genes were functionally categorized using GO terms. In total, 448 CAZyme genes and 21 secondary metabolite biosynthetic gene clusters were identified, which are potentially involved in lignocellulose degradation, cellular redox homeostasis and biosynthesis of bioactive metabolites. Based on ITS sequence alignment, NR annotation, and phylogenetic analysis of single-copy orthologous genes, the strain was confidently identified as A. alternata. This study firstly reports the draft genome of an endophytic A. alternata strain derived from L. chinensis and provides valuable genetic resources for exploring the endophytic lifestyle, stress tolerance, and bioactive metabolite potential of this fungus.

Alternaria

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

Chromosome-level assembly and annotation of the yellow-shelled fish (Barbodes Wynaadensis).

Barbodes wynaadensis, a unique cyprinid species native to Yunnan Province in China, stands out as an allotetraploid (AABB) fish with a complex evolutionary history. Leveraging a multi-platform sequencing strategy combining MGI short-read, PacBio long-read, and Hi-C scaffolding technologies, we assembled the first chromosome-level genome for B. wynaadensis. The final assembled genome spans 1.76&#x2009;Gb in length with a contig N50 of 33.53&#x2009;Mb, demonstrating high assembly continuity. Hi-C scaffolding enabled the reconstruction of 50 pseudochromosomes, representing 99.94% of the total genome assembly. Genome annotation identified 46,121 protein-coding genes, with a functional annotation rate of 99.76%. Repetitive elements constituted 48.26% of the genomic sequences, including lineage-specific expansions of DNA transposons (29.26%) and LTRs (6.36%). This high-quality assembly resolves challenges in polyploid genome reconstruction and provides a critical resource for investigating Cyprinidae evolution, particularly subgenome divergence and adaptation. The dataset also enables practical applications, such as molecular marker development for population monitoring, supporting conservation efforts for this threatened endemic species amid habitat degradation in the Nujiang River basin.

Animals

Chromosome-level genome assembly of Cheilinus chlorourus (Bloch, 1791) (Perciformes: Labridae).

In the classification of marine fish, the Labridae family ranks second in terms of species diversity and plays a vital role in coral reef ecosystems, comprising over 600 species across 82 genera. Despite its significance for ecological and evolutionary studies, genomic research on this group has lagged, resulting in a shortage of data, particularly regarding high-quality chromosome-level genome assemblies. To address this gap, this study focused on Cheilinus chlorourus from the Labridae family and successfully achieved a chromosome-level genome assembly. By integrating Illumina, PacBio, and Hi-C sequencing data, we assembled a genome measuring 940.36&#x2009;Mb, with 926.86&#x2009;Mb (98.56%) of the gene assembly organized into 21 chromosomes. A total of 29,213 protein-coding genes (PCGs) were identified, and 79.93% of these genes were functionally annotated. With this high-quality genome assembly, future investigations into the functional genomics and ecology of C. chlorourus will have a solid scientific foundation.

Animals

Dynamic Protein Structure Paradox: An Integrative Framework for Endpoint-Conditioned Evidentiary Sufficiency in Structure-to-Function Claims.

Accurate coordinates for a represented protein state do not, by themselves, establish activity or any other condition-specific function. This article defines the Dynamic Protein Structure Paradox (DPSP) as the apparent conflict between structural accuracy and functional underdetermination and develops it as an integrative evidentiary assessment framework rather than a new theory or paradigm. The underlying problem has been longstanding, since structural genomics, function annotation, allostery, and disorder research each established that fold does not determine function and that function does not determine fold. DPSP consolidates those results into one endpoint-conditioned rule. Once a measurable endpoint is defined, it assesses four coupled dimensions: relevant-state completeness, context completeness, ensemble or kinetic dependence, and chemical dependence. A rubric rates each dimension as adequate, uncertain, or missing, and a materiality test determines which gaps influence the stated decision. The outcome is one of three mutually exclusive modes of utilization: geometry-led, conditional, or function-measured. The deliverable is a concise evidence statement delineating what the structure supports, which decisive variable remains unmeasured, and what corroboration is necessary. DPSP complements, rather than replaces, existing structural, ensemble, and computational approaches. The framework remains unvalidated, its thresholds are provisional, and the studies necessary to confirm or refute it are specified.

Proteins

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15&#x2009;Gb (contig N50&#x2009;=&#x2009;43.57&#x2009;Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant

Comparative genomics approaches to identify genomic regions associated with the antimicrobial activity of Pseudomonas protegens PBL3.

The environmental bacterium Pseudomonas protegens PBL3 has antagonistic activity against the plant pathogenic bacterium Burkholderia glumae, an important pathogen in rice. The antimicrobial activity of P. protegens PBL3 was found in the bacteria-free secreted fraction (secretome), but the specific molecules, as well as the genetic basis of that activity, have not been identified. In this study, we integrated genomic information with antimicrobial assays on P. protegens PBL3 and additional six Pseudomonas spp. strains, to identify putative genomic regions in P. protegens PBL3 associated with antimicrobial activity. We hypothesized that Pseudomonas spp. strains with antimicrobial activity against B. glumae have conserved genes with P. protegens PBL3 that are absent in strains lacking activity. Comparative genomics analyses with anvi'o and progressiveMauve, and using P. protegens PBL3 as the reference genome, revealed 188 genes uniquely present in antimicrobial-producing strains. Seven of those genes were annotated as biosynthetic gene clusters predicted to encode secondary metabolites; additional genes were grouped into 25 contiguous clusters with functions annotated as secretion, signal transduction, regulation, transport/efflux, carbohydrate metabolism and one with an additional uncharacterized function. Altogether, this study uncovered a complex and multi-functional network of candidate genes, suggesting that the antimicrobial activity in P. protegens PBL3 is not limited to biosynthetic pathways but also involves additional regulatory, metabolic and export modules to synthesize and deploy antimicrobials.

Pseudomonas

Deciphering the genetic background of an industrial 2-ketogluconic acid-producing strain Pseudomonas plecoglossicida JUIM01 using whole-genome sequencing.

2-Ketogluconic acid (2KGA) is an important precursor for the food antioxidant erythorbic acid, currently produced via microbial fermentation using Pseudomonas species. To facilitate the genetic improvement of production strains, the complete genome of an industrial 2KGA producer P. plecoglossicida JUIM01 was sequenced and analyzed. The genome consists of a 5.13-Mb circular chromosome with a GC content of 63.58%, encoding 4,517 predicted proteins. Comprehensive functional annotation identified a putative global regulatory network comprising 75 core regulators, which were classified into six functionally cooperative modules, potentially governing the strain's metabolism and environmental adaptability.&#xa0;We further delineated the genetic determinants hypothetically linked to efficient 2KGA synthesis, including glucose metabolism, fatty acid metabolism, and the oxidative phosphorylation system.&#xa0;These outputs could provide the genomic resource for elucidating high productivity and robustness, and rationally engineering the high-performance chassis cells toward robust 2KGA production.

P. plecoglossicida

Chromosome-Level Genome Assembly and Annotation of the Chinese Lizard Gudgeon (Saurogobio dabryi).

The Chinese lizard gudgeon (Saurogobio dabryi) is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in China, the lack of genomic resources has rendered the genetic breeding and conservation research. Here, we present the first chromosome-level genome assembly of S. dabryi using PacBio HiFi long reads, short reads, and Hi-C sequencing data. The final assembly reaches a total size of 1.09 Gb and Hi-C scaffolding anchors 99.55% of the assembled contigs onto 25 chromosomes, with a scaffold N50 reaching 43.15 Mb. The final genome assembly shows a BUSCO completeness of 98.39%. We annotated 659.55 Mb repetitive sequences and 26,036 protein-coding genes, 99.47% of which are functionally annotated. Comparative phylogenomic analysis clarifies the phylogenetic position of Saurogobio within Gobioninae. This high-quality genome provides a critical genetic basis for exploring cyprinid phylogeny, benthic adaptive evolution, genetic improvement, and conservation efforts of S. dabryi.

Saurogobio dabryi

Comparative transcriptome analysis provides insights into dorso-ventral color pattern formation of Holothuria edulis.

Animal body color patterns are highly diverse and play critical roles in camouflage, intraspecific communication, and environmental adaptation. Holothuria edulis, an important echinoderm inhabiting tropical waters, exhibits a typical dorsoventral dichromatism. This unique body color difference represents a key phenotypic trait for its habitat adaptation; however, the core differential genes regulating this trait remain to be elucidated. In this study, comparative transcriptome sequencing was performed on the dorsal and ventral body wall tissues of H. edulis, leading to the identification of a number of differentially expressed genes (DEGs), followed by GO functional annotation and KEGG pathway enrichment analysis. GO enrichment analysis indicated that the DEGs were significantly enriched in functional categories such as extracellular region, peptidase inhibitor activity, and tetrapyrrole binding. KEGG pathway analysis further revealed significant enrichment of protein digestion and absorption, the TNF signaling pathway, and cholesterol metabolism. Notably, the pigmentation-related gene FMO2 was highly expressed in the dorsal body wall tissue, whereas cyp1a1, ZIC1, Slc7a11, WNT-1, and ADAMTS20 were highly expressed in the ventral body wall tissue. This study identified DEGs and enriched pathways associated with dorsoventral body color differences in H. edulis, providing new insights into the molecular regulatory mechanisms underlying body color pattern formation. From the perspective of aquaculture applications, body color is one of the important traits affecting the quality and market value of sea cucumber products. Elucidating the molecular mechanisms of body color variation can provide a scientific basis for molecular marker-assisted breeding of superior sea cucumber variety.

Animals