PubMed HealthSearch

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Improving the Annotations of JCVI-Syn3a Proteins.

The JCVI-Syn3 organism is a minimal organism derived from Mycoplasma mycoides capri, which is capable of self-replication. While the ancestor has 863 genes, the synthetic progeny has only 473, with 434 of these coding for proteins. Despite initial efforts to understand all functions of the organism, a significant number of these protein-coding genes still have unknown functions, and subsequent studies have been only partially successful in elucidating their roles. In this study, we employ our innovative method PROST to identify homologs and better understand these previously unidentified genes. PROST employs protein language embeddings and enables the identification of remote homologs with as low as 16% sequence identity. PROST successfully finds functionally annotated homologs for 93% of the minimal genome with a high level of accuracy, both confirming previously identified functions, as well as proposing new functions for others. The results of our study can be accessed at https://bit.ly/prost-syn3a .

Molecular Sequence Annotation

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals

Genome-Resolved Functional Profiling of Osteoporosis-Associated Gut Bacteria Highlights Putative Metabolic and Immunogenic Signatures of the Gut-Bone Axis.

The gut microbiota has emerged as a potential regulator of bone metabolism, but the genome-encoded functional repertoire of osteoporosis-associated gut bacteria remains insufficiently characterized. This study performed in silico functional profiling of gut bacterial taxa associated with osteoporosis, low bone mineral density, or comparator bone-related phenotypes. Twenty candidate taxa were selected from evidence in the human microbiome and represented by 26 curated bacterial reference genomes. Genome-wide annotations were used to map predicted gut-bone axis signatures, carbohydrate-active enzyme (CAZyme) repertoires, selected Kyoto Encyclopedia of Genes and Genomes pathways, and gutSMASH-predicted metabolic gene clusters. Functional burdens were normalized as hits per 1000 annotated proteins and integrated into metabolic, immunogenic, CAZyme, KEGG, and metabolic gene cluster profiles. Twelve predicted gut-bone axis signatures were identified, comprising 3337 primary candidate protein hits and a strict high-confidence subset of 2497 hits. Dominant signatures included vitamin B12/cobalamin metabolism, folate/one-carbon metabolism, peptidoglycan/cell-wall biosynthesis, and short-chain fatty acid-related functions. Dialister invisus, Dialister succinatiphilus, Megamonas funiformis, and Megamonas hypermegale showed the strongest normalized predicted gut-bone axis signal. These hypothesis-generating findings prioritize microbial metabolic and immunogenic features for future metagenomic, metabolomic, and experimental validation studies.

Osteoporosis

Chromosome-level genome assembly of Elaeocarpus petiolatus (Elaeocarpaceae).

Elaeocarpus petiolatus is an ecologically and economically important species in tropical and subtropical forests. Despite its significance, the lack of genomic resources has hindered research on the genetic diversity and adaptive traits of E. petiolatus. To address this gap, we present a comprehensive chromosome-level genome assembly of E. petiolatus generated using advanced PacBio high-fidelity (HiFi) long-read sequencing and Hi-C technology. The assembly spans 322.45 Mb, with a scaffold N50 of 20.58 Mb, indicating that 37.11% of the genome is composed of repetitive elements. We identified 25,295 protein-coding genes, of which 96.74% were functionally annotated. This high-quality genome provides a critical resource for understanding the genetic mechanisms underlying environmental adaptability and biosynthesis of bioactive compounds in E. petiolatus, thereby supporting conservation efforts and sustainable forest management. The assembled genome and associated sequencing data are publicly available, facilitating further evolutionary and functional studies on the Elaeocarpaceae family.

Chromosomes, Plant

Chromosome-level genome assembly of the hemiparasitic Taxillus sutchuenensis (Loranthaceae).

Taxillus sutchuenensis, an ecologically and medicinally important hemiparasitic plant that parasitizes diverse woody hosts, was sequenced to generate a high-quality chromosome-level genome assembly. PacBio HiFi long reads, RNA-seq transcriptome data, and Hi-C data were used to assemble a 406.32 Mb genome anchored onto nine pseudo-chromosomes, with a scaffold N50 of 45.59 Mb. The assembly showed high completeness and accuracy, supported by BUSCO (93.6%) and Merqury QV (70.6) assessments. The LTR Assembly Index (LAI) of 13.98 indicated excellent continuity. A total of 21,795 protein-coding genes were predicted, with 94.46% functionally annotated. Repetitive sequences accounted for 50.05% of the genome, primarily LTR retrotransposons. This genome provides a valuable resource for investigating the evolution, functional genomics, and parasitic mechanisms of hemiparasitic plants.

Genome, Plant

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44 Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92 Mb and 29.61 Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals

Annotated genome of the Atlantic dog whelk, Nucella lapillus.

Nucella lapillus is an important player in rocky shore food chains and has been a focal organism of ecological and evolutionary studies for decades. Despite poor dispersal, they have a broad geographic range, which makes them an ideal species to examine isolation by distance and selection across environmental gradients. Here we present the fully annotated genome of N. lapillus generated with Oxford Nanopore Techonology sequencing at ∼37× coverage. The genome assembly is 2.32 Gbp and consists of 2,525 contigs, with an N50 length of 2 Mbp. Repeat annotation identified 2,491 families that cover 67.56% of the genome, which is similar to other gastropods. Despite its large size and high proportion of repeats, the genome is of high quality. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis revealed a score of 96.8%. Functional annotation of the genome produced 45,848 protein-coding genes with a 96.6% BUSCO score. Genomic resources for mollusks lag behind that of other phyla, perhaps because many of their innate characteristics complicate DNA extraction, sequencing, and assembly. This new N. lapillus genome will increase our genomic understanding of the second largest phylum (and the most diverse class within said phylum) and serve as a key resource to advance studies on the organismal biology and population genetics of this iconic species as well as the connection between genomic variation and community-level processes.

Animals

The Complete Chloroplast Genome and the Phylogenetic Analysis of Panicum bisulcatum (Thumb.) (Poaceae).

The chloroplast (cp) genome of Panicum bisulcatum (Thumb.), a significant agricultural weed, was sequenced and characterized to elucidate its genomic architecture, evolutionary dynamics, and phylogenetic relationships. The complete cp genome was assembled as a circular DNA molecule of 138,489 bp, exhibiting a typical quadripartite structure comprising a large single-copy (LSC, 82,260 bp), a small single-copy (SSC, 12,569 bp), and a pair of inverted repeats (IR, 21,830 bp each) regions. It encodes 135 genes, including 89 protein-coding genes, 49 tRNAs, and 8 rRNAs. Functional annotation revealed that most genes are involved in photosynthesis and genetic system. A total of 51 simple sequence repeats (SSRs) and 62 long repeats (LRs) were identified, providing potential molecular markers. Comparative analysis of IR boundaries highlighted both conserved features and species-specific expansion/contraction events among Panicum species. Phylogenomic analysis robustly placed P. bisulcatum within the genus Panicum, showing a closest relationship with P. incomtum and confirming the monophyly of the genus. Furthermore, single nucleotide polymorphism (SNP) analysis with its closest relative, P. incomtum, revealed 4659 SNPs, with a dominance of synonymous substitutions, indicating the action of purifying selection. This study provides the first comprehensive cp genomic resource for P. bisulcatum, which will facilitate future studies in species identification, phylogenetic reconstruction, population genetics, and the development of sustainable management strategies for this weed.

Phylogeny

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62 Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52 Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals

Protective effects of seminal exosomes on cryopreserved sperm via inhibiting oxidative damage.

This study aimed to explore the protective effect of seminal plasma exosomes (SPEs) on human sperm structure and function during cryopreservation and its potential mechanism. The samples were divided into two groups: the control group was treated solely with sperm cryoprotectant before freezing, while the exosome group was supplemented with SPEs. After cryopreservation and thawing, sperm progressive motility, normal morphological rate, and survival rate were evaluated. Furthermore, PKH67 labeling experiments were performed, and oxidative stress markers as well as energy metabolism indicators in sperm were detected. Subsequent mechanism exploration was conducted via proteomic analysis and protein validation assays. This work reveals that adding SPEs at a concentration of 1 or 2 mg/ml effectively improves sperm progressive motility after cryopreservation. After supplementing with SPEs, sperm glucose levels are reduced and mitochondrial membrane potential is enhanced. Simultaneously, SPEs alleviate oxidative stress by decreasing reactive oxygen species (ROS) and DNA fragment index (DFI) while increasing superoxide dismutase (SOD) activity. Functional annotation of proteomics reveals that 14 of the differentially expressed proteins (DEPs) are associated with sperm motility. Enriched metabolic pathways related to sperm motility and sperm protein validation experiments indicate that the expression of MAPK, p-MAPK, and p-JNK proteins in sperm is higher in the Exosome group than in the Control group. This study provides important theoretical support for the application of SPEs in mitigating cryopreservation damage to sperm by enhancing antioxidant capacity. The specific mechanism may be mediated by the MAPK/p-JNK pathway.

Male

Chromosomal-level genome assembly of minute pirate bug Orius nagaii Yasunaga, 1993 (Hemiptera: Anthocoridae).

Species of the genus Orius, diminutive predatory insects that act as natural enemies of other arthropods, are frequently employed in agricultural pest management for controlling various pests, such as thrips, mites, aphids, whiteflies, etc. However, the scarcity of high-quality genomic resources for these predators hinders our comprehension of their population evolution and predation ecology. Consequently, we assembled and annotated a chromosomal-scale genome of Orius nagaii by collating PacBio and Illumina sequencing and Hi-C genomic analysis techniques. The final genome assembly size 152.62 Mb, with scaffold and contig N50 lengths of 11.53 and 2.39 Mb, respectively. It is organized into 12 pairs of autosomes and a pair of XY sex chromosomes. The quality assessment of the genomic data with BUSCO revealed a completeness of 98.5% (n = 1,367). Also, 11,917 protein-coding genes were discovered, with 94.28% of them having functional annotations. The high-quality genome of O. nagaii produced serves as a valuable resource for comprehending the interactions between predatory natural enemies and hosts, along with their evolutionary trajectories.

Animals

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7 Mb with a contig N50 of 39.3 Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38 Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55 Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87 Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97 Mb in size, with a contig N50 of 87.37 Mb, and 99.95% (933.54 Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant

A chromosome-level assembly of the alpine snow alga Chloromonas typhlos.

Chloromonas typhlos is a cosmopolitan alpine snow alga distributed across continents, and its blooming accelerates snow melting by decreasing the amount of snow albedo. To elucidate the genetic traits underlying the adaptation of C. typhlos to the alpine habitat, we combined PacBio sequencing and Hi-C to generate a high-quality chromosome-level genome assembly (contig N50: 1.29 Mb; scaffold N50: 7.23 Mb) with 31 chromosomes and a genome size of 200.86 Mb. Repetitive elements constituted 11.05% of the genome, and 16,133 protein-coding genes were predicted, of which 82% were functionally annotated. This study provides a set of omics resources both for snow algae and the genus Chloromonas.

Snow

In silico analysis of SH3BP2 genomic alterations and expression profiles in CRC.

AIM: Colorectal cancer (CRC) is a widespread health issue that attains high mortality. The adaptor protein SH3BP2 amplification results in metabolic changes, oxidative stress, NK cell activity, and inflammation. The NK cells are capable of destroying tumor cells without prior activation, help prevent metastasis, and have prognostic value. Targeting SH3BP2 to regulate NK cell activity in the TME could enhance CRC-based immunotherapy. MATERIALS AND METHODS: The cancer hallmark tool helps in understanding SH3BP2 hallmark annotation. Utilizing the STRING tool and the KEGG pathway, protein functional enrichment and PPI networking were analyzed. TIMER 2.0 was used for immune cell infiltration correlation analysis, and UALCAN was used for CPTAC-based protein expression profiling. RESULTS AND CONCLUSIONS: The GEO (GSE9348) dataset showed SH3BP2 is upregulated in CRC (log2 fold change = 1.18). GEO, TCGA, and cBioPortal revealed SH3BP2 alterations in CRC cases, potentially aiding immune evasion. Mutations in SH3BP2 influence cancer growth, suppressing tumors or promoting them by activating NF-κB and affecting immune responses through WNT/β-catenin, PI3K, MAPK, and JAK-STAT pathways. Overall, SH3BP2 plays a key role in cancer growth and immune regulation, making it a promising target for CRC therapy. Further experimental validation is needed to demonstrate its diagnostic and therapeutic potency.

Humans