PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

SMART: a web-based tool for the study of genetically mobile domains.

SMART (a Simple Modular Architecture Research Tool) allows the identification and annotation of genetically mobile domains and the analysis of domain architectures (http://SMART.embl-heidelberg.de ). More than 400 domain families found in signalling, extra-cellular and chromatin-associated proteins are detectable. These domains are extensively annotated with respect to phyletic distributions, functional class, tertiary structures and functionally important residues. Each domain found in a non-redundant protein database as well as search parameters and taxonomic information are stored in a relational database system. User interfaces to this database allow searches for proteins containing specific combinations of domains in defined taxa.

Database Management Systems↗

mettannotator: a comprehensive and scalable Nextflow annotation pipeline for prokaryotic assemblies.

SUMMARY: In recent years, there has been a surge in prokaryotic genome assemblies, coming from both isolated organisms and environmental samples. These assemblies often include novel species that are poorly represented in reference databases creating a need for a tool that can annotate both well-described and novel taxa, and can run at scale. Here, we present mettannotator-a comprehensive, scalable Nextflow pipeline for prokaryotic genome annotation that identifies coding and noncoding regions, predicts protein functions, including antimicrobial resistance, and delineates gene clusters. The pipeline summarizes these results in a GFF (General Feature Format) file that can be easily utilized in downstream analysis or visualized using common genome browsers. Here, we show how it works on 200 genomes from 29 prokaryotic phyla, including isolate genomes and known and novel metagenome-assembled genomes, and present metrics on its performance in comparison to other tools. AVAILABILITY AND IMPLEMENTATION: The pipeline is written in Nextflow and Python and published under an open source Apache 2.0 licence. Instructions and source code can be accessed at https://github.com/EBI-Metagenomics/mettannotator. The pipeline is also available on WorkflowHub: https://workflowhub.eu/workflows/1069.

Software↗

BEAUTY: an enhanced BLAST-based search tool that integrates multiple biological information resources into sequence similarity search results.

BEAUTY (BLAST enhanced alignment utility) is an enhanced version of the NCBI's BLAST data base search tool that facilitates identification of the functions of matched sequences. We have created new data bases of conserved regions and functional domains for protein sequences in NCBI's Entrez data base, and BEAUTY allows this information to be incorporated directly into BLAST search results. A Conserved Regions Data Base, containing the locations of conserved regions within Entrez protein sequences, was constructed by (1) clustering the entire data base into families, (2) aligning each family using our PIMA multiple sequence alignment program, and (3) scanning the multiple alignments to locate the conserved regions within each aligned sequence. A separate Annotated Domains Data Base was constructed by extracting the locations of all annotated domains and sites from sequences represented in the Entrez, PROSITE, BLOCKS, and PRINTS data bases. BEAUTY performs a BLAST search of those Entrez sequences with conserved regions and/or annotated domains. BEAUTY then uses the information from the Conserved Regions and Annotated Domains data bases to generate, for each matched sequence, a schematic display that allows one to directly compare the relative locations of (1) the conserved regions, (2) annotated domains and sites, and (3) the locally aligned regions matched in the BLAST search. In addition, BEAUTY search results include World-Wide Web hypertext links to a number of external data bases that provide a variety of additional types of information on the function of matched sequences. This convenient integration of protein families, conserved regions, annotated domains, alignment displays, and World-Wide Web resources greatly enhances the biological informativeness of sequence similarity searches. BEAUTY searches can be performed remotely on our system using the "BCM Search Launcher" World-Wide Web pages (URL is < http:/ /gc.bcm.tmc.edu:8088/ search-launcher/launcher.html > ).

Amino Acid Sequence↗

A chromosome-level reference genome assembly of the Small snakehead (Channa asiatica).

The Small snakehead (Channa asiatica) is an economically important species in both aquaculture and ornamental trade, mainly distributed in South China and Southeast Asia. Despite its significance, limited genomic resources have impeded in-depth genetic studies and breeding programs. In this study, we used PacBio HiFi long-read sequencing, Illumina short-read sequencing, and Hi-C technologies to generate a high-quality chromosome-level genome of the C. asiatica. The final genome spans 659.44&#x2009;Mb, with an impressive 98.18% anchored to 23 chromosomes. Notably, the contig N50 and scaffold N50 are 23.92&#x2009;Mb and 29.61&#x2009;Mb, validated by a BUSCO completeness score of 98.93%. Genome annotation identified 26,603 protein-coding genes, 99.29% of which were confirmed by BUSCO analysis, and 93.68% were functionally annotated. Approximately 27.72% of the genome sequences were classified as repeat elements. This high-fidelity genome assembly provides a robust foundation for advancing molecular breeding, comparative genomics, and evolutionary studies of C. asiatica and related species.

Animals↗

Annotated genome of the Atlantic dog whelk, Nucella lapillus.

Nucella lapillus is an important player in rocky shore food chains and has been a focal organism of ecological and evolutionary studies for decades. Despite poor dispersal, they have a broad geographic range, which makes them an ideal species to examine isolation by distance and selection across environmental gradients. Here we present the fully annotated genome of N. lapillus generated with Oxford Nanopore Techonology sequencing at &#x223c;37&#xd7; coverage. The genome assembly is 2.32 Gbp and consists of 2,525 contigs, with an N50 length of 2 Mbp. Repeat annotation identified 2,491 families that cover 67.56% of the genome, which is similar to other gastropods. Despite its large size and high proportion of repeats, the genome is of high quality. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis revealed a score of 96.8%. Functional annotation of the genome produced 45,848 protein-coding genes with a 96.6% BUSCO score. Genomic resources for mollusks lag behind that of other phyla, perhaps because many of their innate characteristics complicate DNA extraction, sequencing, and assembly. This new N. lapillus genome will increase our genomic understanding of the second largest phylum (and the most diverse class within said phylum) and serve as a key resource to advance studies on the organismal biology and population genetics of this iconic species as well as the connection between genomic variation and community-level processes.

Animals↗

The SBASE protein domain library, release 2.0: a collection of annotated protein sequence segments.

SBASE 2.0 is the second release of SBASE, a collection of annotated protein domain sequences. SBASE entries represent various structural, functional, ligand-binding and topogenic segments of proteins [Pongor, S. et al. (1993) Prot. Eng., in press]. This release contains 34,518 entries provided with standardized names and it is cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalog of protein sequence patterns [Bairoch, A. (1992) Nucl. Acids Res., 20 suppl, 2013-2018]. SBASE can be used for establishing domain homologies using different database-search tools such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], FASTDB [Brutlag et al. (1990) Comp. Appl. Biosci., 6, 237-245] or BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513] which is especially useful in the case of loosely defined domain types for which efficient consensus patterns can not be established. SBASE 2.0 and a set of search and retrieval tools are freely available on request to the authors or by anonymous 'ftp' file transfer from mean value of ftp.icgeb.trieste.it.

Amino Acid Sequence↗

The Complete Chloroplast Genome and the Phylogenetic Analysis of Panicum bisulcatum (Thumb.) (Poaceae).

The chloroplast (cp) genome of Panicum bisulcatum (Thumb.), a significant agricultural weed, was sequenced and characterized to elucidate its genomic architecture, evolutionary dynamics, and phylogenetic relationships. The complete cp genome was assembled as a circular DNA molecule of 138,489 bp, exhibiting a typical quadripartite structure comprising a large single-copy (LSC, 82,260 bp), a small single-copy (SSC, 12,569 bp), and a pair of inverted repeats (IR, 21,830 bp each) regions. It encodes 135 genes, including 89 protein-coding genes, 49 tRNAs, and 8 rRNAs. Functional annotation revealed that most genes are involved in photosynthesis and genetic system. A total of 51 simple sequence repeats (SSRs) and 62 long repeats (LRs) were identified, providing potential molecular markers. Comparative analysis of IR boundaries highlighted both conserved features and species-specific expansion/contraction events among Panicum species. Phylogenomic analysis robustly placed P. bisulcatum within the genus Panicum, showing a closest relationship with P. incomtum and confirming the monophyly of the genus. Furthermore, single nucleotide polymorphism (SNP) analysis with its closest relative, P. incomtum, revealed 4659 SNPs, with a dominance of synonymous substitutions, indicating the action of purifying selection. This study provides the first comprehensive cp genomic resource for P. bisulcatum, which will facilitate future studies in species identification, phylogenetic reconstruction, population genetics, and the development of sustainable management strategies for this weed.

Phylogeny↗

The SBASE protein domain library, release 3.0: a collection of annotated protein sequence segments.

SBASE 3.0 is the third release of SBASE, a collection of annotated protein domain sequences. SBASE entries represent various structural, functional, ligand-binding and topogenic segments of proteins as defined by their publishing authors. SBASE can be used for establishing domain homologies using different database-search tools such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], and BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513] which is especially useful in the case of loosely defined domain types for which efficient consensus patterns can not be established. The present release contains 41,749 entries provided with standardized names and cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalogue of protein sequence patterns. The entries are clustered into 2285 groups using the BLAST algorithm for computing similarity measures. SBASE 3.0 is freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >. Individual records can be retrieved with the gopher server at < icgeb.trieste.it > and with a www-server at < http:@www.icgeb.trieste.it >. Automated searching of SBASE by BLAST can be carried out with the electronic mail server < sbase@icgeb.trieste.it >. Another mail server < domain@hubi.abc.hu > assigns SBASE domain homologies on the basis of SWISS-PROT searches. A comparison of pertinent search strategies is presented.

Amino Acid Sequence↗

High-throughput functional annotation of novel gene products using document clustering.

Gene products differentially expressed in healthy vs. diseased tissues may be considered drug targets since the change in their expression level can be related to the cause and progression of the disease studied. A significant portion of the proteins produced by these genes will be unknown and consequently their function must be characterised. The experimental elucidation of biochemical function must be supported by computational tools which can help predicting the possible function of a given protein from its amino acid sequence. We have designed a high-throughput system which automatically analyses amino acid sequences deduced from differentially represented cDNA clones. The system attempts to assign a biological function to protein sequences by carrying out searches in sequence databanks and by locating functionally relevant motifs in the query sequences. The results delivered by the various prediction methods consist of the annotations of matching sequences and/or motifs, which are free-format texts written by humans and therefore may describe the same concept with synonymous words. It is desirable to present the results in such a way that the annotations describing the same biological function are grouped together. To this end we devised an algorithm that enables the hierarchical clustering of free-format documents based on their contents. The system is capable of detecting and flagging conflicting annotations, and will speed up the interpretation of the function prediction results.

Algorithms↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9&#x2009;Mb assembly is highly continuous (scaffold N50 of 253.58&#x2009;Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals↗

Complementing genomics with proteomics: the membrane subproteome of Pseudomonas aeruginosa PAO1.

With the completion of many genome projects, a shift is now occurring from the acquisition of gene sequence to understanding the role and context of gene products within the genome. The opportunistic pathogen Pseudomonas aeruginosa is one organism for which a genome sequence is now available, including the annotation of open reading frames (ORFs). However, approximately one third of the ORFs are as yet undefined in function. Proteomics can complement genomics, by characterising gene products and their response to a variety of biological and environmental influences. In this study we have established the first two-dimensional gel electrophoresis reference map of proteins from the membrane fraction of P. aeruginosa strain PA01. A total of 189 proteins have been identified and correlated with 104 genes from the P. aeruginosa genome. Annotated membrane proteins could be grouped into three distinct categories: (i) those with functions previously characterised in P. aeruginosa (38%); (ii) those with significant sequence similarity to proteins with assigned function or hypothetical proteins in other organisms (46%); and (iii) those with unknown function (16%). Transmembrane prediction algorithms showed that each identified protein sequence contained at least one membrane-spanning region. Furthermore, the current methodology used to isolate the membrane fraction was shown to be highly specific since no contaminating cytosolic proteins were characterised. Preliminary analysis showed that at least 15 gel spots may be glycosylated in vivo, including three proteins that have not previously been functionally characterised. The reference map of membrane proteins from this organism is now the basis for determining surface molecules associated with antibiotic resistance and efflux, cell-cell signalling and pathogen-host interactions in a variety of P. aeruginosa strains.

Bacterial Proteins↗

Protective effects of seminal exosomes on cryopreserved sperm via inhibiting oxidative damage.

This study aimed to explore the protective effect of seminal plasma exosomes (SPEs) on human sperm structure and function during cryopreservation and its potential mechanism. The samples were divided into two groups: the control group was treated solely with sperm cryoprotectant before freezing, while the exosome group was supplemented with SPEs. After cryopreservation and thawing, sperm progressive motility, normal morphological rate, and survival rate were evaluated. Furthermore, PKH67 labeling experiments were performed, and oxidative stress markers as well as energy metabolism indicators in sperm were detected. Subsequent mechanism exploration was conducted via proteomic analysis and protein validation assays. This work reveals that adding SPEs at a concentration of 1 or 2&#xa0;mg/ml effectively improves sperm progressive motility after cryopreservation. After supplementing with SPEs, sperm glucose levels are reduced and mitochondrial membrane potential is enhanced. Simultaneously, SPEs alleviate oxidative stress by decreasing reactive oxygen species (ROS) and DNA fragment index (DFI) while increasing superoxide dismutase (SOD) activity. Functional annotation of proteomics reveals that 14 of the differentially expressed proteins (DEPs) are associated with sperm motility. Enriched metabolic pathways related to sperm motility and sperm protein validation experiments indicate that the expression of MAPK, p-MAPK, and p-JNK proteins in sperm is higher in the Exosome group than in the Control group. This study provides important theoretical support for the application of SPEs in mitigating cryopreservation damage to sperm by enhancing antioxidant capacity. The specific mechanism may be mediated by the MAPK/p-JNK pathway.

Male↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

Chromosomal-level genome assembly of minute pirate bug Orius nagaii Yasunaga, 1993 (Hemiptera: Anthocoridae).

Species of the genus Orius, diminutive predatory insects that act as natural enemies of other arthropods, are frequently employed in agricultural pest management for controlling various pests, such as thrips, mites, aphids, whiteflies, etc. However, the scarcity of high-quality genomic resources for these predators hinders our comprehension of their population evolution and predation ecology. Consequently, we assembled and annotated a chromosomal-scale genome of Orius nagaii by collating PacBio and Illumina sequencing and Hi-C genomic analysis techniques. The final genome assembly size 152.62&#x2009;Mb, with scaffold and contig N50 lengths of 11.53 and 2.39&#x2009;Mb, respectively. It is organized into 12 pairs of autosomes and a pair of XY sex chromosomes. The quality assessment of the genomic data with BUSCO revealed a completeness of 98.5% (n&#x2009;=&#x2009;1,367). Also, 11,917 protein-coding genes were discovered, with 94.28% of them having functional annotations. The high-quality genome of O. nagaii produced serves as a valuable resource for comprehending the interactions between predatory natural enemies and hosts, along with their evolutionary trajectories.

Animals↗

Near-complete reference genome assembly of Hoya carnosa.

Hoya R. Br. is the largest genus in the tribe Marsdenieae (Apocynaceae), comprising 350-450 species. Hoya species are popular in horticulture for their distinctive floral traits and fragrances, primarily sourced from domestication and mutation breeding. However, the lack of molecular analysis for floral morphological traits has limited their cultivation and application. In this study, we assembled a near-complete reference genome for H. carnosa, the model species of the genus, using PacBio HiFi reads and Hi-C method. The genome size was approximately 465.7&#x2009;Mb with a contig N50 of 39.3&#x2009;Mb. 99.7% of the sequences were anchored to 11 pseudochromosomes, and the assembly achieved a BUSCO score of 98.5%. We predicted 24,309 protein-coding genes, of which 90.2% (21,927) were functionally annotated. This high-quality genome provides a valuable reference for the research of evolution, conservation and molecular breeding in Hoya.

Genome, Plant↗

Chromosomal-level genome assembly of Trypanosoma carassii, the etiologic agent of a recent outbreak of trypanosomiasis in cage-cultured large yellow croaker (Larimichthys crocea) in China.

Trypanosoma carassii, a typical freshwater fish trypanosome, has recently been identified as the etiological agent of a trypanosomiasis outbreak in cage-cultured large yellow croaker (Larimichthys crocea) in China and has been designated as T. c. larimichthys. To date, publicly available genomic data for trypanosomes have been limited to terrestrial species, particularly those of medical importance. Here, we present a chromosome-level genome assembly of T. carassii, the first genome of an aquatic trypanosome, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding technologies. A preliminary genome survey based on Illumina sequencing data estimated the genome size at 56.38&#x2009;Mb with a heterozygosity of 1.17%. The final assembled genome spans 48.55&#x2009;Mb, with contig N50 and scaffold N50 values of 139.15 Kb, and achieves 100.00% BUSCO completeness. Hi-C data resolved the assembly into 34 chromosomes and 9 unanchored scaffolds. Repetitive elements account for 53.29% of the genome (approximately 25.87&#x2009;Mb). A total of 11,584 protein-coding genes were predicted, 95.36% of which were functionally annotated. Synonymous substitution rates analysis of paralogous genes indicates a recent burst of gene duplication, which likely corresponds to a whole-genome duplications. This high-quality genome assembly provides invaluable resources for understanding the evolution and host adaptation of aquatic trypanosomes.

Animals↗

A high-quality chromosome-level genome assembly of apple of Peru (Nicandra physalodes).

Nicandra physalodes, a member of the Solanaceae family, is known for its medicinal potential and strong natural insect-repellent properties, which are mainly attributed to its bioactive withanolides and alkaloids. Despite its ecological and pharmacological significance, genomic information for this species has remained limited. Here, we generated a chromosome-level reference genome for N. physalodes based on PacBio high-fidelity (HiFi) long-read sequencing and Hi-C scaffolding. The assembled genome is 933.97&#x2009;Mb in size, with a contig N50 of 87.37&#x2009;Mb, and 99.95% (933.54&#x2009;Mb) of the sequences anchored to 10 pseudochromosomes. Repetitive elements account for 73.06% of the genome, and 27,925 protein-coding genes were predicted, 97.81% of which were functionally annotated. This genomic resource provides a valuable foundation for investigating the genetic basis of specialized metabolite biosynthesis, insect resistance, and environmental adaptation in N. physalodes, as well as for comparative studies within the Solanaceae family.

Genome, Plant↗