PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “protein function annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Structural insights into adeno-associated virus serotype 5.

The adeno-associated viruses (AAVs) display differential cell binding, transduction, and antigenic characteristics specified by their capsid viral protein (VP) composition. Toward structure-function annotation, the crystal structure of AAV5, one of the most sequence diverse AAV serotypes, was determined to 3.45-Å resolution. The AAV5 VP and capsid conserve topological features previously described for other AAVs but uniquely differ in the surface-exposed HI loop between βH and βI of the core β-barrel motif and have pronounced conformational differences in two of the AAV surface variable regions (VRs), VR-IV and VR-VII. The HI loop is structurally conserved in other AAVs despite amino acid differences but is smaller in AAV5 due to an amino acid deletion. This HI loop is adjacent to VR-VII, which is largest in AAV5. The VR-IV, which forms the larger outermost finger-like loop contributing to the protrusions surrounding the icosahedral 3-fold axes of the AAVs, is shorter in AAV5, creating a smoother capsid surface topology. The HI loop plays a role in AAV capsid assembly and genome packaging, and VR-IV and VR-VII are associated with transduction and antigenic differences, respectively, between the AAVs. A comparison of interior capsid surface charge and volume of AAV5 to AAV2 and AAV4 showed a higher propensity of acidic residues but similar volumes, consistent with comparable DNA packaging capacities. This structure provided a three-dimensional (3D) template for functional annotation of the AAV5 capsid with respect to regions that confer assembly efficiency, dictate cellular transduction phenotypes, and control antigenicity.

Capsid Proteins↗

Chromosome-level genome assembly of Manglietia pachyphylla.

Manglietia pachyphylla, an endangered evergreen tree within the Magnoliaceae family, is renowned for its exceptional ornamental value in landscape horticulture. Despite its classification as a Category II nationally protected plant species in China, the genetic basis of its adaptive traits and conservation priorities remains poorly understood. To address this, we present the first chromosome-scale genome assembly of M. pachyphylla utilizing an integrated approach combining PacBio HiFi long-read and Hi-C chromosome conformation capture sequencing technologies. The assembled genome spans 2.15 Gb (contig N50 = 43.57 Mb), exhibiting a heterozygosity rate of 0.78% and repeat content of 78.64%, predominantly comprising long terminal repeat (LTR) retrotransposons (52.86%). Hi-C scaffolding anchored 99.57% of the assembly to 19 pseudochromosomes, achieving a BUSCO completeness score of 96.4%. Annotation revealed 42,505 putative protein-coding genes, with 84.46% of predicted genes were functionally annotated. Phylogenomic analysis positioned M. pachyphylla and Oyama sieboldii clustered together in a well-supported group. This high-contiguity genome assembly enables future investigations into adaptive evolution, functional genomics, and evidence-based conservation strategies for this endangered species.

Chromosomes, Plant↗

Enzyme function less conserved than anticipated.

The level of sequence similarity that implies similarity in protein structure is well established. Recently, many groups proposed thresholds for similarity in sequence implying similarity in enzymatic function. All previous results suggest the strong conservation of enzymatic function above levels of 50% pairwise sequence identity. Here, I argue that all groups substantially overestimated the conservation of enzyme function because their data sets were either too biased, or too small. An unbiased analysis suggested that less than 30% of the pair fragments above 50% sequence identity have entirely identical EC numbers. Another surprising finding was that even BLAST E-values below 10(-50) did not suffice to automatically transfer enzyme function without errors. As expected, most misclassifications originated from similarities in relatively short regions and/or from transferring annotations for different domains. Both problems cannot be corrected easily by adjusting the thresholds for automatic transfer of genome annotations. A score relating sequence identity to alignment length (distance from HSSP-threshold) outperformed statistical BLAST scores for high sequence similarity. In particular, the distance score allowed error-free transfer of enzyme function for the 10% most similar enzyme pairs. The results illustrated how difficult it is to assess the conservation of protein function and to guarantee error-free genome annotations, in general: sets with millions of pair comparisons might not suffice to arrive at statistically significant conclusions. In practice, the revised detailed estimates for the sequence conservation of enzyme function may provide important benchmarks for everyday sequence analysis and for more cautious automatic genome annotations.

Amino Acid Sequence↗

A gap-free, telomere-to-telomere chromosome-scale genome assembly of the mangrove red snapper, Lutjanus argentimaculatus.

The mangrove red snapper (Lutjanus argentimaculatus) is a commercially important marine fish species in the Indo-Pacific region. Despite its significant economic value for aquaculture, existing genomic resources remain fragmented, limiting the advancement of molecular breeding and functional genomic studies. Here, we present a gap-free, telomere-to-telomere (T2T) genome assembly of L. argentimaculatus, generated using a hybrid approach combining PacBio HiFi, Oxford Nanopore ultra-long reads and Hi-C technology. The resulting assembly comprises exactly 24 scaffolds spanning 1.03 Gb, perfectly matching the haploid chromosome number with a contig N50 of 46.17 Mb. Notably, this assembly resolves all physical gaps present in previous versions, achieving a BUSCO completeness score of 98.2%. Comprehensive genome annotation successfully predicted 23,167 protein-coding genes. Among these, 22,067 genes (95.25%) were functionally annotated across major public databases, including eggNOG, InterPro, and Swiss-Prot. Furthermore, structural analysis successfully identified 19 telomeres and 20 centromeres, validating the chromosomal integrity. This high-fidelity, gap-free reference genome provides a robust foundation for comparative genomics, population genetics, and the genetic improvement of Lutjanidae species.

Animals↗

The Lipase Engineering Database: a navigation and analysis tool for protein families.

The Lipase Engineering Database (LED) (http://www.led.uni-stuttgart.de) integrates information on sequence, structure, and function of lipases, esterases, and related proteins. Sequence data on 806 protein entries are assigned to 38 homologous families, which are grouped into 16 superfamilies with no global sequence similarity between each other. For each family, multisequence alignments are provided with functionally relevant residues annotated. Pre-calculated phylogenetic trees allow navigation inside superfamilies. Experimental structures of 45 proteins are superposed and consistently annotated. The LED has been applied to systematically analyze sequence-structure-function relationships of this vast and diverse enzyme class. It is a useful tool to identify functionally relevant residues apart from the active site residues, and to design mutants with desired substrate specificity.

Amino Acid Sequence↗

The SWISS-PROT protein sequence data bank and its new supplement TREMBL.

SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation (such as the description of the function of a protein, its domain structure, post-translational modifications, variants, etc), a minimal level of redundancy and a high level of integration with other databases. Recent developments of the database include: an increase in the number and scope of model organisms; cross-references to seven additional databases; a variety of new documentation files; the creation of TREMBL, and unannotated supplement to SWISS-PROT. This supplement consists of entries in SWISS-PROT-like format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except CDS already included in SWISS-PROT.

Amino Acid Sequence↗

Mass spectrometric analysis of the editosome and other multiprotein complexes in Trypanosoma brucei.

The composition of the editosome, a multi-protein complex that catalyzes uridine insertion and deletion RNA editing to produce mature mitochondrial mRNAs in trypanosomes, was analyzed by mass spectrometry. The editosomes were isolated by column chromatography, glycerol gradient sedimentation, and monoclonal antibody affinity purifications. At least 16 proteins form the catalytic core of the editosome, and additional associated proteins were identified. Analyses of mitochondrial fractions identified several non-editosome proteins and multi-protein complexes. These studies contribute to the functional annotation of T. brucei genome.

Amino Acid Sequence↗

Chromosomal level genome assembly of medicinal plant Chrysosplenium macrophyllum.

Chrysosplenium macrophyllum Oliv., a perennial herb native to China, is widely used in traditional medicine for its notable therapeutic properties. However, the absence of a reference genome has constrained its full potential for research and application. This study presents the first chromosome-level de novo genome assembly of C. macrophyllum, constructed by integrating long reads from Oxford Nanopore Technologies (ONT), short reads from BGI, and Hi-C data. The final assembly spans 2.55 Gb, with a scaffold N50 of 93.38 Mb, and 83.70% of the genome has been assigned to 22 chromosomes. The mapping rate of the BGI short reads to the genome is approximately 97.94%, and BUSCO analysis reveals that 97.94% of the predicted genes are complete. A total of 62,921 protein-coding genes were predicted, with functional annotations for 93.67% of them. This chromosome-level genome assembly represents an important resource for expanding our understanding of Chrysosplenium species and supports future genomic studies and applications.

Genome, Plant↗

The SBASE domain library: a collection of annotated protein segments.

SBASE is a database of annotated protein domain sequences representing various structural, functional, ligand binding and topogenic segments of proteins. The current release of SBASE contains 27,211 entries which are provided with standardized names in order to facilitate retrieval. SBASE is cross-referenced to the major protein and nucleic acid databanks as well as to the PROSITE catalog of protein sequence patterns [Bairoch, A. (1992) Nucleic Acids Res., 20, Suppl., 2013-2118]. SBASE can be used to establish domain homologies through database search using programs such as FASTA [Lipman and Pearson (1985) Science, 227, 1436-1441], FASTDB [Brutlag et al. (1990) Comp. Appl. Biosci., 6, 237-245] or BLAST3 [Altschul and Lipman (1990) Proc. Natl. Acad. Sci. USA, 87, 5509-5513], which is especially useful in the case of loosely defined domain types for which efficient consensus patterns cannot be established. The use of SBASE is illustrated on the DNA binding protein Brain-4. The database and a set of search and retrieval tools are freely available on request to the authors or by anonymous 'ftp' file transfer from < ftp.icgeb.trieste.it >.

Amino Acid Sequence↗

RNAi screening of uncharacterized genes identifies promising druggable targets in Schistosoma japonicum.

Schistosomiasis affects more than 250 million people worldwide and is one of the neglected tropical diseases. Currently, the treatment of schistosomiasis relies on a single drug-praziquantel-which has led to increasing pressure from drug resistance. Therefore, there is an urgent need to find new treatments. The development of genome sequencing has provided valuable information for understanding the biology of schistosomes. In the genome of Schistosoma japonicum, approximately 11% of the protein-coding sequences are uncharacterized genes (UGs) annotated as "hypothetical protein" or "protein of unknown function." These poorly understood genes have been unjustifiably neglected, although some may be essential for the survival of the parasites and serve as potential drug targets. In this study, we systematically mined the highly expressed UGs in both genders of this parasite throughout key developmental stages in their mammalian host, using our previously published S. japonicum genome and RNA-seq data. By employing in vitro RNA interference (RNAi), we screened 126 UGs that lack homologs in Homo sapiens and identified 8 that are essential for the parasite vitality. We further investigated two UGs, Sjc_0002003 and Sjc_0009272, which resulted in the most severe phenotypes. Fluorescence in situ hybridization demonstrated that both genes were expressed throughout the body without sex bias. Silencing either Sjc_0002003 or Sjc_0009272 reduced the cell proliferation in the body. Furthermore, in vivo RNAi indicated both genes are required for the growth and survival of the parasites in the mammalian host. For Sjc_0002003, we further characterize the underlying molecular cause of the observed phenotype. Through RNA-seq analysis and functional studies, we revealed that silencing Sjc_0002003 reduces the expression of a series of intestinal genes, including Sjc_0007312 (hypothetical protein), Sjc_0008276 (vha-17), Sjc_0002942 (PLA2G15), and Sjc_0003646 (SJCHGC09134 protein), leading to gut dilation. Our work highlights the importance of UGs in schistosomes as promising targets for drug development in the treatment of the schistosomiasis.

Schistosoma japonicum↗

A functional update of the Escherichia coli K-12 genome.

BACKGROUND: Since the genome of Escherichia coli K-12 was initially annotated in 1997, additional functional information based on biological characterization and functions of sequence-similar proteins has become available. On the basis of this new information, an updated version of the annotated chromosome has been generated. RESULTS: The E. coli K-12 chromosome is currently represented by 4,401 genes encoding 116 RNAs and 4,285 proteins. The boundaries of the genes identified in the GenBank Accession U00096 were used. Some protein-coding sequences are compound and encode multimodular proteins. The coding sequences (CDSs) are represented by modules (protein elements of at least 100 amino acids with biological activity and independent evolutionary history). There are 4,616 identified modules in the 4,285 proteins. Of these, 48.9% have been characterized, 29.5% have an imputed function, 2.1% have a phenotype and 19.5% have no function assignment. Only 7% of the modules appear unique to E. coli, and this number is expected to be reduced as more genome data becomes available. The imputed functions were assigned on the basis of manual evaluation of functions predicted by BLAST and DARWIN analyses and by the MAGPIE genome annotation system. CONCLUSIONS: Much knowledge has been gained about functions encoded by the E. coli K-12 genome since the 1997 annotation was published. The data presented here should be useful for analysis of E. coli gene products as well as gene products encoded by other genomes.

Bacterial Proteins↗

Improving the Annotations of JCVI-Syn3a Proteins.

The JCVI-Syn3 organism is a minimal organism derived from Mycoplasma mycoides capri, which is capable of self-replication. While the ancestor has 863 genes, the synthetic progeny has only 473, with 434 of these coding for proteins. Despite initial efforts to understand all functions of the organism, a significant number of these protein-coding genes still have unknown functions, and subsequent studies have been only partially successful in elucidating their roles. In this study, we employ our innovative method PROST to identify homologs and better understand these previously unidentified genes. PROST employs protein language embeddings and enables the identification of remote homologs with as low as 16% sequence identity. PROST successfully finds functionally annotated homologs for 93% of the minimal genome with a high level of accuracy, both confirming previously identified functions, as well as proposing new functions for others. The results of our study can be accessed at https://bit.ly/prost-syn3a .

Molecular Sequence Annotation↗

C. elegans ORFeome version 1.1: experimental verification of the genome annotation and resource for proteome-scale protein expression.

To verify the genome annotation and to create a resource to functionally characterize the proteome, we attempted to Gateway-clone all predicted protein-encoding open reading frames (ORFs), or the 'ORFeome,' of Caenorhabditis elegans. We successfully cloned approximately 12,000 ORFs (ORFeome 1.1), of which roughly 4,000 correspond to genes that are untouched by any cDNA or expressed-sequence tag (EST). More than 50% of predicted genes needed corrections in their intron-exon structures. Notably, approximately 11,000 C. elegans proteins can now be expressed under many conditions and characterized using various high-throughput strategies, including large-scale interactome mapping. We suggest that similar ORFeome projects will be valuable for other organisms, including humans.

Alternative Splicing↗

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75&#x2009;Gb with a contig N50 of 35.0&#x2009;Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5&#x2009;Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

The characterisation of novel secreted Ly-6 proteins from rat urine by the combined use of two-dimensional gel electrophoresis, microbore high performance liquid chromatography and expressed sequence tag data.

A proteomic study of rat urine was undertaken using two-dimensional gel electrophoresis, microbore high performance liquid chromatography, mass spectrometry and N-terminal sequencing. Five known urinary proteins were identified but two novel peptide fragments matched a large number of rat expressed sequence tags (ESTs) from a liver library. By combining protein chemical and nucleotide data, two 101-residue open reading frames with 90% amino acid identity were determined, rat urinary protein 1 (RUP-1) and RUP-2. The data established signal peptide removal and provided evidence for N-glycosylation. A third related sequence, rat spleen protein (RSP-1) was confirmed from EST searches. These three proteins have been submitted to SWISS-PROT as P81827, P81828 and Q9QXN2, respectively. A fourth novel homologue was found in porcine and bovine ESTs from embryo libraries. Alignment with known homologues showed conserved cysteine positions characteristic of a secreted subfamily of Ly-6 proteins. In two cases, antineoplastic urinary protein and caltrin, these homologues have unverified functional annotations. The RUP sequences showed high scoring matches to three unrelated rat mRNAs subsequently established to be chimeric. Two of these share extended sectional identity to RUP-1 but the third may represent another novel Ly-6 homologue. These chimeras have caused serious annotation errors in secondary databases.

Amino Acid Sequence↗

The complete genome sequence of Escherichia coli K-12.

The 4,639,221-base pair sequence of Escherichia coli K-12 is presented. Of 4288 protein-coding genes annotated, 38 percent have no attributed function. Comparison with five other sequenced microbes reveals ubiquitous as well as narrowly distributed gene families; many families of similar genes within E. coli are also evident. The largest family of paralogous proteins contains 80 ABC transporters. The genome as a whole is strikingly organized with respect to the local direction of replication; guanines, oligonucleotides possibly related to replication and recombination, and most genes are so oriented. The genome also contains insertion sequence (IS) elements, phage remnants, and many other patches of unusual composition indicating genome plasticity through horizontal transfer.

Bacterial Proteins↗

SubtiList: the reference database for the Bacillus subtilis genome.

SubtiList is the reference database dedicated to the genome of Bacillus subtilis 168, the paradigm of Gram-positive endospore-forming bacteria. Developed in the framework of the B.subtilis genome project, SubtiList provides a curated dataset of DNA and protein sequences, combined with the relevant annotations and functional assignments. Information about gene functions and products is continuously updated by linking relevant bibliographic references. Recently, sequence corrections arising from both systematic verifications and submissions by individual scientists were included in the reference genome sequence. SubtiList is based on a generic relational data schema and a World Wide Web interface developed for the handling of bacterial genomes, called GenoList. The World Wide Web interface was designed to allow users to easily browse through genome data and retrieve information according to common biological queries. SubtiList also provides more elaborate tools, such as pattern searching, which are tightly connected to the overall browsing system. SubtiList is accessible at http://genolist.pasteur.fr/SubtiList/. Similar bacterial databases are accessible at http://genolist.pasteur.fr/.

Bacillus subtilis↗

The SBASE protein domain library, release 6.0: a collection of annotated protein sequence segments.

The sixth release of the SBASE protein domain library sequences contains 130 703 annotated and crossreferenced entries corresponding to structural, functional, ligand-binding and topogenic segments of proteins. The entries were grouped based on standard names (2312 groups) and futher classified on the basis of the BLAST similarity (2463 clusters). Automated searching with BLAST and a new sequence-plot representation of local domain similarities are available at the WWW-server http://www.icgeb.trieste.it/sbase. A mirror site is at http://sbase.abc.hu/sbase. The database is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it

Amino Acid Sequence↗