PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference protein database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Human liver protein map: a reference database established by microsequencing and gel comparison.

This publication establishes a reference human liver protein map obtained with immobilized pH gradients. By microsequencing, 57 spots or 42 polypeptide chains were identified. By protein map comparison and matching (liver, red blood cell and plasma sample maps), 8 additional proteins were identified. The new polypeptides and previously known proteins are listed in a table and/or labeled on the protein map, thus providing a human liver two-dimensional gel database. This reference map can be used to identify protein spots on other samples such as rectal cancer biopsies.

Amino Acid Sequence

Expanding kinetoplastid genome annotation through protein structure comparison.

Kinetoplastids belong to the Discoba supergroup, an early divergent eukaryotic clade. Although the amount of genomic information on these parasites has grown substantially, assigning gene functions through traditional sequence-based homology methods remains challenging. Recently, significant advancements have been made in in-silico protein structure prediction and algorithms for rapid and precise large-scale protein structure comparisons. In this work, we developed a protein structure-based homology search pipeline (ASC, Annotation by Structural Comparisons) and applied it to transfer biological information to all kinetoplastid proteins available in TriTrypDB, the reference database for this lineage. Our pipeline enabled the assignment of structural similarity to a substantial portion of kinetoplastid proteins, improving current knowledge through annotation transfer. Additionally, we identified structural homologs for representatives of 6,700 uncharacterized proteins across 33 kinetoplastid species, proteins that could not be annotated using existing sequence-based tools and databases. As a result, this approach allowed us to infer potential biological information for a considerable number of kinetoplastid proteins. Among these, we identified structural homologs to ubiquitous eukaryotic proteins that are challenging to detect in kinetoplastid genomes through standard genome annotation pipelines. The results (KASC, Kinetoplastid Annotation by Structural Comparison) are openly accessible to the community at kasc.fcien.edu.uy through a user-friendly, gene-by-gene interface that enables visual inspection of the data.

Kinetoplastida

Combining Annotation Software to Identify Orthologous Genes (CASIO) Provides a New Dataset of Orthologous Genes for Swallowtail Butterflies.

With the massive increase in genomic resources, it is becoming increasingly popular to analyse thousands of loci across many species. However, many of the available genomes are not annotated, which hinders an efficient search for orthologous protein-coding genes. Here, we aim to develop a semi-automated pipeline and compare four genomic annotation methods (BRAKER2, BUSCO, Miniprot and Scipio). Our results highlight the importance of integrating multiple annotation tools to optimise ortholog detection and improve genomic studies. Each annotation method showed different strengths. BRAKER2 annotated a substantial number of genes. BUSCO, despite limitations inherent to its reference database, identified a higher number of orthologs. Miniprot exhibited notable flexibility in accommodating diverse protein datasets, whereas Scipio successfully recovered a considerable set of genes that were not detected by the other tools. The combination of these tools allowed for more comprehensive ortholog detection. Taking advantage of this pipeline, we developed a comprehensive dataset of orthologous genes for swallowtail butterflies (Lepidoptera: Papilionidae), called Papilionidae_odb, which will facilitate future studies, especially for a non-model group with abundant genomic data and few transcriptomic resources. We tested Papilionidae_odb by inferring a robust phylogenetic framework for Leptocircini using 142 complete genomes, which improved branch support for some phylogenetic relationships, although challenges remained in resolving relationships within certain species groups, likely due to rapid radiations. Our results highlight the complementary nature of the annotation methods and suggest that combining these tools can yield more accurate results in genomic research. This approach was implemented in a Snakemake workflow called CASIO (Combining Annotation Software to Identify Orthologous genes) and can easily be applied to other non-model groups to improve genomic datasets in diverse taxa where transcriptomic resources are still limited.

Animals

Workshop on two-dimensional gel protein databases.

A workshop on two-dimensional gel electrophoresis (2-DE) protein database, organized by the Committee on Data for Science and Technology (CODATA) of the International Council of Scientific Unions Task Group on Biological Macromolecules, was held at the CODATA Secretariat in Paris on March 9, 1992. Eleven scientists from eight different countries represented various aspects of 2-DE analysis--namely, cellular protein database development and protein microsequencing methodologies. The purpose of the workshop was to explore means of integrating the rapidly expanding body of information on 2-DE resolved proteins from different laboratories. A major proposal emanating from the workshop was the establishment of an intermediary or "relational" 2-DE gel protein database. This intermediary database, which would catalogue pertinent information on 2-DE resolved proteins (experimental source, 2-DE loci, biological information, etc.) could be an adjunct to, and accessed through, the existing international protein sequence databanks. It would function as a pointer for researchers to the individual 2-DE protein databases where primary and more specialized 2-DE data would be housed.

Databases, Factual

The human keratinocyte two-dimensional gel protein database (update 1992): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional gel database of human keratinocytes currently lists 2980 cellular proteins (2098 isoelectric focusing, IEF; and 882 nonequilibrium pH gradient electrophoresis, NEPHGE) many of which correspond to posttranslational modifications. About 20% of all recorded proteins have been identified (protein name, organelle components, etc.) and they are listed in alphabetical order together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Also, we have listed 145 microsequenced proteins that are recorded in this database. As an aid in localizing the polypeptides we have included blow-ups of the master images (IEF, NEPHGE) displaying all the protein numbers. In the long run, the master keratinocyte database is expected to link protein and DNA sequencing and mapping information (Human Genome Program) and to provide an integrated picture of the expression levels and properties of the thousands of proteins that orchestrate various keratinocyte functions both in health and disease.

Cell Differentiation

Metatranscriptomic analysis of viral sequences associated with Culex nigripalpus at an Alabama aquaculture site.

Mosquitoes associated with aquaculture habitats can harbor diverse viruses, yet the viromes of many locally abundant species remain poorly characterized. At an aquaculture-associated site in Auburn, Alabama, we surveyed mosquito populations and found Culex nigripalpus to be the dominant species collected. To characterize viruses associated with this mosquito, we performed RNA-seq on pooled female Cx. nigripalpus and compared complementary bioinformatic workflows for viral detection and genome recovery. One workflow removed host-associated reads by mapping to the closest available mosquito reference genome prior to assembly, whereas a second workflow used fully de novo assembly and viral database annotation. Additional protein-level filtering, cross-workflow comparison, and comparison of Trinity and rnaSPAdes assemblies were used to prioritize well-supported viral candidates. Across the original analyses, 16 submitted accessions corresponding to 12 collapsed virus/name groups were recovered, including Merida virus, Hubei mosquito virus 5, Zhejiang mosquito virus, Hubei virga-like virus 3, Rinkaby virus, Elemess virus, Qingnian mosquito virus, Serbia narna-like virus 2, XiangYun narna-levi-like virus 8, Ecclesville picorna-like virus, and baculovirus-like fragments. Several candidates were supported across multiple workflows, while others were recovered only under specific analytical conditions, indicating that candidate recovery was influenced by assembly and filtering choices. Selected viral contigs were independently supported by RT-PCR amplification. Overall, these results provide a first characterization of viral sequences associated with Cx. nigripalpus from an Alabama aquaculture-associated site and show that comparison across assembly and filtering strategies helped prioritize the most consistently supported viral candidates.

Animals

Microsequences of 145 proteins recorded in the two-dimensional gel protein database of normal human epidermal keratinocytes.

Microsequencing of proteins recovered from two-dimensional (2-D) gels is being used systematically to identify proteins in the master human keratinocyte 2-D gel database. To date, about 250 protein spots recorded in human 2-D gel databases have been microsequenced and, of these, 145 are recorded in the keratinocyte database under the entry partial amino acid sequence. Coomassie Brilliant Blue-stained protein spots cut from several (up to 40) dry gels were concentrated by elution-concentration gel electrophoresis, electroblotted onto PVDF membranes and digested in situ with trypsin. Eluting peptides were separated by reversed-phase HPLC, collected individually and sequenced. Computer search using the FASTA and TFASTA programs from Genetics Computer Group indicated that 110 of the microsequenced polypeptides shared significant similarity with proteins contained in the PIR, Mipsx or GenEMBL databases. Only 35 polypeptides corresponded to hitherto unknown proteins. Peptide sequences of all 145 proteins are listed together with their coordinates (apparent molecular weight and pI) in the keratinocyte database.

Amino Acid Sequence

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (> 77%), HLA-DMB*01:01 (> 63%), HLA-DOA*01:01 (> 97%) and HLA-DOB*01:01 (> 77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans

High-Resolution Chromosome-Level Genome Assembly and Annotation of Triplophysa stewarti, an Endemic Plateau Loach from the Qinghai-Tibet Plateau.

The bottom-dwelling fish Triplophysa stewarti, endemic to the Qinghai-Tibet Plateau, is a valuable model for studying high-altitude adaptation in aquatic ecosystems. However, the lack of a high-quality reference genome has hindered comparative genomic and evolutionary studies within this genus. Here, we present a chromosome-level genome assembly for T. stewarti, generated using PacBio HiFi long-read sequencing and Hi-C scaffolding. The 697.9 Mb assembly is highly continuous (scaffold N50 of 253.58 Mb) and encompasses 25 chromosomes, representing 92.65% of the genome. BUSCO analysis indicated a 98.4% completeness, supporting the high quality of the assembly. We annotated 28,009 protein-coding genes, with 97.04% being functionally assigned across multiple databases (NR, UniProt, KEGG, GO, Pfam and InterPro). Additionally, repetitive elements constituted 42.47% of the genome, and we identified 52,709 non-coding RNAs. This high-quality reference genome provides a fundamental resource for exploring the adaptive evolution, population structure, and conservation genetics of T. stewarti and related species on the Qinghai-Tibet Plateau.

Animals

Human cerebrospinal fluid protein database: edition 1992.

Two-dimensional electrophoresis maps of human cerebrospinal fluid proteins are presented in the form of labeled images. 931 protein spots are identified in spinal fluid from a normal volunteer. Distinct spots that represent variants of the same protein, especially posttranslational modifications, are estimated to reduce the 931 different spots to < 200 different proteins. 248 spots of 29 protein groups have been identified and are indicated on enlargements of specific gel regions. The distribution of protein abundance, mass, charge and shape characteristics of these normal 931 spinal fluid spots are graphically profiled. Analysis of the shape parameter "vertical height: width ratio" reveals that a ratio > 3.5 correlates with glycoproteins, enabling their identification simply by image analysis. Proteins that are not present on the normal map, but appear in spinal fluid in patients with schizophrenia and Creutzfeldt-Jakob disease are illustrated on additional maps.

Cerebrospinal Fluid Proteins

Environment-specific amino acid substitution tables: tertiary templates and prediction of protein folds.

The local environment of an amino acid in a folded protein determines the acceptability of mutations at that position. In order to characterize and quantify these structural constraints, we have made a comparative analysis of families of homologous proteins. Residues in each structure are classified according to amino acid type, secondary structure, accessibility of the side chain, and existence of hydrogen bonds from the side chains. Analysis of the pattern of observed substitutions as a function of local environment shows that there are distinct patterns, especially for buried polar residues. The substitution data tables are available on diskette with Protein Science. Given the fold of a protein, one is able to predict sequences compatible with the fold (profiles or templates) and potentially to discriminate between a correctly folded and misfolded protein. Conversely, analysis of residue variation across a family of aligned sequences in terms of substitution profiles can allow prediction of secondary structure or tertiary environment.

Amino Acid Sequence

FANTASIA suite: a reproducible and configurable framework for embedding-based functional annotation of proteins.

Embedding-based annotation transfer is increasingly used for protein function inference due to protein language models capture sequence, structural, and functional signals that may extend beyond conventional pairwise similarity. However, systematic application of these approaches requires control over model choice, reference composition, lookup parameters, evidence traceability, and output formats. We developed the FANTASIA suite, a configurable framework for embedding-based functional annotation of proteins. The suite combines a database-backed implementation for reproducible and extensible analyses with a portable flat-file implementation for rapid local annotation and pipeline integration. Using non-model and model-organism proteomes, we show that larger neighbourhood sizes remain practical for proteome-scale analyses and that taxonomy and sequence-identity filtering support leakage-aware benchmarking. We also compare the supported models with baseline methods through external CAFA5 evaluation and provide practical guidance based on empirical evidence variables. FANTASIA provides a controlled, scalable, and reproducible framework for extending functional annotation across the rapidly expanding diversity of sequenced organisms.

Software

Network pharmacology and molecular docking to explore the active compounds and mechanisms of Jerusalem artichoke for treating diabetes.

The effective components and mechanism of Jerusalem artichokes (JAs) in lowering blood glucose were studied through network pharmacology and molecular docking. The active compounds of Jerusalem artichoke were obtained by referring to the literature, and the active compounds were screened. The targets were predicted by the SwissTargetPrediction database, and the disease targets were screened using GeneCard, Disgenet, and OMIM databases. The protein-protein interaction (PPI) network diagram was constructed using the STRING database, and the intersection target was analyzed by gene ontology (GO) biological function and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses using the David database. Finally, molecular docking was verified using AutoDockTools1.5.7 software. After screening, 412 gene targets, 476 disease targets, and 64 intersection targets were identified. The results of GO biological function analysis and KEGG pathway analysis showed that the technology was involved in multiple biological processes and regulatory pathways for hypoglycemia, such as the HIF-1, PI3K-Akt, and AMPK signaling pathways. Molecular docking results showed that Jasmonate, Liquiritigenin and Inulin of JAs had strong binding effects with PPARG and STAT3. JAs exert hypoglycemic effects through multi-component, multi-target and multi-pathway. In summary, this study investigated the hypoglycemic mechanism of JAs using network pharmacology and molecular interconnection technology, and concluded that JAs exert hypoglycemic effects through multiple components, targets, and pathways, which provides a theoretical basis for the study of JAs.

Molecular Docking Simulation

In Situ Hybridization and RT-PCR Detection of Nervous Necrosis Virus in Fourfinger Threadfin, Eleutheronema tetradactylum, in Taiwan.

Between April and July 2020, suspected outbreaks of nervous necrosis virus (NNV) infection were observed in fourfinger threadfin (Eleutheronema tetradactylum) fingerling hatcheries in Pingtung County, southern Taiwan. Affected fish exhibited spiral swimming behaviour and abdominal distension, resulting in mortality rates between 50% and 100%. Histopathological examination showed severe vacuolation in the brain and ocular tissues, with large oval and/or irregular basophilic cytoplasmic inclusion bodies in the brain. Phylogenetic analysis of the viral replicase (RNA1) and capsid protein (RNA2) genes revealed high nucleotide sequence identities among the isolates in this study, with sequence similarity rates of 96.9%-99% for RNA1 and 98.2%-99.0% for RNA2 compared to RGNNV reference strains available in the NCBI GenBank database. This is the first detection of betanodavirus in fourfinger threadfin in Taiwan, using RT-PCR and ISH. A positive correlation between elevated water temperatures and disease severity indicates the need for year-round surveillance and genomic analysis to clarify the epidemiology of FTNNV. The data suggest that infected eggs may facilitate the vertical transmission of Betanodavirus. Crucially, utilising virus-free broodstock, alongside routine health screening and environmental control, is essential for mitigating NNV risks in fourfinger threadfin aquaculture.

Animals

PEELing: an integrated and user-centric platform for spatially resolved proteomics data analysis.

SUMMARY: Molecular compartmentalization is vital for cellular physiology. Spatially resolved proteomics allows biologists to survey protein composition and dynamics with subcellular resolution. Here, we present PEELing, an integrated package and user-friendly web service for analyzing spatially resolved proteomics data. PEELing assesses data quality using curated or user-defined references, performs cutoff analysis to remove contaminants, connects to databases for functional annotation, and generates data visualizations-providing a streamlined and reproducible workflow to explore spatially resolved proteomics data. AVAILABILITY AND IMPLEMENTATION: PEELing and its tutorial are publicly available at https://peeling.janelia.org/ (Zenodo DOI: 10.5281/zenodo.15692517). A Python package of PEELing is available at https://github.com/JaneliaSciComp/peeling/ (Zenodo DOI: 10.5281/zenodo.15692434).

Proteomics

HI-FEVER: a Nextflow pipeline for the high-throughput discovery and annotation of endogenous viral elements.

SUMMARY: Endogenous viral elements (EVEs) offer valuable insights into virus and host evolution, but their detection remains computationally and biologically challenging. We present HI-FEVER, a user-friendly Nextflow pipeline for the discovery of EVEs in eukaryotic host genomes. HI-FEVER is highly parallelizable and customizable, ensuring computational efficiency while allowing researchers to fine-tune parameters to their specific needs. Its output provides a comprehensive analysis of discovered EVEs, including detailed annotations which can provide evolutionary insights. HI-FEVER scales seamlessly to handle millions of viral protein queries across multiple host genomes on both laptops and high-performance computing nodes. AVAILABILITY AND IMPLEMENTATION: The HI-FEVER source code is available on GitHub at https://github.com/PaleovirologyLab/hi-fever. Minimal reference databases, test datasets and benchmarking results are hosted on the Open Science Framework at https://osf.io/y357r. A detailed wiki is available at https://github.com/PaleovirologyLab/hi-fever/wiki, including usage instructions, parameter descriptions, and guidance on interpreting outputs. The pipeline includes a Pixi environment compatible with Conda and Apptainer containerization, and Docker images. HI-FEVER has been tested on Linux, Windows (via WSL2), and macOS (Intel and ARM64).

Software