PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference protein database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics.

MOTIVATION: Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and contain complex patterns, making them inherently challenging to analyze with algorithmic approaches. In terms of the protein sequence landscape, most recent bottom-up MS-based proteomics studies consider either a diverse pool of post-translational modifications, employ large databases-as in metaproteomics or proteogenomics, study multiple isoforms of proteins, include unspecific cleavage sites or even combinations thereof. All this makes peptide and protein identifications challenging. RESULTS: Here, we present a foundation model, called yHydra, that jointly embeds spectra and peptides. This allows us to implement various downstream tasks and search modes in Euclidean space. We implement an open search which allows querying multiple ten-thousands of spectra against millions of peptides. Furthermore, we implement an error-tolerant search for identifying additional proteoforms that are not included in off-the-shelf reference proteomes. Our foundation model provides meaningful embeddings, as we interpret learned peptide embeddings in comparison to the peptide's physico-chemical properties. Hydra's open search, assigns delta masses to each identification which allows to unrestrictedly characterize post-translational modifications. The error-tolerant mode of yHydra can be used as post-processing to existing search engines or as a standalone. yHydra is evaluated on several real life data sets for the identification of modified peptide sequences and shows up to 25% increase in peptide identification at constant false discovery rate compared to the current state-of-the-art. AVAILABILITY AND IMPLEMENTATION: Code is available on Gitlab: https://gitlab.com/dacs-hpi/yHydra, and https://gitlab.com/dacs-hpi/yHydra_train.

Proteomics

Comparison of six microcomputer dietary analysis systems with the USDA Nutrient Data Base for Standard Reference.

We compared the general operating features and nutrient databases of six microcomputer dietary analysis systems. A 3-day food record with 73 food items was entered into each program; nutrient averages were compared with the US Department of Agriculture Nutrient Data Base for Standard Reference (USDA NDB), full version, release 9, for microcomputers. The six programs were found to vary widely in cost, number of foods and nutrients in the database, use of non-USDA data and imputation of data for missing values, number of print/export options, time to analyze the 3-day food record, and overall ease of use. Although all of the microcomputer dietary analysis systems were within 7% of the USDA NDB for energy, protein, total fat, and total carbohydrates, the proportion of other nutrients varying more than 15% from the USDA NDB varied considerably between programs. Variance among programs for 3-day food record nutrient values occurred because of differences in the number of food items included in the database (leading to varying degrees of substitution), the recency of the nutrient data (whether or not the most recent USDA releases had been incorporated), and the number of missing values (the degree to which non-USDA sources or estimated calculations were used to fill in the blanks from the USDA standard). Our results demonstrate that it is important for each dietitian to carefully choose a microcomputer dietary analysis system that is suitable to specific and predetermined needs.

Databases, Factual

Plasma protein map: an update by microsequencing.

The reference plasma protein map, obtained with immobilized pH gradients in the first dimension of two-dimensional electrophoresis, is presented. By microsequencing, more than 40 polypeptide chains were identified. The new polypeptides and previously known proteins are listed in a table and labeled on the protein map, thus providing an update of the human plasma two-dimensional gel database.

Amino Acid Sequence

The molecular mechanism of cuproptosis and research progress in pancreatic diseases.

PURPOSE: Cuproptosis has been proven to be a novel mode of cell death, distinct from other types of cell death such as necrosis, ferroptosis, pyroptosis, and apoptosis. This study aims to systematically review the molecular mechanisms of cuproptosis in recent years and its research progress in pancreatic diseases. METHODS: By searching PubMed and Web of Science databases, 113 key literatures were included for thematic analysis, covering the molecular mechanism of cuproptosis and its role in the occurrence and development of pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cyst, pancreatic injury and pancreatic neuroendocrine tumor. RESULTS: Cuproptosis refers to the accumulation of copper ions in cells, which leads to instability of ferritin and aggregation of acylated proteins, resulting in oxidative stress-related cell death. Recent studies have shown that cuproptosis plays an important role in the occurrence and development of various pancreatic diseases, such as pancreatic cancer, acute and chronic pancreatitis, diabetes, pancreatic cysts, pancreatic injuries and pancreatic neuroendocrine tumor. The inducers of cuproptosis, such as disulfiram, chloroquinolones, and perilla phenols, alleviate pancreatic cancer by promoting cell cuproptosis. Copper chelators such as tetraethylenepentamine and tetrathiomolybdate promote the recovery of pancreatic injury by inhibiting cell cuproptosis. CONCLUSIONS: Cuproptosis plays a crucial role in the pathogenesis of pancreatic diseases. Further research on the cuproptosis pathway may become a potential target for the treatment of pancreatic diseases.

Animals

A new human hypervariable locus (K29) maps to the q37.3 region of chromosome 2 and reveals a fingerprint.

A human genomic library was screened with a 30-base oligomer corresponding to the 5' end of the human calretinin cDNA. A clone that contains a minisatellite composed of 21 imperfect repeats of a 37-bp sequence was isolated. The consensus (GAGGGAGGAACTGGGACGCGTGCATGTTTGCATTCTC) incidentally shares 14 consecutive matches with the oligomer used as a probe, and it was shown that the clone did not belong to the calretinin locus. The minisatellite, named K29, was used as a probe on Southern blots at high stringency. After HaeIII, MboI, or HinfI digestion, it detected a single hypervariable locus, with 65% heterozygosity among Caucasian individuals. The probe used at low stringency revealed a fingerprint, with an average of four bands in addition to the locus-specific pattern. Mendelian inheritance was assessed on pedigrees. The K29 minisatellite was mapped by in situ hybridization to the very end of the long arm of chromosome 2 (2q37.3 band), at close proximity of the Fra2J locus, and is referred to as the D2S88 locus in the genome database.

Base Sequence

HDL particle associated proteins in plasma and cerebrospinal fluid: identification and partial sequencing.

The proteins from plasma HDL particles isolated by immunoaffinity chromatography on anti-apolipoprotein A-1 affinity columns have been analysed and purified by high resolution two dimensional gel electrophoresis. Two of the lipoprotein-associated proteins found in the HDL plasma fraction, previously referred to as NA1 and NA2, have also been found in cerebrospinal fluid. After separation by 2DGE, these two proteins were transferred to PVDF membranes, stained and cut out for N-terminal sequencing. The partial sequences (11 and 13 amino acids) obtained for the two HDL particle associated proteins do not match any of those included in the December 1987 National Biomedical Research Foundation (NBRF) database, and there are no significant sequence similarities.

Amino Acid Sequence

Towards establishing a protein database of Drosophila.

An improved method of high-resolution two-dimensional gel electrophoresis has been used to study the patterns of protein synthesis in wing imaginal discs of late instar larvae of Drosophila melanogaster. A total of one thousand and twenty five labelled polypeptides (787 acidic and 238 basic) have so far been separated and catalogued. For convenience, all these polypeptides have been numbered and their position fixed by its molecular weight and relative mobility. They are indicated on a reference protein map for further studies.

Animals

Genomes of the ex-type strains of Elsinoë mangiferae and E. perseae, the causal agents of scab on mango and avocado.

Elsinoë species are slow-growing, hemibiotrophic to necrotrophic fungi that cause scab diseases on economically important fruit crops. Genome resources for many host-specific species remain limited. We report high-quality draft genome assemblies for the ex-type strains of Elsinoë mangiferae (CBS 226.50) and E. perseae (CBS 406.34), causal agents of mango and avocado scab, respectively. Among 5 approaches tested, a Nanopore-only NextDenovo assembly produced the most contiguous genomes, yielding 24.5 Mb (E. mangiferae) and 25.1 Mb (E. perseae) assemblies with 13 and 18 contigs, respectively, BUSCO completeness scores of ∼94%, and multiple putative telomere-to-telomere chromosomes. Gene prediction identified 9,134 and 9,243 genes, respectively. Functional annotation revealed enrichment of metabolic and regulatory pathways, including those involved in posttranslational modification, protein transport, and secondary metabolism. Carbohydrate-active enzyme repertoires were small but conserved, consistent with stealth pathogenicity strategies and low plant cell wall degradation. Both genomes encoded large secretomes (>850 proteins), diverse protease repertoires (>300 proteins), Ecp2-like effector proteins, and multiple biosynthetic gene clusters, including clusters with similarity to those associated with elsinochrome and ACT-toxin II biosynthesis, some of which may contribute to host-pathogen interactions and disease development. A large fraction of genes lacked functional characterization, suggesting incomplete databases and/or the presence of lineage-specific genes potentially involved in virulence or host adaptation. These genome resources fill critical gaps for underrepresented Elsinoë species and provide taxonomically anchored references essential for diagnostics, comparative genomics, and research into the molecular basis of host specificity and pathogenicity in scab-causing fungi.

Persea

LIPIDAT: a database of lipid phase transition temperatures and enthalpy changes. DMPC data subset analysis.

The systematic study of the mesomorphic phase properties of synthetic and biologically derived lipids began some 30 years ago. In the past decade, interest in this area has grown enormously. As a result, there exists a wealth of information on lipid phase behavior, but unfortunately these data have until now been scattered throughout the literature in a variety of books, proceedings and journals. The data have recently been compiled in a centralized database, LIPIDAT, with a view to providing ready access to the data and to the appropriate literature. LIPIDAT consists of a tabulation of all known mesomorphic and polymorphic phase transition temperatures and enthalpy changes for synthetic and biologically-derived lipids in the dry and in the partially and fully hydrated states. Also included is the effect of pH, and of salt and metal ion concentration and other additives such as proteins, drugs, etc., on the thermodynamic values. The methods used in making the measurements and the experimental conditions are reported. Bibliographic information includes comprehensive literature referencing and list of authors, but does not at the present time include article titles. As of this writing, the database is current through June, 1990 and is approaching 10,000 records in length. Each record contains 28 fields. In this paper we report the contents and present an analysis of LIPIDAT as it refers to fully hydrated 1,2-dimyristoyl-sn-glycero-3-phosphocholine (DMPC). This database subset represents about 7% of all LIPIDAT records. It includes data collected over a 23-year period from 1967 to 1989 and consists of 702 records obtained from 336 articles in 55 different journals. The number of records per year rises steadily beginning in 1971, reaches a maximum of 89 records/year in 1977 and remains relatively constant at 60-70 records/year in the succeeding period. Journals making the greatest contribution to the DMPC subset include Biochimica et Biophysica Acta, Biochemistry, Chemistry and Physics of Lipids and the Biophysical Journal. These four journals account for 71% of the total records in the database subset. The analysis shows that differential scanning calorimetry, electron spin resonance, fluorescence, nuclear magnetic resonance and Raman spectroscopy are the methods most commonly used for DMPC transition temperature determination. An interesting pattern emerges as to the place in time the different methods assume or loose popularity.(ABSTRACT TRUNCATED AT 400 WORDS)

Calorimetry, Differential Scanning

Evaluation of the sequence template method for protein structure prediction. Discrimination of the (beta/alpha)8-barrel fold.

A multiple alignment of five (beta/alpha)8-barrel enzymes has been derived from their structure. The eight beta-strands and eight alpha-helices of the (beta/alpha)8-barrel are correctly aligned and the equivalenced residues in these regions fulfil similar structural roles. Each beta-strand has a central core of usually four residues, two residues contribute side-chains to the barrel core and the other two residues are involved in beta-strand/alpha-helix contacts. However, the fold imposes no constraints on the volumes of the residues at either a local or global level: the volume of the beta-barrel core varies between 1088 A3 in glycolate oxidase and 1571 A3 in taka-amylase. Sequence motifs derived from the multiple alignment were scanned against a database of 124 protein sequences, including 17 (beta/alpha)8-barrel enzymes. The results were evaluated in terms of the discrimination of (beta/alpha)8-barrel sequences and the quality of the alignments obtained. One motif was able to identify the top 12% of high scoring sequences as forming (beta/alpha)8-barrels with 50% accuracy and the bottom 50% of sequences as not being (beta/alpha)8-barrel proteins with 100% accuracy. However, in most instances the alignments were poor. The reasons for this are discussed with reference to the (beta/alpha)8-barrel proteins and the sequence motif method in general.

Alcohol Oxidoreductases

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology

A current genotoxicity database for heterocyclic thermic food mutagens. I. Genetically relevant endpoints.

Cooking, heat processing, or pyrolysis of protein-rich foods induce the formation of a series of structurally related heterocyclic aromatic bases that have been found to be mutagens. The primary genetic assay utilized to detect and isolate these mutagens has been the his reversion assay in Salmonella typhimurium. The classification and nomenclature of these chemicals is revised to reflect recent advances. The findings of short-term tests for genetic injury that have been applied to these agents are presented in a systematic way. Cell-free, bacterial, mammalian cell culture, and in vivo systems are included. Major results, the mutagens tested, and key references are presented in tabular form, with text commentary. Integrated conclusions on the state of current knowledge of the genetic toxicity of thermic food mutagens are presented. Areas in need of further research are defined. Finally, an outline is presented of a suggested path leading to the determination whether normal methods of food preparation and processing constitute a human health hazard.

Animals

Homologous sequences in steroidogenic enzymes, steroid receptors and a steroid binding protein suggest a consensus steroid-binding sequence.

The amino acid sequences of two steroidogenic enzymes, P450c17 (steroid 17 alpha-hydroxylase/17,20 lyse) and P450c21 (steroid 21-hydroxylase), are only 28.9% identical. However, these proteins share a region of 21 amino acids bearing 17 identical residues, which we previously suggested may represent the steroid binding site. We assembled a sequence database of known steroid-binding proteins and searched this with the sequence of this 21 amino acid region. The steroidogenic enzymes, P450c17, P450c21, P450scc (the cholesterol side-chain cleavage enzyme), and P450c11 (steroid 11 beta/18-hydroxylase) share a subregion of 17 amino acids having at least 15 identical residues. Related sequences were identified in a computerized search of the available sequences of steroid hormone receptors and binding proteins. These sequences were invariably found within larger domains previously associated with steroid binding. From these we propose a more general consensus sequence of LPLLL +/- 000KDRE0LKRL +/- PV, where +/- refers to any charged amino acid, and 0 refers to an uncharged amino acid. This consensus sequence predicts 147 or 187 total amino acids in 11 human proteins examined (78.6%). An equivalent degree of sequence identity, 178 of 221 amino acids (80.5%) was found among 13 animal homologs of these human proteins. The ability of this consensus sequence to predict 325 of 408 amino acids (79.7%) strongly suggests this sequence is necessary, if not sufficient, for a steroid binding site in many proteins. Lecithin-cholesterol acetyl transferase, cholesterol ester transfer protein, and steroid sulfatase did not have sequences similar to our consensus sequence.

Amino Acid Sequence

The chromosome-level genome assembly and annotation of the silver-lipped pearl oyster, Pinctada maxima.

The silver-lipped pearl oyster (Pinctada maxima) is a valuable tropical aquaculture species, playing a crucial economic role in the global pearl industry. However, the lack of genomic reference limits our in-depth understanding of this species in genome-based breeding, conservation, evolution and adaptation. Here, annotated chromosome-level reference genome for P. maxima was generated by integrating PacBio long-read sequencing, Illumina short-read sequencing, and Hi-C sequencing data. The total genome size is 1,264.93&#x2009;Mb, with contig N50 and scaffold N50 of 649&#x2009;kb and 89.19&#x2009;Mb, respectively. The majority (97.94%) of the assembled genome was anchored to the 14 chromosomes by Hi-C analysis. The relatively high genome completeness was observed, with 97.38% (metazoa_odb10 database) and 95.26% (mollusca_odb10 database) in BUSCO analysis. Genome annotation revealed approximately 65.46% of the repeat sequences and 26,315 protein-coding genes. Comparative genome analysis revealed 28 expanded and 48 contracted families (p&#x2009;<&#x2009;0.05) in P. maxima, with 3.2% of genes (894) being species-specific. This chromosome-level genome serves as an essential resource for research in evolutionary genomics, phylogenetics, and biomineralization.

Animals

Numerical classification of coding sequences.

DNA sequences coding for protein may be represented by counts of nucleotides or codons. A complete reading frame may be abbreviated by its base count, e.g. A76C158G121T74, or with the corresponding codon table, e.g. (AAA)0(AAC)1(AAG)9 ... (TTT)0. We propose that these numerical designations be used to augment current methods of sequence annotation. Because base counts and codon tables do not require revision as knowledge of function evolves, they are well-suited to act as cross-references, for example to identify redundant GenBank entries. These descriptors may be compared, in place of DNA sequences, to extract homologous genes from large databases. This approach permits rapid searching with good selectivity.

Animals

SCMO: a deep learning model integrating the single-cell resolution TME ecosystem and multi-omics for survival prediction in CRC patients.

BACKGROUND: Colorectal cancer (CRC) remains a leading cause of global cancer mortality, highlighting the need for precise survival prediction to guide clinical decisions. Although tissue-level multi-omics is widely utilized for survival prediction, its limited resolution cannot capture tumor heterogeneity. Single-cell RNA sequencing (scRNA-seq) enables dissection of the tumor microenvironment (TME) at cellular resolution, supporting personalized prognostic assessment. METHODS: We collected 213 CRC scRNA-seq samples and established a CRC-specific TME atlas comprising 339,060 cells. Using this atlas as a reference, we deconvolved bulk RNA-seq data from TCGA-CRC cohort with the EcoTyper algorithm to reconstruct TME features. Clinical, genomic, and transcriptomic data were obtained from the Xena platform; microbial data were sourced from the BIC database. We integrated TME and multi-omics features through a self-normalizing neural network to construct a deep learning model (single-cell resolution TME ecosystem with multi-omics data [SCMO]) for survival prediction. To enhance interpretability, we utilized the Integrated Gradients algorithm and spatial transcriptomic data to analyze multi-omics and TME features. We performed anticancer drug screening with tumor necrosis factor receptor-associated protein 1 (TRAP1), a critical feature according to the Integrated Gradients algorithm, as a potential target. RESULTS: We identified 13 survival-related TME features from the CRC-specific atlas: 12 cell states and one multi-cellular ecosystem. SCMO, which combined TME and multi-omics features, improved survival prediction and outperformed existing methods, achieving a concordance index of 0.762. The SCMO demonstrated robust performance for long-term predictions, achieving areas under the curve (AUCs) of 0.752, 0.772, and 0.869 for 1-, 3-, and 5-year predictions in the training set, with corresponding test set AUCs of 0.639, 0.756, and 0.772. TME features from the SCMO model revealed that ecosystem density increased with CRC malignancy. Multi-omics features included TRAP1 as a potential drug target. Drug screening identified saikosaponin A as a novel TRAP1 inhibitor, and its anticancer activity was validated in vitro. We developed SCMO-Lite, a simplified model incorporating 12 high-attribution-weight multi-omics features, which demonstrated robust risk stratification. CONCLUSIONS: SCMO combines analytical precision with biological interpretability, offering novel insights for oncology survival prediction.

Humans

The REF52 protein database. Methods of database construction and analysis using the QUEST system and characterizations of protein patterns from proliferating and quiescent REF52 cells.

The construction and analysis of protein databases using the QUEST system is described, and the REF52 protein database is presented. A protein database provides the means to store and compare quantitative and descriptive data for up to 2000 proteins from many experiments that employ computer-analyzed two-dimensional gel electrophoresis. The QUEST system provides the tools to manage, analyze, and communicate these data. The REF52 database contains experiments with normal and transformed rat cell lines. In this report, many of the proteins on the REF52 map are identified by name, by subcellular localization, and by mode of post-translational modification. The quantitative experiments analyzed and compared here include 1) a study of the quantitative reproducibility of the analysis system, 2) a study of the clonal reproducibility of REF52 cells, 3) a study of growth-related changes in REF52 cells, and 4) a study of the effects of labeling cells for varying lengths of time. Of the proteins analyzed from REF52 cells, 10% are nuclear, 6% are phosphoproteins, and 4% are mannose-labeled glycoproteins. The mannose-labeled proteins are more prominent in patterns from quiescent cells, while the synthesis of cytoskeletal proteins is generally repressed at quiescence. A small set of proteins, selected for elevated rates of synthesis is generally repressed at quiescence. A small set of proteins, selected for elevated rates of synthesis in quiescent versus proliferating cells includes one of the tropomyosin isoforms, a myosin light chain isoform, and several prominent glycoproteins. These proteins are thought to be characteristic of the differentiated state of untransformed REF52 cells. Proteins induced early versus late after refeeding quiescent cells show very different patterns of growth regulation. These studies lay the foundations of the REF52 database and provide information needed to interpret the experiments with transformed REF52 cells, which are reported in the accompanying paper (Garrels, J., and Franza, B. R., Jr. (1989) J. Biol. Chem. 264, 5299-5312).

Animals

GCRP: Integrated Global Chicken Reference Panel from 11,951 Chicken Genomes.

Chickens are a crucial source of protein for humans and a popular model animal for bird research. Despite the emergence of imputation as a reliable genotyping strategy for large populations, the lack of a high-quality chicken reference panel has hindered progress in chicken genome research. To address this, here we introduce the first phase of the 100K Global Chicken Reference Panel (100K GCRP). Currently, two panels are available: a comprehensive mix panel (CMP) for domestication diversity research and a commercial breed panel (CBP) for breeding broilers specifically. Evaluation of genotype imputation quality showed that CMP had the highest imputation accuracy compared to imputation using existing chicken panels in Animal-SNPAtlas and Animal Genotype Imputation Database (AGIDB), whereas CBP performed stably in the imputation of commercial populations. Additionally, we found that genome-wide association studies using GCRP-imputed data, whether on simulated or real phenotypes, exhibited greater statistical power. In conclusion, our study indicates that the GCRP effectively fills the gap in high-quality reference panels for chickens, providing an effective imputation platform for future genetic and breeding research. The project includes 11,951 samples and provides services for various applications on its website at http://farmrefpanel.com/GCRP/#/.

Animals