PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference protein database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

69 records · Page 4Linked to original sources

Annotation matters: the effect of structural gene annotation on orthology inference.

MOTIVATION: In silico gene annotation, the process of identifying the genes present in a genome, remains a challenging task. As genome assemblies rapidly increase, the corresponding gene models and repertoires often fall short in quality. Despite advances in annotation methods, a lack of community standards means that most published gene annotations result from ad hoc pipelines. As a result, only a few species have nearly complete and accurate gene models. This annotation quality is thought to affect downstream analyses, including orthology inference, often the first step of comparative genomics studies. RESULTS: We show that different annotation methods yield markedly distinct orthology inferences. We compared orthology assignments of gene models obtained by four prominent protein-coding gene model sources: the NCBI Eukaryotic Genome Annotation Pipeline, the Ensembl Gene Annotation System, the UniProt Reference Proteomes, and Augustus 3.4 (an ab initio pipeline). We observe significant discrepancies between sources, namely in the proportion of orthologous genes per genome, the completeness of Hierarchical Orthologous Groups, and the accuracy and recall of the predicted orthologs on a standard orthology benchmark.

Molecular Sequence Annotation

ALADDIN: an integrated tool for computer-assisted molecular design and pharmacophore recognition from geometric, steric, and substructure searching of three-dimensional molecular structures.

ALADDIN is a computer program for the design or recognition of compounds that meet geometric, steric, and substructural criteria. ALADDIN searches a database of three-dimensional structures, marks atoms that meet substructural criteria, evaluates geometric criteria, and prepares a number of files that are input for molecular modification and coordinate generation as well as for molecular graphics. Properties calculated from the three-dimensional structure are described by either properties calculated from the molecule itself or from the molecule as compared to a reference molecule and associated surfaces. ALADDIN was used to design analogues to probe a bioactive conformation of a small molecule and a peptide, to test alternative superposition rules for receptor mapping of the D2 dopamine receptor, to recognize unexpected D2 dopamine agonist activity of existing compounds, and to design compounds to fit a binding site on a protein of known structure. We have found that series designed by ALADDIN show much more subtle variation in shape than do those designed by traditional methods and that compounds can be designed to be very close matches to the objective.

Binding Sites

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins

Characterisation of HIV-1 Gag Cytotoxic T-Lymphocyte Epitopes in the Southern African Region-A Systematic Review.

During early HIV-1 infection, robust Cytotoxic T-lymphocyte (CTL) responses are mostly targeted at immunodominant Gag p24 epitopes to reduce HIV-1 viraemia to a set-point. The aim of this study was to review the current body of knowledge on HIV-1 Gag CTL epitopes in the southern African region where subtype C is prevalent. Peer-reviewed records were obtained from three databases: PubMed Central, Web of Science Core Collection, and Scopus, using the following search terms: HIV subtype C Gag epitopes, and HIV clade C Gag epitopes. The search results were restricted to countries within the southern African region, and only data published in English and between the years 2000-2025 were considered for this review. The search from the three databases produced a total of 2103 peer-reviewed records, and 49 records were included in the review. The majority of studies (58.44%) were conducted in South Africa, followed by Botswana (15.58%), Zambia (10.39%), Malawi (7.79%), Zimbabwe (6.49%) and Angola (1.30%). There were no studies identified from other southern African countries. A total of 60 Gag CTL epitopes were identified, of which 17 (28.33%) were located within the matrix protein (p17), 33 (55.00%) within the capsid protein (p24), and 4 (6.67%) within the Gag polyprotein (p2p7p1p6). The commonly detected immunodominant epitopes were mostly located within the Gag p24 protein; and included TPQDLNTML (TL9, Gag p24 48-56) and TSTLQEQIGW (TW10, Gag p24 108-117) present at 16.00% and 13.3%, respectively. The proportion of HLA-A, B and C allotypes in this systematic review were 18%, 78%, and 4%, respectively. The more common HLA-B allotypes that restrict immunodominant Gag epitopes and facilitate better control of HIV-1 were HLA-B*57, -B*58:01, -B*42:01 and -B*81:01. This systematic review has provided important insights into the description of immunodominant Gag epitopes and HLA-I alleles that contribute to the control of HIV-1 viraemia in the southern African region. It has also exposed that some CTL epitopes identified in the southern African studies are not reported on the Los Alamos HIV database (LANL HIV database). This highlights a need to have this database updated with this information as it is used as a reference for epitopes. This review could provide insights into the design of an epitope-based HIV-1 vaccine that would also be effective in the southern African region.

Humans

Rebuilding flavodoxin from C alpha coordinates: a test study.

The tertiary structure of flavodoxin has been model built from only the X-ray crystallographic alpha-carbon coordinates. Main-chain atoms were generated from a dictionary of backbone structures. Side-chain conformations were initially set according to observed statistical distributions, clashes were resolved with reference to other knowledge-based parameters, and finally, energy minimization was applied. The RMSD of the model was 1.7 A across all atoms to the native structure. Regular secondary structural elements were modeled more accurately than other regions. About 40% of the chi 1 torsional angles were modeled correctly. Packing of side chains in the core was energetically stable but diverged significantly from the native structure in some regions. The modeling of protein structures is increasing in popularity but relatively few checks have been applied to determine the accuracy of the approach. In this work a variety of parameters have been examined. It was found that close contacts, and hydrogen-bonding patterns could identify poorly packed residues. These tests, however, did not indicate which residues had a conformation different from the native structure or how to move such residues to bring them into agreement. To assist in the modeling of interacting side chains a database of known interactions has been prepared.

Flavodoxin

A comprehensive phylogeny of mammalian PRNP gene reveals no influence of prion misfolding propensity on the evolution of this gene.

Prion diseases are invariably fatal neurodegenerative diseases that affect some mammalian species, including humans. These diseases are caused by the misfolding of the cellular prion protein (PrPC) into a pathologic isoform (PrPSc). The prion protein is highly conserved across mammals. However, some species present lower susceptibility to prion diseases than others. This behavior is likely explained by the resistance of these animal species' prion proteins to acquire a pathological conformation. Therefore, the tertiary structure and interspecific variations encoded in the primary structure determine a PrP proneness to misfolding. For this reason, we studied the PRNP gene from a phylogenetic perspective, potentially unveiling evolutionary events related to prion diseases. We generated a database of mammalian PRNP sequences and constructed phylogenetic trees based on nucleotide sequence variations. We aligned 1146 PRNP gene sequences from 901 different mammalian species and built a PRNP gene-based phylogenetic tree. Classical phylogenetic orders tend to maintain their clustering in the PRNP gene tree. Nonetheless, the few differences found may shed some light on potential evolutionary constraints posed by prion disorders. Moreover, this phylogenetic study was combined with an in vitro misfolding study. Protein Misfolding Shaking Amplification (PMSA) was used to evaluate the tendency of many of these proteins to misfold. This comprehensive analysis spanned a wide range of mammalian prion protein sequences and included analysis of different variants with a focus on the human rs1799990 locus (c.385A > G, p.Met129Val). This variant, widely linked to prion disease susceptibility in humans, is explored in the context of its evolutionary origins. All in all, our PRNP gene-based tree, despite showing some topological differences with the reference species tree that could be in some cases related to prion disease susceptibility, is not significantly distinct. Indicating that the proneness of a PrP variant to misfold spontaneously has not shaped the evolution of this gene.

Phylogeny

The conformations and electrostatic potential maps of phorbol esters, teleocidins and ingenols.

Phorbol esters and the structurally dissimilar teleocidins and ingenols bind to and activate protein kinase C (PKC) during the course of tumour promotion. These compounds are referred to as TPA-like tumour promoters (from 12-O-tetradencanoyl phorbol-13-acetate, the most active of the class) and are amongst the most potent tumour promoters known. Despite their structural dissimilarity, all three groups of molecules have been shown to bind to the diacylglycerol site of PKC with high affinity. It is thought that this binding to and consequent activation of PKC is the crucial step in tumour promotion by these compounds. The aim of this work was to provide a description of the binding site by comparing structural features (in particular the electrostatic potential) with the activity of numerous derivatives of the three classes. Initially the description was obtained by consideration of the phorbol derivatives, and then refined using the teleocidins and ingenols. The activity data were collected from a variety of sources and the structures calculated using the semi-empirical MNDO approximation embodied in the MOPAC program. Where possible, the crystal structure was obtained from the Cambridge Crystallographic Database, and used as a starting point for the calculation. In other cases, a preliminary calculation was carried out using the molecular mechanics program AMBER. Electrostatic potentials were calculated and displayed using an in-house program 3D2, while superpositions of molecules were carried out using CHEM-X.

Diterpenes

cDNA cloning of the B cell membrane protein CD22: a mediator of B-B cell interactions.

We have cloned a full-length cDNA for the B cell membrane protein CD22, which is referred to as B lymphocyte cell adhesion molecule (BL-CAM). Using subtractive hybridization techniques, several B lymphocyte-specific cDNAs were isolated. Northern blot analysis with one of the clones, clone 66, revealed expression in normal activated B cells and a variety of B cell lines, but not in normal activated T cells, T cell lines, Hela cells, or several tissues, including brain and placenta. One major transcript of approximately 3.3 kb was found in B cells although several smaller transcripts were also present in low amounts (approximately 2.6, 2.3, and 1.6 kb). Sequence analysis of a full-length cDNA clone revealed an open reading frame of 2,541 bases coding for a predicted protein of 847 amino acids with a molecular mass of 95 kD. The BL-CAM cDNA is nearly identical to a recently isolated cDNA clone for CD22, with the exception of an additional 531 bases in the coding region of BL-CAM. BL-CAM has a predicted transmembrane spanning region and a 140-amino acid intracytoplasmic domain. Search of the National Biological Research Foundation protein database revealed that this protein is a member of the immunoglobulin super family and that it had significant homology with three homotypic cell adhesion proteins: carcinoembryonic antigen (29% identity over 460 amino acids), myelin-associated glycoprotein (27% identity over 425 amino acids), and neural cell adhesion molecule (21.5% over 274 amino acids). Northern blot analysis revealed low-level BL-CAM mRNA expression in unactivated tonsillar B cells, which was rapidly increased after B cell activation with Staphylococcus aureus Cowan strain 1 and phorbol myristate acetate, but not by various cytokines, including interleukin 4 (IL-4), IL-6, and gamma interferon. In situ hybridization with an antisense BL-CAM RNA probe revealed expression in B cell-rich areas in tonsil and lymph node, although the most striking hybridization was in the germinal centers. COS cells transfected with a BL-CAM expression vector were immunofluorescently stained positively with two different CD22 antibodies, each of which recognizes a different epitope. Additionally, both normal tonsil B cells and a B cell line were found to adhere to COS transfected with BL-CAM in the sense but not the antisense direction.(ABSTRACT TRUNCATED AT 400 WORDS)

Adolescent

Precision parameters of methods of analysis required for nutrition labeling. Part I. Major nutrients.

Major components of foods and feeds are fat, protein, and carbohydrates. Fat and protein are determined by direct measurements that are interpreted as the quantity of the constituent. Carbohydrates are usually calculated by difference. For this calculation, values for moisture/solids, ash, and "fiber" are also needed. The readily available collaborative studies for the determination of these major components are reviewed in an attempt to assign precision parameters to validated methods of analysis. When a number of studies for the same analyte, in the same food, by the same method are available, it is seen that the precision parameters among laboratories (standard deviations, SR; relative standard deviations, RSDR) and the ISO maximum tolerable difference functions (repeatability value, r; reproducibility value, R) are not characterized by any conventional distribution. The precision data are best summarized as a median or average parameter and the interval containing the centermost 90% of reported values. Typically, the precision of methods of analysis can be expressed as a function of concentration only, independent of analyte, matrix, and method. The average RSDR value from each collaborative data set can then be used as the numerator in a ratio containing, as the denominator, the value calculated from the Horwitz equation: RSDR = 2 exp (1 - 0.5 log C) where C is the concentration as a decimal fraction. A series of ratios consistently above 1, and especially above 2, probably indicates that a method is unacceptable with respect to precision. By this criterion, only the protein (Kjeldahl) determination is unqualifiedly acceptable with a 90% interval for RSDR of 1 to 3% at C values above about 0.01 (1 g/100 g). Fat, moisture/solids, and ash are acceptable down to limiting concentrations in the region of 1 to 5 g/100 g, if a test portion large enough to provide at least 50 mg of weighable residue or volatiles is specified. Measurements of individual carbohydrates and fiber-related analytes have unexpectedly poor precisions among laboratories. The variability, although high, may still be suitable for nutrition labeling. Reliability of analyses for the control of labeling of the primary nutrients must be achieved through quality assurance programs that require strict adherence to the directions of empirical methods and the use of suitable reference materials for absolute methods.

Databases, Bibliographic

Molecular cloning and characterization of iojap (ij), a pattern striping gene of maize.

Iojap (ij) is a recessive striped mutant of maize affecting the development of plastids in a local and position-dependent manner on the leaves. The ij-affected plastids are transmitted to some of the progeny even when the function of the nuclear gene is restored. Developmental defects during embryogenesis and leaf proliferation are other phenotypic characteristics of ij. The extent of striping and the degree of developmental arrest in ij depend upon genetic background. To understand the diverse and unique phenotypic expression of ij, a transposon tagging experiment has been conducted using Robertson's Mutator (Mu). A new ij mutant was obtained from crosses of the reference allele of (ij-ref) to Mu lines. Subsequent genetic and molecular studies showed that the mutant carried a new ij allele (ij-mum1) from the Mu lines and contained a Mu1 element that cosegregated with the iojap phenotype. A 6.0 kb EcoRI genomic DNA fragment containing the Mu1 element was cloned. ij-ref is unstable, and revertants (Ij-Rev) have been obtained. Using the flanking DNA from the genomic clone as a probe, DNA polymorphisms were detected between ij-ref and these revertants. Further, transcripts were restored to the normal level in Ij-Rev seedlings. Comparison of genomic DNA clones from ij-ref, ij-mum1 and Ij indicated that the ij-ref allele contained 1.5 kb of additional DNA related to a transposable element, Ds. Germinal and somatic revertant alleles were derived by excision of this 1.5 kb element from ij-ref. The structure of the Ij gene and the DNA sequence of its transcribed region were determined. The Ij gene encodes a 24.8 kDa protein that showed no significant sequence similarity with proteins listed in databases.

Alleles

Analysis of the molecular mechanism underlying di(2-ethylhexyl) phthalate-induced bladder carcinogenesis via network toxicology and molecular docking approaches: An observational study.

This study aims to investigate the toxicity of di(2-ethylhexyl) phthalate (DEHP) and the potential molecular mechanisms of DEHP-induced bladder cancer (BLCA) using network toxicology and molecular docking strategies. The toxicity of DEHP was assessed using Prox-II software, and potential targets for DEHP-induced BLCA were identified by integrating data from ChEMBL database, Search Tool for Interactions of Chemicals, SwissTargetPrediction, GeneCards, Therapeutic Target Database, Online Mendelian Inheritance in Man, and The Cancer Genome Atlas. STRING database and Cytoscape were employed to construct target networks and determine core targets. The expression levels of core targets were analyzed using R. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathway enrichment analyses were performed on potential and core targets. Molecular docking was carried out using CB-Dock 2 to verify the interactions between DEHP and core targets. A total of 105 potential targets related to DEHP-induced BLCA were identified, from which 7 core targets were selected: cyclin-dependent kinase 1, interleukin 6, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, cyclin B2, and B-cell lymphoma 2. IL-6 and B-cell lymphoma 2 showed downregulated expression in tumor tissues, while cyclin-dependent kinase 1, cyclin-dependent kinase 2, cyclin B1, Erb-B2 receptor tyrosine kinase 2, and cyclin B2 were upregulated. Gene Ontology and Kyoto Encyclopedia of Genes and Genomes enrichment analyses indicated that these targets were enriched in cell signaling and cancer-related pathways. Molecular docking confirmed that DEHP interacts with these core targets. DEHP may promote the development of BLCA by interacting with key proteins and signaling pathways. This study provides a theoretical basis for understanding the molecular mechanisms of DEHP-induced BLCA and offers references for future prevention and treatment strategies.

Diethylhexyl Phthalate

Quantile distributions of amino acid usage in protein classes.

A comparative study of the compositional properties of various protein sets from both cellular and viral organisms is presented. Invariants and contrasts of amino acid usages have been discerned for different protein function classes and for different species using robust statistical methods based on quantile distributions and stochastic ordering relationships. In addition, a quantitative criterion to assess amino acid compositional extremes relative to a reference protein set is proposed and applied. Invariants of amino acid usage relate mainly to the central range of quantile distributions, whereas contrasts occur mainly in the tails of the distributions, especially contrasts between eukaryote and prokaryote species. Influences from genomic constraint are evident, for example, in the arginine:lysine ratios and the usage frequencies of residues encoded by G + C-rich versus A + T-rich codon types. The structurally similar amino acids, glutamate versus aspartate and phenylalanine versus tyrosine, show stochastic dominance relationships for most species protein sets favoring glutamate and phenylalanine respectively. The quantile distribution of hydrophobic amino acid usages in prokaryote data dominates the corresponding quantile distribution in human data. In contrast, glutamate, cysteine, proline and serine usages in human proteins dominate the corresponding quantile distributions in Escherichia coli. E. coli dominates human in the use of basic residues, but no dominance ordering applies to acidic residues. The discussion centers on commonalities and anomalies of the amino acid compositional spectrum in relation to species, function, cellular localization, biochemical and steric attributes, complexity of the amino acid biosynthetic pathway, amino acid relative abundances and founder effects.

Amino Acids

Large-scale functional annotation establishes a reference framework for human LRRK2 variants.

Pathogenic variants in leucine-rich repeat kinase 2 (LRRK2)1are among the most frequent monogenic causes of Parkinson's disease (PD)2 and act through a gain-of-function mechanism of increased kinase activity. LRRK2-targeted therapies are in clinical development, but interpretation of the rapidly expanding catalogue of rare LRRK2 variants remains a barrier to translation. Here, we present functionally annotated data on >350 LRRK2 coding variants using a standardized cellular assay with Rab10 phosphorylation as a readout of kinase activity and integrated these data with curated genetic and clinical annotations from the Movement Disorders Society Genetic Mutation Database (MDSGene). Variants differed in activation magnitude, ranging from modest increases (e.g., p.G2019S) to strongly activating substitutions such as p.Y1699C or p.L1795F. Activating variants occurred across the full length of LRRK2, although the largest effects clustered within the ROC-COR regulatory hub, where structural analysis identified subdomains forming an allosteric scaffold controlling kinase output. All known/established pathogenic variants showed increased activity, whereas benign and likely benign variants remained within the wild-type range. Functional effect sizes correlated with pathway activation in patient-derived immune cells, altogether providing a framework for ACMG-based variant interpretation in which kinase activation can support PS3 functional evidence for reclassification of variants.

Protein phosphorylation

Comparison on Major Gene Mutations Related to Rifampicin and Isoniazid Resistance between Beijing and Non-Beijing Strains of Mycobacterium tuberculosis: A Systematic Review and Bayesian Meta-Analysis.

Objective: The Beijing strain of Mycobacterium tuberculosis (MTB) is controversially presented as the predominant genotype and is more drug resistant to rifampicin and isoniazid compared to the non-Beijing strain. We aimed to compare the major gene mutations related to rifampicin and isoniazid drug resistance between Beijing and non-Beijing genotypes, and to extract the best evidence using the evidence-based methods for improving the service of TB control programs based on genetics of MTB. Method: Literature was searched in Google Scholar, PubMed and CNKI Database. Data analysis was conducted in R software. The conventional and Bayesian random-effects models were employed for meta-analysis, combining the examinations of publication bias and sensitivity. Results: Of the 8785 strains in the pooled studies, 5225 were identified as Beijing strains and 3560 as non-Beijing strains. The maximum and minimum strain sizes were 876 and 55, respectively. The mutations prevalence of rpoB, katG, inhA and oxyR-ahpC in Beijing strains was 52.40% (2738/5225), 57.88% (2781/4805), 12.75% (454/3562) and 6.26% (108/1724), respectively, and that in non-Beijing strains was 26.12% (930/3560), 28.65% (834/2911), 10.67% (157/1472) and 7.21% (33/458), separately. The pooled posterior value of OR for the mutations of rpoB was 2.72 ((95% confidence interval (CI): 1.90, 3.94) times higher in Beijing than in non-Beijing strains. That value for katG was 3.22 (95% CI: 2.12, 4.90) times. The estimate for inhA was 1.41 (95% CI: 0.97, 2.08) times higher in the non-Beijing than in Beijing strains. That for oxyR-ahpC was 1.46 (95% CI: 0.87, 2.48) times. The principal patterns of the variants for the mutations of the four genes were rpoB S531L, katG S315T, inhA-15C > T and oxyR-ahpC intergenic region. Conclusion: The mutations in rpoB and katG genes in Beijing are significantly more common than that in non-Beijing strains of MTB. We do not have sufficient evidence to support that the prevalence of mutations of inhA and oxyR-ahpC is higher in non-Beijing than in Beijing strains, which provides a reference basis for clinical medication selection.

Isoniazid

[Genetic and phenotypic analysis of three children with Neurodevelopmental disorders due to variants of DEAF1 gene].

OBJECTIVE: To explore the genetic characteristics and clinical phenotypes of three children with novel DEAF1 gene variants. METHODS: Three children who were referred to Capital Children's Medical Center Affiliated to Capital Medical University between January 2018 and December 2025 were selected as study subjects and underwent whole exome sequencing (WES). Candidate variants were verified by Sanger sequencing, and their pathogenicity was evaluated based on the guidelines from American College of Medical Genetics and Genomics (ACMG). A systematic search of databases including PubMed and CNKI was conducted to compile previously reported cases of DEAF1 variants for clinical phenotype comparison. For the non-canonical splice site variant c.870+5G>C located in the intronic region, wild-type and mutant minigene reporter vectors were constructed and transfected into HeLa and 293T cells, respectively, and the splicing patterns were analyzed by RT-PCR and Sanger sequencing. This study was approved by the Medical Ethics Committee of Capital Institute of Pediatrics (Ethics No.: SHERLL 2020001). RESULTS: All three children were found to have carried de novo heterozygous variants of the DEAF1 gene, including two missense variants (c.764G>A, c.641T>C) in the important SAND domain and a splice site variant (c.870+5G>C) in a non-canonical splicing region. The c.764G>A and c.870+5G>C variants were unreported previously. All children had presented with intellectual developmental delay, and two were accompanied by autism spectrum disorder, and two had epilepsy and sleep disorders. In vitro minigene splicing assay showed that the c.870+5G>C variant can lead to abnormal splicing. CONCLUSIONS: This study reported three children with novel DEAF1 variants, two of which have not been previously described, thereby enriched the mutational spectrum of the DEAF1 gene. In vitro functional assay combined with the clinical manifestations of the patients confirmed the pathogenicity of the non-canonical splice site variant in the intronic region.

Humans