PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Protein”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Cluster of fibronectin type III repeats found in the human major histocompatibility complex class III region shows the highest homology with the repeats in an extracellular matrix protein, tenascin.

Walking and sequencing a genome portion centromeric of CYP21B in the human MHC class III region disclosed a cluster of fibronectin type III repeats in an approximately 50-kb DNA segment. Fibronectin type III repeats are known to consist of ca. 90 amino acid residues and exist in a wide range of protein species. Homology searches in protein databases showed that the repeats found had the highest homology with the repeats of human tenascin, an extracellular matrix protein. One cDNA sequence located immediately centromeric of CYP21B, the 3' portion of which is transcribed by the opposite strand of CYP21B, was found also to have six type III repeats followed by a fibrinogen domain. Pairwise homology comparison of these repeats in the MHC locus with those of human tenascin showed a general parallelism in their gene organization, indicating that the newly found repeats are elements of certain tenascin-like genea.

Amino Acid Sequence↗

MPtopo: A database of membrane protein topology.

The reliability of the transmembrane (TM) sequence assignments for membrane proteins (MPs) in standard sequence databases is uncertain because the vast majority are based on hydropathy plots. A database of MPs with dependable assignments is necessary for developing new computational tools for the prediction of MP structure. We have therefore created MPtopo, a database of MPs whose topologies have been verified experimentally by means of crystallography, gene fusion, and other methods. Tests using MPtopo strongly validated four existing MP topology-prediction algorithms. MPtopo is freely available over the internet and can be queried by means of an SQL-based search engine.

Algorithms↗

Protein sequence similarity searches using patterns as seeds.

Protein families often are characterized by conserved sequence patterns or motifs. A researcher frequently wishes to evaluate the significance of a specific pattern within a protein, or to exploit knowledge of known motifs to aid the recognition of greatly diverged but homologous family members. To assist in these efforts, the pattern-hit initiated BLAST (PHI-BLAST) program described here takes as input both a protein sequence and a pattern of interest that it contains. PHI-BLAST searches a protein database for other instances of the input pattern, and uses those found as seeds for the construction of local alignments to the query sequence. The random distribution of PHI-BLAST alignment scores is studied analytically and empirically. In many instances, the program is able to detect statistically significant similarity between homologous proteins that are not recognizably related using traditional single-pass database search methods. PHI-BLAST is applied to the analysis of CED4-like cell death regulators, HS90-type ATPase domains, archaeal tRNA nucleotidyltransferases and archaeal homologs of DnaG-type DNA primases.

Adenosine Triphosphatases↗

A comparison of position-specific score matrices based on sequence and structure alignments.

Sequence comparison methods based on position-specific score matrices (PSSMs) have proven a useful tool for recognition of the divergent members of a protein family and for annotation of functional sites. Here we investigate one of the factors that affects overall performance of PSSMs in a PSI-BLAST search, the algorithm used to construct the seed alignment upon which the PSSM is based. We compare PSSMs based on alignments constructed by global sequence similarity (ClustalW and ClustalW-pairwise), local sequence similarity (BLAST), and local structure similarity (VAST). To assess performance with respect to identification of conserved functional or structural sites, we examine the accuracy of the three-dimensional molecular models predicted by PSSM-sequence alignments. Using the known structures of those sequences as the standard of truth, we find that model accuracy varies with the algorithm used for seed alignment construction in the pattern local-structure (VAST) > local-sequence (BLAST) > global-sequence (ClustalW). Using structural similarity of query and database proteins as the standard of truth, we find that PSSM recognition sensitivity depends primarily on the diversity of the sequences included in the alignment, with an optimum around 30-50% average pairwise identity. We discuss these observations, and suggest a strategy for constructing seed alignments that optimize PSSM-sequence alignment accuracy and recognition sensitivity.

Algorithms↗

Sequence complexity of disordered protein.

Intrinsic disorder refers to segments or to whole proteins that fail to self-fold into fixed 3D structure, with such disorder sometimes existing in the native state. Here we report data on the relationships among intrinsic disorder, sequence complexity as measured by Shannon's entropy, and amino acid composition. Intrinsic disorder identified in protein crystal structures, and by nuclear magnetic resonance, circular dichroism, and prediction from amino acid sequence, all exhibit similar complexity distributions that are shifted to lower values compared to, but significantly overlapping with, the distribution for ordered proteins. Compared to sequences from ordered proteins, these variously characterized intrinsically disordered segments and proteins, and also a collection of low-complexity sequences, typically have obviously higher levels of protein-specific subsets of the following amino acids: R, K, E, P, and S, and lower levels of subsets of the following: C, W, Y, I, and V. The Swiss Protein database of sequences exhibits significantly higher amounts of both low-complexity and predicted-to-be-disordered segments as compared to a non-redundant set of sequences from the Protein Data Bank, providing additional data that nature is richer in disordered and low-complexity segments compared to the commonness of these features in the set of structurally characterized proteins.

Artificial Intelligence↗

Improved sensitivity proteomics by postharvest alkylation and radioactive labelling of proteins.

We describe approaches to improve the detection of proteins by postharvest alkylation and subsequent radioactive labeling with either [3H]iodoacetamide or 125I. Database protein sequence analysis suggested that cysteine is not suitable for detection of the entire proteome, but that cysteine alkylating reagents can increase the number of proteins able to be detected by iodination chemistry. Proteins were alkylated with beta-(4-hydroxyphenyl)ethyl iodoacetamide, or with 1,5-l-AEDANS (the Hudson Weber reagent). Subsequent iodination using the Iodo-Gen system was found to be most efficient. The enhanced sensitivity obtainable by using these approaches is expected to be sufficient for visualization of the lowest copy number proteins from human cells, such as from clinical samples. However, we argue that significantly improved methods of protein separation will be necessary to resolve the large number of proteins expected to be detectable with this sensitivity.

Acetamides↗

Two-dimensional electrophoretic analysis of human breast carcinoma proteins: mapping of proteins that bind to the SH3 domain of mixed lineage kinase MLK2.

MLK2, a member of the mixed lineage kinase (MLK) family of protein kinases, first reported by Dorow et al. (Eur. J. Biochem. 1993, 213, 701-710), comprises several distinct structural domains including an src homology-3 (SH3) domain, a kinase catalytic domain, a unique domain containing two leucine zipper motifs, a polybasic sequence, and a cdc42/rac interactive binding motif. Each of these domains has been shown in other systems to be associated with a specific type of protein interaction in the regulation of cellular signal transduction. To study the role of MLK2 in recruiting specific substrates, we constructed a recombinant cDNA encoding the N-terminal 100 amino acids of MLK2 (MLK2N), including the SH3 domain (residues 23-77), fused to glutathione S-transferase. This fusion protein was expressed in Escherichia coli, purified using gluthathione-Sepharose affinity chromatography and employed in an affinity approach to isolate MLK2-SH3 domain binding proteins from lysates of 35S-labelled MDA-MB231 human breast tumour cells. Electrophoretic analysis of bound proteins revealed that two low-abundance proteins with a molecular weights (Mr) of approximately 31,500 and approximately 34,000, bound consistently to the MLK2N protein. To establish accurately the Mt / isoelectric point (pI) loci of these MLK2-SH3 domain binding proteins, a number of abundant proteins in a two-dimensional electrophoresis (2-DE) master gel were identified to serve as triangulation marker points. Proteins were identified by (i) direct Edman degradation following electroblotting onto polyvinylidene difluoride (PVDF) membranes, (ii) Edman degradation of peptides generated by in-gel proteolysis and fractionation by rapid (approximately 12 min) microbore column (2.1 mm ID) reversed-phase high performance liquid chromatography (HPLC), (iii) mass spectrometric methods including peptide-mass fingerprinting and electrospray (ESI)-mass spectrometry (MS)-MS utilizing capillary (0.2-0.3 mm ID) column chromatography, or (iv) immunoblot analysis. Using this information, a preliminary 2-DE protein database for the human breast carcinoma cell line MDA-MB231, comprising 21 identified proteins, has been constructed and can be accessed via the World Wide Web (URL address: http:(/)/ www.ludwig.edu.au/www/jpsl/jpslhome.htm l).

Amino Acid Sequence↗

Changes in protein expression during melanoma differentiation determined by computer analysis of 2-D gels.

Cytodifferentiation in many melanocytic cells is regulated through the adenylate cyclase-cAMP pathway. To analyse the molecular changes associated with this process we have compared the proteins produced by two closely related cell lines which, though derived from a single cell line, respond very differently to modulation of this signalling pathway. The human melanoma cell line DX3 shows little change in in vitro characteristics following treatment with cAMP elevating agents; in contrast the more malignant DX3 LT5.1 variant, derived from the DX3 parental line, shows pronounced dendrification, decreased proliferation and a reduction in metastatic capacity after similar treatment. The two cell lines were treated with phosphodiesterase inhibitors for 5 days and then processed for two-dimensional gel characterization using an immobilized pH gradient for the IEF dimension. Proteins were detected by silver staining the gels and protein intensities were digitized using a laser densitometer. Two-dimensional gel patterns were edited, matched and a melanoma protein database of 637 spots constructed using PDQUEST software on an Orion 1/05 computer. Eleven proteins were lost and four new proteins were detected in both cell lines following treatment. Twenty-two proteins were present in DX3 LT5.1 after treatment but not in untreated lines or treated DX3. These differentially expressed proteins may be associated with the observed changes in differentiation patterns and metastasis. Our results illustrate the resolving power of this technique and suggest potential applications to the study of cellular differentiation.

Animals↗

Complete nucleotide sequence of a Selenomonas ruminantium plasmid and definition of a region necessary for its replication in Escherichia coli.

A plasmid from Selenomonas ruminantium subspecies lactilytica has been subcloned in Escherichia coli K-12 and completely sequenced. Three open reading frames (ORFs) of 909, 801, and 549 bp were identified and the complete sequence was analyzed by comparison with DNA and protein databases. No significant deoxynucleotide or amino acid sequence homology with other published genes or proteins was detected. The plasmid was shown to replicate independently in E. coli K-12 by a DNA polymerase I-dependent mechanism and deletion analysis defined the DNA sequence responsible for this phenotype.

Amino Acid Sequence↗

Functional analysis of alcS, a gene of the alc cluster in Aspergillus nidulans.

The ethanol utilization pathway (alc system) of Aspergillus nidulans requires two structural genes, alcA and aldA, which encode the two enzymes (alcohol dehydrogenase and aldehyde dehydrogenase, respectively) allowing conversion of ethanol into acetate via acetyldehyde, and a regulatory gene, alcR, encoding the pathway-specific autoregulated transcriptional activator. The alcR and alcA genes are clustered with three other genes that are also positively regulated by alcR, although they are dispensable for growth on ethanol. In this study, we characterized alcS, the most abundantly transcribed of these three genes. alcS is strictly co-regulated with alcA, and encodes a 262-amino acid protein. Sequence comparison with protein databases detected a putative conserved domain that is characteristic of the novel GPR1/FUN34/YaaH membrane protein family. It was shown that the AlcS protein is located in the plasma membrane. Deletion or overexpression of alcS did not result in any obvious phenotype. In particular, AlcS does not appear to be essential for the transport of ethanol, acetaldehyde or acetate. Basic Local Alignment Search Tool analysis against the A. nidulans genome led to the identification of two novel ethanol- and ethylacetate-induced genes encoding other members of the GPR1/FUN34/YaaH family, AN5226 and AN8390.

Alcohol Dehydrogenase↗

A new modular protein of Cryptosporidium parvum, with ricin B and LCCL domains, expressed in the sporozoite invasive stage.

The recombinant SA35 peptide has been described as an antigenic portion of a larger Cryptosporidium parvum protein. We identified and characterized the encoding Cpa135 gene and the entire protein, Cpa135. The Cpa135 gene was found to consist of a single exon of 4671 bp, and the mRNA transcribed in the sporozoites was identified. The predicted 1556 amino-acid protein showed the presence of domains which are widely conserved also in other unrelated phylogenetic groups (i.e. a ricin B and a LCCL motif). Comparison of Cpa135 sequence with genomic and protein databases revealed many related genes in other apicomplexan species and high homology with CCP2 protein from Plasmodium yoelii and Plasmodium berghei. The Cpa135 protein was identified and localized by using a monoclonal antibody (Mab) directed against the SA35 antigen (anti-SA35). In oocyst-sporozoite lysate, the anti-SA35 MAb recognized a 135 kDa protein that forms a protein complex larger than 200 kDa, which is mediated by disulfide bridges. Cpa135 synthesis was up-regulated during the excystation process. After host-cell invasion, Cpa135 gene expression was undetectable up to 48 h, whereas mRNA synthesis was newly observed at 72 h post-infection. The Cpa135 protein was localized in the apical complex, and it was found to be secreted by sporozoites during their gliding. Cpa135 persisted during the intracellular stages of the parasite, and it defined the boundaries of the parasitophorous vacuole in the infected cells. The unique array of domains and the homology with other apicomplexan proteins indicate that the Cpa135 protein is representative of a new family of proteins.

Amino Acid Motifs↗

The distribution of bmpB, a gene encoding a 29.7 kDa lipoprotein with homology to MetQ, in Brachyspira hyodysenteriae and related species.

The distribution of the bmpB gene encoding BmpB, a 29.7 kDa outer membrane lipoprotein of the intestinal spirochaete Brachyspira hyodysenteriae, was investigated. Using PCR, the gene was detected in all the 48 strains of B. hyodysenteriae examined and in Brachyspira innocens strain B256T, but not in 11 other strains of B. innocens nor in 42 strains of other Brachyspira spp. The gene was sequenced from B. innocens strain B256T and from 11 strains of B. hyodysenteriae. The B. hyodysenteriae genes shared 97.9-100% nucleotide sequence similarity and had 97.5-99.5% similarity with the gene of B. innocens strain B256T. Southern hybridisation indicated that bmpB was present on a 1.9 kb HindIII fragment of the B. hyodysenteriae genome and on a 3.1 kb fragment of the B. innocens B256T genome. The B. innocens lipoprotein did not react in Western blots with monoclonal antibody BJL/SH1 that reacts with the B. hyodysenteriae lipoprotein. The difference in binding with the monoclonal antibody may reside in the replacement of a serine residue with a tyrosine residue at base position 210 in the lipoprotein from B. innocens B256T. Comparison of the BmpB amino acid sequence with sequences in the SWISS-PROT protein database indicated that it has 33.9-39.9% similarity with the d-methionine binding proteins (MetQ) of a number of pathogenic bacterial species. The bmpB gene was confirmed to be the same as a gene of B. hyodysenteriae that was recently designated "blpA".

Amino Acid Sequence↗

Preliminary profile of the Cryptosporidium parvum genome: an expressed sequence tag and genome survey sequence analysis.

Cryptosporidium parvum is a protozoan enteropathogen that infects humans and animals and causes a pronounced diarrheal disease that can be life-threatening in immunocompromised hosts. No specific chemo- or immunotherapies exist to treat cryptosporidiosis and little molecular information is available to guide development of such therapies. To accelerate gene discovery and identify genes encoding potential drug and vaccine targets we constructed sporozoite cDNA and genomic DNA sequencing libraries from the Iowa isolate of C. parvum and determined approximately 2000 sequence tags by single-pass sequencing of random clones. Together, the 567 expressed sequence tags (ESTs) and 1507 genome survey sequences (GSSs) totaled one megabase (1 mb) of unique genomic sequence indicating that approximately 10% of the 10.4 mb C. parvum genome has been sequence tagged in this gene discovery expedition. The tags were used to search the public nucleic acid and protein databases via BLAST analyses, and 180 ESTs (32%) and 277 GSSs (18%) exhibited similarity with database sequences at smallest sum probabilities P(N)< or =10(-8). Some tags encoded proteins with clear therapeutic potential including S-adenosylhomocysteine hydrolase, histone deacetylase, polyketide/fatty-acid synthases, various cyclophilins, thrombospondin-related cysteine-rich protein and ATP-binding-cassette transporters. Several anonymous ESTs encoded proteins predicted to contain signal peptides or multiple transmembrane spanning segments suggesting they were destined for membrane-bound compartments, the cell surface or extracellular secretion. One-hundred four simple sequence repeats were identified within the nonredundant sequence tag collection with (TAA)(> or =6)/(TTA)(> or =6) and (TA)(> or = 10)/(AT)(> or =10 ) being the most prevalent, occurring 40 and 15 times, respectively. Various cellular RNAs and their genes were also identified including the small and large ribosomal RNAs, five tRNAs, the U2 small nuclear RNA, and the small and large virus-like, double-stranded RNAs. This investigation has demonstrated that survey sequencing is an efficient procedure for gene discovery and genome characterization and has identified and sequence tagged many C. parvum genes encoding potential therapeutic targets.

Amino Acid Sequence↗

Location proteomics: a systems approach to subcellular location.

Systems Biology requires comprehensive systematic data on all aspects and levels of biological organization and function. In addition to information on the sequence, structure, activities and binding interactions of all biological macromolecules, the creation of accurate predictive models of cell behaviour will require detailed information on the distribution of those molecules within cells and the ways in which those distributions change over the cell cycle and in response to mutations or external stimuli. Current information on subcellular location in protein databases is limited to unstructured text descriptions or sets of terms assigned by human curators. These entries do not permit basic operations that are common to other biological databases, such as measurement of the degree of similarity between the distributions of two proteins, and they are not able to fully capture the complexity of protein patterns that can be observed. The field of location proteomics seeks to provide automated, objective high-resolution descriptions of protein location patterns within cells. Methods have been developed to group proteins into statistically indistinguishable location patterns using automated analysis of fluorescence microscope images. The resulting clusters, or location families, are analogous to clusters found for other domains, such as protein sequence families. Preliminary work suggests the feasibility of expressing each unique pattern as a generative model that can be incorporated into comprehensive models of cell behaviour.

Animals↗

Osmostress response in Bacillus subtilis: characterization of a proline uptake system (OpuE) regulated by high osmolarity and the alternative transcription factor sigma B.

Exogenously provided proline has been shown to serve as an osmoprotectant in Bacillus subtilis. Uptake of proline is under osmotic control and functions independently of the known transport systems for the osmoprotectant glycine betaine. We cloned the structural gene (opuE) for this proline transport system and constructed a chromosomal opuE mutant by marker replacement. The resulting B. subtilis strain was entirely deficient in osmoregulated proline transport activity and was no longer protected by exogenously provided proline, attesting to the central importance of OpuE for proline uptake in high-osmolarity environments. The transport characteristics and growth properties of the opuE mutant revealed the presence of a second proline transport activity in B. subtilis. DNA sequence analysis of the opuE region showed that the OpuE transporter (492 residues) consists of a single integral membrane protein. Database searches indicated that OpuE is a member of the sodium/solute symporter family, comprising proteins from both prokaryotes and eukaryotes that obligatorily couple substrate uptake to Na+ symport. The highest similarity was detected to the PutP proline permeases, which are used in Escherichia coli, Salmonella typhimurium and Staphylococcus aureus for the acquisition of proline as a carbon and nitrogen source, but not for osmoprotective purposes. An elevation of the osmolarity of the growth medium by either ionic or non-ionic osmolytes resulted in a strong increase in the OpuE-mediated proline uptake. This osmoregulated proline transport activity was entirely dependent on de novo protein synthesis, suggesting a transcriptional control mechanism. Primer extension analysis revealed the presence of two osmoregulated and tightly spaced opuE promoters. The activity of one of these promoters was dependent on sigma A and the second promoter was controlled by the general stress transcription factor sigma B.

Amino Acid Sequence↗

Sequence determinants of amyloid fibril formation.

The establishment of rules that link sequence and amyloid feature is critical for our understanding of misfolding diseases. To this end, we have performed a saturation mutagenesis analysis on the de novo-designed amyloid peptide STVIIE (1). The positional scanning mutagenesis has revealed that there is a position dependence on mutation of amyloid fibril formation and that both very tolerant and restrictive positions to mutation can be found within an amyloid sequence. In this system, mutations that accelerate beta-sheet polymerization do not always lead to an increase of amyloid products. On the contrary, abundant fibrils are typically found for mutants that polymerize slowly. From these experiments, we have extracted a sequence pattern to identify amyloidogenic stretches in proteins. The pattern has been validated experimentally. In silico sequence scanning of amyloid proteins also supports the pattern. Analysis of protein databases has shown that highly amyloidogenic sequences matching the pattern are less frequent in proteins than innocuous amino acid combinations and that, if present, they are surrounded by amino acids that disrupt their aggregating capability (amyloid breakers). This study provides the potential for a proteome-wide scanning to detect fibril-forming regions in proteins, from which molecules can be designed to prevent and/or disrupt this process.

Amino Acid Motifs↗

Sisyphus and prediction of protein structure.

The problem of predicting protein structure from the sequence remains fundamentally unsolved despite more than three decades of intensive research effort. However, new and promising methods in three-dimensional (3D), 2D and 1D prediction have reopened the field. Mean-force-potentials derived from the protein databases can distinguish between correct and incorrect models (3D). Inter-residue contacts (2D) can be detected by analysis of correlated mutations, albeit with low accuracy. Secondary structure, solvent accessibility and transmembrane helices (1D) can be predicted with significantly improved accuracy using multiple sequence alignments. Some of these new prediction methods have proven accurate and reliable enough to be useful in genome analysis, and in experimental structure determination. Moreover, the new generation of theoretical methods is increasingly influencing experiments in molecular biology.

Computers↗

Eimeria refractile body proteins contain two potentially functional characteristics: transhydrogenase and carbohydrate transport.

cDNA encoding an immunogenic protein from partially sporulated oocysts of Eimeria acervulina was cloned and used to search for the homologous counterpart in Eimeria tenella. Monospecific antibodies were raised against the recombinant expression product. Using these antibodies, the parasite proteins were found to be localised in the refractile bodies. The derived amino-acid sequences were compared by computer using the SWISSPROT protein database. In addition to high homology between the Eimeria species, extensive similarity was found with pyridine nucleotide transhydrogenase from Escherichia coli. Comparison with the sugar signature database also resulted in a possible sugar binding domain present only in the Eimeria proteins. It is possible that the corresponding parasite proteins play a role in the recently discovered mannitol cycle of Eimeria.

Amino Acid Sequence↗