PubMed Health⌕ Search

Biomedical subjects

Leszek Rychlewski

Publications and source records attributed to Leszek Rychlewski.

At least 19 recordsLinked to original sources

Support-vector-machine classification of linear functional motifs in proteins.

Our algorithm predicts short linear functional motifs in proteins using only sequence information. Statistical models for short linear functional motifs in proteins are built using the database of short sequence fragments taken from proteins in the current release of the Swiss-Prot database. Those segments are confirmed by experiments to have single-residue post-translational modification. The sensitivities of the classification for various types of short linear motifs are in the range of 70%. The query protein sequence is dissected into short overlapping fragments. All segments are represented as vectors. Each vector is then classified by a machine learning algorithm (Support Vector Machine) as potentially modifiable or not. The resulting list of plausible post-translational sites in the query protein is returned to the user. We also present a study of the human protein kinase C family as a biological application of our method.

Databases, Genetic↗

Fluorescent Cell Chip a new in vitro approach for immunotoxicity screening.

The Fluorescent Cell Chip (FCC) has been developed specifically for immunotoxicity screening of chemical compounds. This in vitro test is based on a panel of genetically modified reporter cell lines that regulate the expression of fluorescent protein in the same way as they regulate expression of cytokines. Thus, changes in fluorescence intensity represent changes in cytokine expression. Consequently, this technique conforms to efficiency expected from high throughput screening assay. In a pre-validation effort we analyzed 46 compounds. The experimental protocol employed five reporter cell lines derived from murine EL-4 T cells. Reporter cells were exposed to tested chemicals on a 96 well plate and analyzed for EGFP-mediated fluorescence using automated flow cytometric assay. Tested compounds reproducibly generated compound-specific patterns of changes in fluorescence that allows for the hierarchical clustering of their expected activities based on pattern similarity analysis. Resultant classification revealed correlation with available in vivo immunotoxicity data. In conclusion, FCC is a new promising approach for in vitro screening of chemicals for their immunotoxicity.

Animals↗

Molecular modeling of phosphorylation sites in proteins using a database of local structure segments.

A new bioinformatics tool for molecular modeling of the local structure around phosphorylation sites in proteins has been developed. Our method is based on a library of short sequence and structure motifs. The basic structural elements to be predicted are local structure segments (LSSs). This enables us to avoid the problem of non-exact local description of structures, caused by either diversity in the structural context, or uncertainties in prediction methods. We have developed a library of LSSs and a profile--profile-matching algorithm that predicts local structures of proteins from their sequence information. Our fragment library prediction method is publicly available on a server (FRAGlib), at http://ffas.ljcrf.edu/Servers/frag.html . The algorithm has been applied successfully to the characterization of local structure around phosphorylation sites in proteins. Our computational predictions of sequence and structure preferences around phosphorylated residues have been confirmed by phosphorylation experiments for PKA and PKC kinases. The quality of predictions has been evaluated with several independent statistical tests. We have observed a significant improvement in the accuracy of predictions by incorporating structural information into the description of the neighborhood of the phosphorylated site. Our results strongly suggest that sequence information ought to be supplemented with additional structural context information (predicted with our segment similarity method) for more successful predictions of phosphorylation sites in proteins.

Amino Acid Sequence↗

FFAS03: a server for profile--profile sequence alignments.

The FFAS03 server provides a web interface to the third generation of the profile-profile alignment and fold-recognition algorithm of fold and function assignment system (FFAS) [L. Rychlewski, L. Jaroszewski, W. Li and A. Godzik (2000), Protein Sci., 9, 232-241]. Profile-profile algorithms use information present in sequences of homologous proteins to amplify the patterns defining the family. As a result, they enable detection of remote homologies beyond the reach of other methods. FFAS, initially developed in 2000, is consistently one of the best ranked fold prediction methods in the CAFASP and LiveBench competitions. It is also used by several fold-recognition consensus methods and meta-servers. The FFAS03 server accepts a user supplied protein sequence and automatically generates a profile, which is then compared with several sets of sequence profiles of proteins from PDB, COG, PFAM and SCOP. The profile databases used by the server are automatically updated with the latest structural and sequence information. The server provides access to the alignment analysis, multiple alignment, and comparative modeling tools. Access to the server is open for both academic and commercial researchers. The FFAS03 server is available at http://ffas.burnham.org.

Algorithms↗

Identification of novel restriction endonuclease-like fold families among hypothetical proteins.

Restriction endonucleases and other nucleic acid cleaving enzymes form a large and extremely diverse superfamily that display little sequence similarity despite retaining a common core fold responsible for cleavage. The lack of significant sequence similarity between protein families makes homology inference a challenging task and hinders new family identification with traditional sequence-based approaches. Using the consensus fold recognition method Meta-BASIC that combines sequence profiles with predicted protein secondary structure, we identify nine new restriction endonuclease-like fold families among previously uncharacterized proteins and predict these proteins to cleave nucleic acid substrates. Application of transitive searches combined with gene neighborhood analysis allow us to confidently link these unknown families to a number of known restriction endonuclease-like structures and thus assign folds to the uncharacterized proteins. Finally, our method identifies a novel restriction endonuclease-like domain in the C-terminus of RecC that is not detected with structure-based searches of the existing PDB database.

Amino Acid Sequence↗

Protein domain of unknown function DUF1023 is an alpha/beta hydrolase.

Pfam family DUF1023 consists entirely of uncharacterized proteins generated by sequencing the genomes of Actinobacteria (Bateman A., et al., Nucleic Acids Res. 2004;32 Database issue:D138-141.) Utilizing sequence similarity detection methods, we infer homology between DUF1023 and alpha/beta hydrolases. DUF1023 proteins conserve the core secondary structures in alpha/beta hydrolase fold, and share similar catalytic machinery as that of alpha/beta hydrolases. We predict DUF1023 spatial structure and deduce that they function as hydrolases utilizing catalytic Ser-His-Asp triad with the serine as a nucleophile.

Animals↗

Practical lessons from protein structure prediction.

Despite recent efforts to develop automated protein structure determination protocols, structural genomics projects are slow in generating fold assignments for complete proteomes, and spatial structures remain unknown for many protein families. Alternative cheap and fast methods to assign folds using prediction algorithms continue to provide valuable structural information for many proteins. The development of high-quality prediction methods has been boosted in the last years by objective community-wide assessment experiments. This paper gives an overview of the currently available practical approaches to protein structure prediction capable of generating accurate fold assignment. Recent advances in assessment of the prediction quality are also discussed.

Algorithms↗

AutoMotif server: prediction of single residue post-translational modifications in proteins.

UNLABELLED: The AutoMotif Server allows for identification of post-translational modification (PTM) sites in proteins based only on local sequence information. The local sequence preferences of short segments around PTM residues are described here as linear functional motifs (LFMs). Sequence models for all types of PTMs are trained by support vector machine on short-sequence fragments of proteins in the current release of Swiss-Prot database (phosphorylation by various protein kinases, sulfation, acetylation, methylation, amidation, etc.). The accuracy of the identification is estimated using the standard leave-one-out procedure. The sensitivities for all types of short LFMs are in the range of 70%. AVAILABILITY: The AutoMotif Server is available free for academic use at http://automotif.bioinfo.pl/

Algorithms↗

LiveBench-8: the large-scale, continuous assessment of automated protein structure prediction.

We present the results of the evaluation of the latest LiveBench-8 experiment. These results provide a snapshot view of the state of the art in automated protein structure prediction, just before the 2004 CAFASP-4/CASP-6 experiments begin. The last CAFASP/CASP experiments demonstrated that automated meta-predictors entail a significant advance in the field, already challenging most human expert predictors. LiveBench-8 corroborates the superior performance of meta-predictors, which are able to produce useful predictions for over one-half of the test targets. More importantly, LiveBench-8 identifies a handful of recently developed autonomous (nonmeta) servers that perform at the very top, suggesting that further progress in the individual methods has recently been obtained.

Automation↗

A support vector machine approach to the identification of phosphorylation sites.

We describe a bioinformatics tool that can be used to predict the position of phosphorylation sites in proteins based only on sequence information. The method uses the support vector machine (SVM) statistical learning theory. The statistical models for phosphorylation by various types of kinases are built using a dataset of short (9-amino acid long) sequence fragments. The sequence segments are dissected around post-translationally modified sites of proteins that are on the current release of the Swiss-Prot database, and that were experimentally confirmed to be phosphorylated by any kinase. We represent them as vectors in a multidimensional abstract space of short sequence fragments. The prediction method is as follows. First, a given query protein sequence is dissected into overlapping short segments. All the fragments are then projected into the multidimensional space of sequence fragments via a collection of different representations. Those points are classified with pre-built statistical models (the SVM method with linear, polynomial and radial kernel functions) either as phosphorylated or inactive ones. The resulting list of plausible sites for phosphorylation by various types of kinases in the query protein is returned to the user. The efficiency of the method for each type of phosphorylation is estimated using leave-one-out tests and presented here. The sensitivities of the models can reach over 70%, depending on the type of kinase. The additional information from profile representations of short sequence fragments helps in gaining a higher degree of accuracy in some phosphorylation types. The further development of an automatic phosphorylation site annotation predictor based on our algorithm should yield a significant improvement when using statistical algorithms in order to quantify the results.

Algorithms↗

Lead toxicity through the leadzyme.

Lead is one of the most dangerous toxic agents for all living organisms. In humans, elevated levels of lead have been linked to a number of disorders for which various molecular mechanisms have been proposed. However, none of them has been fully understood. It has also been known for several years that at micromolar concentrations lead can bind a unique RNA motif and catalyze a site-specific hydrolysis of the polyribonucleotide chain. This motif, called leadzyme, may be one of the major targets for lead within the cell, and it can cleave various cellular RNAs. A search of GenBank revealed the sequences that can potentially fold into the structure containing the leadzyme motif and that they are rather common in eukaryotic genomes. We found that the domain occurs with a high frequency in human mRNA sequences. Thus, the leadzyme nucleolytic properties should be considered as a possible mechanism for destruction of RNA within a cell. In particular, targeting of the RNA scaffold of ribosomes or spliceosomes may explain lead-mediated toxicity leading to cell death.

Humans↗

Structure prediction, evolution and ligand interaction of CHASE domain.

Cytokinins are plant hormones involved in the essential processes of plant growth and development. They bind with receptors known as CRE1/WOL/AHK4, AHK2, and AHK3, which possess histidine kinase activity. Recently, the sensor domain cyclases/histidine kinases associated sensory extracellular (CHASE) was identified in those proteins but little is known about its structure and interaction with ligands. Distant homology detection methods developed in our laboratory and molecular phylogeny enabled the prediction of the structure of the CHASE domain as similar to the photoactive yellow protein-like sensor domain. We have identified the active site pocket and amino acids that are involved in receptor-ligand interactions. We also show that fold evolution of cytokinin receptors is very important for a full understanding of the signal transduction mechanism in plants.

Amino Acid Sequence↗

Integrated web service for improving alignment quality based on segments comparison.

BACKGROUND: Defining blocks forming the global protein structure on the basis of local structural regularity is a very fruitful idea, extensively used in description, and prediction of structure from only sequence information. Over many years the secondary structure elements were used as available building blocks with great success. Specially prepared sets of possible structural motifs can be used to describe similarity between very distant, non-homologous proteins. The reason for utilizing the structural information in the description of proteins is straightforward. Structural comparison is able to detect approximately twice as many distant relationships as sequence comparison at the same error rate. RESULTS: Here we provide a new fragment library for Local Structure Segment (LSS) prediction called FRAGlib which is integrated with a previously described segment alignment algorithm SEA. A joined FRAGlib/SEA server provides easy access to both algorithms, allowing a one stop alignment service using a novel approach to protein sequence alignment based on a network matching approach. The FRAGlib used as secondary structure prediction achieves only 73% accuracy in Q3 measure, but when combined with the SEA alignment, it achieves a significant improvement in pairwise sequence alignment quality, as compared to previous SEA implementation and other public alignment algorithms. The FRAGlib algorithm takes approximately 2 min. to search over FRAGlib database for a typical query protein with 500 residues. The SEA service align two typical proteins within circa approximately 5 min. All supplementary materials (detailed results of all the benchmarks, the list of test proteins and the whole fragments library) are available for download on-line at http://ffas.ljcrf.edu/darman/results/. CONCLUSIONS: The joined FRAGlib/SEA server will be a valuable tool both for molecular biologists working on protein sequence analysis and for bioinformaticians developing computational methods of structure prediction and alignment of proteins.

Computational Biology↗

Detecting distant homology with Meta-BASIC.

Meta-BASIC (http://basic.bioinfo.pl) is a novel sensitive approach for recognition of distant similarity between proteins based on consensus alignments of meta profiles. Specifically, Meta-BASIC compares sequence profiles combined with predicted secondary structure by utilizing several scoring systems and alignment algorithms. In our benchmarking tests, Meta-BASIC outperforms many individual servers, including fold recognition servers, and it can compete with meta predictors that base their strength on the structural comparison of models. In addition, Meta-BASIC, which enables detection of very distant relationships even if the tertiary structure for the reference protein is not known, has a high-throughput capability. This new method is applied to 860 PfamA protein families with unknown function (DUF) and provides many novel structure-functional assignments available on-line at http://basic.bioinfo.pl/duf.pl. Detailed discussion is provided for two of the most interesting assignments. DUF271 and DUF431 are predicted to be a nucleotide-diphospho-sugar transferase and an alpha/beta-knot SAM-dependent RNA methyltransferase, respectively.

Algorithms↗

BOF: a novel family of bacterial OB-fold proteins.

Using top-of-the-line fold recognition methods, we assigned an oligonucleotide/oligosaccharide-binding (OB)-fold structure to a family of previously uncharacterized hypothetical proteins from several bacterial genomes. This novel family of bacterial OB-fold (BOF) proteins present in a number of pathogenic strains encompasses sequences of unknown function from DUF388 (in Pfam database) and COG3111. The BOF proteins can be linked evolutionarily to other members of the OB-fold nucleic acid-binding superfamily (anticodon-binding and single strand DNA-binding domains), although they probably lack nucleic acid-binding properties as implied by the analysis of the potential binding site. The presence of conserved N-terminal predicted signal peptide indicates that BOF family members localize in the periplasm where they may function to bind proteins, small molecules, or other typical OB-fold ligands. As hypothesized for the distantly related OB-fold containing bacterial enterotoxins, the loss of nucleotide-binding function and the rapid evolution of the BOF ligand-binding site may be associated with the presence of BOF proteins in mobile genetic elements and their potential role in bacterial pathogenicity.

Amino Acid Sequence↗

The PDB-Preview database: a repository of in-silico models of 'on-hold' PDB entries.

UNLABELLED: The PDB-Preview database is a dynamic web repository of in-silico predicted three-dimensional (3D) models of experimentally determined structures that are deposited into the PDB but are not yet publicly released, and are kept 'on-hold'. The PDB-Preview database is automatically generated on a weekly basis by the bioinfo.pl meta-server, which uses top-of-the-line fold-recognition methods. The PDB-Preview provides biologists with preliminary fold assignments well before the experimentally determined 3D structures are released. AVAILABILITY: http://bioinfo.pl/PDB-Preview/.

Computer Simulation↗

Protein structure prediction for the male-specific region of the human Y chromosome.

The complete sequence of the male-specific region of the human Y chromosome (MSY) has been determined recently; however, detailed characterization for many of its encoded proteins still remains to be done. We applied state-of-the-art protein structure prediction methods to all 27 distinct MSY-encoded proteins to provide better understanding of their biological functions and their mechanisms of action at the molecular level. The results of such large-scale structure-functional annotation provide a comprehensive view of the MSY proteome, shedding light on MSY-related processes. We found that, in total, at least 60 domains are encoded by 27 distinct MSY genes, of which 42 (70%) were reliably mapped to currently known structures. The most challenging predictions include the unexpected but confident 3D structure assignments for three domains identified here encoded by the USP9Y, UTY, and BPY2 genes. The domains with unknown 3D structures that are not predictable with currently available theoretical methods are established as primary targets for crystallographic or NMR studies. The data presented here set up the basis for additional scientific discoveries in human biology of the Y chromosome, which plays a fundamental role in sex determination.

Amino Acid Sequence↗