PubMed Health⌕ Search

Biomedical subjects

Michal Linial

Publications and source records attributed to Michal Linial.

At least 19 recordsLinked to original sources

Micropeptides encoded by lncRNAs associated with cancer progression reveal novel immunogenic epitopes.

MOTIVATION: Long non-coding RNAs (lncRNAs) regulate gene expression, chromatin organization, and cellular signaling. Recent studies indicate that ∼20% of the ∼36 000 human lncRNA genes harbor small open reading frames (sORFs) capable of producing micropeptides (MPs), whose functions remain largely unknown. Whether these peptides contribute to the cancer immunopeptidome is largely unexplored. RESULTS: We systematically analyzed lncRNAs with strong experimental and computational evidence of MP-encoding potential (∼13% of the initial MP collection). Using The Cancer Genome Atlas (TCGA), we identified 2606 high-confidence lncRNA-derived MPs encoded by 647 genes across 16 cancer types. We then focused on 501 MPs from 124 lncRNA genes whose expression changes significantly across tumor stages and metastatic transitions, representing cancer transitional lncRNAs (Tr-lncRNAs). Dipeptide composition and conservation analyses showed that these MPs differ from a size-matched human coding proteome, supporting their potential as neoantigens. All possible 9-mer peptides were evaluated for predicted binding to prevalent European HLA class I alleles. Approximately 60% of Tr-lncRNA genes and 184 (37%) of derived peptides exhibited strong predicted HLA binding. Peptides from XIST, PCAT7, PVT1, HAND2-AS1 showed broad HLA coverage. Notably, TTN-AS1, encoded an MP (79 aa) generated 33 predicted distinct epitopes spanning all 27 HLA alleles. Our analysis identifies lncRNA-derived MPs as a previously underexplored source of potential cancer neoantigens, highlighting their promise as biomarkers and targets for immunotherapy. AVAILABILITY: Data, code and supplementary materials are available in https://doi.org/10.5281/zenodo.20167452 and GitHub: https://github.com/stavzok1/lncrna_peptide_analysis.

Humans↗

On the state of protein function prediction: a report on the fourth CAFA challenge.

BACKGROUND: The Critical Assessment of Functional Annotation (CAFA) is a community effort held to understand the field of computational protein function prediction. Every three years, since 2010, the organizers initiate an experiment to collect function predictions on a large set of proteins and then evaluate the performance of predicting methods on a subset of proteins that have accumulated experimental annotations between the submission deadline and the evaluation time. CAFA provides an independent and rigorous assessment of the current state of the art, thus leveling the playing field, highlighting successes, revealing bottlenecks, and offering a forum for the exchange of ideas in protein science. Here, we report the results of the fourth CAFA experiment (CAFA4). RESULTS: CAFA4 featured the participation of 148 methods from 70 research groups on a total of 46,205 unique proteins over a 5-year annotation accumulation phase, the longest in any CAFA. In a comparison across CAFA2-CAFA4 methods, the prediction of Gene Ontology (GO) terms has clearly improved across all three GO aspects and traditional evaluation settings. While not achieving the first rank, several CAFA2 and CAFA3 methods featured in the top ten methods in many evaluations, suggesting that earlier methods still hold relevance. The performance is weaker in the newly introduced "partial knowledge" evaluation category (proteins with experimental annotations before submission deadline that gained additional annotations in the same GO aspect during the annotation accumulation phase), highlighting the need for a new class of methods. The rankings of the methods were stable over the years in traditional evaluation settings, but less so in the new partial knowledge evaluation. Overall, the field continues to progress with some influx of new participants. Sustained efforts will be necessary to substantially advance it.

Journal Article↗

EVEREST: a collection of evolutionary conserved protein domains.

Protein domains are subunits of proteins that recur throughout the protein world. There are many definitions attempting to capture the essence of a protein domain, and several systems that identify protein domains and classify them into families. EVEREST, recently described in Portugaly et al. (2006) BMC Bioinformatics, 7, 277, is one such system that performs the task automatically, using protein sequence alone. Herein we describe EVEREST release 2.0, consisting of 20,029 families, each defined by one or more HMMs. The current EVEREST database was constructed by scanning UniProt 8.1 and all PDB sequences (total over 3,000,000 sequences) with each of the EVEREST families. EVEREST annotates 64% of all sequences, and covers 59% of all residues. EVEREST is available at http://www.everest.cs.huji.ac.il/. The website provides annotations given by SCOP, CATH, Pfam A and EVEREST. It allows for browsing through the families of each of those sources, graphically visualizing the domain organization of the proteins in the family. The website also provides access to analyzes of relationships between domain families, within and across domain definition systems. Users can upload sequences for analysis by the set of EVEREST families. Finally an advanced search form allows querying for families matching criteria regarding novelty, phylogenetic composition and more.

Amino Acid Sequence↗

Synaptic proteins as multi-sensor devices of neurotransmission.

Neuronal communication is tightly regulated in time and space. Following neuronal activation, an electrical signal triggers neurotransmitter (NT) release at the active zone. The process starts by the signal reaching the synapse followed by a fusion of the synaptic vesicle (SV) and diffusion of the released NT in the synaptic cleft. The NT then binds to the appropriate receptor and induces a membrane potential change at the target cell membrane. The entire process is controlled by a fairly small set of synaptic proteins, collectively called SYCONs. The biochemical features of SYCONs underlie the properties of NT release. SYCONs are characterized by their ability to detect and respond to changes in environmental signals. For example, consider synaptotagmin I (Syt1), a prototype of a protein family with over 20 gene and variants in mammals. Syt1 is a specific example of a multi-sensor device with a large repertoire of discrete states. Several of these states are stimulated by a local concentration of signaling molecules such as Ca2+. The ability of this protein to sense signaling molecules and to adopt multiple biochemical states is shared by other SYCONs such as the synapsins (Syns). Specific biochemical states of Syns determine the accessibility of SV for NT release. Each of these states is defined by a specific alternative spliced variant with a unique profile of phosphorylation modified sites. The plasticity of the synapse is a direct reflection of SYCON's multiple biochemical states. State transitions occurs in a wide range of time scales, and therefore these molecules need to cope with events that last milliseconds (i.e., exocytosis in fast responding synapses) and with events that can carry on for many minutes (i.e., organization of SV pools). We suggest that SYCONs are optimized throughout evolution as multi-sensor devices. A full repertoire of the switches leading to alternation of protein states and a detailed characterization of protein-protein network within the synapse is critical for the development of a dynamic model of synaptic transmission.

Animals↗

ProtoBee: hierarchical classification and annotation of the honey bee proteome.

The recently sequenced genome of the honey bee (Apis mellifera) has produced 10,157 predicted protein sequences, calling for a computational effort to extract biological insights from them. We have applied an unsupervised hierarchical protein-clustering method, which was previously used in the ProtoNet system, to nearly 200,000 proteins consisting of the predicted honey bee proteins, the SWISS-PROT protein database, and the complete set of proteins of the mouse (Mus musculus) and the fruit fly (Drosophila melanogaster). The hierarchy produced by this method has been entitled ProtoBee. In ProtoBee, the proteins are hierarchically organized into 18,936 separate tree hierarchies, each representing a protein functional family. By using the mouse and Drosophila complete proteomes as reference, we are able to highlight functional groups of putative gene-loss events, putative novel proteins of unique functionality, and bee-specific paralogs. We have studied some of the ProtoBee findings and suggest their biological relevance. Examples include novel opsin genes and intriguing nuclear matches of mitochondrial genes. The organization of bee sequences into functional clusters suggests a natural way of automatically inferring functional annotation. Following this notion, we were able to assign functional annotation to about 70% of the sequences. ProtoBee is available at http://www.protobee.cs.huji.ac.il.

Animals↗

Apoptotic cell thrombospondin-1 and heparin-binding domain lead to dendritic-cell phagocytic and tolerizing states.

Apoptotic cells were shown to induce dendritic cell immune tolerance. We applied a proteomic approach to identify molecules that are secreted from apoptotic monocytes, and thus may mediate engulfment and immune suppression. Supernatants of monocytes undergoing apoptosis were collected and compared using sodium dodecyl sulfate-polyacrylamide gel electrophoresis (SDS-PAGE), and differentially expressed proteins were identified using tandem mass spectrometry. Thrombospondin-1 (TSP-1) and its cleaved 26-kDa heparin-binding domain (HBD) were identified. We show that TSP-1 is expressed upon induction of monocyte apoptosis in a caspase-dependent pattern and the HBD is cleaved by chymotrypsin-like serine protease. We further show that CD29, CD36, CD47, CD51, and CD91 simultaneously participate in engulfment induction and generation of an immature dendritic cell (iDC) tolerogenic and phagocytic state. We conclude that apoptotic cell TSP-1, and notably its HBD, creates a signalosome in iDCs to improve engulfment and to tolerate engulfed material prior to the interaction with apoptotic cells.

Antigens, CD↗

Novel unsupervised feature filtering of biological data.

MOTIVATION: Many methods have been developed for selecting small informative feature subsets in large noisy data. However, unsupervised methods are scarce. Examples are using the variance of data collected for each feature, or the projection of the feature on the first principal component. We propose a novel unsupervised criterion, based on SVD-entropy, selecting a feature according to its contribution to the entropy (CE) calculated on a leave-one-out basis. This can be implemented in four ways: simple ranking according to CE values (SR); forward selection by accumulating features according to which set produces highest entropy (FS1); forward selection by accumulating features through the choice of the best CE out of the remaining ones (FS2); backward elimination (BE) of features with the lowest CE. RESULTS: We apply our methods to different benchmarks. In each case we evaluate the success of clustering the data in the selected feature spaces, by measuring Jaccard scores with respect to known classifications. We demonstrate that feature filtering according to CE outperforms the variance method and gene-shaving. There are cases where the analysis, based on a small set of selected features, outperforms the best score reported when all information was used. Our method calls for an optimal size of the relevant feature set. This turns out to be just a few percents of the number of genes in the two Leukemia datasets that we have analyzed. Moreover, the most favored selected genes turn out to have significant GO enrichment in relevant cellular processes.

Artificial Intelligence↗

EVEREST: automatic identification and classification of protein domains in all protein sequences.

BACKGROUND: Proteins are comprised of one or several building blocks, known as domains. Such domains can be classified into families according to their evolutionary origin. Whereas sequencing technologies have advanced immensely in recent years, there are no matching computational methodologies for large-scale determination of protein domains and their boundaries. We provide and rigorously evaluate a novel set of domain families that is automatically generated from sequence data. Our domain family identification process, called EVEREST (EVolutionary Ensembles of REcurrent SegmenTs), begins by constructing a library of protein segments that emerge in an all vs. all pairwise sequence comparison. It then proceeds to cluster these segments into putative domain families. The selection of the best putative families is done using machine learning techniques. A statistical model is then created for each of the chosen families. This procedure is then iterated: the aforementioned statistical models are used to scan all protein sequences, to recreate a library of segments and to cluster them again. RESULTS: Processing the Swiss-Prot section of the UniProt Knoledgebase, release 7.2, EVEREST defines 20,230 domains, covering 85% of the amino acids of the Swiss-Prot database. EVEREST annotates 11,852 proteins (6% of the database) that are not annotated by Pfam A. In addition, in 43,086 proteins (20% of the database), EVEREST annotates a part of the protein that is not annotated by Pfam A. Performance tests show that EVEREST recovers 56% of Pfam A families and 63% of SCOP families with high accuracy, and suggests previously unknown domain families with at least 51% fidelity. EVEREST domains are often a combination of domains as defined by Pfam or SCOP and are frequently sub-domains of such domains. CONCLUSION: The EVEREST process and its output domain families provide an exhaustive and validated view of the protein domain world that is automatically generated from sequence data. The EVEREST library of domain families, accessible for browsing and download at 1, provides a complementary view to that provided by other existing libraries. Furthermore, since it is automatic, the EVEREST process is scalable and we will run it in the future on larger databases as well. The EVEREST source files are available for download from the EVEREST web site.

Cluster Analysis↗

Functional grouping based on signatures in protein termini.

The two ends of each protein are known as the amino (N-) and carboxyl (C-) termini. Short signatures in a protein's termini often carry vital cellular function. No systematic research has been conducted to address the importance of short signatures (3 to 10 amino acids) in protein termini at the proteomic level. Specifically, it is unknown whether such signatures are evolutionarily conserved, and if so, whether this conservation confers shared biological functions. Current signature detection methods fail to detect such short signatures due to inadequate statistical scores. The findings presented in this study strongly support the notion that functional significance of protein sets may be captured by short signatures at their termini. A positional search method was applied to over one million proteins from the UniProt database. The result is a collection of about a thousand significant signature groups (SIGs) that include previously identified as well as many novel signatures in protein termini. These SIGs represent protein sets with minimal or no overall sequence similarity excepting the similarity at their termini. The most significant SIGs are assigned by their strong correspondence to functional annotations derived from external databases such as Gene Ontology. Each of the SIGs is associated with the statistical significance of its functional association. These SIGs provide a valuable source for testing previously overlooked signatures in protein termini and allow for the investigation of the role played by such signatures throughout evolution. The SIGs archive and advanced search options are available at http://www.proteus.cs.huji.ac.il.

Amino Acid Sequence↗

Functional annotation prediction: all for one and one for all.

In an era of rapid genome sequencing and high-throughput technology, automatic function prediction for a novel sequence is of utter importance in bioinformatics. While automatic annotation methods based on local alignment searches can be simple and straightforward, they suffer from several drawbacks, including relatively low sensitivity and assignment of incorrect annotations that are not associated with the region of similarity. ProtoNet is a hierarchical organization of the protein sequences in the UniProt database. Although the hierarchy is constructed in an unsupervised automatic manner, it has been shown to be coherent with several biological data sources. We extend the ProtoNet system in order to assign functional annotations automatically. By leveraging on the scaffold of the hierarchical classification, the method is able to overcome some frequent annotation pitfalls.

Algorithms↗

The secrets of a functional synapse--from a computational and experimental viewpoint.

BACKGROUND: Neuronal communication is tightly regulated in time and in space. The neuronal transmission takes place in the nerve terminal, at a specialized structure called the synapse. Following neuronal activation, an electrical signal triggers neurotransmitter (NT) release at the active zone. The process starts by the signal reaching the synapse followed by a fusion of the synaptic vesicle and diffusion of the released NT in the synaptic cleft; the NT then binds to the appropriate receptor, and as a result, a potential change at the target cell membrane is induced. The entire process lasts for only a fraction of a millisecond. An essential property of the synapse is its capacity to undergo biochemical and morphological changes, a phenomenon that is referred to as synaptic plasticity. RESULTS: In this survey, we consider the mammalian brain synapse as our model. We take a cell biological and a molecular perspective to present fundamental properties of the synapse:(i) the accurate and efficient delivery of organelles and material to and from the synapse; (ii) the coordination of gene expression that underlies a particular NT phenotype; (iii) the induction of local protein expression in a subset of stimulated synapses. We describe the computational facet and the formulation of the problem for each of these topics. CONCLUSION: Predicting the behavior of a synapse under changing conditions must incorporate genomics and proteomics information with new approaches in computational biology.

Animals↗

Is GAS1 a co-receptor for the GDNF family of ligands?

Glial-cell-line-derived neurotrophic factor (GDNF) is a survival and maintenance factor for dopamine-containing neurons and motoneurons. GDNF belongs to a family of structurally related factors that includes neurturin (NRTN), artemin (ARTN) and persephin (PSPN). An initial step in the activation of signaling via the GDNF family of ligands (GFLs) is their binding to their cognate co-receptor GFR alpha. GAS1, an apparently unrelated protein, exhibits homology to GFR alpha and thus we hypothesize that GAS1 can serve as an alternative receptor for GFLs. The functional similarity between GFR alpha and GAS1 extends to their role in embryogenesis, differentiation and glia maintenance, and is substantiated by overlap in their expression profile, subcellular localization and structural details. We propose that the relative expression and localization of the two remote receptors, GFR alpha and GAS1, on the membranes of neuronal and glial cells determines whether these cells survive or undergo apoptotic death.

Amino Acid Sequence↗

ProTeus: identifying signatures in protein termini.

ProTeus (PROtein TErminUS) is a web-based tool for the identification of short linear signatures in protein termini. It is based on a position-based search method for revealing short signatures in termini of all proteins. The initial step in ProTeus development was to collect all signature groups (SIGs) based on their relative positions at the termini. The initial set of SIGs went through a sequential process of inspection and removal of SIGs, which did not meet the attributed statistical thresholds. The SIGs that were found significant represent protein sets with minimal or no overall sequence similarity besides the similarity found at the termini. These SIGs were archived and are presented at ProTeus. The SIGs are sorted by their strong correspondence to functional annotation from external databases such as GO. ProTeus provides rich search and visualization tools for evaluating the quality of different SIGs. A search option allows the identification of terminal signatures in new sequences. ProTeus (ver 1.2) is available at http://www.proteus.cs.huji.ac.il.

Databases, Protein↗

ProTarget: automatic prediction of protein structure novelty.

ProTarget is a Web-based tool for the automatic prediction of fold novelty. It offers the structural genomics community a method for target selection by providing an online analysis of any new or pre-existing sequence for its relationship to any previously solved three-dimensional structure. ProTarget takes as input an amino acid sequence. Regions of this sequence that exhibit high similarity to an existing PDB (Protein Data Bank) sequence are removed, leaving one or more subsequences. Each of these subsequences is then analyzed against a clustering of the protein space to determine the likelihood of its representing a new structural superfamily. This likelihood is derived from the distance in the clustering between the (sub)sequence and sequences that have known structures. The output of ProTarget is a graphical visualization of the protein of interest together with the likelihood that a protein sequence represents a novel structural superfamily. ProTarget is updated regularly and currently covers over 160 000 protein sequences from the SwissProt and PDB databases. ProTarget is available at http://www.protarget.cs.huji.ac.il.

Algorithms↗

Automatic detection of false annotations via binary property clustering.

BACKGROUND: Computational protein annotation methods occasionally introduce errors. False-positive (FP) errors are annotations that are mistakenly associated with a protein. Such false annotations introduce errors that may spread into databases through similarity with other proteins. Generally, methods used to minimize the chance for FPs result in decreased sensitivity or low throughput. We present a novel protein-clustering method that enables automatic separation of FP from true hits. The method quantifies the biological similarity between pairs of proteins by examining each protein's annotations, and then proceeds by clustering sets of proteins that received similar annotation into biological groups. RESULTS: Using a test set of all PROSITE signatures that are marked as FPs, we show that the method successfully separates FPs in 69% of the 327 test cases supplied by PROSITE. Furthermore, we constructed an extensive random FP simulation test and show a high degree of success in detecting FP, indicating that the method is not specifically tuned for PROSITE and performs well on larger scales. We also suggest some means of predicting in which cases this approach would be successful. CONCLUSION: Automatic detection of FPs may greatly facilitate the manual validation process and increase annotation sensitivity. With the increasing number of automatic annotations, the tendency of biological properties to be clustered, once a biological similarity measure is introduced, may become exceedingly helpful in the development of such automatic methods.

Algorithms↗

ProtoNet 4.0: a hierarchical classification of one million protein sequences.

ProtoNet is an automatic hierarchical classification of the protein sequence space. In 2004, the ProtoNet (version 4.0) presents the analysis of over one million proteins merged from SwissProt and TrEMBL databases. In addition to rich visualization and analysis tools to navigate the clustering hierarchy, we incorporated several improvements that allow a simplified view of the scaffold of the proteins. An unsupervised, biologically valid method that was developed resulted in a condensation of the ProtoNet hierarchy to only 12% of the clusters. A large portion of these clusters was automatically assigned high confidence biological names according to their correspondence with functional annotations. ProtoNet is available at: http://www.protonet.cs.huji.ac.il.

Animals↗

Families of membranous proteins can be characterized by the amino acid composition of their transmembrane domains.

SUMMARY: In eukaryotes, membranous proteins account for 20-30% of the proteome. Most of these proteins contain one or more transmembrane (TM) domains. These are short segments that transverse the bilayer lipid membrane. Various properties of the TM domains, such as their number, their topology and their arrangement within the membrane, are closely related to the protein's cellular functions. The properties of the TM domains also determine the cellular targeting and localization of these proteins. It is not known, however, whether the information encoded by TM domains suffices for the purpose of classifying proteins into their functional families. This is the question we address here. We introduce an algorithm that creates a profile of each functional family of membranous proteins based only on the amino acid composition of their TM domains. This is complemented by a classifier program for each such family (to determine whether a given protein belongs to it or not). We find that in most instances TM domains contain enough information to allow an accurate discrimination of approximately 80% sensitivity and approximately 90% specificity among unrelated polytopic functional families with the same number of TM domains. SUPPLEMENTARY INFORMATION: Available at www.protonet.cs.huji.ac.il/TM/

Algorithms↗

A functional hierarchical organization of the protein sequence space.

BACKGROUND: It is a major challenge of computational biology to provide a comprehensive functional classification of all known proteins. Most existing methods seek recurrent patterns in known proteins based on manually-validated alignments of known protein families. Such methods can achieve high sensitivity, but are limited by the necessary manual labor. This makes our current view of the protein world incomplete and biased. This paper concerns ProtoNet, a automatic unsupervised global clustering system that generates a hierarchical tree of over 1,000,000 proteins, based solely on sequence similarity. RESULTS: In this paper we show that ProtoNet correctly captures functional and structural aspects of the protein world. Furthermore, a novel feature is an automatic procedure that reduces the tree to 12% its original size. This procedure utilizes only parameters intrinsic to the clustering process. Despite the substantial reduction in size, the system's predictive power concerning biological functions is hardly affected. We then carry out an automatic comparison with existing functional protein annotations. Consequently, 78% of the clusters in the compressed tree (5,300 clusters) get assigned a biological function with a high confidence. The clustering and compression processes are unsupervised, and robust. CONCLUSIONS: We present an automatically generated unbiased method that provides a hierarchical classification of all currently known proteins.

Algorithms↗