PubMed Health⌕ Search

Biomedical subjects

Lei Lin

Publications and source records attributed to Lei Lin.

10 recordsLinked to original sources

Effect of example weights on prediction of protein-protein interactions.

Protein-protein interactions (PPIs) prediction is an important issue in biology. Recently many computational methods have been proposed to determine PPIs. However, there is no golden standard dataset for these methods now. Furthermore, there exists different quality among training examples and the quality is always ignored by the current methods. In the condition of low-quality examples, the system should tolerate the data noise. Example weighting strategy is used in this paper to build a robust system and solve the problem of data noise. Training examples are investigated and a new example selecting/using strategy is proposed. Training example weighting method based on confidence is proposed. Different weight setting strategies are discussed and the corresponding results are given in the experiment. A new model integrating example weighting strategy, attraction-repulsion (AR) weight model, is proposed. Experimental results on Saccharomyces cerevisiae demonstrate that the new model outperforms the original AR model in the ROC score measure by over 8%. Furthermore, the example weighting strategy is applied to another domain-based PPIs prediction method, maximum likelihood estimation (MLE) method, and the modified MLE method obtains better performance than the original MLE method. At same time, our examples weighting strategy can be applied to any other training example based PPIs prediction methods.

Algorithms↗

Novel knowledge-based mean force potential at the profile level.

BACKGROUND: The development and testing of functions for the modeling of protein energetics is an important part of current research aimed at understanding protein structure and function. Knowledge-based mean force potentials are derived from statistical analyses of interacting groups in experimentally determined protein structures. Current knowledge-based mean force potentials are developed at the atom or amino acid level. The evolutionary information contained in the profiles is not investigated. Based on these observations, a class of novel knowledge-based mean force potentials at the profile level has been presented, which uses the evolutionary information of profiles for developing more powerful statistical potentials. RESULTS: The frequency profiles are directly calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into binary profiles with a probability threshold. As a result, the protein sequences are represented as sequences of binary profiles rather than sequences of amino acids. Similar to the knowledge-based potentials at the residue level, a class of novel potentials at the profile level is introduced. We develop four types of profile-level statistical potentials including distance-dependent, contact, Phi/Psi dihedral angle and accessible surface statistical potentials. These potentials are first evaluated by the fold assessment between the correct and incorrect models generated by comparative modeling from our own and other groups. They are then used to recognize the native structures from well-constructed decoy sets. Experimental results show that all the knowledge-base mean force potentials at the profile level outperform those at the residue level. Significant improvements are obtained for the distance-dependent and accessible surface potentials (5-6%). The contact and Phi/Psi dihedral angle potential only get a slight improvement (1-2%). Decoy set evaluation results show that the distance-dependent profile-level potentials even outperform other atom-level potentials. We also demonstrate that profile-level statistical potentials can improve the performance of threading. CONCLUSION: The knowledge-base mean force potentials at the profile level can provide better discriminatory ability than those at the residue level, so they will be useful for protein structure prediction and model refinement.

Algorithms↗

Domain boundary prediction based on profile domain linker propensity index.

Successful prediction of protein domain boundaries provides valuable information not only for the computational structure prediction of multi-domain proteins but also for the experimental structure determination. In this work, a novel index at the profile level is presented, namely, the profile domain linker propensity index (PDLI), which uses the evolutionary information of profiles for domain linker prediction. The frequency profiles are directly calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into binary profiles with a probability threshold. PDLI is then obtained by the frequencies of binary profiles in domain linkers as compared to those in domains. A smooth and normalized numeric profile is generated for any amino acid sequences from which the domain linkers can be predicted. Testing on the Structural Classification of Proteins (SCOP) database and CASP6 targets shows that PDLI outperforms other indexes at the amino acid level.

Computational Biology↗

[Separation of sitafloxacin epimers by capillary electrophoresis].

Sitafloxacin epimers were separated by capillary zone electrophoresis using gamma-cyclodextrin (gamma-CD) and D-phenylalanine (D-Phe) as chiral selector. The effects of the concentrations of gamma-CD, D-Phe, Cu2+ and pH of buffer were investigated. An uncoated fused-silica capillary of 50 microm i.d. and 60 cm (effective length 52.5 cm) was used. The capillary temperature was maintained at 25 degrees C. Samples were injected under a pressure of 7 kPa for 5 s and separated at 15 kV. A baseline separation of sitafloxacin epimers was achieved with a background electrolyte of 10 mmol/L KH2PO4-K2HPO4(pH 4.5), 10 mmol/L CuSO4, 20 mmol/L gamma-CD and 10 mmol/L D-Phe. The linear range for sitafloxacin was 32 -400 mg/L (0.996). The relative standard deviations (RSDs) of migration time and peak area were less than 1.9% and 3.8% respectively. This method can be applied in qualitative and quantitative analysis for sitafloxacin epimers.

Electrophoresis, Capillary↗

Application of latent semantic analysis to protein remote homology detection.

MOTIVATION: Remote homology detection between protein sequences is a central problem in computational biology. The discriminative method such as the support vector machine (SVM) is one of the most effective methods. Many of the SVM-based methods focus on finding useful representations of protein sequence, using either explicit feature vector representations or kernel functions. Such representations may suffer from the peaking phenomenon in many machine-learning methods because the features are usually very large and noise data may be introduced. Based on these observations, this research focuses on feature extraction and efficient representation of protein vectors for SVM protein classification. RESULTS: In this study, a latent semantic analysis (LSA) model, which is an efficient feature extraction technique from natural language processing, has been introduced in protein remote homology detection. Several basic building blocks of protein sequences have been investigated as the 'words' of 'protein sequence language', including N-grams, patterns and motifs. Each protein sequence is taken as a 'document' that is composed of bags-of-word. The word-document matrix is constructed first. The LSA is performed on the matrix to produce the latent semantic representation vectors of protein sequences, leading to noise-removal and smart description of protein sequences. The latent semantic representation vectors are then evaluated by SVM. The method is tested on the SCOP 1.53 database. The results show that the LSA model significantly improves the performance of remote homology detection in comparison with the basic formalisms. Furthermore, the performance of this method is comparable with that of the complex kernel methods such as SVM-LA and better than that of other sequence-based methods such as PSI-BLAST and SVM-pairwise.

Algorithms↗

Compact continuous-wave blue lasers by direct frequency doubling of laser diodes with periodically poled lithium niobate waveguide crystals.

A compact continuous-wave blue laser has been demonstrated by direct frequency doubling of a laser diode with a periodically poled lithium niobate (PPLN) waveguide crystal. The optimum PPLN temperature is near 28 degrees C, and the dependence of waveguide crystals on crystal temperature is less sensitive than that of bulk crystals. A total of 14.8 mW of 488-nm laser power has been achieved.

Journal Article↗

A seqlet-based maximum entropy Markov approach for protein secondary structure prediction.

A novel method for predicting the secondary structures of proteins from amino acid sequence has been presented. The protein secondary structure seqlets that are analogous to the words in natural language have been extracted. These seqlets will capture the relationship between amino acid sequence and the secondary structures of proteins and further form the protein secondary structure dictionary. To be elaborate, the dictionary is organism-specific. Protein secondary structure prediction is formulated as an integrated word segmentation and part of speech tagging problem. The word-lattice is used to represent the results of the word segmentation and the maximum entropy model is used to calculate the probability of a seqlet tagged as a certain secondary structure type. The method is markovian in the seqlets, permitting efficient exact calculation of the posterior probability distribution over all possible word segmentations and their tags by viterbi algorithm. The optimal segmentations and their tags are computed as the results of protein secondary structure prediction. The method is applied to predict the secondary structures of proteins of four organisms respectively and compared with the PHD method. The results show that the performance of this method is higher than that of PHD by about 3.9% Q3 accuracy and 4.6% SOV accuracy. Combining with the local similarity protein sequences that are obtained by BLAST can give better prediction. The method is also tested on the 50 CASP5 target proteins with Q3 accuracy 78.9% and SOV accuracy 77.1%. A web server for protein secondary structure prediction has been constructed which is available at http://www.insun.hit.edu.cn:81/demos/biology/index.html.

Algorithms↗

Transendothelial movement of liposomes in vitro mediated by cancer cells, neutrophils or histamine.

A two-chamber culture system has been used to examine the ability of small liposomes to cross an endothelial cell barrier in response to various stimuli. Transendothelial transit of liposomes was almost negligible in the presence of intact, healthy endothelial cells (EC). Addition of histamine induced a concentration-dependent increase in the movement of liposomes across the EC monolayer. In the presence of polymorphonuclear neutrophils (PMNs), migrating in response to a chemotactic gradient of N-Formil-Met-Leu-Phe (fMLP), both liposomes and IgG crossed EC monolayer by a paracellular pathway, largely independent of an association with the PMNs. The presence of cancer cell, growing in the lower chamber or the presence of cancer cell-conditioned media, also resulted in the passage of liposome across the EC. We conclude that EC monolayers are sufficiently disrupted by several physiologically relevant stimuli to allow for the transendothelial passage of liposomes. These results have important implications for the therapeutic use of liposome in the treatment of cancer or other inflammatory processes.

Animals↗

Site-specific 32P-labeling of cytokines, monoclonal antibodies, and other protein substrates for quantitative assays and therapeutic application.

Radiolabeled proteins are used in a variety of laboratory applications as well as in radioimmunotherapy. This review focuses on methods that utilize genetic engineering to introduce exogenous phosphorylation sites into proteins. Protein kinase substrate sites can be introduced into target proteins to serve as tags for several purposes. Because many protein kinases, each preferring a unique consensus sequence, are well characterized, the essential structure and function of the target protein can be effectively preserved through judicious selection and design of the phosphate incorporation site. After phosphorylation, these proteins are often indistinguishable from the parent molecules in assays of functional or biological activity. This convenient approach permits incorporation of 32P, 33P, 35S, or nonradioactive 31P, and is rapid, efficient, and safe. Most importantly 32P labeling of monoclonal antibodies or other therapeutic protein candidates has several significant advantages over radioiodination or chemical conjugation of heavy metal isotopes.

Amino Acid Sequence↗