PubMed Health⌕ Search

Biomedical subjects

Youhe Gao

Publications and source records attributed to Youhe Gao.

13 recordsLinked to original sources

A high efficiency strategy for binding property characterization of peptide-binding domains.

A large proportion of protein-protein interactions is mediated by families of peptide-binding domains. Comprehensive characterization of each of these domains is critical for understanding the mechanisms and networks of protein interaction at the domain level. However, existing methods are all based on large scale screenings for each domain that are inefficient to deal with hundreds of members in major domain families. We developed a systematic strategy for efficient binding property characterization of peptide-binding domains based on high throughput validation screening of a specialized candidate ligand library using yeast two-hybrid mating array. Its outstanding feature is that the overall efficiency is dramatically improved compared with that of traditional screening, and it will be higher as the system cycles. PDZ domain family was first used to test the strategy. Five PDZ domains were rapidly characterized. Broader binding properties were identified compared with other methods, including novel recognition specificities that provided the basis for major revision of conventional PDZ classification. Several novel interactions were discovered, serving as significant clues for further functional investigation. This strategy can be easily extended to a variety of peptide-binding domains as a powerful tool for comprehensive analysis of domain binding property in proteomic scale.

Animals↗

An integrated machine learning system to computationally screen protein databases for protein binding peptide ligands.

A fairly large set of protein interactions is mediated by families of peptide binding domains, such as Src homology 2 (SH2), SH3, PDZ, major histocompatibility complex, etc. To identify their ligands by experimental screening is not only labor-intensive but almost futile in screening low abundance species due to the suppression by high abundance species. An ideal way of studying protein-protein interactions is to use high throughput computational approaches to screen protein sequence databases to direct the validating experiments toward the most promising peptides. Predictors with only good cross-validation were not good enough to screen protein databases. In the current study we built integrated machine learning systems using three novel coding methods and screened the Swiss-Prot and GenBank protein databases for potential ligands of 10 SH3 and three PDZ domains. A large fraction of predictions has already been experimentally confirmed by other independent research groups, indicating a satisfying generalization capability for future applications in identifying protein interactions.

Amino Acid Motifs↗

Microwave-assisted protein preparation and enzymatic digestion in proteomics.

The combinations of gel electrophoresis or LC and mass spectrometry are two popular approaches for large scale protein identification. However, the throughput of both approaches is limited by the speed of the protein digestion process. Present research into fast protein enzymatic digestion has been focused mainly on known proteins, and it is unclear whether these results can be extrapolated to complex protein mixtures. In this study microwave technology was used to develop a fast protein preparation and enzymatic digestion method for protein mixtures. The protein mixtures in solution or in gel were prepared and digested by microwave-assisted protein enzymatic digestion, which rapidly produces peptide fragments. The peptide fragments were further analyzed by capillary LC and ESI-ion trap-MS or MALDI-TOF-MS. The technique was optimized using bovine serum albumin and then applied to human urinary proteins and yeast lysate. The method enabled preparation and digestion of protein mixtures in solution (human urinary proteins) or in gel (yeast lysate) in 6 or 25 min, respectively. Equivalent (in-solution) or better (in-gel) digestion efficiency was obtained using microwave-assisted protein enzymatic digestion compared with the standard overnight digestion method. This new application of microwave technology to protein mixture preparation and enzymatic digestion will hasten the application of proteomic techniques to biological and clinical research.

Chromatography, Liquid↗

Concanavalin A-captured glycoproteins in healthy human urine.

Both the urinary proteome and its glycoproteome can reflect human health status, and more directly, functions of kidney and urinary tracts. Because the high abundance protein albumin is not N-glycosylated, the urine N-glycoprotein enrichment procedure could deplete it, and urine proteome could thus provide a more detailed protein profile in addition to glycosylation information especially when albuminuria occurs in some kidney diseases. In terms of describing the details of urinary proteins, the urine glycoproteome is even a better choice than the proteome itself. Pooled urine samples from healthy volunteers were collected and acetone-precipitated for proteins. N-Linked glycoproteins enriched with concanavalin A affinity purification were separated and analyzed by SDS-PAGE-reverse phase LC/MS/MS or two-dimensional LC/MS/MS. A total of 225 urinary proteins were identified based on two-hit criteria with reliability over 97% for each peptide. Among these proteins, 94 were identified in previous urine proteome works, 150 were annotated as glycoproteins in Swiss-Prot, and 43 were predicted as glycoproteins by NetNGlyc 1.0. A number of known biomarkers and disease-related glycoproteins were identified. Because changes in protein quantity or the glycosylation status can lead to changes in the concanavalin A-captured glycoprotein profile, specific urine glycoproteome patterns might be observed for specific pathological conditions as multiplex urinary biomarkers. Knowledge of the urine glycoproteome is important in understanding kidney and body function.

Concanavalin A↗

Human urine proteome analysis by three separation approaches.

The urinary proteome is known to be a valuable field of study related to organ functions. There have been several extensive urine proteome studies. However, the overlapping rate among different studies is relatively low. Whether the low overlapping rate was caused by different sample sources, preparation, separation and identification methods is unknown. Moreover, low molecular mass (<10 kDa) proteins have not been studied extensively. In this report, male and female pooled urine samples were collected from healthy volunteers. The urinary proteins were acetone precipitated, separated and identified by three approaches, 1-DE plus 1-D LC/MS/MS, direct 1-D LC/MS/MS and 2-D LC/MS/MS. 1-D tricine gels were used to separate low molecular mass proteins. The tandem mass spectra of positive identifications were quality controlled both by manual validation and using advanced mass spectrum scanner software. A total of 226 urinary proteins were identified; 171 proteins were identified by proteomics approach for the first time, including 4 male-specific proteins. Twelve low molecular mass proteins were identified. Most urinary proteins had a molecular mass between 30 and 60 kDa and a pI between 4 and 10. The apparent molecular masses of many proteins were different from theoretical ones, which indicated their post-translational modification and degradation. The effects of sample preparation, separation and identification methods on the overlapping rate of different experiments are discussed.

Adult↗

An analysis of protein abundance suppression in data dependent liquid chromatography and tandem mass spectrometry with tryptic peptide mixtures of five known proteins.

Reverse phase liquid chromatography (RPLC) has been widely used in proteomics research for peptide separation. When protein samples are separated by RPLC and identified with electrospray ion trap mass spectrometry (ESI-MS), the signals of high-abundance proteins may suppress those of low-abundance proteins, a phenomenon known as abundance suppression. To what degree the abundance suppression correlates to the number of tryptic peptides in the high-abundance proteins has not been carefully investigated. We tried to answer this question by studying the mixtures digested from five known proteins. The numbers of identified tryptic peptides (longer than five amino acids) of the five proteins ranged from 12 to 47. Four different peptide mixtures with 10- to 100-fold abundance differences of five known proteins were separated by RPLC and identified by ESI-MS. Our results showed that abundance suppression was related to the tryptic peptide numbers in the high-abundance protein. Within a 100-fold protein abundance difference range, tryptic peptide number in the low-abundance proteins could be suppressed up to seven times by high-abundance proteins. The procedure we suggest here can help to identify low-abundance proteins co-purified with their high-abundance binding protein. The result can also help to identify specific high-abundance proteins for removal by immunoaffinity.

Animals↗

Comparative proteome analysis of breast cancer and normal breast.

Breast cancer is a leading cause of death for women. The underlying molecular mechanism is still not well understood. In this study, two-dimensional gel electrophoresis combined with mass spectrometry was used to analyze changes in the proteome of infiltrating ductal carcinoma compared to normal breast tissue. Ten sets of two-dimensional gels per experimental condition were analyzed and more than 500 spots each were detected. This revealed 39 spots for which expression in breast cancer cells were reproducibly altered more than twofold compared to normal controls (p < 0.01). These spots represented 25 different proteins after identification using the database search after mass spectrometry, comprising cell defense proteins, enzymes involved in glycolytic energy metabolism and homeostasis, protein folding and structural proteins, proteins involved in cytoskeleton and cell motility, and proteins involved in other functions. In addition, 28 nondifferentially expressed proteins with different functions were also mapped and identified, which might help to establish a two-dimensional gel electrophoresis reference map of human breast cancer. Our study shows that proteomics offers a powerful methodology to detect the proteins that show different expression patterns in breast cancer tissue and may provide an accurate molecular classification. The differentially expressed proteins may be used as potential candidate markers for diagnostic purposes or for determination of tumor sensitivity to therapy. The functional implications of the identified proteins are discussed.

Biomarkers, Tumor↗

A method for generation of arbitrary peptide libraries using genomic DNA.

Random peptide libraries can be constructed either by in vitro synthesis of random peptides, or through translation of DNA sequences from synthetic random oligonucleotides. Here we describe an alternative way of making arbitrary peptide libraries with high diversity that can be used in screening as random peptide libraries. Genomic DNA digested with a frequent-cutting restriction enzyme recognizing four nucleotides will theoretically consist of small DNA pieces with average length of 256 nucleotides, and on average around 107 fragments can be generated from a genome of 3 x 109 bases. A peptide library translated from these fragments will have sufficient diversity for some protein interaction screening experiments. Moreover, the same genome digested with a different four-cutter enzyme or ligated into different reading frames will result in different nonoverlapping libraries. A series of such libraries could be generated with genomic DNAs from different species. In this study, human genomic DNA was digested with four-cutter restriction enzymes DpnII and Tsp509I, respectively, and cloned into yeast expression vector pGADT7 to generate arbitrary peptide libraries. These libraries were used in yeast two-hybrid assays to screen for binding motifs of the PDZ domain containing protein synectin. Our results showed that in addition to various native carboxy-terminal tails, synectin could also bind to many artificial ones, some of which contained a consensus sequence--(S/T)XC-COOH.

Adaptor Proteins, Signal Transducing↗

AMASS: software for automatically validating the quality of MS/MS spectrum from SEQUEST results.

Time-consuming and experience-dependent manual validations of tandem mass spectra are usually applied to SEQUEST results. This inefficient method has become a significant bottleneck for MS/MS data processing. Here we introduce a program AMASS (advanced mass spectrum screener), which can filter the tandem mass spectra of SEQUEST results by measuring the match percentage of high-abundant ions and the continuity of matched fragment ions in b, y series. Compared with Xcorr and DeltaCn filter, AMASS can increase the number of positives and reduce the number of negatives in 22 datasets generated from 18 known protein mixtures. It effectively removed most noisy spectra, false interpretations, and about half of poor fragmentation spectra, and AMASS can work synergistically with Rscore filter. We believe the use of AMASS and Rscore can result in a more accurate identification of peptide MS/MS spectra and reduce the time and energy for manual validation.

Algorithms↗

RScore: a peptide randomicity score for evaluating tandem mass spectra.

RScore, a new criterion of randomicity for evaluating tandem mass (MS/MS) spectra, is described. RScore is defined as the relative quality in cross-correlation and matched intensity percentage of a potentially positive peptide to those of other possible candidates for the same spectrum. By utilizing RScore combined with less stringent SEQUEST score filters, the number of true positive peptides can be increased and the number of false positives in datasets from a known protein mixture can be reduced compared with current SEQUEST parameters used alone. This algorithm is simple and adds little overheads to SEQUEST computation.

Algorithms↗

Construction of a non-redundant human SH2 domain database.

Domain database is essential for domain property research. Eliminating redundant information in database query is very important for database quality. Here we report the manual construction of a non-redundant human SH2 domain database. There are 119 human SH2 domains in 110 SH2-containing proteins. Human SH2s were aligned with ClustalX, and a homologous tree was generated. In this tree, proteins with similar known function were classified into the same group. Some proteins in the same group have been reported to have similar binding motifs experimentally. The tree might provide clues about possible functions of hypothetical proteins for further experimental verification.

Amino Acid Sequence↗

A systematical analysis of tryptic peptide identification with reverse phase liquid chromatography and electrospray ion trap mass spectrometry.

In this study we systematically analyzed the elution condition of tryptic peptides and the characteristics of identified peptides in reverse phase liquid chromatography and electrospray tandem mass spectrometry (RPLC-MS/MS) analysis. Following protein digestion with trypsin, the peptide mixture was analyzed by on-line RPLC-MS/MS. Bovine serum albumin (BSA) was used to optimize acetonitrile (ACN) elution gradient for tryptic peptides, and Cytochrome C was used to retest the gradient and the sensitivity of LC-MS/MS. The characteristics of identified peptides were also analyzed. In our experiments, the suitable ACN gradient is 5% to 30% for tryptic peptide elution and the sensitivity of LC-MS/MS is 50 fmol. Analysis of the tryptic peptides demonstrated that longer (more than 10 amino acids) and multi-charge state (+2, +3) peptides are likely to be identified, and the hydropathicity of the peptides might not be related to whether it is more likely to be identified or not. The number of identified peptides for a protein might be used to estimate its loading amount under the same sample background. Moreover, in this study the identified peptides present three types of redundancy, namely identification, charge, and sequence redundancy, which may repress low abundance protein identification.

Acetonitriles↗

Proline- and arginine-rich peptides constitute a novel class of allosteric inhibitors of proteasome activity.

Substrate-specific inhibition of the proteasome has been unachievable despite great interest in proteasome inhibitors as drugs. Recent studies demonstrated that PR39, a natural proline- and arginine-rich antibacterial peptide, stimulates angiogenesis and inhibits inflammatory responses by specifically blocking degradation of IkappaBalpha and HIF-1alpha by the proteasome. However, molecular events involved in the PR39-proteasome interaction have not been elucidated. Here we show that PR39 is a noncompetitive and reversible inhibitor of the proteasome function. This effect is achieved by a unique allosteric mechanism allowing for specific inhibition of degradation of selected proteins without affecting total proteasome-dependent proteolysis. Atomic force microscopy (AFM) studies demonstrate that 20S and 26S proteasomes treated with PR39 or its derivatives exhibit serious perturbations in their structure and their normal allosteric movements. These effects are universal for proteasomes from yeast to human. The shortest functional sequence derived from PR39 still showing the allosteric inhibitory effect consists of eleven NH(2)-terminal residues containing essential three NH(2)-terminal arginines. The noncompetitive and reversible in vitro action of PR39 and its truncated derivatives is matched by the ability of the peptides to induce angiogenesis in vivo. We postulate that PR39 changes conformational dynamics of the proteasomes by interactions with the noncatalytic subunit alpha7 in a way that prevents the enzyme from cleaving the substrates of unique structural constraints.

Allosteric Site↗