PubMed Health⌕ Search

Biomedical subjects

Yudong Cai

Publications and source records attributed to Yudong Cai.

7 recordsLinked to original sources

LCORL and STC2 Variants Increase Body Size and Growth Rate in Cattle and Other Animals.

Natural variants can significantly improve growth traits in livestock and serve as safe targets for gene editing, thus being applied in animal molecular design breeding. However, such safe and large-effect mutations are severely lacking. Using ancestral recombination graphs, we investigated recent selection signatures in beef cattle breeds, pinpointing sweep-driving variants in the LCORL and STC2 loci with notable effects on body size and growth rate. The ACT-to-A frameshift mutation in LCORL occurs mainly in central-European cattle, and stimulates growth. Remarkably, convergent truncating mutations were also found in commercial breeds of sheep, goats, pigs, horses, dogs, rabbits, and chickens. In the STC2 gene, we identified a missense mutation (A60P) located within the conserved region across vertebrates. We validated the two natural mutations in gene-edited mouse models, where both variants in homozygous carriers significantly increase the average weight by 11%. Our findings provide insights into a seemingly recurring gene target of body size enhancing truncating mutations across domesticated species, and offer valuable targets for gene editing-based breeding in animals.

Animals↗

Demonstration of two novel methods for predicting functional siRNA efficiency.

BACKGROUND: siRNAs are small RNAs that serve as sequence determinants during the gene silencing process called RNA interference (RNAi). It is well know that siRNA efficiency is crucial in the RNAi pathway, and the siRNA efficiency for targeting different sites of a specific gene varies greatly. Therefore, there is high demand for reliable siRNAs prediction tools and for the design methods able to pick up high silencing potential siRNAs. RESULTS: In this paper, two systems have been established for the prediction of functional siRNAs: (1) a statistical model based on sequence information and (2) a machine learning model based on three features of siRNA sequences, namely binary description, thermodynamic profile and nucleotide composition. Both of the two methods show high performance on the two datasets we have constructed for training the model. CONCLUSION: Both of the two methods studied in this paper emphasize the importance of sequence information for the prediction of functional siRNAs. The way of denoting a bio-sequence by binary system in mathematical language might be helpful in other analysis work associated with fixed-length bio-sequence.

Algorithms↗

Predicting O-glycosylation sites in mammalian proteins by using SVMs.

O-glycosylation is one of the most important, frequent and complex post-translational modifications. This modification can activate and affect protein functions. Here, we present three support vector machines models based on physical properties, 0/1 system, and the system combining the above two features. The prediction accuracies of the three models have reached 0.82, 0.85 and 0.85, respectively. The accuracies of the three SVMs methods were evaluated by 'leave-one-out' cross validation. This approach provides a useful tool to help identify the O-glycosylation sites in mammalian proteins. An online prediction web server is available at http://www.biosino.org/Oglyc.

Animals↗

Using Bagging classifier to predict protein domain structural class.

Classification and prediction of protein domain structural class is one of the important topics in the molecular biology. We introduce the Bagging (Bootstrap aggregating), one of the bootstrap methods, for classifying and predicting protein structural classes. By a bootstrap aggregating procedure, the Bagging can improve a weak classifier, for instance the random tree method, to a significant step towards optimality. In this research, it is demonstrated that the Bagging performed at least as well as LogitBoost and Support vector machines in predicting the structural classes for a given protein domain dataset by 10 cross-validation test, which indicate that the Bagging method is promising and anticipated that it could be potentially further improved on predicting protein structural classes as well as other bio-macromolecular attributes, if the bagging method and other existing methods can be effectively complemented with each other.

Algorithms↗

Predicting rRNA-, RNA-, and DNA-binding proteins from primary structure with support vector machines.

In the post-genome era, the prediction of protein function is one of the most demanding tasks in the study of bioinformatics. Machine learning methods, such as the support vector machines (SVMs), greatly help to improve the classification of protein function. In this work, we integrated SVMs, protein sequence amino acid composition, and associated physicochemical properties into the study of nucleic-acid-binding proteins prediction. We developed the binary classifications for rRNA-, RNA-, DNA-binding proteins that play an important role in the control of many cell processes. Each SVM predicts whether a protein belongs to rRNA-, RNA-, or DNA-binding protein class. Self-consistency and jackknife tests were performed on the protein data sets in which the sequences identity was < 25%. Test results show that the accuracies of rRNA-, RNA-, DNA-binding SVMs predictions are approximately 84%, approximately 78%, approximately 72%, respectively. The predictions were also performed on the ambiguous and negative data set. The results demonstrate that the predicted scores of proteins in the ambiguous data set by RNA- and DNA-binding SVM models were distributed around zero, while most proteins in the negative data set were predicted as negative scores by all three SVMs. The score distributions agree well with the prior knowledge of those proteins and show the effectiveness of sequence associated physicochemical properties in the protein function prediction. The software is available from the author upon request.

Amino Acid Sequence↗

Carbon-carbon bond formation by radical addition-fragmentation reactions of O-alkylated enols.

Alpha-tert-butoxystyrene [H2C=C(OBut)Ph] reacts with alpha-bromocarbonyl or alpha-bromosulfonyl compounds [R1R2C(Br)EWG; EWG =-C(O)X or -S(O2)X] to bring about replacement of the bromine atom by the phenacyl group and give R1R2C(EWG)CH2C(O)Ph. These reactions take place in refluxing benzene or cyclohexane with dilauroyl peroxide or azobis(isobutyronitrile) as initiator and proceed by a radical-chain mechanism that involves addition of the relatively electrophilic radical R1R2(EWG)C* to the styrene. This is followed by beta-scission of the derived alpha-tert-butoxybenzylic adduct radical to give But*, which then abstracts bromine from the organic halide to complete the chain. Alpha-1-adamantoxystyrene reacts similarly with R1R2C(Br)EWG, at higher temperature in refluxing octane using di-tert-amyl peroxide as initiator, and gives phenacylation products in generally higher yields than are obtained using alpha-tert-butoxystyrene. Simple iodoalkanes, which afford relatively nucleophilic alkyl radicals, can also be successfully phenacylated using alpha-1-adamantoxystyrene. O-Alkyl O-(tert-butyldimethylsilyl) ketene acetals H2C=C(OR)OTBS, in which R is a secondary or tertiary alkyl group, react in an analogous fashion with organic halides of the type R1R2C(Br)EWG to give the carboxymethylation products R1R2C(EWG)CH2CO2Me, after conversion of the first-formed silyl ester to the corresponding methyl ester. The silyl ketene acetals also undergo radical-chain reactions with electron-poor alkenes to bring about alkylation-carboxymethylation of the latter. For example, phenyl vinyl sulfone reacts with H2C=C(OBut)OTBS to afford ButCH2CH(SO2Ph)CH2CO2Me via an initial silyl ester. In a more complex chain reaction, involving rapid ring opening of the cyclopropyldimethylcarbinyl radical, the ketene acetal H2C=C(OCMe2C3H5-cyclo)OTBS reacts with two molecules of N-methyl- or N-phenyl-maleimide to bring about [3 + 2] annulation of one molecule of the maleimide, and then to link the bicyclic moiety thus formed to the second molecule of the maleimide via an alkylation-carboxymethylation reaction.

Alkenes↗

Information-theoretic analysis of protein sequences shows that amino acids self-cluster.

We analyse for each of 20 amino acids X the statistics of spacings between consecutive occurrences of X within the well-characterized Saccharomyces cerevisiae genome. The occurrences of amino acids may exhibit near random, clustered or smoothed out behaviour, like one-dimensional stochastic processes along the protein chain. If amino acids are distributed randomly within a sequence, then they follow a Poisson process, and a histogram of the number of observations of each gap size would asymptotically follow a negative exponential distribution. The novelty of the present approach lies in the use of differential geometric methods to quantify information on sequencing of amino acids and groups of amino acids, via the sequences of intervals between their occurrences. The differential geometry arises from an information-theoretic distance function on the two-dimensional space of stochastic processes subordinate to gamma distributions-which latter include the random process as a special case. We find that maximum-likelihood estimates of parametric statistics show that all 20 amino acids tend to cluster, some substantially. In other words, the frequencies of short gap lengths tend to be higher and the variance of the gap lengths is greater than expected by chance. This may be because localizing amino acids with the same properties may favour secondary structure formation or transmembrane domains. Gap sizes of 1 or 2 are generally disfavoured, 1 strongly so. The only exceptions to this are Gln and Ser, as a result of poly(Gln) or poly(Ser) sequences. There are preferences for gaps of 4 and 7 that can be attributed to alpha -helices. In particular, a favoured gap of 7 for Leu is found in coiled coils. Our method contributes to the characterization of whole sequences by extracting and quantifying stable stochastic features.

Amino Acids↗