PubMed Health⌕ Search

Biomedical subjects

Yuanyuan Xiao

Publications and source records attributed to Yuanyuan Xiao.

7 recordsLinked to original sources

Pilot study of allele-specific multi-InDel markers for the detection of extremely unbalanced DNA mixtures.

Mixtures are common in forensic casework, and they represent one of the most challenging types of biological evidence. Traditional short tandem repeat analyses are often associated with limitations when dealing with extremely unbalanced mixtures because alleles from minor contributors can easily be masked by those of major contributors. Consequently, researchers have developed new technologies and methods for improving the analysis of mixtures, spanning upstream DNA extraction and downstream software analysis. Among these, strategies combining allele-specific amplification with compound markers have drawn particular interest because of their ability to selectively detect minor contributors in complex mixtures. In this study, we screened multi-InDels across the entire genome, designed allele-specific primers compatible with the capillary electrophoresis platform, and further explored their potential in unbalanced DNA mixtures and cell-free fetal DNA (cffDNA). Ultimately, a set comprising 10 multi-InDels was developed, and this included two groups of primers that separately amplified the long alleles (L primer set) and short alleles (S primer set). The results demonstrated that each primer pair could detect the minor component at a 1:1000 mixture ratio, whereas the L and S primer sets successfully detected the minor contributors at mixture ratios of 1:200 and 1:500, respectively. Furthermore, in the cffDNA analysis, 60 of 78 informative markers were successfully detected, with the complete detection of all informative markers achieved in 18 mother-child reference pairs. Overall, allele-specific amplification-based multi-InDel markers enabled the sensitive detection of minor contributors, providing a potential strategy for the analysis of unbalanced two-person mixtures.

Allelic-specific amplification↗

Analysis of a splice array experiment elucidates roles of chromatin elongation factor Spt4-5 in splicing.

Splicing is an important process for regulation of gene expression in eukaryotes, and it has important functional links to other steps of gene expression. Two examples of these linkages include Ceg1, a component of the mRNA capping enzyme, and the chromatin elongation factors Spt4-5, both of which have recently been shown to play a role in the normal splicing of several genes in the yeast Saccharomyces cerevisiae. Using a genomic approach to characterize the roles of Spt4-5 in splicing, we used splicing-sensitive DNA microarrays to identify specific sets of genes that are mis-spliced in ceg1, spt4, and spt5 mutants. In the context of a complex, nested, experimental design featuring 22 dye-swap array hybridizations, comprising both biological and technical replicates, we applied five appropriate statistical models for assessing differential expression between wild-type and the mutants. To refine selection of differential expression genes, we then used a robust model-synthesizing approach, Differential Expression via Distance Synthesis, to integrate all five models. The resultant list of differentially expressed genes was then further analyzed with regard to select attributes: we found that highly transcribed genes with long introns were most sensitive to spt mutations. QPCR confirmation of differential expression was established for the limited number of genes evaluated. In this paper, we showcase splicing array technology, as well as powerful, yet general, statistical methodology for assessing differential expression, in the context of a real, complex experimental design. Our results suggest that the Spt4-Spt5 complex may help coordinate splicing with transcription under conditions that present kinetic challenges to spliceosome assembly or function.

Chromatin↗

Prediction of genomewide conserved epitope profiles of HIV-1: classifier choice and peptide representation.

Identification of peptides binding to Major Histocompatibility Complex (MHC) molecules is important for accelerating vaccine development and improving immunotherapy. Accordingly, a wide variety of prediction methods have been applied in this context. In this paper, we introduce (tree-based) ensemble classifiers for such problems and contrast their predictive performance with forefront existing methods for both MHC class I and class II molecules. In addition, we investigate the impact of differing peptide representation schemes on performance. Finally, classifier predictions are used to conduct genomewide scans of a diverse collection of HIV-1 strains, enabling assessment of epitope conservation. We investigated all combinations of six classification methods (classification trees, artificial neural networks, support vector machines, as well as the more recently devised ensemble methods (bagging, random forests, boosting) with four peptide representation schemes (amino acid sequence, select biophysical properties, select quantitative structure-activity relationship (QSAR) descriptors, and the combination of the latter two) in predicting peptide binding to an MHC class I molecule (HLA-A2) and MHC class II molecule (HLA-DR4). Our results show that the ensemble methods are consistently more accurate than the other three alternatives. Furthermore, they are robust with respect to parameter tuning. Among the four representation schemes, the amino acid sequence representation gave consistently (across classifiers) best results. This finding obviates the need for feature selection strategies incurred by use of biophysical and/or QSAR properties. We obtained, and aligned, a diverse set of 32 HIV-1 genomes and pursued genomewide HLA-DR4 epitope profiling by querying with respect to classifier predictions, as obtained under each of the four peptide representation schemes. We validated those epitopes conserved across strains against known T-cell epitopes. Once again, amino acid sequence representation was at least as effective as using properties. Assessment of novel epitope predictions awaits experimental verification.

Journal Article↗

Stepwise normalization of two-channel spotted microarrays.

Intensities measurements of spotted microarrays embody many undesirable systematic variations. Very commonly, varying amounts and types of such variations are observed in different arrays. Although various normalization methods have been proposed to remove such systematic effects, it has not been well studied how to assess or select the most appropriate method for different arrays and data sets. To address this issue, we present a novel normalization technique, STEPNORM, for data-dependent and adaptive normalization of two-channel spotted microarrays. STEPNORM performs a stepwise interrogation of a range of different normalization models and selects the appropriate method based on formal model selection criteria. In addition, we evaluate the effectiveness of STEPNORM and other commonly used normalization methods utilizing a set of specially constructed splicing arrays.

Journal Article↗

Identifying differentially expressed genes from microarray experiments via statistic synthesis.

MOTIVATION: A common objective of microarray experiments is the detection of differential gene expression between samples obtained under different conditions. The task of identifying differentially expressed genes consists of two aspects: ranking and selection. Numerous statistics have been proposed to rank genes in order of evidence for differential expression. However, no one statistic is universally optimal and there is seldom any basis or guidance that can direct toward a particular statistic of choice. RESULTS: Our new approach, which addresses both ranking and selection of differentially expressed genes, integrates differing statistics via a distance synthesis scheme. Using a set of (Affymetrix) spike-in datasets, in which differentially expressed genes are known, we demonstrate that our method compares favorably with the best individual statistics, while achieving robustness properties lacked by the individual statistics. We further evaluate performance on one other microarray study.

Algorithms↗

Plasticity of gene expression in injured human dorsal root ganglia revealed by GeneChip oligonucleotide microarrays.

Root avulsion from the spinal cord occurs in brachial plexus lesions. It is the practice to repair such injuries by transferring an intact neighbouring nerve to the distal stump of the damaged nerve; avulsed dorsal root ganglia (DRG) are removed to enable nerve transfer. Such avulsed adult human cervical DRG ( [Formula: see text] ) obtained at surgery were compared to controls, for the first time, using GeneChip oligonucleotide arrays. We report 91 genes whose expression levels are clearly altered by the injury. This first study provides a global assessment of the molecular events or "gene switches" as a consequence of DRG injuries, as the tissues represent a wide range of surgical delay, from 1 to 100 days. A number of these genes are novel with respect to sensory ganglia, while others are known to be involved in neurotransmission, trophism, cytokine functions, signal transduction, myelination, transcription regulation, and apoptosis. Cluster analysis showed that genes involved in the same functional groups are largely positioned close to each other. This study represents an important step in identifying new genes and molecular mechanisms in human DRG, with potential therapeutic relevance for nerve repair and relief of chronic neuropathic pain.

Adult↗

Assessment of differential gene expression in human peripheral nerve injury.

BACKGROUND: Microarray technology is a powerful methodology for identifying differentially expressed genes. However, when thousands of genes in a microarray data set are evaluated simultaneously by fold changes and significance tests, the probability of detecting false positives rises sharply. In this first microarray study of brachial plexus injury, we applied and compared the performance of two recently proposed algorithms for tackling this multiple testing problem, Significance Analysis of Microarrays (SAM) and Westfall and Young step down adjusted p values, as well as t-statistics and Welch statistics, in specifying differential gene expression under different biological states. RESULTS: Using SAM based on t statistics, we identified 73 significant genes, which fall into different functional categories, such as cytokines / neurotrophin, myelin function and signal transduction. Interestingly, all but one gene were down-regulated in the patients. Using Welch statistics in conjunction with SAM, we identified an additional set of up-regulated genes, several of which are engaged in transcription and translation regulation. In contrast, the Westfall and Young algorithm identified only one gene using a conventional significance level of 0.05. CONCLUSION: In coping with multiple testing problems, Family-wise type I error rate (FWER) and false discovery rate (FDR) are different expressions of Type I error rates. The Westfall and Young algorithm controls FWER. In the context of this microarray study, it is, seemingly, too conservative. In contrast, SAM, by controlling FDR, provides a promising alternative. In this instance, genes selected by SAM were shown to be biologically meaningful.

Journal Article↗