PubMed Health⌕ Search

Biomedical subjects

Biao Xing

Publications and source records attributed to Biao Xing.

5 recordsLinked to original sources

A causal inference approach for constructing transcriptional regulatory networks.

MOTIVATION: Transcriptional regulatory networks specify the interactions among regulatory genes and between regulatory genes and their target genes. Discovering transcriptional regulatory networks helps us to understand the underlying mechanism of complex cellular processes and responses. METHOD: This paper describes a causal inference approach for constructing transcriptional regulatory networks using gene expression data, promoter sequences and information on transcription factor (TF) binding sites. The method first identifies active TFs in each individual experiment using a feature selection approach. TFs are viewed as "treatments" and gene expression levels as "responses". For every TF and gene pair, a marginal structural model is built to estimate the causal effect of the TF on the expression level of the gene. The model parameters can be estimated using the G-computation procedure or the IPTW estimator. The P-value associated with the causal parameter in each of these models is used to measure how strongly a TF regulates a gene. These results are further used to infer the overall regulatory network structures. RESULTS: Our analysis of yeast data suggests that the method is capable of identifying significant transcriptional regulatory interactions and the corresponding regulatory networks. AVAILABILITY: The software is under development.

Algorithms↗

A method to estimate the variance of an endpoint from an on-going blinded trial.

Blinded estimation of variance allows for changing the sample size without compromising the integrity of the trial. Some of the methods that estimate the variance in a blinded manner either make untenable assumptions or are only applicable to two-treatment trials. We propose a new method for continuous endpoints that makes minimal assumptions. The method uses the enrollment order of subjects and the randomization block size to estimate the variance. It can be applied to normal or non-normal data, trials with two or more treatments, equal or unequal allocation schemes, fixed or random randomization block sizes, and single or multi-centre trials. The variance estimator is unbiased and performs best when the randomization block size is the smallest. Simulation results suggest that for many commonly used randomization block sizes the proposed estimator is expected to perform well. The proposed method is used to estimate the variance of the endpoint for two trials and is shown to perform well by comparison with its unblinded counterpart.

Analysis of Variance↗

A statistical method for constructing transcriptional regulatory networks using gene expression and sequence data.

Transcriptional regulation is one of the most important means of gene regulation. Uncovering transcriptional regulatory networks helps us to understand the complex cellular process. In this paper, we describe a statistical approach for constructing transcriptional regulatory networks using data of gene expression, promoter sequence, and transcription factor binding sites. Our simulation studies show that the overall and false positive error rates in the estimated transcriptional regulatory networks are expected to be small if the systematic noise in the constructed feature matrix is small. Our analysis based on 658 microarray experiments on yeast gene expression programs and 46 transcription factors suggests that the method is capable of identifying significant transcriptional regulatory interactions and uncovering the corresponding regulatory network structures.

Computational Biology↗

Supervised detection of regulatory motifs in DNA sequences.

Identification of transcription factor binding sites (regulatory motifs) is a major interest in contemporary biology. We propose a new likelihood based method, COMODE, for identifying structural motifs in DNA sequences. Commonly used methods (e.g. MEME, Gibbs motif sampler) model binding sites as families of sequences described by a position weight matrix (PWM) and identify PWMs that maximize the likelihood of observed sequence data under a simple multinomial mixture model. This model assumes that the positions of the PWM correspond to independent multinomial distributions with four cell probabilities. We address supervising the search for DNA binding sites using the information derived from structural characteristics of protein-DNA interactions. We extend the simple multinomial mixture model to a constrained multinomial mixture model by incorporating constraints on the information content profiles or on specific parameters of the motif PWMs. The parameters of this extended model are estimated by maximum likelihood using a nonlinear constraint optimization method. Likelihood-based cross-validation is used to select model parameters such as motif width and constraint type. The performance of COMODE is compared with existing motif detection methods on simulated data that incorporate real motif examples from Saccharomyces cerevisiae. The proposed method is especially effective when the motif of interest appears as a weak signal in the data. Some of the transcription factor binding data of Lee et al. (2002) were also analyzed using COMODE and biologically verified sites were identified.

Journal Article↗

Rank regression in stability analysis.

Stability data are often collected to determine the shelf life of certain characteristics of a pharmaceutical product, for example, a drug's potency over time. Statistical approaches such as the linear regression models are considered as appropriate to analyze the stability data. However, most of these regression models in both theory and practice rely heavily on their underlying parametric assumptions, such as normality of the continuous characteristics or their transformations. In this article, we propose and study some rank-based regression procedures for the stability data when the linear regression models are semiparametric with unspecified error structure. Numerical studies including Monte Carlo simulations and practical example are demonstrated with the proposed procedures as well.

Drug Stability↗