PubMed Health⌕ Search

Biomedical subjects

Yinglei Lai

Publications and source records attributed to Yinglei Lai.

10 recordsLinked to original sources

A statistical method for estimating the proportion of differentially expressed genes.

Microarrays have been widely used to identify differentially expressed genes. One related problem is to estimate the proportion of differentially expressed genes. For some complex diseases, the amount of differentially expressed genes may be relatively small and these genes may only have subtly differential expressions. For these microarray data, it is generally difficult to efficiently estimate the proportion of differentially expressed genes. In this study, I propose a likelihood-based method coupled with an expectation-maximization (E-M) algorithm for estimating the proportion of differentially expressed genes. The proposed method has favorable performances if either (i) the P values of differentially expressed genes are homogeneously distributed or (ii) the proportion of differentially expressed genes is relatively small. In both of these situations, I showed through simulations that the proposed method gave satisfactory performances when it was compared to other existing methods. As applications, these methods were applied to two microarray gene expression data sets generated from different platforms.

Algorithms↗

List of lists-annotated (LOLA): a database for annotation and comparison of published microarray gene lists.

Microarray profiling of RNA expression is a powerful tool that generates large lists of transcripts that are potentially relevant to a disease or treatment. However, because the lists of changed transcripts are embedded in figures and tables, they are typically inaccessible for search engines. Due to differences in gene nomenclatures, the lists are difficult to compare between studies. LOLA (Lists of Lists Annotated) is an internet-based database for comparing gene lists from microarray studies or other genomic-scale methods. It serves as a common platform to compare and reannotate heterogeneous gene lists from different microarray platforms or different genomic methodologies such as serial analysis of gene expression (SAGE) or proteomics. LOLA () provides researchers with a means to store, annotate, and compare gene lists produced from different studies or different analyses of the same study. It is especially useful in identifying potentially "high interest" genes which are reported as significant across multiple studies and species. Its application to the fields of stem cell, cancer, and aging research is demonstrated by comparing published papers.

Animals↗

Serum protein markers for early detection of ovarian cancer.

Early diagnosis of epithelial ovarian cancer (EOC) would significantly decrease the morbidity and mortality from this disease but is difficult in the absence of physical symptoms. Here, we report a blood test, based on the simultaneous quantization of four analytes (leptin, prolactin, osteopontin, and insulin-like growth factor-II), that can discriminate between disease-free and EOC patients, including patients diagnosed with stage I and II disease, with high efficiency (95%). Microarray analysis was used initially to determine the levels of 169 proteins in serum from 28 healthy women, 18 women newly diagnosed with EOC, and 40 women with recurrent disease. Evaluation of proteins that showed significant differences in expression between controls and cancer patients by ELISA assays yielded the four analytes. These four proteins then were evaluated in a blind cross-validation study by using an additional 106 healthy females and 100 patients with EOC (24 stage I/II and 76 stage III/IV). Upon sample decoding, the results were analyzed by using three different classification algorithms and a binary code methodology. The four-analyte test was further validated in a blind binary code study by using 40 additional serum samples from normal and EOC cancer patients. No single protein could completely distinguish the cancer group from the healthy controls. However, the combination of the four analytes exhibited the following: sensitivity 95%, positive predictive value (PPV) 95%, specificity 95%, and negative predictive value (NPV) 94%, a considerable improvement on current methodology.

Aged↗

A statistical method to detect chromosomal regions with DNA copy number alterations using SNP-array-based CGH data.

Single nucleotide polymorphism (SNP) arrays were used to detect chromosomal regions with DNA copy number alterations. Current statistical methods for microarray-based comparative genomic hybridization (array-CGH) analysis generally assume certain relationships among adjacent markers on the same chromosome, and these assumptions may be questionable. For an SNP-array-based CGH study, multiple normal reference SNP arrays were collected. In order to utilize these normal reference SNP arrays, we derived an empirical distribution of signal ratios for each SNP marker. With an assumed threshold value for the overall error rate control and the defined signal ratio ranges for chromosomal amplification and deletion, we proposed a procedure to identify chromosomal alteration regions based on several bootstrapped one-sample t-tests and the false discovery rate control. When we have multiple arrays for different individuals with the same disease, our method can also be used to detect SNP markers for chromosomal alteration regions that are common among these individuals. We applied our method to a published SNP array data set for breast carcinoma cell lines. For an individual with breast cancer, numerous chromosomal alteration regions were identified. Compared to results of previous studies, our method identified more chromosomal alteration regions, with some being implicated in the literature to harbor genes associated with breast cancer. For multiple cancer arrays, our results suggested the existence of common chromosomal alteration regions. However, a high proportion of false positives also indicated that genetic variations among different individuals with breast cancer can be present.

Breast Neoplasms↗

A statistical method for identifying differential gene-gene co-expression patterns.

MOTIVATION: To understand cancer etiology, it is important to explore molecular changes in cellular processes from normal state to cancerous state. Because genes interact with each other during cellular processes, carcinogenesis related genes may form differential co-expression patterns with other genes in different cell states. In this study, we develop a statistical method for identifying differential gene-gene co-expression patterns in different cell states. RESULTS: For efficient pattern recognition, we extend the traditional F-statistic and obtain an Expected Conditional F-statistic (ECF-statistic), which incorporates statistical information of location and correlation. We also propose a statistical method for data transformation. Our approach is applied to a microarray gene expression dataset for prostate cancer study. For a gene of interest, our method can select other genes that have differential gene-gene co-expression patterns with this gene in different cell states. The 10 most frequently selected genes, include hepsin, GSTP1 and AMACR, which have recently been proposed to be associated with prostate carcinogenesis. However, genes GSTP1 and AMACR cannot be identified by studying differential gene expression alone. By using tumor suppressor genes TP53, PTEN and RB1, we identify seven genes that also include hepsin, GSTP1 and AMACR. We show that genes associated with cancer may have differential gene-gene expression patterns with many other genes in different cell states. By discovering such patterns, we may be able to identify carcinogenesis related genes.

Algorithms↗

Sampling distribution for microsatellites amplified by PCR: mean field approximation and its applications to genotyping.

Due to microsatellite mutations during PCR, stutter patterns may appear in the final PCR product, which hinder us from accurate genotyping microsatellite markers. The existing methods for microsatellite stutter pattern deconvolution required large amount of data. A mathematical model for microsatellite mutations during PCR and an estimation method based on mean field approximation for branching processes have recently been developed. In this paper, we study the asymptotic behaviors for mean field approximation when experiments are started from a large number of molecules, and we derive an upper bound for the approximation error when experiments are started from a finite number of molecules. Based on the theories of mean field approximation and Bayesian statistics, we develop a novel method for microsatellite stutter pattern deconvolution.

Gene Amplification↗

Microsatellite mutations during the polymerase chain reaction: mean field approximations and their applications.

We develop a novel mathematical model for microsatellite mutations during polymerase chain reaction (PCR). Based on the model, we study the first- and second-order moments of the number of repeat units in a randomly chosen molecule after n PCR cycles and their corresponding mean field approximations. We give upper bounds for the approximation errors and show that the approximation errors are small when the mutation rate is low. Based on the theoretical results, we develop a moment estimation method to estimate the mutation rate per-repeat-unit per PCR cycle and the probability of expansion when mutations occur. Simulation studies show that the moment estimation method can accurately recover the true mutation rate and probability of expansion. Finally, the method is applied to experimental data from single-molecule PCR experiments.

Humans↗

The relationship between microsatellite slippage mutation rate and the number of repeat units.

Microsatellite markers are widely used for genetic studies, but the relationship between microsatellite slippage mutation rate and the number of repeat units remains unclear. In this study, microsatellite distributions in the human genome are collected from public sequence databases. We observe that there is a threshold size for slippage mutations. We consider a model of microsatellite mutation consisting of point mutations and single stepwise slippage mutations. From two sets of equations based on two stochastic processes and equilibrium assumptions, we estimate microsatellite slippage mutation rates without assuming any relationship between microsatellite slippage mutation rate and the number of repeat units. We use the least squares method with constraints to estimate expansion and contraction mutation rates. The estimated slippage mutation rate increases exponentially as the number of repeat units increases. When slippage mutations happen, expansion occurs more frequently for short microsatellites and contraction occurs more frequently for long microsatellites. Our results agree with the length-dependent mutation pattern observed from experimental data, and they explain the scarcity of long microsatellites.

Algorithms↗

Taq DNA polymerase slippage mutation rates measured by PCR and quasi-likelihood analysis: (CA/GT)n and (A/T)n microsatellites.

During microsatellite polymerase chain reaction (PCR), insertion-deletion mutations produce stutter products differing from the original template by multiples of the repeat unit length. We analyzed the PCR slippage products of (CA)n and (A)n tracts cloned in a pUC18 vector. Repeat numbers varied from two to 14 (CA)n and four to 12 (A)n. Data was generated on approximately 10 single molecules for each clone type using two rounds of nested PCR. The size and peak areas of the products were obtained by capillary electrophoresis. A quasi- likelihood approach to the analysis of the data estimated the mutation rate/repeat/PCR cycle. The rate for (CA)n tracts was 3.6 x 10(-3) with contractions 14 times greater than expansions. For (A)n tracts the rate was 1.5 x 10(-2) and contractions outnumbered expansions by 5-fold. The threshold for detecting 'stutter' products was computed to be four repeats for (CA)n and eight repeats for (A)n or approximately 8 bp in both cases. A comparison was made between the computationally and experimentally derived threshold values. The threshold and expansion to contraction ratios are explained on the basis of the active site structure of Taq DNA polymerase and models of the energetics of slippage events, respectively.

Base Sequence↗

The mutation process of microsatellites during the polymerase chain reaction.

We build a mathematical model for the mutation process of microsatellites during polymerase chain reaction (PCR) using the theory of branching processes. Based on the model, we develop a method to estimate the mutation rate of microsatellites per PCR cycle and the probability of expansion by maximizing a quasi-likelihood of the observed data. We show by simulations that the proposed estimation method can accurately recover the relationship between the mutation rate and number of repeat units. The theoretical basis for the proposed method is also given. We apply the method to experimental data on poly-A and poly-CA repeats.

Animals↗