PubMed Health⌕ Search

Biomedical subjects

Erliang Zeng

Publications and source records attributed to Erliang Zeng.

4 recordsLinked to original sources

Clustering genes using gene expression and text literature data.

Clustering of gene expression data is a standard technique used to identify closely related genes. In this paper, we develop a new clustering algorithm, MSC (Multi-Source Clustering), to perform exploratory analysis using two or more diverse sources of data. In particular, we investigate the problem of improving the clustering by integrating information obtained from gene expression data with knowledge extracted from biomedical text literature. In each iteration of algorithm MSC, an EM-type procedure is employed to bootstrap the model obtained from one data source by starting with the cluster assignments obtained in the previous iteration using the other data sources. Upon convergence, the two individual models are used to construct the final cluster assignment. We compare the results of algorithm MSC for two data sources with the results obtained when the clustering is applied on the two sources of data separately. We also compare it with that obtained using the feature level integration method that performs the clustering after simply concatenating the features obtained from the two data sources. We show that the z-scores of the clustering results from MSC are better than that from the other methods. To evaluate our clusters better, function enrichment results are presented using terms from the Gene Ontology database. Finally, by investigating the success of motif detection programs that use the clusters, we show that our approach integrating gene expression data and text data reveals clusters that are biologically more meaningful than those identified using gene expression data alone.

Artificial Intelligence↗

Detection of rifampin-resistant Mycobacterium tuberculosis strains by using a specialized oligonucleotide microarray.

DNA microarray represents one of the major advances in diagnostic sequencing of polymerase chain reaction (PCR) products. Until now, arrays have been relatively expensive, complex to perform, and difficult to interpret, limiting their wide application in the clinical laboratory. A moderate-density oligonucleotide microarray that can rapidly identify Mycobacterium tuberculosis rifampin-resistant strains was developed. The method is based on the detection of point mutations and other rearrangements in the rpoB gene region determining rifampin resistance. Rifampin resistance was determined by hybridizing fluorescently labeled, amplified genetic material generated from bacterial colonies to the array. Fifty-three rifampin-resistant M. tuberculosis and 15 rifampin-susceptible M. tuberculosis were tested and results were concordant with those based on culture drug susceptibility testing and sequencing. Rifampin-resistant clinical isolates were detected in as little as 1.5 hours after PCR amplification with visual results. It is demonstrated that oligonucleotide microarray is an efficient, specialized technique to implement and can be used as a rapid method for detecting rifampin resistance to complement standard culture-based method.

Antibiotics, Antitubercular↗

Determining a detectable threshold of signal intensity in cDNA microarray based on accumulated distribution.

In microarray data mining, one of the key problems is how to handle weak signals. Based on a bent piecewise linear accumulated distribution generally found in the microarray data, a new detectable threshold finding method is proposed to filter genes with unreliable information in this paper. More reliable and reproducible data is produced for the subsequent data mining.

DNA, Complementary↗

Mutations in the rpoB gene of multidrug-resistant Mycobacterium tuberculosis isolates from China.

Mutations in the 81-bp rifampin resistance determining region (RRDR) and mutation V176F locating at the beginning of the ropB gene were analyzed by DNA sequencing of 86 Mycobacterium tuberculosis clinical isolates (72 resistant and 14 sensitive) from different parts of China. Sixty-five mutations of 22 distinct kinds, 21 point mutations, and 1 insertion were found in 65 of 72 resistant isolates. The most common mutations were in codons 531 (41%), 526 (40%), and 516 (4%). Mutations were not found in seven (10%) of the resistant isolates. Six new alleles within the RRDR, along with five novel mutations outside the RRDR, are reported. None of isolates contained the V176 mutation.

Alleles↗