PubMed Health⌕ Search

Biomedical subjects

Roger Perkins

Publications and source records attributed to Roger Perkins.

15 recordsLinked to original sources

GOFFA: gene ontology for functional analysis--a FDA gene ontology tool for analysis of genomic and proteomic data.

BACKGROUND: Gene Ontology (GO) characterizes and categorizes the functions of genes and their products according to biological processes, molecular functions and cellular components, facilitating interpretation of data from high-throughput genomics and proteomics technologies. The most effective use of GO information is achieved when its rich and hierarchical complexity is retained and the information is distilled to the biological functions that are most germane to the phenomenon being investigated. RESULTS: Here we present a FDA GO tool named Gene Ontology for Functional Analysis (GOFFA). GOFFA first ranks GO terms in the order of prevalence for a list of selected genes or proteins, and then it allows the user to interactively select GO terms according to their significance and specific biological complexity within the hierarchical structure. GOFFA provides five interactive functions (Tree view, Terms View, Genes View, GO Path and GO TreePrune) to analyze the GO data. Among the five functions, GO Path and GO TreePrune are unique. The GO Path simultaneously displays the ranks that order GOFFA Tree Paths based on statistical analysis. The GO TreePrune provides a visual display of a reduced GO term set based on a user's statistical cut-offs. Therefore, the GOFFA visual display can provide an intuitive depiction of the most likely relevant biological functions. CONCLUSION: With GOFFA, the user can dynamically interact with the GO data to interpret gene expression results in the context of biological plausibility, which can lead to new discoveries or identify new hypotheses. AVAILABILITY: GOFFA is available through ArrayTrack softwarehttp://edkb.fda.gov/webstart/arraytrack/.

Genomics↗

Gene expression profile exploration of a large dataset on chronic fatigue syndrome.

OBJECTIVE: To gain understanding of the molecular basis of chronic fatigue syndrome (CFS) through gene expression analysis using a large microarray data set in conjunction with clinically administrated questionnaires. METHOD: Data from the Wichita (KS, USA) CFS Surveillance Study was used, comprising 167 participants with two self-report questionnaires (multidimensional fatigue inventory [MFI] and Zung depression scale [Zung]), microarray data, empiric classification, and others. Microarray data was analyzed using bioinformatics tools from ArrayTrack. RESULTS: Correspondence analysis was applied to the MFI questionnaire to select the 23 samples having either the most or the least fatigue, and to the Zung questionnaire to select the 26 samples having either the most or least depression; ten samples were common, resulting in a total of 39 samples. The MFI and Zung-based CFS/non-CFS (NF) classifications on the 39 samples were consistent with the empiric classification. Two differentially-expressed gene lists were determined, 188 fatigue-related genes and 164 depression-related genes, which shared 24 common genes and involved 11 common pathways. Principal component analysis based on 24 genes clearly separates 39 samples with respect to their likelihood to be CFS. Most of the 24 genes are not previously reported for CFS, yet their functions are consistent with the prevailing model of CFS, such as immune response, apoptosis, ion channel activity, signal transduction, cell-cell signaling, regulation of cell growth and neuronal activity. Hierarchical cluster analysis was performed based on 24 genes to classify 128 (=167-39) unassigned samples. Several of the 11 identified common pathways are supported by earlier findings for CFS, such as cytokine-cytokine receptor interaction and neuroactive ligand-receptor interaction. Importantly, most of the 11 common pathways are interrelated, suggesting complex biological mechanisms associated with CFS. CONCLUSION: Bioinformatics is critical in this study to select definitive sample groups, analyze gene expression data and gain insight into biological mechanisms. The 24 identified common genes and 11 common pathways could be important in future studies of CFS at the molecular level.

Adult↗

Decision forest analysis of 61 single nucleotide polymorphisms in a case-control study of esophageal cancer; a novel method.

BACKGROUND: Systematic evaluation and study of single nucleotide polymorphisms (SNPs) made possible by high throughput genotyping technologies and bioinformatics promises to provide breakthroughs in the understanding of complex diseases. Understanding how the millions of SNPs in the human genome are involved in conferring susceptibility or resistance to disease, or in rendering a drug efficacious or toxic in the individual is a major goal of the relatively new fields of pharmacogenomics. Esophageal squamous cell carcinoma is a high-mortality cancer with complex etiology and progression involving both genetic and environmental factors. We examined the association between esophageal cancer risk and patterns of 61 SNPs in a case-control study for a population from Shanxi Province in North Central China that has among the highest rates of esophageal squamous cell carcinoma in the world. METHODS: High-throughput Masscode mass spectrometry genotyping was done on genomic DNA from 574 individuals (394 cases and 180 age-frequency matched controls). SNPs were chosen from among genes involving DNA repair enzymes, and Phase I and Phase II enzymes. We developed a novel adaptation of the Decision Forest pattern recognition method named Decision Forest for SNPs (DF-SNPs). The method was designated to analyze the SNP data. RESULTS: The classifier in separating the cases from the controls developed with DF-SNPs gave concordance, sensitivity and specificity, of 94.7%, 99.0% and 85.1%, respectively; suggesting its usefulness for hypothesizing what SNPs or combinations of SNPs could be involved in susceptibility to esophageal cancer. Importantly, the DF-SNPs algorithm incorporated a randomization test for assessing the relevance (or importance) of individual SNPs, SNP types (Homozygous common, heterozygous and homozygous variant) and patterns of SNP types (SNP patterns) that differentiate cases from controls. For example, we found that the different genotypes of SNP GADD45B E1122 are all associated with cancer risk. CONCLUSION: The DF-SNPs method can be used to differentiate esophageal squamous cell carcinoma cases from controls based on individual SNPs, SNP types and SNP patterns. The method could be useful to identify potential biomarkers from the SNP data and complement existing methods for genotype analyses.

Algorithms↗

Quality control and quality assessment of data from surface-enhanced laser desorption/ionization (SELDI) time-of flight (TOF) mass spectrometry (MS).

BACKGROUND: Proteomic profiling of complex biological mixtures by the ProteinChip technology of surface-enhanced laser desorption/ionization time-of-flight (SELDI-TOF) mass spectrometry (MS) is one of the most promising approaches in toxicological, biological, and clinic research. The reliable identification of protein expression patterns and associated protein biomarkers that differentiate disease from health or that distinguish different stages of a disease depends on developing methods for assessing the quality of SELDI-TOF mass spectra. The use of SELDI data for biomarker identification requires application of rigorous procedures to detect and discard low quality spectra prior to data analysis. RESULTS: The systematic variability from plates, chips, and spot positions in SELDI experiments was evaluated using biological and technical replicates. Systematic biases on plates, chips, and spots were not found. The reproducibility of SELDI experiments was demonstrated by examining the resulting low coefficient of variances of five peaks presented in all 144 spectra from quality control samples that were loaded randomly on different spots in the chips of six bioprocessor plates. We developed a method to detect and discard low quality spectra prior to proteomic profiling data analysis, which uses a correlation matrix to measure the similarities among SELDI mass spectra obtained from similar biological samples. Application of the correlation matrix to our SELDI data for liver cancer and liver toxicity study and myeloma-associated lytic bone disease study confirmed this approach as an efficient and reliable method for detecting low quality spectra. CONCLUSION: This report provides evidence that systematic variability between plates, chips, and spots on which the samples were assayed using SELDI based proteomic procedures did not exist. The reproducibility of experiments in our studies was demonstrated to be acceptable and the profiling data for subsequent data analysis are reliable. Correlation matrix was developed as a quality control tool to detect and discard low quality spectra prior to data analysis. It proved to be a reliable method to measure the similarities among SELDI mass spectra and can be used for quality control to decrease noise in proteomic profiling data prior to data analysis.

Female↗

Development of public toxicogenomics software for microarray data management and analysis.

A robust bioinformatics capability is widely acknowledged as central to realizing the promises of toxicogenomics. Successful application of toxicogenomic approaches, such as DNA microarray, inextricably relies on appropriate data management, the ability to extract knowledge from massive amounts of data and the availability of functional information for data interpretation. At the FDA's National Center for Toxicological Research (NCTR), we are developing a public microarray data management and analysis software, called ArrayTrack. ArrayTrack is Minimum Information About a Microarray Experiment (MIAME) supportive for storing both microarray data and experiment parameters associated with a toxicogenomics study. A quality control mechanism is implemented to assure the fidelity of entered expression data. ArrayTrack also provides a rich collection of functional information about genes, proteins and pathways drawn from various public biological databases for facilitating data interpretation. In addition, several data analysis and visualization tools are available with ArrayTrack, and more tools will be available in the next released version. Importantly, gene expression data, functional information and analysis methods are fully integrated so that the data analysis and interpretation process is simplified and enhanced. ArrayTrack is publicly available online and the prospective user can also request a local installation version by contacting the authors.

Databases, Genetic↗

Multiclass Decision Forest--a novel pattern recognition method for multiclass classification in microarray data analysis.

The wealth of knowledge imbedded in gene expression data from DNA microarrays portends rapid advances in both research and clinic. Turning the prodigious and noisy data into knowledge is a challenge to the field of bioinformatics, and development of classifiers using supervised learning techniques is the primary methodological approach for clinical application using gene expression data. In this paper, we present a novel classification method, multiclass Decision Forest (DF), that is the direct extension of the two-class DF previously developed in our lab. Central to DF is the synergistic combining of multiple heterogenic but comparable decision trees to reach a more accurate and robust classification model. The computationally inexpensive multiclass DF algorithm integrates gene selection and model development, and thus eliminates the bias of gene preselection in crossvalidation. Importantly, the method provides several statistical means for assessment of prediction accuracy, prediction confidence, and diagnostic capability. We demonstrate the method by application to gene expression data for 83 small round blue-cell tumors (SRBCTs) samples belonging to one of four different classes. Based on 500 runs of 10-fold crossvalidation, tumor prediction accuracy was approximately 97%, sensitivity was approximately 95%, diagnostic sensitivity was approximately 91%, and diagnostic accuracy was approximately 99.5%. Among 25 genes selected to distinguish tumor class, 12 have functional information in the literature implicating their involvement in cancer. The four types of SRBCTs samples are also distinguishable in a clustering analysis based on the expression profiles of these 25 genes. The results demonstrated that the multiclass DF is an effective classification method for analysis of gene expression data for the purpose of molecular diagnostics.

Carcinoma, Small Cell↗

Using decision forest to classify prostate cancer samples on the basis of SELDI-TOF MS data: assessing chance correlation and prediction confidence.

Class prediction using "omics" data is playing an increasing role in toxicogenomics, diagnosis/prognosis, and risk assessment. These data are usually noisy and represented by relatively few samples and a very large number of predictor variables (e.g., genes of DNA microarray data or m/z peaks of mass spectrometry data). These characteristics manifest the importance of assessing potential random correlation and overfitting of noise for a classification model based on omics data. We present a novel classification method, decision forest (DF), for class prediction using omics data. DF combines the results of multiple heterogeneous but comparable decision tree (DT) models to produce a consensus prediction. The method is less prone to overfitting of noise and chance correlation. A DF model was developed to predict presence of prostate cancer using a proteomic data set generated from surface-enhanced laser deposition/ionization time-of-flight mass spectrometry (SELDI-TOF MS). The degree of chance correlation and prediction confidence of the model was rigorously assessed by extensive cross-validation and randomization testing. Comparison of model prediction with imposed random correlation demonstrated biologic relevance of the model and the reduction of overfitting in DF. Furthermore, two confidence levels (high and low confidences) were assigned to each prediction, where most misclassifications were associated with the low-confidence region. For the high-confidence prediction, the model achieved 99.2% sensitivity and 98.2% specificity. The model also identified a list of significant peaks that could be useful for biomarker identification. DF should be equally applicable to other omics data such as gene expression data or metabolomic data. The DF algorithm is available upon request.

Decision Support Techniques↗

Assessment of prediction confidence and domain extrapolation of two structure-activity relationship models for predicting estrogen receptor binding activity.

Quantitative structure-activity relationship (QSAR) methods have been widely applied in drug discovery, lead optimization, toxicity prediction, and regulatory decisions. Despite major advances in algorithms and software, QSAR models have inherent limitations associated with a size and chemical-structure diversity of the training set, experimental error, and many characteristics of structure representation and correlation algorithms. Whereas excellent fit to the training data may be readily attainable, often models fail to predict accurately chemicals that are outside their domain of applicability. A QSAR's utility and, in the case of regulatory decisions, justification for usage increasingly depend on the ability to quantify a model's potential for predicting unknown chemicals with some known degree of certainty. It is never possible to predict an unknown chemical with absolute certainty. Here we report on two QSAR models based on different data sets for classification of chemicals according to their ability to bind to the estrogen receptor. The models were developed by using a novel QSAR method, Decision Forest, which combines the results of multiple heterogeneous but comparable Decision Tree models to produce a consensus prediction. We used an extensive cross-validation process to define an applicability domain for model predictions based on two quantitative measures: prediction confidence and domain extrapolation. Together, these measures quantify the accuracy of each prediction within and outside of the training domain. Despite being based on large and diverse training sets, both QSAR models had poor accuracy for chemicals within the domain of low confidence, whereas good accuracy was obtained for those within the domain of high confidence. For prediction in the high confidence domain, accuracy was inversely proportional to the degree of domain extrapolation. The model with a larger training set of 1,092, compared with 232 for the other, was more accurate in predicting chemicals at larger domain extrapolation, and could be particularly useful for rapidly prioritizing potential endocrine disruptors from large chemical universe.

Animals↗

Study of 202 natural, synthetic, and environmental chemicals for binding to the androgen receptor.

A number of environmental and industrial chemicals are reported to possess androgenic or antiandrogenic activities. These androgenic endocrine disrupting chemicals may disrupt the endocrine system of humans and wildlife by mimicking or antagonizing the functions of natural hormones. The present study developed a low cost recombinant androgen receptor (AR) competitive binding assay that uses no animals. We validated the assay by comparing the protocols and results from other similar assays, such as the binding assay using prostate cytosol. We tested 202 natural, synthetic, and environmental chemicals that encompass a broad range of structural classes, including steroids, diethylstilbestrol and related chemicals, antiestrogens, flutamide derivatives, bisphenol A derivatives, alkylphenols, parabens, alkyloxyphenols, phthalates, siloxanes, phytoestrogens, DDTs, PCBs, pesticides, organophosphate insecticides, and other chemicals. Some of these chemicals are environmentally persistent and/or commercially important, but their AR binding affinities have not been previously reported. To the best of our knowledge, these results represent the largest and most diverse data set publicly available for chemical binding to the AR. Through a careful structure-activity relationship (SAR) examination of the data set in conjunction with knowledge of the recently reported ligand-AR crystal structures, we are able to define the general structural requirements for chemical binding to AR. Hydrophobic interactions are important for AR binding. The interaction between ligand and AR at the 3- and 17-positions of testosterone and R1881 found in other chemical classes are discussed in depth. The SAR studies of ligand binding characteristics for AR are compared to our previously reported results for estrogen receptor binding.

Androgen Receptor Antagonists↗

ArrayTrack--supporting toxicogenomic research at the U.S. Food and Drug Administration National Center for Toxicological Research.

The mapping of the human genome and the determination of corresponding gene functions, pathways, and biological mechanisms are driving the emergence of the new research fields of toxicogenomics and systems toxicology. Many technological advances such as microarrays are enabling this paradigm shift that indicates an unprecedented advancement in the methods of understanding the expression of toxicity at the molecular level. At the National Center for Toxicological Research (NCTR) of the U.S. Food and Drug Administration, core facilities for genomic, proteomic, and metabonomic technologies have been established that use standardized experimental procedures to support centerwide toxicogenomic research. Collectively, these facilities are continuously producing an unprecedented volume of data. NCTR plans to develop a toxicoinformatics integrated system (TIS) for the purpose of fully integrating genomic, proteomic, and metabonomic data with the data in public repositories as well as conventional (Italic)in vitro(/Italic) and (Italic)in vivo(/Italic) toxicology data. The TIS will enable data curation in accordance with standard ontology and provide or interface a rich collection of tools for data analysis and knowledge mining. In this article the design, practical issues, and functions of the TIS are discussed through presenting its prototype version, ArrayTrack, for the management and analysis of DNA microarray data. ArrayTrack is logically constructed of three linked components: a) a library (LIB) that mirrors critical data in public databases; b) a database (MicroarrayDB) that stores microarray experiment information that is Minimal Information About a Microarray Experiment (MIAME) compliant; and c) tools (TOOL) that operate on experimental and public data for knowledge discovery. Using ArrayTrack, we can select an analysis method from the TOOL and apply the method to selected microarray data stored in the MicroarrayDB; the analysis results can be linked directly to gene information in the LIB.

Databases, Factual↗

Quantitative structure-activity relationship methods: perspectives on drug discovery and toxicology.

Quantitative structure-activity relationships (QSARs) attempt to correlate chemical structure with activity using statistical approaches. The QSAR models are useful for various purposes including the prediction of activities of untested chemicals. Quantitative structure-activity relationships and other related approaches have attracted broad scientific interest, particularly in the pharmaceutical industry for drug discovery and in toxicology and environmental science for risk assessment. An assortment of new QSAR methods have been developed during the past decade, most of them focused on drug discovery. Besides advancing our fundamental knowledge of QSARs, these scientific efforts have stimulated their application in a wider range of disciplines, such as toxicology, where QSARs have not yet gained full appreciation. In this review, we attempt to summarize the status of QSAR with emphasis on illuminating the utility and limitations of QSAR technology. We will first review two-dimensional (2D) QSAR with a discussion of the availability and appropriate selection of molecular descriptors. We will then proceed to describe three-dimensional (3D) QSAR and key issues associated with this technology, then compare the relative suitability of 2D and 3D QSAR for different applications. Given the recent technological advances in biological research for rapid identification of drug targets, we mention several examples in which QSAR approaches are employed in conjunction with improved knowledge of the structure and function of the target receptor. The review will conclude by discussing statistical validation of QSAR models, a topic that has received sparse attention in recent years despite its critical importance.

Drug Design↗

Structure-activity relationship approaches and applications.

New techniques and software have enabled ubiquitous use of structure-activity relationships (SARs) in the pharmaceutical industry and toxicological sciences. We review the status of SAR technology by using examples to underscore the advances as well as the unique technical challenges. Applying SAR involves two steps: Characterization of the chemicals under investigation, and application of chemometric approaches to explore data patterns or to establish the relationships between structure and activity. We describe generally but not exhaustively the SAR methodologies popular use in toxicology, including representation of chemical structure, and chemometric techniques where models are both unsupervised and supervised. The utility of SAR technology is most evident when supervised methods are used to predict toxicity of untested chemicals based only on chemical structure. Such models can predict on both an ordinal scale (e.g., active vs inactive) or a continuouis scale (e.g., median lethal dose [LD50] dose). The reader is also referred to a companion paper in this issue that discusses quantitative structure-activity relationship (QSAR) methods that have advanced markedly over the past decade.

Forecasting↗

Prediction of estrogen receptor binding for 58,000 chemicals using an integrated system of a tree-based model with structural alerts.

A number of environmental chemicals, by mimicking natural hormones, can disrupt endocrine function in experimental animals, wildlife, and humans. These chemicals, called "endocrine-disrupting chemicals" (EDCs), are such a scientific and public concern that screening and testing 58,000 chemicals for EDC activities is now statutorily mandated. Computational chemistry tools are important to biologists because they identify chemicals most important for in vitro and in vivo studies. Here we used a computational approach with integration of two rejection filters, a tree-based model, and three structural alerts to predict and prioritize estrogen receptor (ER) ligands. The models were developed using data for 232 structurally diverse chemicals (training set) with a 10(6) range of relative binding affinities (RBAs); we then validated the models by predicting ER RBAs for 463 chemicals that had ER activity data (testing set). The integrated model gave a lower false negative rate than any single component for both training and testing sets. When the integrated model was applied to approximately 58,000 potential EDCs, 80% (approximately 46,000 chemicals) were predicted to have negligible potential (log RBA < -4.5, with log RBA = 2.0 for estradiol) to bind ER. The ability to process large numbers of chemicals to predict inactivity for ER binding and to categorically prioritize the remainder provides one biologic measure to prioritize chemicals for entry into more expensive assays (most chemicals have no biologic data of any kind). The general approach for predicting ER binding reported here may be applied to other receptors and/or reversible binding mechanisms involved in endocrine disruption.

Animals↗

Decision forest: combining the predictions of multiple independent decision tree models.

The techniques of combining the results of multiple classification models to produce a single prediction have been investigated for many years. In earlier applications, the multiple models to be combined were developed by altering the training set. The use of these so-called resampling techniques, however, poses the risk of reducing predictivity of the individual models to be combined and/or over fitting the noise in the data, which might result in poorer prediction of the composite model than the individual models. In this paper, we suggest a novel approach, named Decision Forest, that combines multiple Decision Tree models. Each Decision Tree model is developed using a unique set of descriptors. When models of similar predictive quality are combined using the Decision Forest method, quality compared to the individual models is consistently and significantly improved in both training and testing steps. An example will be presented for prediction of binding affinity of 232 chemicals to the estrogen receptor.

Algorithms↗