PubMed Health⌕ Search

Biomedical subjects

E E Schadt

Publications and source records attributed to E E Schadt.

9 recordsLinked to original sources

Novel integrative genomics strategies to identify genes for complex traits.

Forward genetics is a common approach to dissecting complex traits like common human diseases. The ultimate aim of this approach was the identification of genes that are causal for disease or other phenotypes of interest. However, the forward genetics approach is by definition restricted to the identification of genes that have incurred mutations over the course of evolution or that incurred mutations as a result of chemical mutagenesis, and that as a result lead to disease or to variations in other phenotypes of interest. Genes that harbour no such mutations, but that play key roles in parts of the biological network that lead to disease, are systematically missed by this class of approaches. Recently, a class of novel integrative genomics approaches has been devised to elucidate the complexity of common human diseases by intersecting genotypic, molecular profiling, and clinical data in segregating populations. These novel approaches take a more holistic view of biological systems and leverage the vast network of gene-gene interactions, in combination with DNA variation data, to establish causal relationships among molecular profiling traits and between molecular profiling and disease (or other classic phenotypes). A number of novel genes for disease phenotypes have been identified as a result of these approaches, highlighting the utility of integrating orthogonal sources of data to get at the underlying causes of disease.

Animals↗

Genetic inheritance of gene expression in human cell lines.

Combining genetic inheritance information, for both molecular profiles and complex traits, is a promising strategy not only for detecting quantitative trait loci (QTLs) for complex traits but for understanding which genes, pathways, and biological processes are also under the influence of a given QTL. As a primary step in determining the feasibility of such an approach in humans, we present the largest survey to date, to our knowledge, of the heritability of gene-expression traits in segregating human populations. In particular, we measured expression for 23,499 genes in lymphoblastoid cell lines for members of 15 Centre d'Etude du Polymorphisme Humain (CEPH) families. Of the total set of genes, 2,340 were found to be expressed, of which 31% had significant heritability when a false-discovery rate of 0.05 was used. QTLs were detected for 33 genes on the basis of at least one P value <.000005. Of these, 13 genes possessed a QTL within 5 Mb of their physical location. Hierarchical clustering was performed on the basis of both Pearson correlation of gene expression and genetic correlation. Both reflected biologically relevant activity taking place in the lymphoblastoid cell lines, with greater coherency represented in Kyoto Encyclopedia of Genes and Genomes database (KEGG) pathways than in Gene Ontology database pathways. However, more pathway coherence was observed in KEGG pathways when clustering was based on genetic correlation than when clustering was based on Pearson correlation. As more expression data in segregating populations are generated, viewing clusters or networks based on genetic correlation measures and shared QTLs will offer potentially novel insights into the relationship among genes that may underlie complex traits.

Cell Line↗

An integrative genomics approach to the reconstruction of gene networks in segregating populations.

The reconstruction of genetic networks in mammalian systems is one of the primary goals in biological research, especially as such reconstructions relate to elucidating not only common, polygenic human diseases, but living systems more generally. Here we propose a novel gene network reconstruction algorithm, derived from classic Bayesian network methods, that utilizes naturally occurring genetic variations as a source of perturbations to elucidate the network. This algorithm incorporates relative transcript abundance and genotypic data from segregating populations by employing a generalized scoring function of maximum likelihood commonly used in Bayesian network reconstruction problems. The utility of this novel algorithm is demonstrated via application to liver gene expression data from a segregating mouse population. We demonstrate that the network derived from these data using our novel network reconstruction algorithm is able to capture causal associations between genes that result in increased predictive power, compared to more classically reconstructed networks derived from the same data.

11-beta-Hydroxysteroid Dehydrogenases↗

A new paradigm for drug discovery: integrating clinical, genetic, genomic and molecular phenotype data to identify drug targets.

Application of statistical genetics approaches to variations in mRNA transcript abundances in segregating populations can be used to identify genes and pathways associated with common human diseases. The combination of this genetic information with gene expression and clinical trait data can also be used to identify subtypes of a disease and the genetic loci specific to each subtype. Here we highlight results from some of our recent work in this area and further explore the many possibilities that exist in employing a more comprehensive genetics and functional genomics approach to the functional annotation of genomes, and in applying such methods to the validation of targets for complex traits in the drug discovery process.

Animals↗

Experimental annotation of the human genome using microarray technology.

The most important product of the sequencing of a genome is a complete, accurate catalogue of genes and their products, primarily messenger RNA transcripts and their cognate proteins. Such a catalogue cannot be constructed by computational annotation alone; it requires experimental validation on a genome scale. Using 'exon' and 'tiling' arrays fabricated by ink-jet oligonucleotide synthesis, we devised an experimental approach to validate and refine computational gene predictions and define full-length transcripts on the basis of co-regulated expression of their exons. These methods can provide more accurate gene numbers and allow the detection of mRNA splice variants and identification of the tissue- and disease-specific conditions under which genes are expressed. We apply our technique to chromosome 22q under 69 experimental condition pairs, and to the entire human genome under two experimental conditions. We discuss implications for more comprehensive, consistent and reliable genome annotation, more efficient, full-length complementary DNA cloning strategies and application to complex diseases.

Algorithms↗

Feature extraction and normalization algorithms for high-density oligonucleotide gene expression array data.

Algorithms for performing feature extraction and normalization on high-density oligonucleotide gene expression arrays, have not been fully explored, and the impact these algorithms have on the downstream analysis is not well understood. Advances in such low-level analysis methods are essential to increase the sensitivity and specificity of detecting whether genes are present and/or differentially expressed. We have developed and implemented a number of algorithms for the analysis of expression array data in a software application, the DNA-Chip Analyzer (dChip). In this report, we describe the algorithms for feature extraction and normalization, and present validation data and comparison results with some of the algorithms currently in use.

Algorithms↗

Analyzing high-density oligonucleotide gene expression array data.

We have developed methods and identified problems associated with the analysis of data generated by high-density, oligonuceotide gene expression arrays. Our methods are aimed at accounting for many of the sources of variation that make it difficult, at times, to realize consistent results. We present here descriptions of some of these methods and how they impact the analysis of oligonucleotide gene expression array data. We will discuss the process of recognizing the "spots" (or features) on the Affymetrix GeneChip(R) probe arrays, correcting for background and intensity gradients in the resulting images, scaling/normalizing an array to allow array-to-array comparisons, monitoring probe performance with respect to hybridization efficiency, and assessing whether a gene is present or differentially expressed. Examples from the analyses of gene expression validation data are presented to contrast the different methods applied to these types of data.

Base Sequence↗

Computational advances in maximum likelihood methods for molecular phylogeny.

We have developed a generalization of Kimura's Markov chain model for base substitution at a single nucleotide site. This generalized model incorporates more flexible transition rates and consequently allows irreversible as well as reversible chains. Because the model embodies just the right amount of symmetry, it permits explicit calculation of finite-time transition probabilities and equilibrium distributions. The model also meshes well with maximum likelihood methods for phylogenetic analysis. Quick calculation of likelihoods and their derivatives can be carried out by adapting Baum's forward and backward algorithms from the theory of hidden Markov chains. Analysis of HIV sequence data illustrates the speed of the algorithms on trees with many contemporary taxa. Analysis of some of Lake's data on the origin of the eukaryotic nucleus contrasts the reversible and irreversible versions of the model.

Algorithms↗