PubMed Health⌕ Search

Biomedical subjects

Isaac S Kohane

Publications and source records attributed to Isaac S Kohane.

18 recordsLinked to original sources

Joint, multifaceted genomic analysis enables diagnosis of diverse, ultra-rare monogenic presentations.

Genomics for rare disease diagnosis has advanced at a rapid pace due to our ability to perform in-depth analyses on individual patients with ultra-rare diseases. The increasing sizes of ultra-rare disease cohorts internationally newly enables cohort-wide analyses for new discoveries, but well-calibrated statistical genetics approaches for jointly analyzing these patients are still under development. The Undiagnosed Diseases Network (UDN) brings multiple clinical, research and experimental centers under the same umbrella across the United States to facilitate and scale case-based diagnostic analyses. Here, we present the first joint analysis of whole genome sequencing data of UDN patients across the network. We introduce new, well-calibrated statistical methods for prioritizing disease genes with de novo recurrence and compound heterozygosity. We also detect pathways enriched with candidate and known diagnostic genes. Our computational analysis, coupled with a systematic clinical review, recapitulated known diagnoses and revealed new disease associations. We further release a software package, RaMeDiES, enabling automated cross-analysis of deidentified sequenced cohorts for new diagnostic and research discoveries. Gene-level findings and variant-level information across the cohort are available in a public-facing browser ( https://dbmi-bgm.github.io/udn-browser/ ). These results show that case-level diagnostic efforts should be supplemented by a joint genomic analysis across cohorts.

Humans↗

Minimal haplotype tagging.

The high frequency of single-nucleotide polymorphisms (SNPs) in the human genome presents an unparalleled opportunity to track down the genetic basis of common diseases. At the same time, the sheer number of SNPs also makes unfeasible genome-wide disease association studies. The haplotypic nature of the human genome, however, lends itself to the selection of a parsimonious set of SNPs, called haplotype tagging SNPs (htSNPs), able to distinguish the haplotypic variations in a population. Current approaches rely on statistical analysis of transmission rates to identify htSNPs. In contrast to these approximate methods, this contribution describes an exact, analytical, and lossless method, called BEST (Best Enumeration of SNP Tags), able to identify the minimum set of SNPs tagging an arbitrary set of haplotypes from either pedigree or independent samples. Our results confirm that a small proportion of SNPs is sufficient to capture the haplotypic variations in a population and that this proportion decreases exponentially as the haplotype length increases. We used BEST to tag the haplotypes of 105 genes in an African-American and a European-American sample. An interesting finding of this analysis is that the vast majority (95%) of the htSNPs in the European-American sample is a subset of the htSNPs of the African-American sample. This result seems to provide further evidence that a severe bottleneck occurred during the founding of Europe and the conjectured "Out of Africa" event.

Algorithms↗

Reproducibility of gene expression across generations of Affymetrix microarrays.

BACKGROUND: The development of large-scale gene expression profiling technologies is rapidly changing the norms of biological investigation. But the rapid pace of change itself presents challenges. Commercial microarrays are regularly modified to incorporate new genes and improved target sequences. Although the ability to compare datasets across generations is crucial for any long-term research project, to date no means to allow such comparisons have been developed. In this study the reproducibility of gene expression levels across two generations of Affymetrix GeneChips (HuGeneFL and HG-U95A) was measured. RESULTS: Correlation coefficients were computed for gene expression values across chip generations based on different measures of similarity. Comparing the absolute calls assigned to the individual probe sets across the generations found them to be largely unchanged. CONCLUSION: We show that experimental replicates are highly reproducible, but that reproducibility across generations depends on the degree of similarity of the probe sets and the expression level of the corresponding transcript.

Calibration↗

Gene expression profiling of Duchenne muscular dystrophy skeletal muscle.

The primary cause of Duchenne muscular dystrophy (DMD) is a mutation in the dystrophin gene, leading to absence of the corresponding protein, disruption of the dystrophin-associated protein complex, and substantial changes in skeletal muscle pathology. Although the primary defect is known and the histological pathology well documented, the underlying molecular pathways remain in question. To clarify these pathways, we used expression microarrays to compare individual gene expression profiles for skeletal muscle biopsies from DMD patients and unaffected controls. We have previously published expression data for the 12,500 known genes and full-length expressed sequence tags (ESTs) on the Affymetrix HG-U95Av2 chips. Here we present comparative expression analysis of the 50,000 EST clusters represented on the remainder of the Affymetrix HG-U95 set. Individual expression profiles were generated for biopsies from 10 DMD patients and 10 unaffected control patients. Two methods of statistical analysis were used to interpret the resulting data (t-test analysis to determine the statistical significance of differential expression and geometric fold change analysis to determine the extent of differential expression). These analyses identified 183 probe sets (59 of which represent known genes) that differ significantly in expression level between unaffected and disease muscle. This study adds to our knowledge of the molecular pathways that are altered in the dystrophic state. In particular, it suggests that signaling pathways might be substantially involved in the disease process. It also highlights a large number of unknown genes whose expression is altered and whose identity therefore becomes important in understanding the pathogenesis of muscular dystrophy.

Biopsy↗

PGAGENE: integrating quantitative gene-specific results from the NHLBI programs for genomic applications.

SUMMARY: PGAGENE is a web-based gene-specific genomic data search engine, which allows users to search over 5.9 million pieces of collective genetic and genomic data from the NHLBI supported Programs for Genomic Applications. This data includes microarray measurements, SNPs, and mutations, and data may be found using symbols, parts of gene names or products, Affymetrix probe IDs, GenBank accession numbers, UniGene IDs, dbSNP IDs, and others. The PGAGENE indexing agent periodically maps all publicly available gene-specific PGA data onto LocusLink using dynamically generated cross-referencing tables.

Base Sequence↗

Expression profiling reveals altered satellite cell numbers and glycolytic enzyme transcription in nemaline myopathy muscle.

The nemaline myopathies (NMs) are a clinically and genetically heterogeneous group of disorders characterized by nemaline rods and skeletal muscle weakness. Mutations in five sarcomeric thin filament genes have been identified. However, the molecular consequences of these mutations are unknown. Using Affymetrix oligonucleotide microarrays, we have analyzed the expression patterns of >21,000 genes and expressed sequence tags in skeletal muscles of 12 NM patients and 21 controls. Multiple complementary approaches were used for data analysis, including geometric fold analysis, two-tailed unequal variance t test, hierarchical clustering, relevance network, and nearest-neighbor analysis. We report the identification of high satellite cell populations in NM and the significant down-regulation of transcripts for key enzymes of glucose and glycogen metabolism as well as a possible regulator of fatty acid metabolism, UCP3. Interestingly, transcript level changes of multiple genes suggest possible changes in Ca(2+) homeostasis. The increased expression of multiple structural proteins was consistent with increased fibrosis. This comprehensive study of downstream molecular consequences of NM gene mutations provides insights in the cellular events leading to the NM phenotype.

Adult↗

The value of parental report for diagnosis and management of dehydration in the emergency department.

STUDY OBJECTIVES: We define the predictive value of parents' computer-based report for history and physical signs of dehydration for a primary outcome of percentage of dehydration (fluid deficit) and 2 secondary outcomes: clinically important acidosis and hospital admission. We also sought to compare the reports of physical signs related to dehydration made by parents and nurses. METHODS: We performed a prospective observational trial in an urban pediatric emergency department. A convenience sample of parents completed a computer-based interview covering historical details and physical signs (ill appearance, sunken fontanelle, sunken eyes, decreased tears, dry mouth, cool extremities, and weak cry) related to dehydration. Nurses independently completed an assessment of physical signs for enrolled children. The primary outcome was the degree of dehydration (fluid deficit), which was defined as the percentage difference between initial ED weight and stable final weight after the illness. Secondary outcomes included clinically important acidosis (defined as a serum CO(2) value of </=15 mEq/L) and hospital admission. RESULTS: One hundred thirty-two parent-child dyads comprised the final sample. Parent-reported data manifested higher sensitivity (range 73% to 100%) than specificity (range 0% to 49%) for the prediction of dehydration of 5% or greater. Likelihood ratios (LRs) near zero (<0.1) suggest that a normal history of fluid intake and urine output reduced the likelihood of significant dehydration. Parental report of a normal tearing state reduced the likelihood of significant dehydration and clinically important acidosis (negative LRs of 0.4 and 0.1, respectively). Two physical signs reported by parents, sunken fontanelle and decreased tears, were associated with hospital admission (positive LR of 3.4 and 4.0, respectively). CONCLUSION: Parents' report of history and observations for children captured through computer-based interview demonstrates predictive value for relevant outcomes in dehydration.

Adult↗

Gene homology resources on the World Wide Web.

As the amount of information available to biologists increases exponentially, data analysis becomes progressively more challenging. Sequence homology has been a traditional tool in the researchers' armamentarium; it is a very versatile instrument and can be employed to assist in numerous tasks, from establishing the function of a gene to determination of the evolutionary development of an organism. Consequently, numerous specialized tools have been established in the public domain (most commonly, the World Wide Web) to help investigators use sequence homology in their research. These homology databases differ both in techniques they use to compare sequences as well as in the size of the unit of analysis, which can be the whole gene, a domain, or a motif. In this paper, we aim to present a systematic review of the inner details of the most commonly used databases as well as to offer guidelines for their use.

Animals↗

Gene expression comparison of biopsies from Duchenne muscular dystrophy (DMD) and normal skeletal muscle.

The primary cause of Duchenne muscular dystrophy (DMD) is a mutation in the dystrophin gene leading to the absence of the corresponding RNA transcript and protein. Absence of dystrophin leads to disruption of the dystrophin-associated protein complex and substantial changes in skeletal muscle pathology. Although the histological pathology of dystrophic tissue has been well documented, the underlying molecular pathways remain poorly understood. To examine the pathogenic pathways and identify new or modifying factors involved in muscular dystrophy, expression microarrays were used to compare individual gene expression profiles of skeletal muscle biopsies from 12 DMD patients and 12 unaffected control patients. Two separate statistical analysis methods were used to interpret the resulting data: t test analysis to determine the statistical significance of differential expression and geometric fold change analysis to determine the extent of differential expression. These analyses identified 105 genes that differ significantly in expression level between unaffected and DMD muscle. Many of the differentially expressed genes reflect changes in histological pathology. For instance, immune response signals and extracellular matrix genes are overexpressed in DMD muscle, an indication of the infiltration of inflammatory cells and connective tissue. Significantly more genes are overexpressed than are underexpressed in dystrophic muscle, with dystrophin underexpressed, whereas other genes encoding muscle structure and regeneration processes are overexpressed, reflecting the regenerative nature of the disease.

Adult↗

Cluster analysis of gene expression dynamics.

This article presents a Bayesian method for model-based clustering of gene expression dynamics. The method represents gene-expression dynamics as autoregressive equations and uses an agglomerative procedure to search for the most probable set of clusters given the available data. The main contributions of this approach are the ability to take into account the dynamic nature of gene expression time series during clustering and a principled way to identify the number of distinct clusters. As the number of possible clustering models grows exponentially with the number of observed time series, we have devised a distance-based heuristic search procedure able to render the search process feasible. In this way, the method retains the important visualization capability of traditional distance-based clustering and acquires an independent, principled measure to decide when two series are different enough to belong to different clusters. The reliance of this method on an explicit statistical representation of gene expression dynamics makes it possible to use standard statistical techniques to assess the goodness of fit of the resulting model and validate the underlying assumptions. A set of gene-expression time series, collected to study the response of human fibroblasts to serum, is used to identify the properties of the method.

Bayes Theorem↗

Visualization and evaluation of clusters for exploratory analysis of gene expression data.

Clustering algorithms have been shown to be useful to explore large-scale gene expression profiles. Visualization and objective evaluation of clusters are two important considerations when users are selecting different clustering algorithms, but they are often overlooked. The developments of a framework and software tools that implement comprehensive data visualization and objective measures of cluster quality are crucial. In this paper, we describe a theoretical framework and formalizations for consistently developing clustering algorithms. A new clustering algorithm was developed within the proposed framework. We demonstrate that a theoretically sound principle can be uniformly applied to the developments of cluster-optimization function, comprehensive data-visualization strategy, and objective cluster-evaluation measures as well as actual implementation of the principle. Cluster consistency and quality measures of the algorithm are rigorously evaluated against those of popular clustering algorithms for gene expression data analysis (K-means and self-organizing maps), in four data sets, yielding promising results.

Algorithms↗

Comparing expression profiles of genes with similar promoter regions.

MOTIVATION: Gene regulatory elements are often predicted by seeking common sequences in the promoter regions of genes that are clustered together based on their expression profiles. We consider the problem in the opposite direction: we seek to find the genes that have similar promoter regions and determine the extent to which these genes have similar expression profiles. RESULTS: We use the data sets from experiments on Saccharomyces cerevisiae. Our similarity measure for the promoter regions is based on the set of common mapped or putative transcription factor binding sites and other regulatory elements in the upstream region of the genes, as contained in the Saccharomyces cerevisiae Promoter Database. We pair up the genes with high similarity scores and compare their expression levels in time-course experiment data. We find that genes with similar promoter regions on the average have significantly higher correlation, but it can vary widely depending on the genes. This confirms that the presence of similar regulatory elements often does not correspond to similarity in expression profiles and indicates that finding transcription factor binding sites or other regulatory elements starting with the expression patterns may be limited in many cases. Regardless of the correlation, the degree to which the profiles agree under different experimental conditions can be examined to derive hypotheses concerning the role of common regulatory elements. Overall, we find that considering the relationship between the promoter regions and the expression profiles starting with the regulatory elements is a difficult but useful process that can provide valuable insights.

Databases, Nucleic Acid↗

Analysis of matched mRNA measurements from two different microarray technologies.

MOTIVATION: [corrected] The existence of several technologies for measuring gene expression makes the question of cross-technology agreement of measurements an important issue. Cross-platform utilization of data from different technologies has the potential to reduce the need to duplicate experiments but requires corresponding measurements to be comparable. METHODS: A comparison of mRNA measurements of 2895 sequence-matched genes in 56 cell lines from the standard panel of 60 cancer cell lines from the National Cancer Institute (NCI 60) was carried out by calculating correlation between matched measurements and calculating concordance between cluster from two high-throughput DNA microarray technologies, Stanford type cDNA microarrays and Affymetrix oligonucleotide microarrays. RESULTS: In general, corresponding measurements from the two platforms showed poor correlation. Clusters of genes and cell lines were discordant between the two technologies, suggesting that relative intra-technology relationships were not preserved. GC-content, sequence length, average signal intensity, and an estimator of cross-hybridization were found to be associated with the degree of correlation. This suggests gene-specific, or more correctly probe-specific, factors influencing measurements differently in the two platforms, implying a poor prognosis for a broad utilization of gene expression measurements across platforms.

Cluster Analysis↗

Linking gene expression data with patient survival times using partial least squares.

There is an increasing need to link the large amount of genotypic data, gathered using microarrays for example, with various phenotypic data from patients. The classification problem in which gene expression data serve as predictors and a class label phenotype as the binary outcome variable has been examined extensively, but there has been less emphasis in dealing with other types of phenotypic data. In particular, patient survival times with censoring are often not used directly as a response variable due to the complications that arise from censoring. We show that the issues involving censored data can be circumvented by reformulating the problem as a standard Poisson regression problem. The procedure for solving the transformed problem is a combination of two approaches: partial least squares, a regression technique that is especially effective when there is severe collinearity due to a large number of predictors, and generalized linear regression, which extends standard linear regression to deal with various types of response variables. The linear combinations of the original variables identified by the method are highly correlated with the patient survival times and at the same time account for the variability in the covariates. The algorithm is fast, as it does not involve any matrix decompositions in the iterations. We apply our method to data sets from lung carcinoma and diffuse large B-cell lymphoma studies to verify its effectiveness.

Algorithms↗

Newborn screening program practices in the United States: notification, research, and consent.

OBJECTIVE: To define current practice among US newborn screening programs for notification of results, research, and consenting procedures. METHODS: A telephone survey of all US newborn screening program supervisors. RESULTS: All 51 programs participated. All states reported abnormal results to the infant's physician, and some also reported to the hospital and parents. Cases with abnormal results were tracked to different endpoints but usually (92.1%) at least until a follow-up appointment was made. A total of 66.6% of programs can communicate with programs in other states; 9.8% enable families to suppress reporting of results to the infant's physician. No state has a mechanism for parents to prevent results from entering the medical record. Parents or physicians who request results are often authenticated by providing their name (52.9%). Many programs (45.1%) report only to physicians and require just their name (43.5%), an identification number (17.4%), a letter (26.1%), or a parent's signature (26.1%). A total of 70.6% retain residual blood samples; of these, only 8.3% store them completely devoid of patient identifiers. A total of 49.0% of programs aggregate data for research. In 16.0% of these, the data are publicly available. In 24.0%, researchers obtain approval at their own institution; in 24.0%, researchers obtain approval through the state laboratory Institutional Review Board. In 74.5% of programs, parents are notified but not asked for consent before collection of the sample; 19.6% neither notify parents nor obtain consent before screening. CONCLUSIONS: There is wide variation in practice among the US newborn screening programs. Because the programs collectively manage a comprehensive nationwide genomic databank, careful consideration of how information technology and high-throughput genomic analysis are used will be essential to allow progress in clinical care, public health, and research while protecting individual privacy.

Communicable Disease Control↗

Accessing genomic data through XML-based remote procedure calls.

As the amount of data in public genomic databases grows, interoperability among them is becoming an increasingly critical feature. The ability for automated systems to mine and integrate data will be crucial to extracting knowledge from sources of data whose volume far exceeds the capabilities of human researchers. The currently dominant paradigm of presenting information as Web pages and using hyperlinks to describe relationships between pieces of information favors usability, but makes interoperability and automated data exchange more difficult. In this paper we describe how SNPper, a web-based system for the retrieval and analysis of Single Nucleotide Polymorphisms (SNPs), was augmented with a Remote Procedure Call interface, allowing client applications to query our program for SNP data and to receive the response as an XML document. Data represented in this form can be easily parsed by the requesting program, and thus reused for other applications. In this paper we describe the implementation of the interface and we show examples of its usage in a number of existing applications.

Databases, Genetic↗

An unsupervised self-optimizing gene clustering algorithm.

We have devised a gene-clustering algorithm that is completely unsupervised in that no parameters need be set by the user, and the clustering of genes is self-optimizing to yield the set of clusters that minimizes within-cluster distance and maximizes between-cluster distance. This algorithm was implemented in Java, and tested on a randomly selected 200-gene subset of 3000 genes from cell-cycle data in S. cerevisiae. AlignACE was used to evaluate the resulting optimized cluster set for upstream cis-regulons. The optimized cluster set was found to be of comparable quality to cluster sets obtained by two established methods (complete linkage and k-means), even when provided with only a small, randomly selected subset of the data (200 vs 3000 genes), and with absolutely no supervision. MAP and specificity scores of the highest ranking motifs identified in the largest clusters were comparable.

Algorithms↗

The contributions of biomedical informatics to the fight against bioterrorism.

A comprehensive and timely response to current and future bioterrorist attacks requires a data acquisition, threat detection, and response infrastructure with unprecedented scope in time and space. Fortunately, biomedical informaticians have developed and implemented architectures, methodologies, and tools at the local and the regional levels that can be immediately pressed into service for the protection of our populations from these attacks. These unique contributions of the discipline of biomedical informatics are reviewed here.

Bioterrorism↗