PubMed Health⌕ Search

Biomedical subjects

Marco F Ramoni

Publications and source records attributed to Marco F Ramoni.

18 recordsLinked to original sources

Serum proteome profiling detects myelodysplastic syndromes and identifies CXC chemokine ligands 4 and 7 as markers for advanced disease.

Myelodysplastic syndromes (MDS) are among the most frequent hematologic malignancies. Patients have a short survival and often progress to acute myeloid leukemia. The diagnosis of MDS can be difficult; there is a paucity of molecular markers, and the pathophysiology is largely unknown. Therefore, we conducted a multicenter study investigating whether serum proteome profiling may serve as a noninvasive platform to discover novel molecular markers for MDS. We generated serum proteome profiles from 218 individuals by MS and identified a profile that distinguishes MDS from non-MDS cytopenias in a learning sample set. This profile was validated by testing its ability to predict MDS in a first independent validation set and a second, prospectively collected, independent validation set run 5 months apart. Accuracy was 80.5% in the first and 79.0% in the second validation set. Peptide mass fingerprinting and quadrupole TOF MS identified two differential proteins: CXC chemokine ligands 4 (CXCL4) and 7 (CXCL7), both of which had significantly decreased serum levels in MDS, as confirmed with independent antibody assays. Western blot analyses of platelet lysates for these two platelet-derived molecules revealed a lack of CXCL4 and CXCL7 in MDS. Subtype analyses revealed that these two proteins have decreased serum levels in advanced MDS, suggesting the possibility of a concerted disturbance of transcription or translation of these chemokines in advanced MDS.

Biomarkers↗

GO PaD: the Gene Ontology Partition Database.

Gene Ontology (GO) has been widely used to infer functional significance associated with sets of genes in order to automate discoveries within large-scale genetic studies. A level in GO's direct acyclic graph structure is often assumed to be indicative of its terms' specificities, although other work has suggested this assumption does not hold. Unfortunately, quantitative analysis of biological functions based on nodes at the same level (as is common in gene enrichment analysis tools) can lead to incorrect conclusions as well as missed discoveries due to inefficient use of available information. This paper addresses these using an informational theoretic approach encoded in the GO Partition Database that guarantees to maximize information for gene enrichment analysis. The GO Partition Database was designed to feature ontology partitions with GO terms of similar specificity. The GO partitions comprise varying numbers of nodes and present relevant information theoretic statistics, so researchers can choose to analyze datasets at arbitrary levels of specificity. The GO Partition Database, featuring GO partition sets for functional analysis of genes from human and 10 other commonly studied organisms with a total of 131,972 genes, is available on the internet at: bcl.med.harvard.edu/proj/gopart. The site also includes an online tutorial.

Computer Graphics↗

Regulation of myogenic progenitor proliferation in human fetal skeletal muscle by BMP4 and its antagonist Gremlin.

Skeletal muscle side population (SP) cells are thought to be "stem"-like cells. Despite reports confirming the ability of muscle SP cells to give rise to differentiated progeny in vitro and in vivo, the molecular mechanisms defining their phenotype remain unclear. In this study, gene expression analyses of human fetal skeletal muscle demonstrate that bone morphogenetic protein 4 (BMP4) is highly expressed in SP cells but not in main population (MP) mononuclear muscle-derived cells. Functional studies revealed that BMP4 specifically induces proliferation of BMP receptor 1a-positive MP cells but has no effect on SP cells, which are BMPR1a-negative. In contrast, the BMP4 antagonist Gremlin, specifically up-regulated in MP cells, counteracts the stimulatory effects of BMP4 and inhibits proliferation of BMPR1a-positive muscle cells. In vivo, BMP4-positive cells can be found in the proximity of BMPR1a-positive cells in the interstitial spaces between myofibers. Gremlin is expressed by mature myofibers and interstitial cells, which are separate from BMP4-expressing cells. Together, these studies propose that BMP4 and Gremlin, which are highly expressed by human fetal skeletal muscle SP and MP cells, respectively, are regulators of myogenic progenitor proliferation.

Bone Morphogenetic Protein 4↗

Melanoma cell adhesion molecule is a novel marker for human fetal myogenic cells and affects myoblast fusion.

Myoblast fusion is a highly regulated process that is important during muscle development and myofiber repair and is also likely to play a key role in the incorporation of donor cells in myofibers for cell-based therapy. Although several proteins involved in muscle cell fusion in Drosophila are known, less information is available on the regulation of this process in vertebrates, including humans. To identify proteins that are regulated during fusion of human myoblasts, microarray studies were performed on samples obtained from human fetal skeletal muscle of seven individuals. Primary muscle cells were isolated, expanded, induced to fuse in vitro, and gene expression comparisons were performed between myoblasts and early or late myotubes. Among the regulated genes, melanoma cell adhesion molecule (M-CAM) was found to be significantly downregulated during human fetal muscle cell fusion. M-CAM expression was confirmed on activated myoblasts, both in vitro and in vivo, and on myoendothelial cells (M-CAM(+) CD31(+)), which were positive for the myogenic markers desmin and MyoD. Lastly, in vitro functional studies using M-CAM RNA knockdown demonstrated that inhibition of M-CAM expression enhances myoblast fusion. These studies identify M-CAM as a novel marker for myogenic progenitors in human fetal muscle and confirm that downregulation of this protein promotes myoblast fusion.

Adult↗

A Bayesian dynamic model for influenza surveillance.

The severe acute respiratory syndrome (SARS) epidemic, the growing fear of an influenza pandemic and the recent shortage of flu vaccine highlight the need for surveillance systems able to provide early, quantitative predictions of epidemic events. We use dynamic Bayesian networks to discover the interplay among four data sources that are monitored for influenza surveillance. By integrating these different data sources into a dynamic model, we identify in children and infants presenting to the pediatric emergency department with respiratory syndromes an early indicator of impending influenza morbidity and mortality. Our findings show the importance of modelling the complex dynamics of data collected for influenza surveillance, and suggest that dynamic Bayesian networks could be suitable modelling tools for developing epidemic surveillance systems.

Bayes Theorem↗

SELDI-TOF MS of quadruplicate urine and serum samples to evaluate changes related to storage conditions.

Proteomic profiling with SELDI-TOF MS has facilitated the discovery of disease-specific protein profiles. However, multicenter studies are often hindered by the logistics required for prompt deep-freezing of samples in liquid nitrogen or dry ice within the clinic setting prior to shipping. We report high concordance between MS profiles within sets of quadruplicate split urine and serum samples deep-frozen at 0, 2, 6, and 24 h after sample collection. Gage R&R results confirm that deep-freezing times are not a statistically significant source of SELDI-TOF MS variability for either blood or urine.

Albumins↗

Automation, parallelism, and robotics for proteomics.

The speed of the human genome project (Lander, E. S., Linton, L. M., Birren, B., Nusbaum, C. et al., Nature 2001, 409, 860-921) was made possible, in part, by developments in automation of sequencing technologies. Before these technologies, sequencing was a laborious, expensive, and personnel-intensive task. Similarly, automation and robotics are changing the field of proteomics today. Proteomics is defined as the effort to understand and characterize proteins in the categories of structure, function and interaction (Englbrecht, C. C., Facius, A., Comb. Chem. High Throughput Screen. 2005, 8, 705-715). As such, this field nicely lends itself to automation technologies since these methods often require large economies of scale in order to achieve cost and time-saving benefits. This article describes some of the technologies and methods being applied in proteomics in order to facilitate automation within the field as well as in linking proteomics-based information with other related research areas.

Animals↗

The gene expression profile in refractory periodontitis patients.

BACKGROUND: There are no specific bacterial profiles or diagnostic tests capable of identifying refractory periodontitis patients before a treatment regimen is initiated. Therefore, in this high-risk cohort of patients who do not respond appropriately, host factors that might be partly under genetic control may play a crucial role in their susceptibility. Specifically, we tested the hypothesis that patients with refractory periodontitis have multiple upregulated and/or downregulated genes that might be important in influencing clinical risk. METHODS: Oral subepithelial connective tissues were harvested aseptically from seven refractory periodontitis and seven periodontally well-maintained patients. An RNA isolation kit was used to isolate total RNA from tissue samples that had been stabilized in the RNA stabilizing reagent. The isolated total RNA was then subjected to gene expression profiling using the microarray to measure gene expression levels. The retrieved data were analyzed with a computer program for the differential analysis of gene expression microarray experiments. In addition, real-time polymerase chain reaction (PCR) analysis was performed on selected samples to confirm the microarray data's gene expression patterns. RESULTS: A total of 68 upregulated and six downregulated genes were identified that were differentially expressed at least two-fold out of 22,283 genes we analyzed. The selected model provided a 93% intrinsic validation along with a 93% extrinsic validation. To validate the microarray data, five upregulated genes (lactotransferrin [LTF], matrix metalloproteinase-1 [MMP-1], MMP-3, interferon induced-15 [IFI-15], and Homo sapiens hypothetical protein MGC5566) and two downregulated genes (keratin 2A [KRT2A] and desmocollin-1 [DSC-1]) were randomly selected for further analysis by real-time PCR. The relative RNA expression level of these genes measured by real-time PCR was similar to those measured by microarrays. CONCLUSION: The combined use of microarray technology with the computer program for the differential analysis of gene expression microarray experiments provided a set of candidate genes that may serve as novel therapeutic intervention points and improved diagnostic and screening procedures for high-risk individuals.

Aged↗

Genetic dissection and prognostic modeling of overt stroke in sickle cell anemia.

Sickle cell anemia (SCA) is a paradigmatic single gene disorder caused by homozygosity with respect to a unique mutation at the beta-globin locus. SCA is phenotypically complex, with different clinical courses ranging from early childhood mortality to a virtually unrecognized condition. Overt stroke is a severe complication affecting 6-8% of individuals with SCA. Modifier genes might interact to determine the susceptibility to stroke, but such genes have not yet been identified. Using Bayesian networks, we analyzed 108 SNPs in 39 candidate genes in 1,398 individuals with SCA. We found that 31 SNPs in 12 genes interact with fetal hemoglobin to modulate the risk of stroke. This network of interactions includes three genes in the TGF-beta pathway and SELP, which is associated with stroke in the general population. We validated this model in a different population by predicting the occurrence of stroke in 114 individuals with 98.2% accuracy.

Anemia, Sickle Cell↗

Factors affecting automated syndromic surveillance.

OBJECTIVE: The increased threat of bioterroristic attacks and epidemic events requires the development of accurate and timely outbreak detection systems for early identification of anomalies in public health data. MATERIAL AND METHODS: We propose an automated outbreak detection system based on syndromic data. This system uses an autoregressive model with seasonal components to monitor, online, the daily counts of chief complaints for respiratory syndromes at the emergency department of two major metropolitan hospitals. We evaluate this system by estimating the false positive rate in real data under the assumption that there were no outbreaks of disease, and the true positive rate in real baseline data in which we injected stochastically simulated outbreaks of different shape and size. We then use directed graphical models to account for the effect of exogenous factors on the detection performance of the system. RESULTS: Our study shows that for a week-long outbreak, our model has an overall 84.8% true detection accuracy across all shapes of outbreaks, while the outbreak size influences the earliness to detection. The false and true positive rates are also associated with the exogenous factors and knowledge about these factors can help to improve the detection accuracy. CONCLUSION: This study suggests that the integration of multiple data sources can significantly improve the detection accuracy of syndromic surveillance systems.

Automation↗

Optimization and evaluation of surface-enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF MS) with reversed-phase protein arrays for protein profiling.

Surface-enhanced laser desorption/ionization (SELDI) time-of-flight mass spectrometry with protein arrays has facilitated the discovery of disease-specific protein profiles in serum. Such results raise hopes that protein profiles may become a powerful diagnostic tool. To this end, reliable and reproducible protein profiles need to be generated from many samples, accurate mass peak heights are necessary, and the experimental variation of the profiles must be known. We adapted the entire processing of protein arrays to a robotics system, thus improving the intra-assay coefficients of variation (CVs) from 45.1% to 27.8% (p<0.001). In addition, we assessed up to 16 technical replicates, and demonstrated that analysis of 2-4 replicates significantly increases the reliability of the protein profiles. A recent report on limited long-term reproducibility seemed to concord with our initial inter-assay CVs, which varied widely and reached up to 56.7%. However, we discovered that the inter-assay CV is strongly dependent on the drying time before application of the matrix molecule. Therefore, we devised a standardized drying process and demonstrated that our optimized SELDI procedure generates reliable and long-term reproducible protein profiles with CVs ranging from 25.7% to 32.6%, depending on the signal-to-noise ratio threshold used.

Lasers↗

Gene expression signature with independent prognostic significance in epithelial ovarian cancer.

PURPOSE: Currently available clinical and molecular prognostic factors provide an imperfect assessment of prognosis for patients with epithelial ovarian cancer (EOC). In this study, we investigated whether tumor transcription profiling could be used as a prognostic tool in this disease. METHODS: Tumor tissue from 68 patients was profiled with oligonucleotide microarrays. Samples were randomly split into training and validation sets. A three-step training procedure was used to discover a statistically significant Kaplan-Meier split in the training set. The resultant prognostic signature was then tested on an independent validation set for confirmation. RESULTS: In the training set, a 115-gene signature referred to as the Ovarian Cancer Prognostic Profile (OCPP) was identified. When applied to the validation set, the OCPP distinguished between patients with unfavorable and favorable overall survival (median, 30 months v not yet reached, respectively; log-rank P = .004). The signature maintained independent prognostic value in multivariate analysis, controlling for other known prognostic factors such as age, stage, grade, and debulking status. The hazard ratio for death in the unfavorable OCPP group was 4.8 (P = .021 by Cox proportional hazards analysis). CONCLUSION: The OCPP is an independent prognostic determinant of outcome in EOC. The use of gene profiling may ultimately permit identification of EOC patients appropriate for investigational treatment approaches, based on a low likelihood of achieving prolonged survival with standard first-line platinum-based therapy.

Adult↗

Identification of a transcriptional profile associated with in vitro invasion in non-small cell lung cancer cell lines.

Although much has been learned about basic mechanisms of cell invasion, the genes whose expression is required for this process by malignant cell lines have remained obscure. We assessed invasion through Matrigel using EGF as a chemoattractant and gene expression profiles using oligonucleotide microarrays for 22 non-small cell lung cancer cell lines. The expression of 22 genes were significantly correlated (p < 0.001) with the measured invasion index. Cluster analysis demonstrated that gene expression profiles classify the cell lines into low and high invasive subgroups. Considering invasiveness as a dichotomous variable, Bayesian analysis was used to identify genes that have the highest probability of being differentially expressed between the high and low invasion groups. This analysis identified 16 genes whose expression was associated with invasiveness. "Leave one out" cross validation was 91% accurate. Nine genes were identified in both correlation and Bayesian analyses. Seven of the nine genes were negatively associated with invasion and four of those genes are plasma membrane proteins. The two genes with the highest inverse association with invasion, TACSTD1 and CLDN3, are involved with cell adhesion and cell-cell interactions, respectively. Interestingly, the gene with the highest positive association with invasion, SERPINE1 (PAI-1), is a protease inhibitor. These and the other genes identified by both analyses represent targets for further study to assess their importance in non-small cell lung cancer invasion and metastasis.

Adenocarcinoma↗

Bayesian approach to discovering pathogenic SNPs in conserved protein domains.

The success rate of association studies can be improved by selecting better genetic markers for genotyping or by providing better leads for identifying pathogenic single nucleotide polymorphisms (SNPs) in the regions of linkage disequilibrium with positive disease associations. We have developed a novel algorithm to predict pathogenic single amino acid changes, either nonsynonymous SNPs (nsSNPs) or missense mutations, in conserved protein domains. Using a Bayesian framework, we found that the probability of a microbial missense mutation causing a significant change in phenotype depended on how much difference it made in several phylogenetic, biochemical, and structural features related to the single amino acid substitution. We tested our model on pathogenic allelic variants (missense mutations or nsSNPs) included in OMIM, and on the other nsSNPs in the same genes (from dbSNP) as the nonpathogenic variants. As a result, our model predicted pathogenic variants with a 10% false-positive rate. The high specificity of our prediction algorithm should make it valuable in genetic association studies aimed at identifying pathogenic SNPs.

Algorithms↗

Robust transmission/disequilibrium test for incomplete family genotypes.

Several solutions have been proposed to extend the transmission disequilibrium test (TDT) to include cases with missing parental genotype. However, completion of the missing parental genotype may bias the test if the underlying missing data mechanism is informative. Furthermore, all these solutions resolve the problem of missing parental genotype, while offspring with missing genotypes are typically ignored. We propose here an extension to the TDT, called robust TDT (rTDT), able to handle incomplete genotypes on both parents and children and that does not rest on any assumption about the missing data mechanism. rTDT returns minimum and maximum values of TDT that are consistent with all the possible completions of the missing data. We also show that, in some situations, rTDT can achieve both greater power and greater significance than the popular TDT analysis of incomplete data. rTDT is applied to a database of markers of susceptibility to Crohn's disease and it shows that only 2 of the 11 markers originally associated with the phenotype do not depend on assumptions about the missing data mechanism.

Crohn Disease↗

Expression profiling and identification of novel genes involved in myogenic differentiation.

Skeletal muscle differentiation is a complex, highly coordinated process that relies on precise temporal gene expression patterns. To better understand this cascade of transcriptional events, we used expression profiling to analyze gene expression in a 12-day time course of differentiating C2C12 myoblasts. Cluster analysis specific for time-ordered microarray experiments classified 2895 genes and ESTs with variable expression levels between proliferating and differentiating cells into 22 clusters with distinct expression patterns during myogenesis. Expression patterns for several known and novel genes were independently confirmed by real-time quantitative RT-PCR and/or Western blotting and immunofluorescence. MyoD and MEF family members exhibited unique expression kinetics that were highly coordinated with cell-cycle withdrawal regulators. Among genes with peak expression levels during cell cycle withdrawal were Vcam1, Itgb3, Itga5, Vcl, as well as Ptger4, a gene not previously associated with the process of myogenesis. One interesting uncharacterized transcript that is highly induced during myogenesis encodes several immunoglobulin repeats with sequence similarity to titin, a large sarcomeric protein. These data sets identify many additional uncharacterized transcripts that may play important functions in muscle cell proliferation and differentiation and provide a baseline for comparison with C2C12 cells expressing various mutant genes involved in myopathic disorders.

Animals↗

Minimal haplotype tagging.

The high frequency of single-nucleotide polymorphisms (SNPs) in the human genome presents an unparalleled opportunity to track down the genetic basis of common diseases. At the same time, the sheer number of SNPs also makes unfeasible genome-wide disease association studies. The haplotypic nature of the human genome, however, lends itself to the selection of a parsimonious set of SNPs, called haplotype tagging SNPs (htSNPs), able to distinguish the haplotypic variations in a population. Current approaches rely on statistical analysis of transmission rates to identify htSNPs. In contrast to these approximate methods, this contribution describes an exact, analytical, and lossless method, called BEST (Best Enumeration of SNP Tags), able to identify the minimum set of SNPs tagging an arbitrary set of haplotypes from either pedigree or independent samples. Our results confirm that a small proportion of SNPs is sufficient to capture the haplotypic variations in a population and that this proportion decreases exponentially as the haplotype length increases. We used BEST to tag the haplotypes of 105 genes in an African-American and a European-American sample. An interesting finding of this analysis is that the vast majority (95%) of the htSNPs in the European-American sample is a subset of the htSNPs of the African-American sample. This result seems to provide further evidence that a severe bottleneck occurred during the founding of Europe and the conjectured "Out of Africa" event.

Algorithms↗

Cluster analysis of gene expression dynamics.

This article presents a Bayesian method for model-based clustering of gene expression dynamics. The method represents gene-expression dynamics as autoregressive equations and uses an agglomerative procedure to search for the most probable set of clusters given the available data. The main contributions of this approach are the ability to take into account the dynamic nature of gene expression time series during clustering and a principled way to identify the number of distinct clusters. As the number of possible clustering models grows exponentially with the number of observed time series, we have devised a distance-based heuristic search procedure able to render the search process feasible. In this way, the method retains the important visualization capability of traditional distance-based clustering and acquires an independent, principled measure to decide when two series are different enough to belong to different clusters. The reliance of this method on an explicit statistical representation of gene expression dynamics makes it possible to use standard statistical techniques to assess the goodness of fit of the resulting model and validate the underlying assumptions. A set of gene-expression time series, collected to study the response of human fibroblasts to serum, is used to identify the properties of the method.

Bayes Theorem↗