PubMed Health⌕ Search

Biomedical subjects

Chiara Romualdi

Publications and source records attributed to Chiara Romualdi.

11 recordsLinked to original sources

MIDAW: a web tool for statistical analysis of microarray data.

MIDAW (microarray data analysis web tool) is a web interface integrating a series of statistical algorithms that can be used for processing and interpretation of microarray data. MIDAW consists of two main sections: data normalization and data analysis. In the normalization phase the simultaneous processing of several experiments with background correction, global and local mean and variance normalization are carried out. The data analysis section allows graphical display of expression data for descriptive purposes, estimation of missing values, reduction of data dimension, discriminant analysis and identification of marker genes. The statistical results are organized in dynamic web pages and tables, where the transcript/gene probes contained in a specific microarray platform can be linked (according to user choice) to external databases (GenBank, Entrez Gene, UniGene). Tutorial files help the user throughout the statistical analysis to ensure that the forms are filled out correctly. MIDAW has been developed using Perl and PHP and it uses R/Bioconductor languages and routines. MIDAW is GPL licensed and freely accessible at http://muscle.cribi.unipd.it/midaw/. Perl and PHP source codes are available from the authors upon request.

Algorithms↗

Novel genes, possibly relevant for molecular diagnosis or therapy of human rhabdomyosarcoma, detected by genomic expression profiling.

Transcriptional profiles of an alveolar rhabdomyosarcoma (RMS) and of a RMS cell line were reconstructed by a computational and statistical approach. Expression data of 29,963 genes in 11 adult human healthy tissues and in 37 tumour tissues were analysed for comparison. We identified 202 genes differentially expressed in at least one RMS sample, as compared with normal skeletal muscle. Among them, 107 resulted specifically overexpressed in RMS, but in no tumour affecting other tissues. Cluster analysis applied to expression data detected a series of genes presumably co-expressed with genes encoding known tumour markers and/or reportedly involved in genesis or development of rhabdomyosarcoma. This study succeeded in identifying a number of genes, which become candidates for in vitro study, thus facilitating discovery of novel tumour markers or targets for drug therapy.

Adult↗

A leukemia-enriched cDNA microarray platform identifies new transcripts with relevance to the biology of pediatric acute lymphoblastic leukemia.

BACKGROUND AND OBJECTIVES: Microarray gene expression profiling has been widely applied to characterize hematologic malignancies, has attributed a molecular signature to leukemia subclasses and has allowed new subclasses to be distinguished. We set out to use microarray technology to identify novel genes relevant for leukemogenesis. To this end we used a unique leukemia-enriched cDNA microarray platform. DESIGN AND METHODS: The systematic sequencing of cDNA libraries of normal and leukemic bone marrow allowed us to increase the number of genes to yield a new release of a previously generated cDNA microarray. Using this platform we analyzed the expression profiles of 4,670 genes in bone marrow samples from 18 pediatric patients with acute lymphoblastic leukemia (ALL). RESULTS: Expression profiling consistently distinguished the leukemia patients into three groups, those with T-ALL, B-ALL and B-ALL with MLL/AF4 rearrangement, in agreement with the clinical classification. Our platform identified 30 genes that best discriminate these three subtypes. Using mini-array technology these 30 genes were validated in another cohort of 17 patients. In particular we identified two novel genes not previously reported: endomucin (EMCN) and ubiquitin specific protease 33 (USP33) that appear to be over-expressed in B-ALL relative to their expression in T-ALL. INTERPRETATION AND CONCLUSIONS: Microarray technology not only allows the distinction between disease subclasses but also offers a chance to identify new genes involved in leukemogenesis. Our approach of using a unique platform has proven to be fruitful in identifying new genes and we suggest exploration of other malignancies using this approach.

Adolescent↗

RAP: a new computer program for de novo identification of repeated sequences in whole genomes.

MOTIVATION: DNA repeats are a common feature of most genomic sequences. Their de novo identification is still difficult despite being a crucial step in genomic analysis and oligonucleotides design. Several efficient algorithms based on word counting are available, but too short words decrease specificity while long words decrease sensitivity, particularly in degenerated repeats. RESULTS: The Repeat Analysis Program (RAP) is based on a new word-counting algorithm optimized for high resolution repeat identification using gapped words. Many different overlapping gapped words can be counted at the same genomic position, thus producing a better signal than the single ungapped word. This results in better specificity both in terms of low-frequency detection, being able to identify sequences repeated only once, and highly divergent detection, producing a generally high score in most intron sequences. AVAILABILITY: The program is freely available for non-profit organizations, upon request to the authors. CONTACT: giorgio.valle@unipd.it SUPPLEMENTARY INFORMATION: The program has been tested on the Caenorhabditis elegans genome using word lengths of 12, 14 and 16 bases. The full analysis has been implemented in the UCSC Genome Browser and is accessible at http://genome.cribi.unipd.it.

Algorithms↗

Improved detection of differentially expressed genes in microarray experiments through multiple scanning and image integration.

The variability of results in microarray technology is in part due to the fact that independent scans of a single hybridised microarray give spot images that are not quite the same. To solve this problem and turn it to our advantage, we introduced the approach of multiple scanning and of image integration of microarrays. To this end, we have developed specific software that creates a virtual image that statistically summarises a series of consecutive scans of a microarray. We provide evidence that the use of multiple imaging (i) enhances the detection of differentially expressed genes; (ii) increases the image homogeneity; and (iii) reveals false-positive results such as differentially expressed genes that are detected by a single scan but not confirmed by successive scanning replicates. The increase in the final number of differentially expressed genes detected in a microarray experiment with this approach is remarkable; 50% more for microarrays hybridised with targets labelled by reverse transcriptase, and 200% more for microarrays developed with the tyramide signal amplification (TSA) technique. The results have been confirmed by semi-quantitative RT-PCR tests.

False Negative Reactions↗

Pattern recognition in gene expression profiling using DNA array: a comparative study of different statistical methods applied to cancer classification.

Large-scale parallel measurements of the expression of many thousands genes are now available with high-density array made with collections of cDNA fragments, or oligonucleotide corresponding to different transcripts. These technologies have been applied to cancer investigations since the availability of such a large number of markers makes DNA array a powerful diagnostic tool for tumour and patient classification. Over the last two years, a series of computational tools have been developed for the analysis of different aspects of gene profiling. Our work tries to compare a series of supervised statistical techniques on the basis of their ability to correctly classify different types of tumours. A simulation approach was initially used to control the huge source of variation among and between patients, and to evaluate the ability of algorithms to classify tumours in relation to different types of experimental variables. Different techniques for reduction of data dimension were then added to the discriminant analysis and compared according to their ability to capture the main genetic information. The simulation results have been tested by applying the selected classification algorithms to two experimental microarray datasets of human cancers, and by measuring the correspondent rates of misclassification. Our analyses identify in these datasets a series of genes principally involved in tumour characterization. The functional role of these discriminant transcripts is discussed.

Algorithms↗

TRAIT (TRAnscript Integrated Table): a knowledgebase of human skeletal muscle transcripts.

TRAIT is a knowledgebase integrating information on transcripts with related data from genome, proteins, ortholog genes and diseases. It was initially built as a system to manage an EST-based gene discovery project on human skeletal muscle, which yielded over 4500 independent sequence clusters. Transcripts are annotated using automatic as well as manual procedures, linking known transcripts to public databases and unknown transcripts to tables of predicted features. Data are stored in a MySQL database. Complex queries are automatically built by means of a user-friendly web interface that allows the concurrent selection of many fields such as ontology, expression level, map position and protein domains. The results are parsed by the system and returned in a ranked order, in respect to the number of satisfied criteria.

Database Management Systems↗

IDEG6: a web tool for detection of differentially expressed genes in multiple tag sampling experiments.

Here we present a novel web tool for the statistical analysis of gene expression data in multiple tag sampling experiments. Differentially expressed genes are detected by using six different test statistics. Result tables, linked to the GenBank, UniGene, or LocusLink database, can be browsed or searched in different ways. Software is freely available at the site: http://telethon.bio.unipd.it/bioinfo/IDEG6_form/, together with additional information on statistical methodologies.

Database Management Systems↗

Gene expression profiling in dysferlinopathies using a dedicated muscle microarray.

We have performed expression profiling to define the molecular changes in dysferlinopathy using a novel dedicated microarray platform made with 3'-end skeletal muscle cDNAs. Eight dysferlinopathy patients, defined by western blot, immunohistochemistry and mutation analysis, were investigated with this technology. In a first experiment RNAs from different limb-girdle muscular dystrophy type 2B patients were pooled and compared with normal muscle RNA to characterize the general transcription pattern of this muscular disorder. Then the expression profiles of patients with different clinical traits were independently obtained and hierarchical clustering was applied to discover patient-specific gene variations. MHC class I genes and genes involved in protein biosynthesis were up-regulated in relation to muscle histopathological features. Conversely, the expression of genes codifying the sarcomeric proteins titin, nebulin and telethonin was down-regulated. Neither calpain-3 nor caveolin, a sarcolemmal protein interacting with dysferlin, was consistently reduced. There was a major up-regulation of proteins interacting with calcium, namely S100 calcium-binding proteins and sarcolipin, a sarcoplasmic calcium regulator.

Adolescent↗

Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matrices.

MOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp

Algorithms↗

Patterns of human diversity, within and among continents, inferred from biallelic DNA polymorphisms.

Previous studies have reported that about 85% of human diversity at Short Tandem Repeat (STR) and Restriction Fragment Length Polymorphism (RFLP) autosomal loci is due to differences between individuals of the same population, whereas differences among continental groups account for only 10% of the overall genetic variance. These findings conflict with popular notions of distinct and relatively homogeneous human races, and may also call into question the apparent usefulness of ethnic classification in, for example, medical diagnostics. Here, we present new data on 21 Alu insertions in 32 populations. We analyze these data along with three other large, globally dispersed data sets consisting of apparently neutral biallelic nuclear markers, as well as with a beta-globin data set possibly subject to selection. We confirm the previous results for the autosomal data, and find a higher diversity among continents for Y-chromosome loci. We also extend the analyses to address two questions: (1) whether differences between continental groups, although small, are nevertheless large enough to confidently assign individuals to their continent on the basis of their genotypes; (2) whether the observed genotypes naturally cluster into continental or population groups when the sample source location is ignored. Using a range of statistical methods, we show that classification errors are at best around 30% for autosomal biallelic polymorphisms and 27% for the Y chromosome. Two data sets suggest the existence of three and four major groups of genotypes worldwide, respectively, and the two groupings are inconsistent. These results suggest that, at random biallelic loci, there is little evidence, if any, of a clear subdivision of humans into biologically defined groups.

Alleles↗