PubMed Health⌕ Search

Biomedical subjects

Giorgio Valle

Publications and source records attributed to Giorgio Valle.

10 recordsLinked to original sources

Pattern recognition in gene expression profiling using DNA array: a comparative study of different statistical methods applied to cancer classification.

Large-scale parallel measurements of the expression of many thousands genes are now available with high-density array made with collections of cDNA fragments, or oligonucleotide corresponding to different transcripts. These technologies have been applied to cancer investigations since the availability of such a large number of markers makes DNA array a powerful diagnostic tool for tumour and patient classification. Over the last two years, a series of computational tools have been developed for the analysis of different aspects of gene profiling. Our work tries to compare a series of supervised statistical techniques on the basis of their ability to correctly classify different types of tumours. A simulation approach was initially used to control the huge source of variation among and between patients, and to evaluate the ability of algorithms to classify tumours in relation to different types of experimental variables. Different techniques for reduction of data dimension were then added to the discriminant analysis and compared according to their ability to capture the main genetic information. The simulation results have been tested by applying the selected classification algorithms to two experimental microarray datasets of human cancers, and by measuring the correspondent rates of misclassification. Our analyses identify in these datasets a series of genes principally involved in tumour characterization. The functional role of these discriminant transcripts is discussed.

Algorithms↗

TRAIT (TRAnscript Integrated Table): a knowledgebase of human skeletal muscle transcripts.

TRAIT is a knowledgebase integrating information on transcripts with related data from genome, proteins, ortholog genes and diseases. It was initially built as a system to manage an EST-based gene discovery project on human skeletal muscle, which yielded over 4500 independent sequence clusters. Transcripts are annotated using automatic as well as manual procedures, linking known transcripts to public databases and unknown transcripts to tables of predicted features. Data are stored in a MySQL database. Complex queries are automatically built by means of a user-friendly web interface that allows the concurrent selection of many fields such as ontology, expression level, map position and protein domains. The results are parsed by the system and returned in a ranked order, in respect to the number of satisfied criteria.

Database Management Systems↗

Human MYO18B, a novel unconventional myosin heavy chain expressed in striated muscles moves into the myonuclei upon differentiation.

We have characterized a novel unconventional myosin heavy chain, named MYO18B, that appears to be expressed mainly in human cardiac and skeletal muscles and, at lower levels, in testis. MYO18B transcript is detected in all types of striated muscles but at much lower levels compared to class II sarcomeric myosins, and it is up regulated after in vitro differentiation of myoblasts into myotubes. Phylogenetic analysis shows that this myosin belongs to the recently identified class XVIII, however, unlike the other member of this class, it seems to be unique to Vertebrate since it contains two large amino acid domains of unknown function at the N and C-termini. Immunolocalization of MYO18B protein in skeletal muscle cells shows that this myosin heavy chain is located in the cytoplasm of undifferentiated myoblasts. After in vitro differentiation into myotubes, a fraction of this protein is accumulated in a subset of myonuclei. This nuclear localization was confirmed by immunofluorescence experiments on primary cardiomyocytes and adult muscle sections. In the cytoplasm MYO18B shows a punctate staining, both in cardiac and skeletal fibers. In some cases, cardiomyocytes show a partial sarcomeric pattern of MYO18B alternating that of alpha-actinin-2. In skeletal muscle the cytoplasmic MYO18B results much more evident in the fast type fibers.

Animals↗

Simple consensus procedures are effective and sufficient in secondary structure prediction.

We have analyzed the performance of majority voting on minimal combination sets of three state-of-the-art secondary structure prediction methods in order to obtain a consensus prediction. Using three large benchmark sets from the EVA server, our results show a significant improvement in the average Q3 prediction accuracy of up to 1.5 percentage points by consensus formation. The application of an additional trivial filtering procedure for predicted secondary structure elements that are too short, does not significantly affect the prediction accuracy. Our analysis also provides valuable insight into the similarity of the results of the prediction methods that we combine as well as the higher confidence in consistently predicted secondary structure.

Computational Biology↗

Gene expression profiling in dysferlinopathies using a dedicated muscle microarray.

We have performed expression profiling to define the molecular changes in dysferlinopathy using a novel dedicated microarray platform made with 3'-end skeletal muscle cDNAs. Eight dysferlinopathy patients, defined by western blot, immunohistochemistry and mutation analysis, were investigated with this technology. In a first experiment RNAs from different limb-girdle muscular dystrophy type 2B patients were pooled and compared with normal muscle RNA to characterize the general transcription pattern of this muscular disorder. Then the expression profiles of patients with different clinical traits were independently obtained and hierarchical clustering was applied to discover patient-specific gene variations. MHC class I genes and genes involved in protein biosynthesis were up-regulated in relation to muscle histopathological features. Conversely, the expression of genes codifying the sarcomeric proteins titin, nebulin and telethonin was down-regulated. Neither calpain-3 nor caveolin, a sarcolemmal protein interacting with dysferlin, was consistently reduced. There was a major up-regulation of proteins interacting with calcium, namely S100 calcium-binding proteins and sarcolipin, a sarcoplasmic calcium regulator.

Adolescent↗

Development and production of an oligonucleotide MuscleChip: use for validation of ambiguous ESTs.

BACKGROUND: We describe the development, validation, and use of a highly redundant 120,000 oligonucleotide microarray (MuscleChip) containing 4,601 probe sets representing 1,150 known genes expressed in muscle and 2,075 EST clusters from a non-normalized subtracted muscle EST sequencing project (28,074 EST sequences). This set included 369 novel EST clusters showing no match to previously characterized proteins in any database. Each probe set was designed to contain 20-32 25 mer oligonucleotides (10-16 paired perfect match and mismatch probe pairs per gene), with each probe evaluated for hybridization kinetics (Tm) and similarity to other sequences. The 120,000 oligonucleotides were synthesized by photolithography and light-activated chemistry on each microarray. RESULTS: Hybridization of human muscle cRNAs to this MuscleChip (33 samples) showed a correlation of 0.6 between the number of ESTs sequenced in each cluster and hybridization intensity. Out of 369 novel EST clusters not showing any similarity to previously characterized proteins, we focused on 250 EST clusters that were represented by robust probe sets on the MuscleChip fulfilling all stringent rules. 102 (41%) were found to be consistently "present" by analysis of hybridization to human muscle RNA, of which 40 ESTs (39%) could be genome anchored to potential transcription units in the human genome sequence. 19 ESTs of the 40 ESTs were furthermore computer-predicted as exons by one or more than three gene identification algorithms. CONCLUSION: Our analysis found 40 transcriptionally validated, genome-anchored novel EST clusters to be expressed in human muscle. As most of these ESTs were low copy clusters (duplex and triplex) in the original 28,000 EST project, the identification of these as significantly expressed is a robust validation of the transcript units that permits subsequent focus on the novel proteins encoded by these genes.

Algorithms↗

Functional profiling of the Saccharomyces cerevisiae genome.

Determining the effect of gene deletion is a fundamental approach to understanding gene function. Conventional genetic screens exhibit biases, and genes contributing to a phenotype are often missed. We systematically constructed a nearly complete collection of gene-deletion mutants (96% of annotated open reading frames, or ORFs) of the yeast Saccharomyces cerevisiae. DNA sequences dubbed 'molecular bar codes' uniquely identify each strain, enabling their growth to be analysed in parallel and the fitness contribution of each gene to be quantitatively assessed by hybridization to high-density oligonucleotide arrays. We show that previously known and new genes are necessary for optimal growth under six well-studied conditions: high salt, sorbitol, galactose, pH 8, minimal medium and nystatin treatment. Less than 7% of genes that exhibit a significant increase in messenger RNA expression are also required for optimal growth in four of the tested conditions. Our results validate the yeast gene-deletion collection as a valuable resource for functional genomics.

Cell Size↗

A two-step strategy for constructing specifically self-subtracted cDNA libraries.

We have developed a new strategy for producing subtracted cDNA libraries that is optimized for connective and epithelial tissues, where a few exceptionally abundant (super-prevalent) RNA species account for a large fraction of the total mRNA mass. Our method consists of a two-step subtraction of the most abundant mRNAs: the first step involves a novel use of oligo-directed RNase H digestion to lower the concentration of tissue-specific, super-prevalent RNAs. In the second step, a highly specific subtraction is achieved through hybridization with probes from a 3'-end ESTs collection. By applying this technique in skeletal muscle, we have constructed subtracted cDNA libraries that are effectively enriched for genes expressed at low levels. We further report on frequent premature termination of transcription in human muscle mitochondria and discuss the importance of this phenomenon in designing subtractive approaches. The tissue-specific collections of cDNA clones generated by our method are particularly well suited for expression profiling.

DNA, Complementary↗

Analysis of 22 deletion breakpoints in dystrophin intron 49.

Over 60% of Duchenne and Becker muscular dystrophies are caused by deletions spanning tens or hundreds of kilobases in the dystrophin gene. The molecular mechanisms underlying the loss of DNA at this genomic locus are not yet understood. By studying the distribution of deletion breakpoints at the genomic level, we have previously shown that intron 49 exhibits a higher relative density of breakpoints than most dystrophin introns. To determine whether the mechanisms leading to deletions in this intron preferentially involve specific sequence elements, we sublocalized 22 deletion endpoints along its length by a polymerase-chain-reaction-based approach and, in particular, analyzed the nucleotide sequences of five deletion junctions. Deletion breakpoints were homogeneously distributed throughout the intron length, and no extensive homology was observed between the sequences adjacent to each breakpoint. However, a short sequence able to curve the DNA molecule was found at or near three breakpoint junctions.

Base Sequence↗

Simplifying amino acid alphabets by means of a branch and bound algorithm and substitution matrices.

MOTIVATION: Protein and DNA are generally represented by sequences of letters. In a number of circumstances simplified alphabets (where one or more letters would be represented by the same symbol) have proved their potential utility in several fields of bioinformatics including searching for patterns occurring at an unexpected rate, studying protein folding and finding consensus sequences in multiple alignments. The main issue addressed in this paper is the possibility of finding a general approach that would allow an exhaustive analysis of all the possible simplified alphabets, using substitution matrices like PAM and BLOSUM as a measure for scoring. RESULTS: The computational approach presented in this paper has led to a computer program called AlphaSimp (Alphabet Simplifier) that can perform an exhaustive analysis of the possible simplified amino acid alphabets, using a branch and bound algorithm together with standard or user-defined substitution matrices. The program returns a ranked list of the highest-scoring simplified alphabets. When the extent of the simplification is limited and the simplified alphabets are maintained above ten symbols the program is able to complete the analysis in minutes or even seconds on a personal computer. However, the performance becomes worse, taking up to several hours, for highly simplified alphabets. AVAILABILITY: AlphaSimp and other accessory programs are available at http://bioinformatics.cribi.unipd.it/alphasimp

Algorithms↗