PubMed Health⌕ Search

Biomedical subjects

S Toppo

Publications and source records attributed to S Toppo.

7 recordsLinked to original sources

Large-scale prediction of protein structure and function from sequence.

The identification of novel drug targets from genomic data involves the large-scale analysis of many protein sequences. Methods for automated structure and function prediction are an essential tool for this purpose. In this review we concentrate on the recent developments in the field of protein structure prediction and how these can be used to gain hints about the function of proteins. The current state-of-the-art is highlighted through recent community-wide experiments aimed at comparing different approaches. For structure prediction this allows the identification of key improvements to increase the crucial sequence to structure alignment needed for accurate models. Function prediction is a rapidly maturing field that is still being benchmarked. Definitions for protein function are presented and available methods, mostly concentrating on functional site descriptors and structural motifs, presented.

Algorithms↗

Characterization of 16 novel human genes showing high similarity to yeast sequences.

The entire set of open reading frames (ORFs) of Saccharomyces cerevisiae has been used to perform systematic similarity searches against nucleic acid and protein databases: with the aim of identifying interesting homologies between yeast and mammalian genes. Many similarities were detected: mostly with known genes. However: several yeast ORFs were only found to match human partial sequence tags: indicating the presence of human transcripts still uncharacterized that have a homologous counterpart in yeast. About 30 such transcripts were further studied and named HUSSY (human sequence similar to yeast). The 16 most interesting are presented in this paper along with their sequencing and mapping data. As expected: most of these genes seem to be involved in basic metabolic and cellular functions (lipoic acid biosynthesis: ribulose-5-phosphate-3-epimerase: glycosyl transferase: beta-transducin: serine-threonine-kinase: ABC proteins: cation transporters). Genes related to RNA maturation were also found (homologues to DIM1: ROK1-RNA-elicase and NFS1). Furthermore: five novel human genes were detected (HUSSY-03: HUSSY-22: HUSSY-23: HUSSY-27: HUSSY-29) that appear to be homologous to yeast genes whose function is still undetermined. More information on this work can be obtained at the website http://grup.bio.unipd.it/hussy

Amino Acid Sequence↗

Sequence and analysis of chromosome 3 of the plant Arabidopsis thaliana.

Arabidopsis thaliana is an important model system for plant biologists. In 1996 an international collaboration (the Arabidopsis Genome Initiative) was formed to sequence the whole genome of Arabidopsis and in 1999 the sequence of the first two chromosomes was reported. The sequence of the last three chromosomes and an analysis of the whole genome are reported in this issue. Here we present the sequence of chromosome 3, organized into four sequence segments (contigs). The two largest (13.5 and 9.2 Mb) correspond to the top (long) and the bottom (short) arms of chromosome 3, and the two small contigs are located in the genetically defined centromere. This chromosome encodes 5,220 of the roughly 25,500 predicted protein-coding genes in the genome. About 20% of the predicted proteins have significant homology to proteins in eukaryotic genomes for which the complete sequence is available, pointing to important conserved cellular functions among eukaryotes.

Arabidopsis↗

A comprehensive, high-resolution genomic transcript map of human skeletal muscle.

We present the Human Muscle Gene Map (HMGM), the first comprehensive and updated high-resolution expression map of human skeletal muscle. The 1078 entries of the map were obtained by merging data retrieved from UniGene with the RH mapping information on 46 novel muscle transcripts, which showed no similarity to any known sequence. In the map, distances are expressed in megabase pairs. About one-quarter of the map entries represents putative novel genes. Genes known to be specifically expressed in muscle account for <4% of the total. The genomic distribution of the map entries confirmed the previous finding that muscle genes are selectively concentrated in chromosomes 17, 19, and X. Five chromosomal regions are suspected to have a significant excess of muscle genes. Present data support the hypothesis that the biochemical and functional properties of differentiated muscle cells may result from the transcription of a very limited number of muscle-specific genes along with the activity of a large number of genes, shared with other tissues, but showing different levels of expression in muscle. [The sequence data described in this paper have been submitted to the EMBL data library under accession nos. F23198-F23242.]

Chromosome Mapping↗

Telethonin, a novel sarcomeric protein of heart and skeletal muscle.

In this paper we describe a novel 19 kDa sarcomeric protein named telethonin. The cDNA sequence discloses an open reading frame of 167 amino acids that does not resemble any known protein. Antibodies against a recombinant telethonin fragment were used for Western blot analysis, confirming the presence of this 19 kDa protein in heart and skeletal muscle and revealing an immunofluorescence pattern typical of sarcomeric proteins, overlapping myosin. The frequency of specific cDNA clones in different libraries indicates that the telethonin transcript is amongst the most abundant in skeletal muscle. In human, telethonin maps at 17q12, adjacent to the phenylethanolamine N-methyltransferase gene.

Amino Acid Sequence↗

Identification of 4370 expressed sequence tags from a 3'-end-specific cDNA library of human skeletal muscle by DNA sequencing and filter hybridization.

A systematic study on the mRNA species expressed in the human skeletal muscle is presented in this paper. To carry on this study, a new method has been developed for the construction of unbiased cDNA libraries specially designed for the production of ESTs corresponding to the 3'-end portion of the mRNAs. The method has been applied to human skeletal muscle, where the analysis of the transcription profile is particularly difficult for the presence of several very abundant transcripts. To detect and quantify high-level mRNAs, the first 1054 ESTs were obtained from randomly selected clones. The 10 most abundant transcripts accounted for > 45% of the clones. Subsequently, these transcripts were identified by filter hybridization, thus making DNA sequencing more productive. Overall, 4370 clones were identified: 3372 by DNA sequencing and 998 by filter hybridization. The number of groups of sequences identifying individual transcripts was relatively low compared with other tissues, resulting in a total of 934 groups out of 4370 ESTs. Of these, 719 groups were represented by only one sequence.

Cloning, Molecular↗

Semi-multiplex PCR technique for screening of abundant transcripts during systematic sequencing of cDNA libraries.

The systematic sequencing of cDNA libraries is an efficient approach for the identification of new genes, but the presence of abundant mRNAs is often a major problem. This paper describes a very simple method of "semi-multiplex PCR" that allows specific identification of such abundant transcripts before DNA sequencing without using nonrepresentative subtracted libraries. The PCR utilizes a series of forward primers specific for abundant transcripts with a pair of universal primers used for template generation. cDNA clones corresponding to abundant mRNAs are then revealed by double bands in agarose gel.

Base Sequence↗