PubMed Health⌕ Search

Biomedical subjects

Francesca D Ciccarelli

Publications and source records attributed to Francesca D Ciccarelli.

18 recordsLinked to original sources

Toward automatic reconstruction of a highly resolved tree of life.

We have developed an automatable procedure for reconstructing the tree of life with branch lengths comparable across all three domains. The tree has its basis in a concatenation of 31 orthologs occurring in 191 species with sequenced genomes. It revealed interdomain discrepancies in taxonomic classification. Systematic detection and subsequent exclusion of products of horizontal gene transfer increased phylogenetic resolution, allowing us to confirm accepted relationships and resolve disputed and preliminary classifications. For example, we place the phylum Acidobacteria as a sister group of delta-Proteobacteria, support a Gram-positive origin of Bacteria, and suggest a thermophilic last universal common ancestor.

Amino Acyl-tRNA Synthetases↗

Very-KIND is a novel nervous system specific guanine nucleotide exchange factor for Ras GTPases.

The kinase non-catalytic c-lobe domain (KIND) evolved from the catalytic protein kinase fold into a potential protein interaction module for signalling proteins. Spir family actin organizers and the non-receptor phosphatase type 13 (PTP type 13) encode a KIND domain in the very N-terminal parts of the proteins. Here we report the characterization and cloning of a third member of the KIND protein family, which we have named very-KIND (VKIND) because of its two KIND domains. Like the other members of the protein family, VKIND has a KIND domain at the N-terminus. A second KIND domain is located in the central part of the protein. The C-terminal half encodes a guanine nucleotide exchange factor motif for Ras-like GTPases (RasGEF) and a RasGEF N-terminal module (RasGEFN). There is only one VKIND gene in the mammalian genomes and up to now we have found the gene only in vertebrates. During mouse embryogenesis the VKIND gene was specifically expressed in the developing nervous system. In adult mice Northern hybridizations revealed high expression only in brain. Low expression could be detected in ovary. In situ hybridizations showed a specific expression of VKIND in neuronal cells of the granular and Purkinje cell layers of the cerebellum.

Amino Acid Sequence↗

Generation and annotation of the DNA sequences of human chromosomes 2 and 4.

Human chromosome 2 is unique to the human lineage in being the product of a head-to-head fusion of two intermediate-sized ancestral chromosomes. Chromosome 4 has received attention primarily related to the search for the Huntington's disease gene, but also for genes associated with Wolf-Hirschhorn syndrome, polycystic kidney disease and a form of muscular dystrophy. Here we present approximately 237 million base pairs of sequence for chromosome 2, and 186 million base pairs for chromosome 4, representing more than 99.6% of their euchromatic sequences. Our initial analyses have identified 1,346 protein-coding genes and 1,239 pseudogenes on chromosome 2, and 796 protein-coding genes and 778 pseudogenes on chromosome 4. Extensive analyses confirm the underlying construction of the sequence, and expand our understanding of the structure and evolution of mammalian chromosomes, including gene deserts, segmental duplications and highly variant regions.

Animals↗

Complex genomic rearrangements lead to novel primate gene function.

Orthologous genes that maintain a single-copy status in a broad range of species may indicate a selection against gene duplication. If this is the case, then duplicates of such genes that do survive may have escaped the dosage control by rapid and sizable changes in their function. To test this hypothesis and to develop a strategy for the identification of novel gene functions, we have analyzed 22 primate-specific intrachromosomal duplications of genes with a single-copy ortholog in all other completely sequenced metazoans. When comparing this set to genes not exposed to the single-copy status constraint, we observed a higher tendency of the former to modify their gene structure, often through complex genomic rearrangements. The analysis of the most dramatic of these duplications, affecting approximately 10% of human Chromosome 2, enabled a detailed reconstruction of the events leading to the appearance of a novel gene family. The eight members of this family originated from the highly conserved nucleoporin RanBP2 by several genetic rearrangements such as segmental duplications, inversions, translocations, exon loss, and domain accretion. We have experimentally verified that at least one of the newly formed proteins has a cellular localization different from RanBP2's, and we show that positive selection did act on specific domains during evolution.

Alternative Splicing↗

The WHy domain mediates the response to desiccation in plants and bacteria.

MOTIVATION: The hypersensitive response (HR) is a process activated by plants after microbial infection. Its main phenotypic effects are both a programmed death of the plant cells near the infection site and a reduction of the microbial proliferation. Although many resistance genes (R genes) associated to HR have been identified, very little is known about the molecular mechanisms activated after their expression. RESULTS: The analysis of the product of one of the R genes, the Hin1 protein, led to the identification of a novel domain, which we named WHy because it is detectable in proteins involved in Water stress and Hypersensitive response. The expression of this domain during both biotic infection and response to desiccation points to a molecular machinery common to these two stress conditions. Moreover, its presence in a restricted number of bacteria suggests a possible use for marking plant pathogenicity. CONTACT: francesca.ciccarelli@embl.de SUPPLEMENTARY INFORMATION: Supplementary data (Figures S1 and S2 and Table S1) and the alignment in clustal format are available at http://www.bork.embl.de/~ciccarel/WHy_add_data.html.

Adaptation, Physiological↗

SMART 4.0: towards genomic data integration.

SMART (Simple Modular Architecture Research Tool) is a web tool (http://smart.embl.de/) for the identification and annotation of protein domains, and provides a platform for the comparative study of complex domain architectures in genes and proteins. The January 2004 release of SMART contains 685 protein domains. New developments in SMART are centred on the integration of data from completed metazoan genomes. SMART now uses predicted proteins from complete genomes in its source sequence databases, and integrates these with predictions of orthology. New visualization tools have been developed to allow analysis of gene intron-exon structure within the context of protein domain structure, and to align these displays to provide schematic comparisons of orthologous genes, or multiple transcripts from the same gene. Other improvements include the ability to query SMART by Gene Ontology terms, improved structure database searching and batch retrieval of multiple entries.

Algorithms↗

RanBP2/Nup358 provides a major binding site for NXF1-p15 dimers at the nuclear pore complex and functions in nuclear mRNA export.

Metazoan NXF1-p15 heterodimers promote the nuclear export of bulk mRNA across nuclear pore complexes (NPCs). In vitro, NXF1-p15 forms a stable complex with the nucleoporin RanBP2/Nup358, a component of the cytoplasmic filaments of the NPC, suggesting a role for this nucleoporin in mRNA export. We show that depletion of RanBP2 from Drosophila cells inhibits proliferation and mRNA export. Concomitantly, the localization of NXF1 at the NPC is strongly reduced and a significant fraction of this normally nuclear protein is detected in the cytoplasm. Under the same conditions, the steady-state subcellular localization of other nuclear or cytoplasmic proteins and CRM1-mediated protein export are not detectably affected, indicating that the release of NXF1 into the cytoplasm and the inhibition of mRNA export are not due to a general defect in NPC function. The specific role of RanBP2 in the recruitment of NXF1 to the NPC is highlighted by the observation that depletion of CAN/Nup214 also inhibits cell proliferation and mRNA export but does not affect NXF1 localization. Our results indicate that RanBP2 provides a major binding site for NXF1 at the cytoplasmic filaments of the NPC, thereby restricting its diffusion in the cytoplasm after NPC translocation. In RanBP2-depleted cells, NXF1 diffuses freely through the cytoplasm. Consequently, the nuclear levels of the protein decrease and export of bulk mRNA is impaired.

Animals↗

The PAM domain, a multi-protein complex-associated module with an all-alpha-helix fold.

BACKGROUND: Multimeric protein complexes have a role in many cellular pathways and are highly interconnected with various other proteins. The characterization of their domain composition and organization provides useful information on the specific role of each region of their sequence. RESULTS: We identified a new module, the PAM domain (PCI/PINT associated module), present in single subunits of well characterized multiprotein complexes, like the regulatory lid of the 26S proteasome, the COP-9 signalosome and the Sac3-Thp1 complex. This module is an around 200 residue long domain with a predicted TPR-like all-alpha-helical fold. CONCLUSIONS: The occurrence of the PAM domain in specific subunits of multimeric protein complexes, together with the role of other all-alpha-helical folds in protein-protein interactions, suggest a function for this domain in mediating transient binding to diverse target proteins.

Amino Acid Sequence↗

Genome evolution reveals biochemical networks and functional modules.

The analysis of completely sequenced genomes uncovers an astonishing variability between species in terms of gene content and order. During genome history, the genes are frequently rear-ranged, duplicated, lost, or transferred horizontally between genomes. These events appear to be stochastic, yet they are under selective constraints resulting from the functional interactions between genes. These genomic constraints form the basis for a variety of techniques that employ systematic genome comparisons to predict functional associations among genes. The most powerful techniques to date are based on conserved gene neighborhood, gene fusion events, and common phylogenetic distributions of gene families. Here we show that these techniques, if integrated quantitatively and applied to a sufficiently large number of genomes, have reached a resolution which allows the characterization of function at a higher level than that of the individual gene: global modularity becomes detectable in a functional protein network. In Escherichia coli, the predicted modules can be bench-marked by comparison to known metabolic pathways. We found as many as 74% of the known metabolic enzymes clustering together in modules, with an average pathway specificity of at least 84%. The modules extend beyond metabolism, and have led to hundreds of reliable functional predictions both at the protein and pathway level. The results indicate that modularity in protein networks is intrinsically encoded in present-day genomes.

Amino Acids↗

Nonsense-mediated mRNA decay in Drosophila: at the intersection of the yeast and mammalian pathways.

The nonsense-mediated mRNA decay (NMD) pathway promotes the rapid degradation of mRNAs containing premature stop codons (PTCs). In Caenorhabditis elegans, seven genes (smg1-7) playing an essential role in NMD have been identified. Only SMG2-4 (known as UPF1-3) have orthologs in Saccharomyces cerevisiae. Here we show that the Drosophila orthologs of UPF1-3, SMG1, SMG5 and SMG6 are required for the degradation of PTC-containing mRNAs, but that there is no SMG7 ortholog in this organism. In contrast, orthologs of SMG5-7 are encoded by the human genome and all three are required for NMD. In human cells, exon boundaries have been shown to play a critical role in defining PTCs. This role is mediated by components of the exon junction complex (EJC). Contrary to expectation, however, we show that the components of the EJC are dispensable for NMD in Drosophila cells. Consistently, PTC definition occurs independently of exon boundaries in Drosophila. Our findings reveal that despite conservation of the NMD machinery, different mechanisms have evolved to discriminate premature from natural stop codons in metazoa.

Amino Acid Sequence↗

The identification of a conserved domain in both spartin and spastin, mutated in hereditary spastic paraplegia.

Multiple sequence alignment has revealed the presence of a sequence domain of approximately 80 amino acids in two molecules, spartin and spastin, mutated in hereditary spastic paraplegia. The domain, which corresponds to a slightly extended version of the recently described ESP domain of unknown function, was also identified in VPS4, SKD1, RPK118, and SNX15, all of which have a well established and consistent role in endosomal trafficking. Recent functional information indicates that spastin is likely to be involved in microtubule interaction. With this new information relating to its likely function, we propose the more descriptive name 'MIT' (contained within microtubule-interacting and trafficking molecules) for the domain and predict endosomal trafficking as the principal functionality of all molecules in which it is present.

Adenosine Triphosphatases↗

A protocol for the update of references to scientific literature in biological databases.

Entries in biological databases are usually linked to scientific references. To generate those links and to keep them up-to-date, database maintainers have to continuously scan the scientific literature to select references that are relevant for each single database entry. The continuous growth of both the corpus of scientific literature and the size of biological databases makes this task very hard. We present a protocol intended to assist the updating of an existing set of literature (abstract) links from a single database entry with new references. It consists of taking the set of MEDLINE neighbour references of the existing linked abstracts and evaluating their relevance according to the existing set of abstracts. To test the applicability of the algorithm, we did a simple benchmark of the system using the references associated with the entries of a protein domain database. Human experts found the references that the algorithm scored highly were more relevant to the database entry than those scored lowly, suggesting that the algorithm was useful.

Abstracting and Indexing↗

NEAT: a domain duplicated in genes near the components of a putative Fe3+ siderophore transporter from Gram-positive pathogenic bacteria.

BACKGROUND: Iron uptake from the host is essential for bacteria that infect animals. To find potential targets for drugs active against pathogenic bacteria, we have searched all completely sequenced genomes of pathogenic bacteria for genes relevant for iron transport. RESULTS: We identified a protein domain that appears in variable copy number in bacterial genes that are usually in the vicinity of a putative Fe3+ siderophore transporter. Accordingly, we have denoted this domain NEAT for 'near transporter'. Most of the bacterial species containing this domain are pathogenic. Sequence features indicate that the domain is anchored to the extracellular side of the membrane. The domain seems to be under high selective pressure for rapid independent duplications that are typical of sequences involved in signaling and binding. CONCLUSIONS: The NEAT domain might be functionally related to iron transport. The taxonomic specificity of this domain and its predicted extracellular position could make it an interesting target for designing new drugs against some highly pathogenic bacteria.

Amino Acid Sequence↗

SPG20 is mutated in Troyer syndrome, an hereditary spastic paraplegia.

Troyer syndrome (TRS) is an autosomal recessive complicated hereditary spastic paraplegia (HSP) that occurs with high frequency in the Old Order Amish. We report mapping of the TRS locus to chromosome 13q12.3 and identify a frameshift mutation in SPG20, encoding spartin. Comparative sequence analysis indicates that spartin shares similarity with molecules involved in endosomal trafficking and with spastin, a molecule implicated in microtubule interaction that is commonly mutated in HSP.

Adenosine Triphosphatases↗

CASH--a beta-helix domain widespread among carbohydrate-binding proteins.

In this article, we describe a novel, widespread domain (CASH) that is shared by many carbohydrate-binding proteins and sugar hydrolases. This domain occurs in more than 1000 proteins distributed among all three kingdoms of life. The CASH domain is characterized by internal repetitions of glycines and hydrophobic residues that correspond to the repetitive units of a predicted or observed right-handed beta-helix structure of the pectate lyase superfamily.

Amino Acid Motifs↗

AMOP, a protein module alternatively spliced in cancer cells.

This article describes a new extracellular domain--AMOP, for adhesion-associated domain in MUC4 and other proteins. This domain occurs in putative cell adhesion molecules and in some splice variants of MUC4. MUC4 splice variants are overexpressed in several tumours; in particular, they are highly expressed in pancreatic carcinomas but not in normal pancreas. The presence of AMOP in cell adhesion molecules could be indicative of a role for this domain in adhesion.

Alternative Splicing↗

A complex prediction: three-dimensional model of the yeast exosome.

We present a model of the yeast exosome based on the bacterial degradosome component polynucleotide phosphorylase (PNPase). Electron microscopy shows the exosome to resemble PNPase but with key differences likely related to the position of RNA binding domains, and to the location of domains unique to the exosome. We use various techniques to reduce the many possible models of exosome subunits based on PNPase to just one. The model suggests numerous experiments to probe exosome function, particularly with respect to subunits making direct atomic contacts and conserved, possibly functional residues within the predicted central pore of the complex.

Amino Acid Sequence↗