PubMed Health⌕ Search

Biomedical subjects

Olivier Poch

Publications and source records attributed to Olivier Poch.

At least 19 recordsLinked to original sources

PeroxisomeDB: a database for the peroxisomal proteome, functional genomics and disease.

Peroxisomes are essential organelles of eukaryotic origin, ubiquitously distributed in cells and organisms, playing key roles in lipid and antioxidant metabolism. Loss or malfunction of peroxisomes causes more than 20 fatal inherited conditions. We have created a peroxisomal database (http://www.peroxisomeDB.org) that includes the complete peroxisomal proteome of Homo sapiens and Saccharomyces cerevisiae, by gathering, updating and integrating the available genetic and functional information on peroxisomal genes. PeroxisomeDB is structured in interrelated sections 'Genes', 'Functions', 'Metabolic pathways' and 'Diseases', that include hyperlinks to selected features of NCBI, ENSEMBL and UCSC databases. We have designed graphical depictions of the main peroxisomal metabolic routes and have included updated flow charts for diagnosis. Precomputed BLAST, PSI-BLAST, multiple sequence alignment (MUSCLE) and phylogenetic trees are provided to assist in direct multispecies comparison to study evolutionary conserved functions and pathways. Highlights of the PeroxisomeDB include new tools developed for facilitating (i) identification of novel peroxisomal proteins, by means of identifying proteins carrying peroxisome targeting signal (PTS) motifs, (ii) detection of peroxisomes in silico, particularly useful for screening the deluge of newly sequenced genomes. PeroxisomeDB should contribute to the systematic characterization of the peroxisomal proteome and facilitate system biology approaches on the organelle.

Animals↗

Identification of a novel BBS gene (BBS12) highlights the major role of a vertebrate-specific branch of chaperonin-related proteins in Bardet-Biedl syndrome.

Bardet-Biedl syndrome (BBS) is primarily an autosomal recessive ciliopathy characterized by progressive retinal degeneration, obesity, cognitive impairment, polydactyly, and kidney anomalies. The disorder is genetically heterogeneous, with 11 BBS genes identified to date, which account for ~70% of affected families. We have combined single-nucleotide-polymorphism array homozygosity mapping with in silico analysis to identify a new BBS gene, BBS12. Patients from two Gypsy families were homozygous and haploidentical in a 6-Mb region of chromosome 4q27. FLJ35630 was selected as a candidate gene, because it was predicted to encode a protein with similarity to members of the type II chaperonin superfamily, which includes BBS6 and BBS10. We found pathogenic mutations in both Gypsy families, as well as in 14 other families of various ethnic backgrounds, indicating that BBS12 accounts for approximately 5% of all BBS cases. BBS12 is vertebrate specific and, together with BBS6 and BBS10, defines a novel branch of the type II chaperonin superfamily. These three genes are characterized by unusually rapid evolution and are likely to perform ciliary functions specific to vertebrates that are important in the pathophysiology of the syndrome, and together they account for about one-third of the total BBS mutational load. Consistent with this notion, suppression of each family member in zebrafish yielded gastrulation-movement defects characteristic of other BBS morphants, whereas simultaneous suppression of all three members resulted in severely affected embryos, possibly hinting at partial functional redundancy within this protein family.

Animals↗

Pitfalls of homozygosity mapping: an extended consanguineous Bardet-Biedl syndrome family with two mutant genes (BBS2, BBS10), three mutations, but no triallelism.

The extensive genetic heterogeneity of Bardet-Biedl syndrome (BBS) is documented by the identification, by classical linkage analysis complemented recently by comparative genomic approaches, of nine genes (BBS1-9) that account cumulatively for about 50% of patients. The BBS genes appear implicated in cilia and basal body assembly or function. In order to find new BBS genes, we performed SNP homozygosity mapping analysis in an extended consanguineous family living in a small Lebanese village. This uncovered an unexpectedly complex pattern of mutations, and led us to identify a novel BBS gene (BBS10). In one sibship of the pedigree, a BBS2 homozygous mutation was identified, while in three other sibships, a homozygous missense mutation was identified in a gene encoding a vertebrate-specific chaperonine-like protein (BBS10). The single patient in the last sibship was a compound heterozygote for the above BBS10 mutation and another one in the same gene. Although triallelism (three deleterious alleles in the same patient) has been described in some BBS families, we have to date no evidence that this is the case in the present family. The analysis of this family challenged linkage analysis based on the expectation of a single locus and mutation. The very high informativeness of SNP arrays was instrumental in elucidating this case, which illustrates possible pitfalls of homozygosity mapping in extended families, and that can be explained by the rather high prevalence of heterozygous carriers of BBS mutations (estimated at one in 50 in Europeans).

Adolescent↗

PromAn: an integrated knowledge-based web server dedicated to promoter analysis.

PromAn is a modular web-based tool dedicated to promoter analysis that integrates distinct complementary databases, methods and programs. PromAn provides automatic analysis of a genomic region with minimal prior knowledge of the genomic sequence. Prediction programs and experimental databases are combined to locate the transcription start site (TSS) and the promoter region within a large genomic input sequence. Transcription factor binding sites (TFBSs) can be predicted using several public databases and user-defined motifs. Also, a phylogenetic footprinting strategy, combining multiple alignment of large genomic sequences and assignment of various scores reflecting the evolutionary selection pressure, allows for evaluation and ranking of TFBS predictions. PromAn results can be displayed in an interactive graphical user interface, PromAnGUI. It integrates all of this information to highlight active promoter regions, to identify among the huge number of TFBS predictions those which are the most likely to be potentially functional and to facilitate user refined analysis. Such an integrative approach is essential in the face of a growing number of tools dedicated to promoter analysis in order to propose hypotheses to direct further experimental validations. PromAn is publicly available at http://bips.u-strasbg.fr/PromAn.

Binding Sites↗

MACSIMS: multiple alignment of complete sequences information management system.

BACKGROUND: In the post-genomic era, systems-level studies are being performed that seek to explain complex biological systems by integrating diverse resources from fields such as genomics, proteomics or transcriptomics. New information management systems are now needed for the collection, validation and analysis of the vast amount of heterogeneous data available. Multiple alignments of complete sequences provide an ideal environment for the integration of this information in the context of the protein family. RESULTS: MACSIMS is a multiple alignment-based information management program that combines the advantages of both knowledge-based and ab initio sequence analysis methods. Structural and functional information is retrieved automatically from the public databases. In the multiple alignment, homologous regions are identified and the retrieved data is evaluated and propagated from known to unknown sequences with these reliable regions. In a large-scale evaluation, the specificity of the propagated sequence features is estimated to be >99%, i.e. very few false positive predictions are made. MACSIMS is then used to characterise mutations in a test set of 100 proteins that are known to be involved in human genetic diseases. The number of sequence features associated with these proteins was increased by 60%, compared to the features available in the public databases. An XML format output file allows automatic parsing of the MACSIM results, while a graphical display using the JalView program allows manual analysis. CONCLUSION: MACSIMS is a new information management system that incorporates detailed analyses of protein families at the structural, functional and evolutionary levels. MACSIMS thus provides a unique environment that facilitates knowledge extraction and the presentation of the most pertinent information to the biologist. A web server and the source code are available at http://bips.u-strasbg.fr/MACSIMS/.

Algorithms↗

Phosphorylation by PKA potentiates retinoic acid receptor alpha activity by means of increasing interaction with and phosphorylation by cyclin H/cdk7.

Nuclear retinoic acid receptors (RARs) work as ligand-dependent heterodimeric RAR/retinoid X receptor transcription activators, which are targets for phosphorylations. The N-terminal activation function (AF)-1 domain of RARalpha is phosphorylated by the cyclin-dependent kinase (cdk) 7/cyclin H complex of the general transcription factor TFIIH and the C-terminal AF-2 domain by the cAMP-dependent protein kinase A (PKA). Here, we report the identification of a molecular pathway by which phosphorylation by PKA propagates cAMP signaling from the AF-2 domain to the AF-1 domain. The first step is the phosphorylation of S369, located in loop 9-10 of the AF-2 domain. This signal is transferred to the cyclin H binding domain (at the N terminus of helix 9 and loop 8-9), resulting in enhanced cyclin H interaction and, thereby, greater amounts of RARalpha phosphorylated at S77 located in the AF-1 domain by the cdk7/cyclin H complex. This molecular mechanism relies on the integrity of the ligand-binding domain and the cyclin H binding surface. Finally, it results in higher DNA-binding efficiency, providing an explanation for how cAMP synergizes with retinoic acid for transcription.

Amino Acid Sequence↗

BBS10 encodes a vertebrate-specific chaperonin-like protein and is a major BBS locus.

Bardet-Biedl syndrome (BBS) is a genetically heterogeneous ciliopathy. Although nine BBS genes have been cloned, they explain only 40-50% of the total mutational load. Here we report a major new BBS locus, BBS10, that encodes a previously unknown, rapidly evolving vertebrate-specific chaperonin-like protein. We found BBS10 to be mutated in about 20% of an unselected cohort of families of various ethnic origins, including some families with mutations in other BBS genes, consistent with oligogenic inheritance. In zebrafish, mild suppression of bbs10 exacerbated the phenotypes of other bbs morphants.

Bardet-Biedl Syndrome↗

The evolutionary origin of peroxisomes: an ER-peroxisome connection.

The peroxisome is an essential eukaryotic organelle, crucial for lipid metabolism and free radical detoxification, development, differentiation, and morphogenesis from yeasts to humans. Loss of peroxisomes invariably leads to fatal peroxisome biogenesis disorders in man. The evolutionary origin of peroxisomes remains unsolved; proposals for either a symbiogenetic or cellular membrane invagination event are unconclusive. To address this question, we have probed with a peroxisomal proteome, an "ensemble" of 19 representative eukaryotic complete genomes. Molecular phylogenetic and sequence comparison tools allowed us to identify four proteins as peroxisomal markers for unequivocal in silico peroxisome detection. We have then detected the Apicomplexa phylum as the first group of organisms devoid of peroxisomes, in the presence of mitochondria. Finally, we deliver evidence against a prokaryotic ancestor of peroxisomes: (1) the peroxisomal membrane is composed of purely eukaryotic bricks and is thus useful to trace the eukaryotes in their evolutionary paths and (2) the peroxisomal matrix protein import system shares mechanistic similarities with the endoplasmic reticulum/proteasome degradation process, indicating a common evolutionary history.

Animals↗

Polyglutamine expansion causes neurodegeneration by altering the neuronal differentiation program.

Huntington's disease (HD) and spinocerebellar ataxia type 7 (SCA7) belong to a group of inherited neurodegenerative diseases caused by polyglutamine (polyQ) expansion in corresponding proteins. Transcriptional alteration is a unifying feature of polyQ disorders; however, the relationship between polyQ-induced gene expression deregulation and degenerative processes remains unclear. R6/2 and R7E mouse models of HD and SCA7, respectively, present a comparable retinal degeneration characterized by progressive reduction of electroretinograph activity and important morphological changes of rod photoreceptors. The retina, which is a simple central nervous system tissue, allows correlating functional, morphological and molecular defects. Taking advantage of comparing polyQ-induced degeneration in two retina models, we combined gene expression profiling and molecular biology techniques to decipher the molecular pathways underlying polyQ expansion toxicity. We show that R7E and R6/2 retinal phenotype strongly correlates with loss of expression of a large cohort of genes specifically involved in phototransduction function and morphogenesis of differentiated rod photoreceptors. Accordingly, three key transcription factors (Nrl, Crx and Nr2e3) controlling rod differentiation genes, hence expression of photoreceptor specific traits, are down-regulated. Interestingly, other transcription factors known to cause inhibitory effects on photoreceptor differentiation when mis-expressed, such as Stat3, are aberrantly re-activated. Thus, our results suggest that independently from the protein context, polyQ expansion overrides the control of neuronal differentiation and maintenance, thereby causing dysfunction and degeneration.

Animals↗

ICDS database: interrupted CoDing sequences in prokaryotic genomes.

Unrecognized frameshifts, in-frame stop codons and sequencing errors lead to Interrupted CoDing Sequence (ICDS) that can seriously affect all subsequent steps of functional characterization, from in silico analysis to high-throughput proteomic projects. Here, we describe the Interrupted CoDing Sequence database containing ICDS detected by a similarity-based approach in 80 complete prokaryotic genomes. ICDS can be retrieved by species browsing or similarity searches via a web interface (http://www-bio3d-igbmc.u-strasbg.fr/ICDS/). The definition of each interrupted gene is provided as well as the ICDS genomic localization with the surrounding sequence. Furthermore, to facilitate the experimental characterization of ICDS, we propose optimized primers for re-sequencing purposes. The database will be regularly updated with additional data from ongoing sequenced genomes. Our strategy has been validated by three independent tests: (i) ICDS prediction on a benchmark of artificially created frameshifts, (ii) comparison of predicted ICDS and results obtained from the comparison of the two genomic sequences of Bacillus licheniformis strain ATCC 14580 and (iii) re-sequencing of 25 predicted ICDS of the recently sequenced genome of Mycobacterium smegmatis. This allows us to estimate the specificity and sensitivity (95 and 82%, respectively) of our program and the efficiency of primer determination.

Bacillus↗

BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.

Multiple sequence alignment is one of the cornerstones of modern molecular biology. It is used to identify conserved motifs, to determine protein domains, in 2D/3D structure prediction by homology and in evolutionary studies. Recently, high-throughput technologies such as genome sequencing and structural proteomics have lead to an explosion in the amount of sequence and structure information available. In response, several new multiple alignment methods have been developed that improve both the efficiency and the quality of protein alignments. Consequently, the benchmarks used to evaluate and compare these methods must also evolve. We present here the latest release of the most widely used multiple alignment benchmark, BAliBASE, which provides high quality, manually refined, reference alignments based on 3D structural superpositions. Version 3.0 of BAliBASE includes new, more challenging test cases, representing the real problems encountered when aligning large sets of complex sequences. Using a novel, semiautomatic update protocol, the number of protein families in the benchmark has been increased and representative test cases are now available that cover most of the protein fold space. The total number of proteins in BAliBASE has also been significantly increased from 1444 to 6255 sequences. In addition, full-length sequences are now provided for all test cases, which represent difficult cases for both global and local alignment programs. Finally, the BAliBASE Web site (http://www-bio3d-igbmc.u-strasbg.fr/balibase) has been completely redesigned to provide a more user-friendly, interactive interface for the visualization of the BAliBASE reference alignments and the associated annotations.

Amino Acid Sequence↗

Cloning, purification and crystallization of a Walker-type Pyrococcus abyssi ATPase family member.

Several ATPase proteins play essential roles in the initiation of chromosomal DNA replication in archaea. Walker-type ATPases are defined by their conserved Walker A and B motifs, which are associated with nucleotide binding and ATP hydrolysis. A family of 28 ATPase proteins with non-canonical Walker A sequences has been identified by a bioinformatics study of comparative genomics in Pyrococcus genomes. A high-throughput structural study on P. abyssi has been started in order to establish the structure of these proteins. 16 genes have been cloned and characterized. Six out of the seven soluble constructs were purified in Escherichia coli and one of them, PABY2304, has been crystallized. X-ray diffraction data were collected from selenomethionine-derivative crystals using synchrotron radiation. The crystals belong to the orthorhombic space group C2, with unit-cell parameters a = 79.41, b = 48.63, c = 108.77 A, and diffract to beyond 2.6 A resolution.

Adenosine Triphosphatases↗

Sequence and comparative genomic analysis of actin-related proteins.

Actin-related proteins (ARPs) are key players in cytoskeleton activities and nuclear functions. Two complexes, ARP2/3 and ARP1/11, also known as dynactin, are implicated in actin dynamics and in microtubule-based trafficking, respectively. ARP4 to ARP9 are components of many chromatin-modulating complexes. Conventional actins and ARPs codefine a large family of homologous proteins, the actin superfamily, with a tertiary structure known as the actin fold. Because ARPs and actin share high sequence conservation, clear family definition requires distinct features to easily and systematically identify each subfamily. In this study we performed an in depth sequence and comparative genomic analysis of ARP subfamilies. A high-quality multiple alignment of approximately 700 complete protein sequences homologous to actin, including 148 ARP sequences, allowed us to extend the ARP classification to new organisms. Sequence alignments revealed conserved residues, motifs, and inserted sequence signatures to define each ARP subfamily. These discriminative characteristics allowed us to develop ARPAnno (http://bips.u-strasbg.fr/ARPAnno), a new web server dedicated to the annotation of ARP sequences. Analyses of sequence conservation among actins and ARPs highlight part of the actin fold and suggest interactions between ARPs and actin-binding proteins. Finally, analysis of ARP distribution across eukaryotic phyla emphasizes the central importance of nuclear ARPs, particularly the multifunctional ARP4.

Actins↗

MAO: a Multiple Alignment Ontology for nucleic acid and protein sequences.

The application of high-throughput techniques such as genomics, proteomics or transcriptomics means that vast amounts of heterogeneous data are now available in the public databases. Bioinformatics is responding to the challenge with new integrated management systems for data collection, validation and analysis. Multiple alignments of genomic and protein sequences provide an ideal environment for the integration of this mass of information. In the context of the sequence family, structural and functional data can be evaluated and propagated from known to unknown sequences. However, effective integration is being hindered by syntactic and semantic differences between the different data resources and the alignment techniques employed. One solution to this problem is the development of an ontology that systematically defines the terms used in a specific domain. Ontologies are used to share data from different resources, to automatically analyse information and to represent domain knowledge for non-experts. Here, we present MAO, a new ontology for multiple alignments of nucleic and protein sequences. MAO is designed to improve interoperation and data sharing between different alignment protocols for the construction of a high quality, reliable multiple alignment in order to facilitate knowledge extraction and the presentation of the most pertinent information to the biologist.

Databases, Genetic↗

vALId: validation of protein sequence quality based on multiple alignment data.

The validation of sequences is essential to perform accurate phylogeny and structure/function analysis. However among the thousands of protein sequences available in the public databases, most have been predicted in silico and have not systematically undergone a quality verification. It has recently become evident that they often contain sequence errors. To address the problem of automatic protein quality control, we have developed vALId, an interactive web interfaced software. Taking advantage of high quality multiple alignments of complete protein sequences (MACS), vALId first warns about the presence of suspicious insertions, deletions (indels) and divergent segments, and second, proposes corrections based on transcripts and genome contigs. In a first evaluation test, hundreds of indels and divergent segments were randomly generated in a manually refined MACS. The sensitivity (Sn) and specificity (Sp) of indel detection were excellent (0.96) while the mean Sn(0.49) and Sp(0.56) of divergent segment delineation depended on the percent identity between sequence neighbors. In a second test, 6195 sequences in 100 MACS corresponding to different functional and structural protein families were analyzed. 65% of the sequences were in silico predictions and 44% of eukaryote predicted proteins were partially incorrect with at least one suspicious indel or divergent segment.

Algorithms↗

Domain architecture of the p62 subunit from the human transcription/repair factor TFIIH deduced by limited proteolysis and mass spectrometry analysis.

TFIIH is a multiprotein complex that plays a central role in both transcription and DNA repair. The subunit p62 is a structural component of the TFIIH core that is known to interact with VP16, p53, Eralpha, and E2F1 in the context of activated transcription, as well as with the endonuclease XPG in DNA repair. We used limited proteolysis experiments coupled to mass spectrometry to define structural domains within the conserved N-terminal part of the molecule. The first domain identified resulted from spontaneous proteolysis and corresponds to residues 1-108. The second domain encompasses residues 186-240, and biophysical characterization by fluorescence studies and NMR analysis indicated that it is at least partially folded and thus may correspond to a structural entity. This module contains a region of high sequence conservation with an invariant FWxxPhiPhi motif (Phi representing either tyrosine or phenylalanine), which was also found in other protein families and could play a key role as a protein-protein recognition module within TFIIH. The approach used in this study is general and can be straightforwardly applied to other multidomain proteins and/or multiprotein assemblies.

Amino Acid Motifs↗