PubMed Health⌕ Search

Biomedical subjects

Alexandre Reymond

Publications and source records attributed to Alexandre Reymond.

At least 19 recordsLinked to original sources

EGASP: the human ENCODE Genome Annotation Assessment Project.

BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a 'reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40% to 50% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.

Alternative Splicing↗

GENCODE: producing a reference annotation for ENCODE.

BACKGROUND: The GENCODE consortium was formed to identify and map all protein-coding genes within the ENCODE regions. This was achieved by a combination of initial manual annotation by the HAVANA team, experimental validation by the GENCODE consortium and a refinement of the annotation based on these experimental results. RESULTS: The GENCODE gene features are divided into eight different categories of which only the first two (known and novel coding sequence) are confidently predicted to be protein-coding genes. 5' rapid amplification of cDNA ends (RACE) and RT-PCR were used to experimentally verify the initial annotation. Of the 420 coding loci tested, 229 RACE products have been sequenced. They supported 5' extensions of 30 loci and new splice variants in 50 loci. In addition, 46 loci without evidence for a coding sequence were validated, consisting of 31 novel and 15 putative transcripts. We assessed the comprehensiveness of the GENCODE annotation by attempting to validate all the predicted exon boundaries outside the GENCODE annotation. Out of 1,215 tested in a subset of the ENCODE regions, 14 novel exon pairs were validated, only two of them in intergenic regions. CONCLUSION: In total, 487 loci, of which 434 are coding, have been annotated as part of the GENCODE reference set available from the UCSC browser. Comparison of GENCODE annotation with RefSeq and ENSEMBL show only 40% of GENCODE exons are contained within the two sets, which is a reflection of the high number of alternative splice forms with unique exons annotated. Over 50% of coding loci have been experimentally verified by 5' RACE for EGASP and the GENCODE collaboration is continuing to refine its annotation of 1% human genome with the aid of experimental validation.

Chromosome Mapping↗

Submicroscopic deletion in patients with Williams-Beuren syndrome influences expression levels of the nonhemizygous flanking genes.

Genomic imbalance is a common cause of phenotypic abnormalities. We measured the relative expression level of genes that map within the microdeletion that causes Williams-Beuren syndrome and within its flanking regions. We found, unexpectedly, that not only hemizygous genes but also normal-copy neighboring genes show decreased relative levels of expression. Our results suggest that not only the aneuploid genes but also the flanking genes that map several megabases away from a genomic rearrangement should be considered possible contributors to the phenotypic variation in genomic disorders.

Cell Line, Transformed↗

Conserved noncoding sequences are selectively constrained and not mutation cold spots.

Noncoding genetic variants are likely to influence human biology and disease, but recognizing functional noncoding variants is difficult. Approximately 3% of noncoding sequence is conserved among distantly related mammals, suggesting that these evolutionarily conserved noncoding regions (CNCs) are selectively constrained and contain functional variation. However, CNCs could also merely represent regions with lower local mutation rates. Here we address this issue and show that CNCs are selectively constrained in humans by analyzing HapMap genotype data. Specifically, new (derived) alleles of SNPs within CNCs are rarer than new alleles in nonconserved regions (P = 3 x 10(-18)), indicating that evolutionary pressure has suppressed CNC-derived allele frequencies. Intronic CNCs and CNCs near genes show greater allele frequency shifts, with magnitudes comparable to those for missense variants. Thus, conserved noncoding variants are more likely to be functional. Allele frequency distributions highlight selectively constrained genomic regions that should be intensively surveyed for functionally important variation.

Conserved Sequence↗

Tandem chimerism as a means to increase protein complexity in the human genome.

The "one-gene, one-protein" rule, coined by Beadle and Tatum, has been fundamental to molecular biology. The rule implies that the genetic complexity of an organism depends essentially on its gene number. The discovery, however, that alternative gene splicing and transcription are widespread phenomena dramatically altered our understanding of the genetic complexity of higher eukaryotic organisms; in these, a limited number of genes may potentially encode a much larger number of proteins. Here we investigate yet another phenomenon that may contribute to generate additional protein diversity. Indeed, by relying on both computational and experimental analysis, we estimate that at least 4%-5% of the tandem gene pairs in the human genome can be eventually transcribed into a single RNA sequence encoding a putative chimeric protein. While the functional significance of most of these chimeric transcripts remains to be determined, we provide strong evidence that this phenomenon does not correspond to mere technical artifacts and that it is a common mechanism with the potential of generating hundreds of additional proteins in the human genome.

Gene Fusion↗

Emergence of young human genes after a burst of retroposition in primates.

The origin of new genes through gene duplication is fundamental to the evolution of lineage- or species-specific phenotypic traits. In this report, we estimate the number of functional retrogenes on the lineage leading to humans generated by the high rate of retroposition (retroduplication) in primates. Extensive comparative sequencing and expression studies coupled with evolutionary analyses and simulations suggest that a significant proportion of recent retrocopies represent bona fide human genes. We estimate that at least one new retrogene per million years emerged on the human lineage during the past approximately 63 million years of primate evolution. Detailed analysis of a subset of the data shows that the majority of retrogenes are specifically expressed in testis, whereas their parental genes show broad expression patterns. Consistently, most retrogenes evolved functional roles in spermatogenesis. Proteins encoded by X chromosome-derived retrogenes were strongly preserved by purifying selection following the duplication event, supporting the view that they may act as functional autosomal substitutes during X-inactivation of late spermatogenesis genes. Also, some retrogenes acquired a new or more adapted function driven by positive selection. We conclude that retroduplication significantly contributed to the formation of recent human genes and that most new retrogenes were progressively recruited during primate evolution by natural and/or sexual selection to enhance male germline function.

Animals↗

A novel TMPRSS3 missense mutation in a DFNB8/10 family prevents proteolytic activation of the protein.

Pathogenic mutations in TMPRSS3, which encodes a transmembrane serine protease, cause non-syndromic deafness DFNB8/10. Missense mutations map in the low density-lipoprotein receptor A (LDLRA), scavenger-receptor cysteine-rich (SRCR), and protease domains of the protein, indicating that all domains are important for its function. TMPRSS3 undergoes proteolytic cleavage and activates the ENaC sodium channel in a Xenopus oocyte model system. To assess the importance of this gene in non-syndromic childhood or congenital deafness in Turkey, we screened for mutations affected members of 25 unrelated Turkish families. The three families with the highest LOD score for linkage to chromosome 21q22.3 were shown to harbor P404L, R216L, or Q398X mutations, suggesting that mutations in TMPRSS3 are a considerable contributor to non-syndromic deafness in the Turkish population. The mutant TMPRSS3 harboring the novel R216L missense mutation within the predicted cleavage site of the protein fails to undergo proteolytic cleavage and is unable to activate ENaC, thus providing evidence that pre-cleavage of TMPRSS3 is mandatory for normal function.

Amino Acid Sequence↗

LKB1 interacts with and phosphorylates PTEN: a functional link between two proteins involved in cancer predisposing syndromes.

Germline mutations of the LKB1 (STK11) tumor suppressor gene lead to Peutz-Jeghers syndrome (PJS) and predisposition to cancer. LKB1 encodes a serine/threonine kinase generally inactivated in PJS patients. We identified the dual phosphatase and tumor suppressor protein PTEN as an LKB1-interacting protein. Several LKB1 point mutations associated with PJS disrupt the interaction with PTEN suggesting that the loss of this interaction might contribute to PJS. Although PTEN and LKB1 are predominantly cytoplasmic and nuclear, respectively, their interaction leads to a cytoplasmic relocalization of LKB1. In addition, we show that PTEN is a substrate of the kinase LKB1 in vitro. As PTEN is a dual phosphatase mutated in autosomal inherited disorders with phenotypes similar to those of PJS (Bannayan-Riley-Ruvalcaba syndrome and Cowden disease), our study suggests a functional link between the proteins involved in different hamartomatous polyposis syndromes and emphasizes the central role played by LKB1 as a tumor suppressor in the small intestine.

AMP-Activated Protein Kinase Kinases↗

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals↗

Comparative gene finding in chicken indicates that we are closing in on the set of multi-exonic widely expressed human genes.

The recent availability of the chicken genome sequence poses the question of whether there are human protein-coding genes conserved in chicken that are currently not included in the human gene catalog. Here, we show, using comparative gene finding followed by experimental verification of exon pairs by RT-PCR, that the addition to the multi-exonic subset of this catalog could be as little as 0.2%, suggesting that we may be closing in on the human gene set. Our protocol, however, has two shortcomings: (i) the bioinformatic screening of the predicted genes, applied to filter out false positives, cannot handle intronless genes; and (ii) the experimental verification could fail to identify expression at a specific developmental time. This highlights the importance of developing methods that could provide a reliable estimate of the number of these two types of genes.

Animals↗

Different mechanisms preclude mutant CLDN14 proteins from forming tight junctions in vitro.

Mutations in claudin 14 (CLDN14) cause nonsyndromic DFNB29 deafness in humans. The analysis of a murine model indicated that this phenotype is associated with degeneration of hair cells, possibly due to cation overload. However, the mechanism linking these alterations to CLDN14 mutations is unknown. To investigate this mechanism, we compared the ability of wild-type and missense mutant CLDN14 to form tight junctions. Ectopic expression in L mouse fibroblasts (LM cells) of wild-type CLDN14 protein induced the formation of tight junctions, while both the c.254T>A (p.V85D) mutant, previously identified in a Pakistani family, and the c.301 G>A (p.G101R) mutant, identified in this study through the screen of 183 Spanish and Greek patients affected with sporadic nonsyndromic deafness, failed to form such junctions. However, the two mutant proteins differed in their ability to localize at the plasma membrane. We further identified hitherto undescribed exons of CLDN14 that are utilized in alternative spliced transcripts. We demonstrated that different mutations of CLDN14 impaired by different mechanisms the ability of the protein to form tight junctions. Our results indicate that the ability of CLDN14 to be recruited to these junctions is crucial for the hearing process.

Alternative Splicing↗

Conserved non-genic sequences - an unexpected feature of mammalian genomes.

Mammalian genomes contain highly conserved sequences that are not functionally transcribed. These sequences are single copy and comprise approximately 1-2% of the human genome. Evolutionary analysis strongly supports their functional conservation, although their potentially diverse, functional attributes remain unknown. It is likely that genomic variation in conserved non-genic sequences is associated with phenotypic variability and human disorders. So how might their function and contribution to human disorders be examined?

Animals↗

Evolutionary comparison provides evidence for pathogenicity of RMRP mutations.

Cartilage-hair hypoplasia (CHH) is a pleiotropic disease caused by recessive mutations in the RMRP gene that result in a wide spectrum of manifestations including short stature, sparse hair, metaphyseal dysplasia, anemia, immune deficiency, and increased incidence of cancer. Molecular diagnosis of CHH has implications for management, prognosis, follow-up, and genetic counseling of affected patients and their families. We report 20 novel mutations in 36 patients with CHH and describe the associated phenotypic spectrum. Given the high mutational heterogeneity (62 mutations reported to date), the high frequency of variations in the region (eight single nucleotide polymorphisms in and around RMRP), and the fact that RMRP is not translated into protein, prediction of mutation pathogenicity is difficult. We addressed this issue by a comparative genomic approach and aligned the genomic sequences of RMRP gene in the entire class of mammals. We found that putative pathogenic mutations are located in highly conserved nucleotides, whereas polymorphisms are located in non-conserved positions. We conclude that the abundance of variations in this small gene is remarkable and at odds with its high conservation through species; it is unclear whether these variations are caused by a high local mutation rate, a failure of repair mechanisms, or a relaxed selective pressure. The marked diversity of mutations in RMRP and the low homozygosity rate in our patient population indicate that CHH is more common than previously estimated, but may go unrecognized because of its variable clinical presentation. Thus, RMRP molecular testing may be indicated in individuals with isolated metaphyseal dysplasia, anemia, or immune dysregulation.

Animals↗

The subcellular localization of the ChoRE-binding protein, encoded by the Williams-Beuren syndrome critical region gene 14, is regulated by 14-3-3.

The Williams-Beuren syndrome (WBS) is a contiguous gene syndrome caused by chromosomal rearrangements at chromosome band 7q11.23. Several endocrine phenotypes, in particular impaired glucose tolerance and silent diabetes, have been described for this clinically complex disorder. The WBSCR14 gene, one of the genes mapping to the WBS critical region, encodes a member of the basic-helix-loop-helix leucine zipper family of transcription factors, which dimerizes with the Max-like protein, Mlx. This heterodimeric complex binds and activates, in a glucose-dependent manner, carbohydrate response element (ChoRE) motifs in the promoter of lipogenic enzymes. We identified five novel WBSCR14-interacting proteins, four 14-3-3 isotypes and NIF3L1, which form a single polypeptide complex in mammalian cells. Phosphatase treatment abrogates the association between WBSCR14 and 14-3-3, as shown previously for multiple 14-3-3 interactors. WBSCR14 is exported actively from the nucleus through a CRM1-dependent mechanism. This translocation is contingent upon the ability to bind 14-3-3. Through this mechanism the 14-3-3 isotypes directly affect the WBSCR14:Mlx complexes, which activate the transcription of lipogenic genes.

14-3-3 Proteins↗

Comparison of human chromosome 21 conserved nongenic sequences (CNGs) with the mouse and dog genomes shows that their selective constraint is independent of their genic environment.

The analysis of conservation between the human and mouse genomes resulted in the identification of a large number of conserved nongenic sequences (CNGs). The functional significance of this nongenic conservation remains unknown, however. The availability of the sequence of a third mammalian genome, the dog, allows for a large-scale analysis of evolutionary attributes of CNGs in mammals. We have aligned 1638 previously identified CNGs and 976 conserved exons (CODs) from human chromosome 21 (Hsa21) with their orthologous sequences in mouse and dog. Attributes of selective constraint, such as sequence conservation, clustering, and direction of substitutions were compared between CNGs and CODs, showing a clear distinction between the two classes. We subsequently performed a chromosome-wide analysis of CNGs by correlating selective constraint metrics with their position on the chromosome and relative to their distance from genes. We found that CNGs appear to be randomly arranged in intergenic regions, with no bias to be closer or farther from genes. Moreover, conservation and clustering of substitutions of CNGs appear to be completely independent of their distance from genes. These results suggest that the majority of CNGs are not typical of previously described regulatory elements in terms of their location. We propose models for a global role of CNGs in genome function and regulation, through long-distance cis or trans chromosomal interactions.

Animals↗

Knobloch syndrome: novel mutations in COL18A1, evidence for genetic heterogeneity, and a functionally impaired polymorphism in endostatin.

Knobloch syndrome (KNO) is an autosomal recessive disorder characterized by high myopia, vitreoretinal degeneration with retinal detachment, and congenital encephalocele. Pathogenic mutations in the COL18A1 gene on 21q22.3 were recently identified in KNO families. Analysis of two unrelated KNO families from Hungary and New Zealand allowed us to confirm the involvement of COL18A1 in the pathogenesis of KNO and to demonstrate the existence of genetic heterogeneity. Two COL18A1 mutations were identified in the Hungarian family: a 1-bp insertion causing a frameshift and a premature in-frame stop codon and an amino acid substitution. This missense variant is located in a conserved amino acid of endostatin, a cleavage product of the carboxy-terminal domain of collagen alpha 1 XVIII. D1437N (D104N in endostatin) likely represents a pathogenic mutation, as we show that the endostatin N104 mutant is impaired in its affinity towards laminin. Linkage to the COL18A1 locus was excluded in the New Zealand family, providing evidence for the existence of a second KNO locus. We named the second unmapped locus for Knobloch syndrome KNO2. Mutation analysis excluded COL15A1, a member of the multiplexin collagen subfamily similar to COL18A1, as being responsible for KNO2.

Amino Acid Sequence↗

The Caenorhabditis elegans ortholog of C21orf80, a potential new protein O-fucosyltransferase, is required for normal development.

Down syndrome (DS), as a phenotypic result of trisomy 21, is the most frequent aneuploidy at birth and the most common known genetic cause of mental retardation. DS is also characterized by other phenotypes affecting many organs, including brain, muscle, heart, limbs, gastrointestinal tract, skeleton, and blood. Any of the human chromosome 21 (Hsa21) genes may contribute to some of the DS phenotypes. To determine which of the Hsa21 genes are involved in DS, the effects of disrupting and overexpressing individual human gene orthologs in model organisms, such as the nematode Caenorhabditis elegans, can be analyzed. Here, we isolated and characterized C21orf80 (human chromosome 21 open reading frame 80), a potential novel protein O-fucosyltransferase gene that encodes three alternatively spliced transcripts. Transient expression of tagged C21orf80 proteins suggests a primary intracellular localization in the Golgi apparatus. To gain insight into the biological role of C21orf80 and its potential role in DS, we isolated its C. elegans ortholog, pad-2, and performed RNA interference (RNAi) and overexpression experiments. pad-2(RNAi) embryos showed failure to undergo normal morphogenesis. Transgenic worms with elevated dosage of pad-2 displayed severe body malformations and abnormal neuronal development. These results show that pad-2 is required for normal development and suggest potential roles for C21orf80 in the pathogenesis of DS.

Amino Acid Sequence↗