PubMed Health⌕ Search

Biomedical subjects

M Mar Albà

Publications and source records attributed to M Mar Albà.

18 recordsLinked to original sources

Highly constrained proteins contain an unexpectedly large number of amino acid tandem repeats.

Single-amino-acid tandem repeats are very common in mammalian proteins but their function and evolution are still poorly understood. Here we investigate how the variability and prevalence of amino acid repeats are related to the evolutionary constraints operating on the proteins. We find a significant positive correlation between repeat size difference and protein nonsynonymous substitution rate in human and mouse orthologous genes. This association is observed for all the common amino acid repeat types and indicates that rapid diversification of repeat structures, involving both trinucleotide slippage and nucleotide substitutions, preferentially occurs in proteins subject to low selective constraints. However, strikingly, we also observe a significant negative correlation between the number of repeats in a protein and the gene nonsynonymous substitution rate, particularly for glutamine, glycine, and alanine repeats. This implies that proteins subject to strong selective constraints tend to contain an unexpectedly high number of repeats, which tend to be well conserved between the two species. This is consistent with a role for selection in the maintenance of a significant number of repeats. Analysis of the codon structure of the sequences encoding the repeats shows that codon purity is associated with high repeat size interspecific variability. Interestingly, polyalanine and polyglutamine repeats associated with disease show very distinctive features regarding the degree of repeat conservation and the protein sequence selective constraints.

Amino Acid Sequence↗

PEAKS: identification of regulatory motifs by their position in DNA sequences.

UNLABELLED: Many DNA functional motifs tend to accumulate or cluster at specific gene locations. These locations can be detected, in a group of gene sequences, as high frequency 'peaks' with respect to a reference position, such as the transcription start site (TSS). We have developed a web tool for the identification of regions containing significant motif peaks. We show, by using different yeast gene datasets, that peak regions are strongly enriched in experimentally-validated motifs and contain potentially important novel motifs. AVAILABILITY: http://genomics.imim.es/peaks

Algorithms↗

Differences in the evolutionary history of disease genes affected by dominant or recessive mutations.

BACKGROUND: Global analyses of human disease genes by computational methods have yielded important advances in the understanding of human diseases. Generally these studies have treated the group of disease genes uniformly, thus ignoring the type of disease-causing mutations (dominant or recessive). In this report we present a comprehensive study of the evolutionary history of autosomal disease genes separated by mode of inheritance. RESULTS: We examine differences in protein and coding sequence conservation between dominant and recessive human disease genes. Our analysis shows that disease genes affected by dominant mutations are more conserved than those affected by recessive mutations. This could be a consequence of the fact that recessive mutations remain hidden from selection while heterozygous. Furthermore, we employ functional annotation analysis and investigations into disease severity to support this hypothesis. CONCLUSION: This study elucidates important differences between dominantly- and recessively-acting disease genes in terms of protein and DNA sequence conservation, paralogy and essentiality. We propose that the division of disease genes by mode of inheritance will enhance both understanding of the disease process and prediction of candidate disease genes in the future.

Animals↗

Mutation patterns of amino acid tandem repeats in the human proteome.

BACKGROUND: Amino acid tandem repeats are found in nearly one-fifth of human proteins. Abnormal expansion of these regions is associated with several human disorders. To gain further insight into the mutational mechanisms that operate in this type of sequence, we have analyzed a large number of mutation variants derived from human expressed sequence tags (ESTs). RESULTS: We identified 137 polymorphic variants in 115 different amino acid tandem repeats. Of these, 77 contained amino acid substitutions and 60 contained gaps (expansions or contractions of the repeat unit). The analysis showed that at least about 21% of the repeats might be polymorphic in humans. We compared the mutations found in different types of amino acid repeats and in adjacent regions. Overall, repeats showed a five-fold increase in the number of gap mutations compared to adjacent regions, reflecting the action of slippage within the repetitive structures. Gap and substitution mutations were very differently distributed between different amino acid repeat types. Among repeats containing gap variants we identified several disease and candidate disease genes. CONCLUSION: This is the first report at a genome-wide scale of the types of mutations occurring in the amino acid repeat component of the human proteome. We show that the mutational dynamics of different amino acid repeat types are very diverse. We provide a list of loci with highly variable repeat structures, some of which may be potentially involved in disease.

Amino Acid Substitution↗

ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.

Information about the genomic coordinates and the sequence of experimentally identified transcription factor binding sites is found scattered under a variety of diverse formats. The availability of standard collections of such high-quality data is important to design, evaluate and improve novel computational approaches to identify binding motifs on promoter sequences from related genes. ABS (http://genome.imim.es/datasets/abs2005/index.html) is a public database of known binding sites identified in promoters of orthologous vertebrate genes that have been manually curated from bibliography. We have annotated 650 experimental binding sites from 68 transcription factors and 100 orthologous target genes in human, mouse, rat or chicken genome sequences. Computational predictions and promoter alignment information are also provided for each entry. A simple and easy-to-use web interface facilitates data retrieval allowing different views of the information. In addition, the release 1.0 of ABS includes a customizable generator of artificial datasets based on the known sites contained in the collection and an evaluation tool to aid during the training and the assessment of motif-finding programs.

Animals↗

Inverse relationship between evolutionary rate and age of mammalian genes.

A large number of genes is shared by all living organisms, whereas many others are unique to some specific lineages, indicating their different times of origin. The availability of a growing number of eukaryotic genomes allows us to estimate which mammalian genes are novel genes and, approximately, when they arose. In this article, we classify human genes into four different age groups and estimate evolutionary rates in human and mouse orthologs. We show that older genes tend to evolve more slowly than newer ones; that is, proteins that arose earlier in evolution currently have a larger proportion of sites subjected to negative selection. Interestingly, this property is maintained when a fraction of the fastest-evolving genes is excluded or when only genes belonging to a given functional class are considered. One way to explain this relationship is by assuming that genes maintain their functional constraints along all their evolutionary history, but the nature of more recent evolutionary innovations is such that the functional constraints operating on them are increasingly weaker. Alternatively, our results would also be consistent with a scenario in which the functional constraints acting on a gene would not need to be constant through evolution. Instead, starting from weak functional constraints near the time of origin of a gene-as supported by mechanisms proposed for the origin of orphan genes-there would be a gradual increase in selective pressures with time, resulting in fewer accepted mutations in older versus more novel genes.

Animals↗

Evolutionary conservation and selection of human disease gene orthologs in the rat and mouse genomes.

BACKGROUND: Model organisms have contributed substantially to our understanding of the etiology of human disease as well as having assisted with the development of new treatment modalities. The availability of the human, mouse and, most recently, the rat genome sequences now permit the comprehensive investigation of the rodent orthologs of genes associated with human disease. Here, we investigate whether human disease genes differ significantly from their rodent orthologs with respect to their overall levels of conservation and their rates of evolutionary change. RESULTS: Human disease genes are unevenly distributed among human chromosomes and are highly represented (99.5%) among human-rodent ortholog sets. Differences are revealed in evolutionary conservation and selection between different categories of human disease genes. Although selection appears not to have greatly discriminated between disease and non-disease genes, synonymous substitution rates are significantly higher for disease genes. In neurological and malformation syndrome disease systems, associated genes have evolved slowly whereas genes of the immune, hematological and pulmonary disease systems have changed more rapidly. Amino-acid substitutions associated with human inherited disease occur at sites that are more highly conserved than the average; nevertheless, 15 substituting amino acids associated with human disease were identified as wild-type amino acids in the rat. Rodent orthologs of human trinucleotide repeat-expansion disease genes were found to contain substantially fewer of such repeats. Six human genes that share the same characteristics as triplet repeat-expansion disease-associated genes were identified; although four of these genes are expressed in the brain, none is currently known to be associated with disease. CONCLUSIONS: Most human disease genes have been retained in rodent genomes. Synonymous nucleotide substitutions occur at a higher rate in disease genes, a finding that may reflect increased mutation rates in the chromosomal regions in which disease genes are found. Rodent orthologs associated with neurological function exhibit the greatest evolutionary conservation; this suggests that rodent models of human neurological disease are likely to most faithfully represent human disease processes. However, with regard to neurological triplet repeat expansion-associated human disease genes, the contraction, relative to human, of rodent trinucleotide repeats suggests that rodent loci may not achieve a 'critical repeat threshold' necessary to undergo spontaneous pathological repeat expansions. The identification of six genes in this study that have multiple characteristics associated with repeat expansion-disease genes raises the possibility that not all human loci capable of facilitating neurological disease by repeat expansion have as yet been identified.

Animals↗

Clustering of genes coding for DNA binding proteins in a region of atypical evolution of the human genome.

Comparison of the human and mouse genomes has revealed that significant variations in evolutionary rates exist among genomic regions and that a large part of this variation is interchromosomal. We confirm in this work, using a large collection of introns, that human chromosome 19 is the one that shows the highest divergence with respect to mouse. To search for other differences among chromosomes, we examine the distribution of gene functions in human and mouse chromosomes using the Gene Ontology definitions. We found by correspondence analysis that among the strongest clusterings of gene functions in human chromosomes is a group of genes coding for DNA binding proteins in chromosome 19. Interestingly, chromosome 19 also has a very high GC content, a feature that has been proposed to promote an opening of the chromatin, thereby facilitating binding of proteins to the DNA helix. In the mouse genome, however, a similar aggregation of genes coding for DNA binding proteins and high GC content cannot be found. This suggests that the distribution of genes coding for DNA binding proteins and the variations of the chromatin accessibility to these proteins are different in the human and mouse genomes. It is likely that the overall high synonymous and intron rates in chromosome 19 are a by-product of the high GC content of this chromosome.

Base Composition↗

Comparative analysis of amino acid repeats in rodents and humans.

Amino acid tandem repeats, also called homopolymeric tracts, are extremely abundant in eukaryotic proteins. To gain insight into the genome-wide evolution of these regions in mammals, we analyzed the repeat content in a large data set of rat-mouse-human orthologs. Our results show that human proteins contain more amino acid repeats than rodent proteins and that trinucleotide repeats are also more abundant in human coding sequences. Using the human species as an outgroup, we were able to address differences in repeat loss and repeat gain in the rat and mouse lineages. In this data set, mouse proteins contain substantially more repeats than rat proteins, which can be at least partly attributed to a higher repeat loss in the rat lineage. The data are consistent with a role for trinucleotide slippage in the generation of novel amino acid repeats. We confirm the previously observed functional bias of proteins with repeats, with overrepresentation of transcription factors and DNA-binding proteins. We show that genes encoding amino acid repeats tend to have an unusually high GC content, and that differences in coding GC content among orthologs are directly related to the presence/absence of repeats. We propose that the different GC content isochore structure in rodents and humans may result in an increased amino acid repeat prevalence in the human lineage.

Animals↗

Interaction of the plant glycine-rich RNA-binding protein MA16 with a novel nucleolar DEAD box RNA helicase protein from Zea mays.

The maize RNA-binding MA16 protein is a developmentally and environmentally regulated nucleolar protein that interacts with RNAs through complex association with several proteins. By using yeast two-hybrid screening, we identified a DEAD box RNA helicase protein from Zea mays that interacted with MA16, which we named Z. maysDEAD box RNA helicase 1 (ZmDRH1). The sequence of ZmDRH1 includes the eight RNA helicase motifs and two glycine-rich regions with arginine-glycine-rich (RGG) boxes at the amino (N)- and carboxy (C)-termini of the protein. Both MA16 and ZmDRH1 were located in the nucleus and nucleolus, and analysis of the sequence determinants for their cellular localization revealed that the region containing the RGG motifs in both proteins was necessary for nuclear/nucleolar localization The two domains of MA16, the RNA recognition motif (RRM) and the RGG, were tested for molecular interaction with ZmDRH1. MA16 specifically interacted with ZmDRH1 through the RRM domain. A number of plant proteins and vertebrate p68/p72 RNA helicases showed evolutionary proximity to ZmDRH1. In addition, like p68, ZmDRH1 was able to interact with fibrillarin. Our data suggest that MA16, fibrillarin, and ZmDRH1 may be part of a ribonucleoprotein complex involved in ribosomal RNA (rRNA) metabolism.

Amino Acid Sequence↗

Identification of patterns in biological sequences at the ALGGEN server: PROMO and MALGEN.

In this paper we present several web-based tools to identify conserved patterns in sequences. In particular we present details on the functionality of PROMO version 2.0, a program for the prediction of transcription factor binding site in a single sequence or in a group of related sequences and, of MALGEN, a tool to visualize sequence correspondences among long DNA sequences. The web tools and associated documentation can be accessed at http://www.lsi.upc.es/~alggen (RESEARCH link).

Animals↗

Detecting cryptically simple protein sequences using the SIMPLE algorithm.

MOTIVATION: Low-complexity or cryptically simple sequences are widespread in protein sequences but their evolution and function are poorly understood. To date methods for the detection of low complexity in proteins have been directed towards the filtering of such regions prior to sequence homology searches but not to the analysis of the regions per se. However, many of these regions are encoded by non-repetitive DNA sequences and may therefore result from selection acting on protein structure and/or function. RESULTS: We have developed a new tool, based on the SIMPLE algorithm, that facilitates the quantification of the amount of simple sequence in proteins and determines the type of short motifs that show clustering above a certain threshold. By modifying the sensitivity of the program simple sequence content can be studied at various levels, from highly organised tandem structures to complex combinations of repeats. We compare the relative amount of simplicity in different functional groups of yeast proteins and determine the level of clustering of the different amino acids in these proteins. AVAILABILITY: The program is available on request or online at http://www.biochem.ucl.ac.uk/bsm/SIMPLE.

Algorithms↗

Identification of new herpesvirus gene homologs in the human genome.

Viruses are intracellular parasites that use many cellular pathways during their replication. Large DNA viruses, such as herpesviruses, have captured a repertoire of cellular genes to block or mimic host immune responses, apoptosis regulation, and cell-cycle control mechanisms. We have conducted a systematic search for all homologs of herpesvirus proteins in the human genome using position-specific scoring matrices representing herpesvirus protein sequence domains, and pair-wise sequence comparisons. The analysis shows that approximately 13% of the herpesvirus proteins have clear sequence similarity to products of the human genome. Different human herpesviruses vary in their numbers of human homologs, indicating distinct rates of gene acquisition in different lineages. Our analysis has identified new families of herpesvirus/human homologs from viruses including human herpesvirus 5 (human cytomegalovirus; HCMV) and human herpesvirus 8 (Kaposi's sarcoma-associated herpesvirus; KSHV), which may play important roles in host-virus interactions.

Amino Acid Sequence↗

Virus bioinformatics: databases and recent applications.

Bioinformatics is now used as an umbrella term for almost all aspects of computational biology. Bioinformatics research will have an impact on all of biology, and virology is not immune from these research methods. Although virology has been slower to embrace bioinformatics this is now changing, particularly in the areas of viral sequences databasing and the systematic identification of viral and host homologous proteins. Here we will review some of these recent advances focusing mainly on the herpesvirus.

Computational Biology↗

Amino acid reiterations in yeast are overrepresented in particular classes of proteins and show evidence of a slippage-like mutational process.

Long amino acid repeats are often observed in eukaryotic proteins. In humans, several neurological disorders are caused by proteins containing abnormally long polyglutamines. However, no systematic analysis has attempted to investigate the relationship between reiterations of particular amino acids and protein function, the possible mechanisms involved in the generation of these regions, or the contribution of selection in restricting their genomic distribution, in a large collection of wild-type proteins. We have used baker's yeast open reading frames to study these questions. The most abundant amino acid repeats found in yeast proteins are repeats of glutamine, asparagine, aspartic acid, glutamic acid, and serine. Different amino acid repeats are concentrated in different classes of proteins. Acidic and polar amino acid repeats are significantly associated with transcription factors and protein kinases, while serine repeats are significantly associated with membrane transporter proteins. In most cases the codon structures encoding the repeats at the gene level show a significant bias toward long tracts of one of the possible codons, suggesting that trinucleotide slippage has played an important role in generating these reiterations. However, many, particularly those encoding serine repeats, do not show evidence of slippage. The distributions of codon repeats within proteins and between coding and noncoding regions of the genome, and of amino acids between proteins with different functions, suggest that repeats of these kinds are subject to strong selection.

Base Sequence↗

Expression and cellular localization of rab28 mRNA and Rab28 protein during maize embryogenesis.

The maize abscisic acid (ABA) responsive gene rab28, has been shown to be ABA-inducible in embryos and vegetative tissues. A polyclonal antiserum was raised against Rab28 protein. Using immunoblotting and immunoprecipitation, the antiserum specifically recognized a protein of about 30 kDa and pl 6 which is in close agreement with the molecular weight and pl predicted by the deduced amino acid sequence. The rab28 gene product accumulated during late embryogenesis. In vegetative tissues, dehydration stress induced rab28 gene expression both in the light and in the dark. The spatial and temporal pattern of rab28 mRNA expression during embryogenesis was investigated by in situ hybridization using digoxigenin-labelled rab28 probes, and the immunochemical localization of Rab28 protein using anti-Rab28 antibodies. Expression of rab28 mRNA is restricted to provascular tissues in young embryos, and at later stages of development the most prevalent accumulation occurred in meristem and in the vascular elements of the plumule, root and scutellum. Using immunoelectron microscopy the Rab28 protein has been located in the nucleolus of different cell types. In light of these results the stress regulation of rab28 and a likely role for this protein during late embryogenesis are discussed.

Antibody Specificity↗

The maize abscisic acid-responsive protein Rab17 is located in the nucleus and interacts with nuclear localization signals.

The maize abscisic acid (ABA)-responsive rab17 mRNA and Rab17 protein distribution in maize embryo tissues was investigated by in situ hybridization and immunocytochemistry. rab17 mRNA and Rab17 protein were found in all cells of embryo tissues. Synthesis of rab17 mRNA occurred initially in the embryo axis. As maturation progressed, rab17 mRNA was detectable in the scutellum and accumulated in axis cells and provascular tissues. However, the response to exogenous ABA differed in various embryo cell types. The Rab17 protein was located in the nucleus and in the cytoplasm, and qualitative differences in the phosphorylation states of the protein were found between the two subcellular compartments. Based on the similar domain arrangements of Rab17 and a nuclear localization signal (NLS) binding phosphoprotein, Nopp140, interaction of Rab17 with NLS peptides was studied. We found specific binding of Rab17 to the wild-type NLS of the SV40 T antigen but not to an import incompetent mutant peptide. Moreover, binding of the NLS peptide to Rab17 was found to be dependent upon phosphorylation. These results suggest that Rab17 may play a role in nuclear protein transport.

Abscisic Acid↗