PubMed Health⌕ Search

Biomedical subjects

Laurence D Hurst

Publications and source records attributed to Laurence D Hurst.

At least 19 recordsLinked to original sources

Protein evolution: causes of trends in amino-acid gain and loss.

Understanding how proteins evolve is important for determining the molecular basis of adaptation, for inferring phylogenies and for engineering novel proteins. It has been suggested that some amino acids were incorporated into the genetic code more recently than others and, after comparing pairs of closely related genomes, Jordan et al. report that 'recent' amino acids are becoming more common. They argue that this process has been going on since the genetic code first evolved to encompass all 20 amino acids. Here we provide evidence that the patterns observed conform with standard, nearly neutral theoretical expectations and require no new explanation. This reinforces the need for caution in the interpretation of results derived from closely related taxa.

Amino Acids↗

Is optimal gene order impossible?

Recent evidence suggests that yeast genes encoding proteins that are present in the same protein complex tend to be linked and to be co-expressed. More generally, we found that genes that are close to each other in the protein interaction network tend to be linked more often than expected and are often co-expressed. Unexpectedly, we found that linked genes in network proximity have unusually high recombination rates. Because high recombination rates are associated with high rates of genome re-organization, our findings might explain why the clustering of genes in proximity in the network is such a weak effect: there could be a co-evolutionary cycle of physical linkage for co-expression, upwards modification of the recombination rate and concomitant break-up of a cluster. Under such a model an "optimal" gene order is never stable.

Evolution, Molecular↗

Preliminary assessment of the impact of microRNA-mediated regulation on coding sequence evolution in mammals.

Despite prior claims to the contrary, several lines of evidence suggest that selection acts on synonymous mutations in mammals. What might be the mechanisms for such selection? Here I attempt to quantify the constraints on the evolution of the coding sequence resulting from regulation of mRNA by microRNAs (miRNAs) that antisense-bind to the coding region of mRNAs. I employ a set of genes recently experimentally verified to be the target of a miRNA, all with putative antisense pairing domains within the coding sequence. Although very small ( approximately 22 nucleotides), 2 of 13 pairing domains show evidence of significantly slow sequence evolution. This, along with evidence that these genes are regulated by the miRNA under consideration, provides the first good candidate domains for intra-CDS pairing of a miRNA in mammals. When analyzed en masse, the putative pairing domains have a significantly reduced rate of synonymous evolution (approximately 35% lower than null). However, given the size and rarity of pairing domains within the coding sequence, the effects that such constraint has on estimates of the mutation rate are small enough to be ignored (probably less than 1% reduction). The pairing sites also have low Ka values and the selection on the synonymous sites is unlikely to lead to misleading reports of localized high Ka/Ks ratios.

3' Untranslated Regions↗

Evidence for variation in abundance of antisense transcripts between multicellular animals but no relationship between antisense transcriptionand organismic complexity.

Given that humans have about the same number of genes as mice and not so many more than worm, what makes us more complex? Antisense transcripts are implicated in many aspects of gene regulation. Is there a functional connection between antisense transcription and organismic complexity, that is, is antisense regulation especially prevalent in humans? We used the same robust protocol to identify antisense transcripts in humans and five other metazoan genomes (mouse, rat, chicken, fruit fly, and nematode), and found that the estimated proportions of genes involved in antisense transcription are highly sensitive to the number of transcripts included in the analysis. By controlling for transcript abundance, we find that the probability that any given transcript is putatively involved in sense-antisense regulation is no higher in humans than in other vertebrates but appears unusually high in flies and especially low in nematodes. Similarly, there is no evidence that the proportion of sense-antisense transcripts is especially higher in humans than other vertebrates in a given subset of transcript sequences such as mRNAs, coding sequences, conserved, or nonconserved transcripts. Although antisense transcription might be enriched in mammalian brains compared with nonbrain tissues, it is no more enriched in human brain than in mouse brain. Overall, therefore, while we see striking variation between multicellular animals in the abundance of antisense transcripts, there is no evidence for a link between antisense transcription and organismic complexity. More particularly, we see no evidence that humans are in any way unusual among the vertebrates in this regard. Instead, our results suggest that antisense transcription might be prevalent in almost all metazoan genomes, nematodes being an unexplained exception.

Animals↗

Evolutionary and physiological importance of hub proteins.

It has been claimed that proteins with more interaction partners (hubs) are both physiologically more important (i.e., less dispensable) and, owing to an assumed high density of binding sites, slow evolving. Not all analyses, however, support these results, probably because of biased and less-than reliable global protein interaction data. Here we provide the first examination of these issues using a comprehensive literature-curated dataset of well-substantiated protein interactions in Saccharomyces cerevisiae. Whereas use of less reliable yeast two-hybrid data alone can reject the possibility that local connectivity correlates with measures of dispensability, in higher quality datasets a relatively robust correlation is observed. In contrast, local connectivity does not correlate with the rate of protein evolution even in reliable datasets. This perhaps surprising lack of correlation with evolutionary rate appears in part to arise from the fact that hub proteins do not have a higher density of residues associated with binding. However, hub proteins do have at least one other set of unusual features, namely rapid turnover and regulation, as manifest in high mRNA decay rates and a large number of phosphorylation sites. This, we suggest, is an adaptation to minimize unwanted activation of pathways that might be mediated by adventitious binding to hubs, were they to actively persist longer than required at any given time point. We conclude that hub proteins are more important for cellular growth rate and under tight regulation but are not slow evolving.

Biological Evolution↗

Co-expressed yeast genes cluster over a long range but are not regularly spaced.

Analysis in yeast of the relationship between a gene's genomic position and its expression profile, derived from chip array data, suggests that both closely linked genes and genes spaced at regular intervals show correlated expression profiles. Unfortunately, yeast arrays are often printed in genomic order. The above results may hence reflect little more than known spatial biases within arrays. To circumvent this problem, we analyse spatially unbiased expression data derived from a large Northern blot study. We find that local domains of co-expressed genes range up to 30 genes (100 kb), and are thus much larger than previously considered. There is, by contrast, no evidence for periodicity of co-expression in yeast. We likewise find no convincing evidence for periodicity in the human or mouse genome. Further, analysis of yeast transcription factor binding data sets suggests that there is currently no statistical evidence for chromosomal periodicity of co-regulation.

Animals↗

Chance and necessity in the evolution of minimal metabolic networks.

It is possible to infer aspects of an organism's lifestyle from its gene content. Can the reverse also be done? Here we consider this issue by modelling evolution of the reduced genomes of endosymbiotic bacteria. The diversity of gene content in these bacteria may reflect both variation in selective forces and contingency-dependent loss of alternative pathways. Using an in silico representation of the metabolic network of Escherichia coli, we examine the role of contingency by repeatedly simulating the successive loss of genes while controlling for the environment. The minimal networks that result are variable in both gene content and number. Partially different metabolisms can thus evolve owing to contingency alone. The simulation outcomes do preserve a core metabolism, however, which is over-represented in strict intracellular bacteria. Moreover, differences between minimal networks based on lifestyle are predictable: by simulating their respective environmental conditions, we can model evolution of the gene content in Buchnera aphidicola and Wigglesworthia glossinidia with over 80% accuracy. We conclude that, at least for the particular cases considered here, gene content of an organism can be predicted with knowledge of its distant ancestors and its current lifestyle.

Biological Evolution↗

Hearing silence: non-neutral evolution at synonymous sites in mammals.

Although the assumption of the neutral theory of molecular evolution - that some classes of mutation have too small an effect on fitness to be affected by natural selection - seems intuitively reasonable, over the past few decades the theory has been in retreat. At least in species with large populations, even synonymous mutations in exons are not neutral. By contrast, in mammals, neutrality of these mutations is still commonly assumed. However, new evidence indicates that even some synonymous mutations are subject to constraint, often because they affect splicing and/or mRNA stability. This has implications for understanding disease, optimizing transgene design, detecting positive selection and estimating the mutation rate.

Animals↗

Stratus not altocumulus: a new view of the yeast protein interaction network.

Systems biology approaches can reveal intermediary levels of organization between genotype and phenotype that often underlie biological phenomena such as polygenic effects and protein dispensability. An important conceptualization is the module, which is loosely defined as a cohort of proteins that perform a dedicated cellular task. Based on a computational analysis of limited interaction datasets in the budding yeast Saccharomyces cerevisiae, it has been suggested that the global protein interaction network is segregated such that highly connected proteins, called hubs, tend not to link to each other. Moreover, it has been suggested that hubs fall into two distinct classes: "party" hubs are co-expressed and co-localized with their partners, whereas "date" hubs interact with incoherently expressed and diversely localized partners, and thereby cohere disparate parts of the global network. This structure may be compared with altocumulus clouds, i.e., cotton ball-like structures sparsely connected by thin wisps. However, this organization might reflect a small and/or biased sample set of interactions. In a multi-validated high-confidence (HC) interaction network, assembled from all extant S. cerevisiae interaction data, including recently available proteome-wide interaction data and a large set of reliable literature-derived interactions, we find that hub-hub interactions are not suppressed. In fact, the number of interactions a hub has with other hubs is a good predictor of whether a hub protein is essential or not. We find that date hubs are neither required for network tolerance to node deletion, nor do date hubs have distinct biological attributes compared to other hubs. Date and party hubs do not, for example, evolve at different rates. Our analysis suggests that the organization of global protein interaction network is highly interconnected and hence interdependent, more like the continuous dense aggregations of stratus clouds than the segregated configuration of altocumulus clouds. If the network is configured in a stratus format, cross-talk between proteins is potentially a major source of noise. In turn, control of the activity of the most highly connected proteins may be vital. Indeed, we find that a fluctuation in steady-state levels of the most connected proteins is minimized.

Computational Biology↗

Unusual linkage patterns of ligands and their cognate receptors indicate a novel reason for non-random gene order in the human genome.

BACKGROUND: Prior to the sequencing of the human genome it was typically assumed that, tandem duplication aside, gene order is for the most part random. Numerous observers, however, highlighted instances in which a ligand was linked to one of its cognate receptors, with some authors suggesting that this may be a general and/or functionally important pattern, possibly associated with recombination modification between epistatically interacting loci. Here we ask whether ligands are more closely linked to their receptors than expected by chance. RESULTS: We find no evidence that ligands are linked to their receptors more closely than expected by chance. However, in the human genome there are approximately twice as many co-occurrences of ligand and receptor on the same human chromosome as expected by chance. Although a weak effect, the latter might be consistent with a past history of block duplication. Successful duplication of some ligands, we hypothesise, is more likely if the cognate receptor is duplicated at the same time, so ensuring appropriate titres of the two products. CONCLUSION: While there is an excess of ligands and their receptors on the same human chromosome, this cannot be accounted for by classical models of non-random gene order, as the linkage of ligands/receptors is no closer than expected by chance. Alternative hypotheses for non-random gene order are hence worth considering.

Animals↗

Comparisons of dN/dS are time dependent for closely related bacterial genomes.

The ratio of non-synonymous (dN) to synonymous (dS) changes between taxa is frequently computed to assay the strength and direction of selection. Here we note that for comparisons between closely related strains and/or species a second parameter needs to be considered, namely the time since divergence of the two sequences under scrutiny. We demonstrate that a simple time lag model provides a general, parsimonious explanation of the extensive variation in the dN/dS ratio seen when comparing closely related bacterial genomes. We explore this model through simulation and comparative genomics, and suggest a role for hitch-hiking in the accumulation of non-synonymous mutations. We also note taxon-specific differences in the change of dN/dS over time, which may indicate variation in selection, or in population genetics parameters such as population size or the rate of recombination. The effect of comparing intra-species polymorphism and inter-species substitution, and the problems associated with these concepts for asexual prokaryotes, are also discussed. We conclude that, because of the critical effect of time since divergence, inter-taxa comparisons are only possible by comparing trajectories of dN/dS over time and it is not valid to compare taxa on the basis of single time points.

Bacteria↗

Evidence for purifying selection against synonymous mutations in mammalian exonic splicing enhancers.

Silent sites in mammals have classically been assumed to be free from selective pressures. Consequently, the synonymous substitution rate (Ks) is often used as a proxy for the mutation rate. Although accumulating evidence demonstrates that the assumption is not valid, the mechanism by which selection acts remain unclear. Recent work has revealed that the presence of exonic splicing enhancers (ESEs) in coding sequence might influence synonymous evolution. ESEs are predominantly located near intron-exon junctions, which may explain the reduced single-nucleotide polymorphism (SNP) density in these regions. Here we show that synonymous sites in putative ESEs evolve more slowly than the remaining exonic sequence. Differential mutabilities of ESEs do not appear to explain this difference. We observe that substitution frequency at fourfold synonymous sites decreases as one approaches the ends of exons, consistent with the existing SNP data. This gradient is at least in part explained by ESEs being more abundant near junctions. Between-gene variation in Ks is hence partly explained by the proportion of the gene that acts as an ESE. Given the relative abundance of ESEs and the reduced rates of synonymous divergence within them, we estimate that constraints on synonymous evolution within ESEs causes the true mutation rate to be underestimated by not more than approximately 8%. We also find that Ks outside of ESEs is much lower in alternatively spliced exons than in constitutive exons, implying that other causes of selection on synonymous mutations exist. Additionally, selection on ESEs appears to affect nonsynonymous sites and may explain why amino acid usage near intron-exon junctions is nonrandom.

Alternative Splicing↗

Evidence for a preferential targeting of 3'-UTRs by cis-encoded natural antisense transcripts.

Although both the 5'- and 3'-untranslated regions (5'- and 3'-UTRs) of eukaryotic mRNAs may play a crucial role in posttranscriptional gene regulation, we observe that cis-encoded natural antisense RNAs have a striking preferential complementarity to the 3'-UTRs of their target genes in mammalian (human and mouse) genomes. A null neutral model, evoking differences in the rate of 3'-UTR and 5'-UTR extension, could potentially explain high rates of 3'-to-3' overlap compared with 5'-to-5' overlap. However, employing a simulation model we show that this null model probably cannot explain the finding that 3'-to-3' overlapping pairs have a much higher probability (>5 times) of conservation in both mouse and human genomes with the same overlapping pattern than do 5'-to-5' overlaps. Furthermore, it certainly cannot explain the finding that overlapping pairs seen in both genomes have a significantly higher probability of having co-expression and inverse expression (i.e. characteristic of sense-antisense regulation) than do overlapping pairs seen in only one of the two species. We infer that the function of many 3'-to-3' overlaps is indeed antisense regulation. These findings underscore the preference for, and conservation of, 3'-UTR-targeted antisense regulation, and the importance of 3'-UTRs in gene regulation.

3' Untranslated Regions↗

The small introns of antisense genes are better explained by selection for rapid transcription than by "genomic design".

Several models have been proposed to explain why expression parameters of a gene might be related to the size of the gene's introns. These include the idea that an energetic cost of transcription should favor smaller introns in highly expressed genes (the "economy selection" argument) and that tissue-specific genes reside in genomic locations with complex chromatin level control requiring large amounts of noncoding DNA (the "genomic design" hypothesis). We recently proposed a modification of the economy model arguing that, for some genes, the time that expression takes is more important than the energetic cost, such that some weakly but rapidly expressed genes might also have small introns. We suggested that antisense genes might be such a class and showed that the data appear to be consistent with this. We now reexamine this model to ask (a) whether the effects described were owing solely to the fact that antisense genes are often noncoding RNA and (b) whether we can confidently reject the "genomic design" model as an explanation for the facts. We show that the effects are not specific to noncoding RNAs and that the predictions of the "genomic design" model for the most part are not upheld.

Antisense Elements (Genetics)↗

Evidence for selection on synonymous mutations affecting stability of mRNA secondary structure in mammals.

BACKGROUND: In mammals, contrary to what is usually assumed, recent evidence suggests that synonymous mutations may not be selectively neutral. This position has proven contentious, not least because of the absence of a viable mechanism. Here we test whether synonymous mutations might be under selection owing to their effects on the thermodynamic stability of mRNA, mediated by changes in secondary structure. RESULTS: We provide numerous lines of evidence that are all consistent with the above hypothesis. Most notably, by simulating evolution and reallocating the substitutions observed in the mouse lineage, we show that the location of synonymous mutations is non-random with respect to stability. Importantly, the preference for cytosine at 4-fold degenerate sites, diagnostic of selection, can be explained by its effect on mRNA stability. Likewise, by interchanging synonymous codons, we find naturally occurring mRNAs to be more stable than simulant transcripts. Housekeeping genes, whose proteins are under strong purifying selection, are also under the greatest pressure to maintain stability. CONCLUSION: Taken together, our results provide evidence that, in mammals, synonymous sites do not evolve neutrally, at least in part owing to selection on mRNA stability. This has implications for the application of synonymous divergence in estimating the mutation rate.

Amino Acid Substitution↗

Gametophytic selection in Arabidopsis thaliana supports the selective model of intron length reduction.

Why do highly expressed genes have small introns? This is an important issue, not least because it provides a testing ground to compare selectionist and neutralist models of genome evolution. Some argue that small introns are selectively favoured to reduce the costs of transcription. Alternatively, large introns might permit complex regulation, not needed for highly expressed genes. This "genome design" hypothesis evokes a regionalized model of control of expression and hence can explain why intron size covaries with intergene distance, a feature also consistent with the hypothesis that highly expressed genes cluster in genomic regions with high deletion rates. As some genes are expressed in the haploid stage and hence subject to especially strong purifying selection, the evolution of genes in Arabidopsis provides a novel testing ground to discriminate between these possibilities. Importantly, controlling for expression level, genes that are expressed in pollen have shorter introns than genes that are expressed in the sporophyte. That genes flanking pollen-expressed genes have average-sized introns and intergene distances argues against regional mutational biases and genomic design. These observations thus support the view that selection for efficiency contributes to the reduction in intron length and provide the first report of a molecular signature of strong gametophytic selection.

Arabidopsis↗

Comparative evolutionary analysis of VPS33 homologues: genetic and functional insights.

VPS33B protein is a homologue of the yeast class C vacuolar protein sorting protein Vps33p that is involved in the biogenesis and function of vacuoles. Vps33p homologues contain a Sec1 domain and belong to the family of Sec1/Munc18 (SM) proteins that regulate fusion of membrane-bound organelles and interact with other vps proteins and also SNARE proteins that execute membrane fusion in all cells. We demonstrated recently that mutations in VPS33B cause ARC syndrome (MIM 208085), a lethal multisystem disease. In contrast, mutations in other Vps33p homologues result in different phenotypes, e.g. a mutation in Drosophila melanogaster car gene causes the carnation eye colour mutant and inactivation of mouse Vps33a causes buff hypopigmentation phenotype. In mammals two Vps33p homologues (e.g. VPS33A and VPS33B in humans) have been identified. As comparative genome analysis can provide novel insights into gene evolution and function, we performed nucleotide and protein sequence comparisons of Vps33 homologues in different species to define their inter-relationships and evolution. In silico analysis (a) identified two homologues of yeast Vps33p in the worm, fly, zebrafish, rodent and human genomes, (b) suggested that Carnation is an orthologue of VPS33A rather than VPS33B and (c) identified conserved candidate functional domains within VPS33B. We have shown previously that wild-type VPS33B induced perinuclear clustering of late endosomes and lysosomes in human renal cells. Consistent with the predictions of comparative analysis: (a) VPS33B induced significantly more clustering than VPS33A in a renal cell line, (b) a putative fly VPS33B homologue but not Carnation protein also induced clustering and (c) the ability to induce clustering in renal cells was linked to two evolutionary conserved domains within VPS33B. One domain was present in VPS33B but not VPS33A homologues and the other was one of three regions predicted to form a t-SNARE binding site in VPS33B. In contrast, VPS33A induced significantly more clustering of melanosomes in melanoma cells than VPS33B. These investigations are consistent with the hypothesis that there are two functional classes of Vps33p homologues in all multicellular organisms and that the two classes reflect the evolution of organelle/tissue-specific functions.

Amino Acid Sequence↗