PubMed Health⌕ Search

Biomedical subjects

Sean D Hooper

Publications and source records attributed to Sean D Hooper.

9 recordsLinked to original sources

Identification of tightly regulated groups of genes during Drosophila melanogaster embryogenesis.

Time-series analysis of whole-genome expression data during Drosophila melanogaster development indicates that up to 86% of its genes change their relative transcript level during embryogenesis. By applying conservative filtering criteria and requiring 'sharp' transcript changes, we identified 1534 maternal genes, 792 transient zygotic genes, and 1053 genes whose transcript levels increase during embryogenesis. Each of these three categories is dominated by groups of genes where all transcript levels increase and/or decrease at similar times, suggesting a common mode of regulation. For example, 34% of the transiently expressed genes fall into three groups, with increased transcript levels between 2.5-12, 11-20, and 15-20 h of development, respectively. We highlight common and distinctive functional features of these expression groups and identify a coupling between downregulation of transcript levels and targeted protein degradation. By mapping the groups to the protein network, we also predict and experimentally confirm new functional associations.

Analysis of Variance↗

Medusa: a simple tool for interaction graph analysis.

SUMMARY: Medusa is a Java application for visualizing and manipulating graphs of interaction, such as data from the STRING database. It features an intuitive user interface developed with the help of biologists. Medusa is optimized for accessing protein interaction data from STRING, but can be used for any type of graph from any scientific field.

Algorithms↗

Systematic association of genes to phenotypes by genome and literature mining.

One of the major challenges of functional genomics is to unravel the connection between genotype and phenotype. So far no global analysis has attempted to explore those connections in the light of the large phenotypic variability seen in nature. Here, we use an unsupervised, systematic approach for associating genes and phenotypic characteristics that combines literature mining with comparative genome analysis. We first mine the MEDLINE literature database for terms that reflect phenotypic similarities of species. Subsequently we predict the likely genomic determinants: genes specifically present in the respective genomes. In a global analysis involving 92 prokaryotic genomes we retrieve 323 clusters containing a total of 2,700 significant gene-phenotype associations. Some clusters contain mostly known relationships, such as genes involved in motility or plant degradation, often with additional hypothetical proteins associated with those phenotypes. Other clusters comprise unexpected associations; for example, a group of terms related to food and spoilage is linked to genes predicted to be involved in bacterial food poisoning. Among the clusters, we observe an enrichment of pathogenicity-related associations, suggesting that the approach reveals many novel genes likely to play a role in infectious diseases.

Bacteria↗

STRING: known and predicted protein-protein associations, integrated and transferred across organisms.

A full description of a protein's function requires knowledge of all partner proteins with which it specifically associates. From a functional perspective, 'association' can mean direct physical binding, but can also mean indirect interaction such as participation in the same metabolic pathway or cellular process. Currently, information about protein association is scattered over a wide variety of resources and model organisms. STRING aims to simplify access to this information by providing a comprehensive, yet quality-controlled collection of protein-protein associations for a large number of organisms. The associations are derived from high-throughput experimental data, from the mining of databases and literature, and from predictions based on genomic context analysis. STRING integrates and ranks these associations by benchmarking them against a common reference set, and presents evidence in a consistent and intuitive web interface. Importantly, the associations are extended beyond the organism in which they were originally described, by automatic transfer to orthologous protein pairs in other organisms, where applicable. STRING currently holds 730,000 proteins in 180 fully sequenced organisms, and is available at http://string.embl.de/.

Databases, Protein↗

Environments shape the nucleotide composition of genomes.

To test the impact of environments on genome evolution, we analysed the relative abundance of the nucleotides guanine and cytosine ('GC content') of large numbers of sequences from four distinct environmental samples (ocean surface water, farm soil, an acidophilic mine drainage biofilm and deep-sea whale carcasses). We show that the GC content of complex microbial communities seems to be globally and actively influenced by the environment. The observed nucleotide compositions cannot be easily explained by distinct phylogenetic origins of the species in the environments; the genomic GC content may change faster than was previously thought, and is also reflected in the amino-acid composition of the proteins in these habitats.

Amino Acids↗

Duplication is more common among laterally transferred genes than among indigenous genes.

BACKGROUND: Recent developments in the understanding of paralogous evolution have prompted a focus not only on obviously advantageous genes, but also on genes that can be considered to have a weak or sporadic impact on the survival of the organism. Here we examine the duplicative behavior of a category of genes that can be considered to be mostly transient in the genome, namely laterally transferred genes. Using both a compositional method and a gene-tree approach, we identify a number of proposed laterally transferred genes and study their nucleotide composition and frequency of duplication. RESULTS: It is found that duplications are significantly overrepresented among potential laterally transferred genes compared to the indigenous ones. Furthermore, the GC3 distribution of potential laterally transferred genes was found to be largely uniform in some genomes, suggesting an import from a broad range of donors. CONCLUSIONS: The results are discussed not in a context of strongly optimized established genes, but rather of genes with weak or ancillary functions. The importance of duplication may therefore depend on the variability and availability of weak genes for which novel functions may be discovered. Therefore, lateral transfer may accelerate the evolutionary process of duplication by bringing foreign genes that have mainly weak or no function into the genome.

Bacillus↗

On the nature of gene innovation: duplication patterns in microbial genomes.

Gene duplication is considered a major force in gene family expansion and gene innovation. As gene copies assume novel functions, they must avoid periods of neutrality or be deleted from the genome. Current opinions state that copies avoid neutrality through gene dosage effects. These copies are therefore selected from an early stage. This study concentrates on the flow of copies from recent duplication to gene innovation. We have studied 21 microbial genomes using amino acid divergence to describe paralog evolution in the long-term perspective. Five of these were studied in closer detail using nucleotide divergence for a shorter perspective. It was found that rates of duplication and deletion are high, with only a small fraction of duplications retained and apparently selected. This leads to a steady accumulation of paralogs, which seems to be of a similar magnitude in most of the genomes. Furthermore, it is found that genes of high expression level, as measured by their codon bias, are strongly underrepresented among the most recent duplications. Based on these and other observations, it is suggested that gene innovation is driven by amplification of weak, ancillary functions rather than strong, established functions.

DNA Transposable Elements↗

Detection of genes with atypical nucleotide sequence in microbial genomes.

Along the gene, nucleotides in various codon positions tend to exert a slight but observable influence on the nucleotide choice at neighboring positions. Such context biases are different in different organisms and can be used as genomic signatures. In this paper, we will focus specifically on the dinucleotide composed of a third codon position nucleotide and its succeeding first position nucleotide. Using the 16 possible dinucleotide combinations, we calculate how well individual genes conform to the observed mean dinucleotide frequencies of an entire genome, forming a distance measure for each gene. It is found that genes from different genomes can be separated with a high degree of accuracy, according to these distance values. In particular, we address the problem of recent horizontal gene transfer, and how imported genes may be evaluated by their poor assimilation to the host's context biases. By concentrating on the third- and succeeding first position nucleotides, we eliminate most spurious contributions from codon usage and amino-acid requirements, focusing mainly on mutational effects. Since imported genes are expected to converge only gradually to genomic signatures, it is possible to question whether a gene present in only one of two closely related organisms has been imported into one organism or deleted in the other. Striking correlations between the proposed distance measure and poor homology are observed when Escherichia coli genes are compared to Salmonella typhi, indicating that sets of outlier genes in E. coli may contain a high number of genes that have been imported into E. coli, and not deleted in S. typhi.

Bacteria↗

Gene Import or Deletion: A Study of the Different Genes in Escherichia coli Strains K12 and O157:H7.

By comparing two strains of Escherichia coli (K12 and O157:H7) with an outgroup of Salmonella and Klebsiella species and analyzing the sets of genes which are present or absent in either of the three groups, we study the gene history of K12, in particular, since the respective divergences of these bacteria. Furthermore, by using a compositional method based on context bias, we evaluate not only recently imported genes but also deleted genes. In addition, we examine recent gene duplications in the two E. coli strains. It is found that turnover of DNA is high in E. coli and, more importantly, that turnover is highest for genes of low GC content. Although levels of import are high, most of the imported genes seem to be "junk" or have poorly understood functions. Nevertheless, selected genes do persist, and may even define some E. coli strains as pathogenic. Our results support the conclusion that some of the pathogenic islands in O157:H7 are likely to have been imported in recent time.

Escherichia coli↗