PubMed Health⌕ Search

Biomedical subjects

Simon Rayner

Publications and source records attributed to Simon Rayner.

3 recordsLinked to original sources

Genome wide prediction of protein function via a generic knowledge discovery approach based on evidence integration.

BACKGROUND: The automation of many common molecular biology techniques has resulted in the accumulation of vast quantities of experimental data. One of the major challenges now facing researchers is how to process this data to yield useful information about a biological system (e.g. knowledge of genes and their products, and the biological roles of proteins, their molecular functions, localizations and interaction networks). We present a technique called Global Mapping of Unknown Proteins (GMUP) which uses the Gene Ontology Index to relate diverse sources of experimental data by creation of an abstraction layer of evidence data. This abstraction layer is used as input to a neural network which, once trained, can be used to predict function from the evidence data of unannotated proteins. The method allows us to include almost any experimental data set related to protein function, which incorporates the Gene Ontology, to our evidence data in order to seek relationships between the different sets. RESULTS: We have demonstrated the capabilities of this method in two ways. We first collected various experimental datasets associated with yeast (Saccharomyces cerevisiae) and applied the technique to a set of previously annotated open reading frames (ORFs). These ORFs were divided into training and test sets and were used to examine the accuracy of the predictions made by our method. Then we applied GMUP to previously un-annotated ORFs and made 1980, 836 and 1969 predictions corresponding to the GO Biological Process, Molecular Function and Cellular Component sub-categories respectively. We found that GMUP was particularly successful at predicting ORFs with functions associated with the ribonucleoprotein complex, protein metabolism and transportation. CONCLUSION: This study presents a global and generic gene knowledge discovery approach based on evidence integration of various genome-scale data. It can be used to provide insight as to how certain biological processes are implemented by interaction and coordination of proteins, which may serve as a guide for future analysis. New data can be readily incorporated as it becomes available to provide more reliable predictions or further insights into processes and interactions.

Algorithms↗

ORF-FINDER: a vector for high-throughput gene identification.

We have developed a simple and efficient system (ORF-FINDER) for selecting open reading frames (ORFs) from randomly fragmented genomic DNA fragments. The ORF-FINDER vectors are plasmids that contain a translational start site out of frame with respect to the gene for green fluorescent protein (GFP). Insertion of DNA fragments that bring the initiating ATG in frame with GFP and that contain no stop codons (that is, ORFs) results in the expression of ORF-GFP fusion proteins. In addition, we have developed software (GeneWorks and GenomeAnalyzer) to predict the optimal insert size for maximizing the number of gene-coding ORFs and minimizing unintentionally selected non-coding ORFs. To demonstrate the feasibility of using the ORF-FINDER system to screen genomes for ORFs, we cloned yeast genomic DNA and succeeded in enriching for ORFs by 25-fold. Furthermore, we have shown that the vector can effectively isolate ORFs from the more complex genomes of eukaryotic parasites. We envision that ORF-FINDER will have several applications including genome sequencing projects, gene building from oligonucleotides and construction of expression libraries enriched for ORFs.

Animals↗

A scalable high-throughput chemical synthesizer.

A machine that employs a novel reagent delivery technique for biomolecular synthesis has been developed. This machine separates the addressing of individual synthesis sites from the actual process of reagent delivery by using masks placed over the sites. Because of this separation, this machine is both cost-effective and scalable, and thus the time required to synthesize 384 or 1536 unique biomolecules is very nearly the same. Importantly, the mask design allows scaling of the number of synthesis sites without the addition of new valving. Physical and biological comparisons between DNA made on a commercially available synthesizer and this unit show that it produces DNA of similar quality.

DNA↗