PubMed Health⌕ Search

Biomedical subjects

Stephen Rudd

Publications and source records attributed to Stephen Rudd.

At least 19 recordsLinked to original sources

Separation of sequences from host-pathogen interface using triplet nucleotide frequencies.

The identification of genes involved in host-pathogen interactions is important for the elucidation of mechanisms of disease resistance and host susceptibility. A traditional way to classify the origin of genes sampled from a pool of mixed cDNA is through sequence similarity to known genes from either the pathogen or host organism or other closely related species. This approach does not work when the identified sequence has no close homologues in the sequence databases. In our previous studies, we classified genes using their codon frequencies. This method, however, explicitly required the prediction of CDS regions and thus could not be applied to sequences composed from the non-coding regions of genes. In this study, we show that the use of sliding-window triplet frequencies extends the application of the algorithm to both coding and non-coding sequences and also increases the prediction accuracy of a Support Vector Machine classifier from 95.6+/-0.3 to 96.5+/-0.2. Thus the use of the triplet frequencies increased the prediction accuracy of the new method by more than 20% compared to our previous approach. A functional analysis of sequences detected gene families having significantly higher or lower probability to be correctly classified compared to the average accuracy of the method is described. The server to perform classification of EST sequences using triplet frequencies is available at (URL: http://mips.gsf.de/proj/est3).

Algorithms↗

Construction, database integration, and application of an Oenothera EST library.

Coevolution of cellular genetic compartments is a fundamental aspect in eukaryotic genome evolution that becomes apparent in serious developmental disturbances after interspecific organelle exchanges. The genus Oenothera represents a unique, at present the only available, resource to study the role of the compartmentalized plant genome in diversification of populations and speciation processes. An integrated approach involving cDNA cloning, EST sequencing, and bioinformatic data mining was chosen using Oenothera elata with the genetic constitution nuclear genome AA with plastome type I. The Gene Ontology system grouped 1621 unique gene products into 17 different functional categories. Application of arrays generated from a selected fraction of ESTs revealed significantly differing expression profiles among closely related Oenothera species possessing the potential to generate fertile and incompatible plastid/nuclear hybrids (hybrid bleaching). Furthermore, the EST library provides a valuable source of PCR-based polymorphic molecular markers that are instrumental for genotyping and molecular mapping approaches.

Cell Nucleus↗

Genetic screen for signal peptides in Hydra reveals novel secreted proteins and evidence for non-classical protein secretion.

We have screened a Hydra cDNA library for sequences encoding N-terminal signal peptides using the yeast invertase secretion vector pSUC [Jacobs et al., 1997. A genetic selection for isolating cDNAs encoding secreted proteins. Gene 198, 289-296]. We isolated and sequenced 907 positive clones; 88% encoded signal peptides; 12% lacked signal peptides. By searching the Hydra EST database we identified full-length sequences for the selected clones. These encoded 37 known proteins with signal peptides and 40 novel Hydra-specific proteins with signal peptides. Localization of two signal peptide-containing sequences, VEGF and ferritin, to the secretory pathway was confirmed with GFP fusion proteins. In addition, we isolated 105 clones which lacked signal peptides but which supported invertase secretion from yeast. Isolation of plasmids from these clones and retransformation in invertase-negative yeast cells confirmed the phenotype. A GFP fusion protein of one such clone encoding the foot morphogen pedibin was localized to the cytoplasm in transfected Hydra cells and did not enter the ER/Golgi secretory pathway. Secretion of pedibin and other proteins lacking signal peptides appears to occur by a non-classical protein secretion route.

Amino Acid Sequence↗

A wing expressed sequence tag resource for Bicyclus anynana butterflies, an evo-devo model.

BACKGROUND: Butterfly wing color patterns are a key model for integrating evolutionary developmental biology and the study of adaptive morphological evolution. Yet, despite the biological, economical and educational value of butterflies they are still relatively under-represented in terms of available genomic resources. Here, we describe an Expression Sequence Tag (EST) project for Bicyclus anynana that has identified the largest available collection to date of expressed genes for any butterfly. RESULTS: By targeting cDNAs from developing wings at the stages when pattern is specified, we biased gene discovery towards genes potentially involved in pattern formation. Assembly of 9,903 ESTs from a subtracted library allowed us to identify 4,251 genes of which 2,461 were annotated based on BLAST analyses against relevant gene collections. Gene prediction software identified 2,202 peptides, of which 215 longer than 100 amino acids had no homology to any known proteins and, thus, potentially represent novel or highly diverged butterfly genes. We combined gene and Single Nucleotide Polymorphism (SNP) identification by constructing cDNA libraries from pools of outbred individuals, and by sequencing clones from the 3' end to maximize alignment depth. Alignments of multi-member contigs allowed us to identify over 14,000 putative SNPs, with 316 genes having at least one high confidence double-hit SNP. We furthermore identified 320 microsatellites in transcribed genes that can potentially be used as genetic markers. CONCLUSION: Our project was designed to combine gene and sequence polymorphism discovery and has generated the largest gene collection available for any butterfly and many potential markers in expressed genes. These resources will be invaluable for exploring the potential of B. anynana in particular, and butterflies in general, as models in ecological, evolutionary, and developmental genetics.

Animals↗

Spatiotemporal expression control correlates with intragenic scaffold matrix attachment regions (S/MARs) in Arabidopsis thaliana.

Scaffold/matrix attachment regions (S/MARs) are essential for structural organization of the chromatin within the nucleus and serve as anchors of chromatin loop domains. A significant fraction of genes in Arabidopsis thaliana contains intragenic S/MAR elements and a significant correlation of S/MAR presence and overall expression strength has been demonstrated. In this study, we undertook a genome scale analysis of expression level and spatiotemporal expression differences in correlation with the presence or absence of genic S/MAR elements. We demonstrate that genes containing intragenic S/MARs are prone to pronounced spatiotemporal expression regulation. This characteristic is found to be even more pronounced for transcription factor genes. Our observations illustrate the importance of S/MARs in transcriptional regulation and the role of chromatin structural characteristics for gene regulation. Our findings open new perspectives for the understanding of tissue- and organ-specific regulation of gene expression.

Arabidopsis↗

Gene expression and metabolite profiling of Populus euphratica growing in the Negev desert.

BACKGROUND: Plants growing in their natural habitat represent a valuable resource for elucidating mechanisms of acclimation to environmental constraints. Populus euphratica is a salt-tolerant tree species growing in saline semi-arid areas. To identify genes involved in abiotic stress responses under natural conditions we constructed several normalized and subtracted cDNA libraries from control, stress-exposed and desert-grown P. euphratica trees. In addition, we identified several metabolites in desert-grown P. euphratica trees. RESULTS: About 14,000 expressed sequence tag (EST) sequences were obtained with a good representation of genes putatively involved in resistance and tolerance to salt and other abiotic stresses. A P. euphratica DNA microarray with a uni-gene set of ESTs representing approximately 6,340 different genes was constructed. The microarray was used to study gene expression in adult P. euphratica trees growing in the desert canyon of Ein Avdat in Israel. In parallel, 22 selected metabolites were profiled in the same trees. CONCLUSION: Of the obtained ESTs, 98% were found in the sequenced P. trichocarpa genome and 74% in other Populus EST collections. This implies that the P. euphratica genome does not contain different genes per se, but that regulation of gene expression might be different and that P. euphratica expresses a different set of genes that contribute to adaptation to saline growth conditions. Also, all of the five measured amino acids show increased levels in trees growing in the more saline soil.

Desert Climate↗

Maintenance of ancestral complexity and non-metazoan genes in two basal cnidarians.

Cnidarians are among the simplest extant animals; however EST analyses reveal that they have a remarkably high level of genetic complexity. In this article, we show that the full diversity of metazoan signaling pathways is represented in this phylum, as are antagonists previously known only in chordates. Many of the cnidarian ESTs match genes previously known only in non-animal kingdoms. At least some of these represent ancient genes lost by all bilaterians examined so far, rather than genes gained by recent lateral gene transfer.

Animals↗

Eclair--a web service for unravelling species origin of sequences sampled from mixed host interfaces.

The identification of the genes that participate at the biological interface of two species remains critical to our understanding of the mechanisms of disease resistance, disease susceptibility and symbiosis. The sequencing of complementary DNA (cDNA) libraries prepared from the biological interface between two organisms provides an inexpensive way to identify the novel genes that may be expressed as a cause or consequence of compatible or incompatible interactions. Sequence classification and annotation of species origin typically use an orthology-based approach and require access to large portions of either genome, or a close relative. Novel species- or clade-specific sequences may have no counterpart within existing databases and remain ambiguous features. Here we present a web-service, Eclair, which utilizes support vector machines for the classification of the origin of expressed sequence tags stemming from mixed host cDNA libraries. In addition to providing an interface for the classification of sequences, users are presented with the opportunity to train a model to suit their preferred species pair. Eclair is freely available at http://eclair.btk.fi.

Artificial Intelligence↗

Analysis of the floral transcriptome uncovers new regulators of organ determination and gene families related to flower organ differentiation in Gerbera hybrida (Asteraceae).

Development of composite inflorescences in the plant family Asteraceae has features that cannot be studied in the traditional model plants for flower development. In Gerbera hybrida, inflorescences are composed of morphologically different types of flowers tightly packed into a flower head (capitulum). Individual floral organs such as pappus bristles (sepals) are developmentally specialized, stamens are aborted in marginal flowers, petals and anthers are fused structures, and ovaries are located inferior to other floral organs. These specific features have made gerbera a rewarding target of comparative studies. Here we report the analysis of a gerbera EST database containing 16,994 cDNA sequences. Comparison of the sequences with all plant peptide sequences revealed 1656 unique sequences for gerbera not identified elsewhere within the plant kingdom. Based on the EST database, we constructed a cDNA microarray containing 9000 probes and have utilized it in identification of flower-specific genes and abundantly expressed marker genes for flower scape, pappus, stamen, and petal development. Our analysis revealed several regulatory genes with putative functions in flower-organ development. We were also able to associate a number of abundantly and specifically expressed genes with flower-organ differentiation. Gerbera is an outcrossing species, for which genetic approaches to gene discovery are not readily amenable. However, reverse genetics with the help of gene transfer has been very informative. We demonstrate here the usability of the gerbera microarray as a reliable new tool for identifying novel genes related to specific biological questions and for large-scale gene expression analysis.

Asteraceae↗

openSputnik--a database to ESTablish comparative plant genomics using unsaturated sequence collections.

The public expressed sequence tag collections are continually being enriched with high-quality sequences that represent an ever-expanding range of taxonomically diverse plant species. While these sequence collections provide biased insight into the populations of expressed genes available within individual species and their associated tissues, the information is conceivably of wider relevance in a comparative context. When we consider the available expressed sequence tag (EST) collections of summer 2004, most of the major plant taxonomic clades are at least superficially represented. Investigation of the five million available plant ESTs provides a wealth of information that has applications in modelling the routes of plant genome evolution and the identification of lineage-specific genes and gene families. Over four million ESTs from over 50 distinct plant species have been collated within an EST analysis pipeline called openSputnik. The ESTs were resolved down into approximately one million unigene sequences. These have been annotated using orthology-based annotation transfer from reference plant genomes and using a variety of contemporary bioinformatics methods to assign peptide, structural and functional attributes. The openSputnik database is available at http://sputnik.btk.fi.

Cluster Analysis↗

PlantMarkers--a database of predicted molecular markers from plants.

Molecular markers are required in a broad spectrum of gene screening approaches, ranging from gene-mapping within traditional 'forward'-genetics approaches through QTL identification studies to genotyping and haplotyping studies. As we enter the post-genomics era, the need for genetic markers does not diminish, even in the species with fully sequenced genomes. PlantMarkers is a genetic marker database that contains a comprehensive pool of predicted molecular markers. We have adopted contemporary techniques to identify putative single nucleotide polymorphism (SNP), simple sequence repeat (SSR) and conserved orthologue set markers. A systematic approach to identify as broad a range of putative markers has been undertaken by screening the available openSputnik unigene consensus sequences from over 50 plant species. A web presence at http://markers.btk.fi provides functionality so that a user may search for species-specific markers on the basis of many specific criteria not limited to non-synonymous SNPs segregating between different varieties or measured polymorphic SSRs. Feedback forms are provided with all sequence entries to enable inclusion of, for example, map location for markers validated by the research community.

Base Sequence↗

A 6374 unigene set corresponding to low abundance transcripts expressed following fertilization in Solanum chacoense Bitt, and characterization of 30 receptor-like kinases.

In order to characterize regulatory genes that are expressed in ovule tissues after fertilization we have undertaken an EST sequencing project in Solanum chacoense, a self-incompatible wild potato species. Two cDNA libraries made from ovule tissues covering embryo development from zygote to late torpedo-stage were constructed and plated at high density on nylon membranes. To decrease EST redundancy and enrich for transcripts corresponding to weakly expressed genes a self-probe subtraction method was used to select the colonies harboring the genes to be sequenced. 7741 good sequences were obtained and, from these, 6374 unigenes were isolated. Thus, the self-probe subtraction resulted in a strong enrichment in singletons, a decrease in the number of clones per contigs, and concomitantly, an enrichment in the total number of unigenes obtained (82%). To gain insights into signal transduction events occurring during embryo development all the receptor-like kinases (or protein receptor kinases) were analyzed by quantitative real-time RT-PCR. Interestingly, 28 out of the 30 RLK isolated were predominantly expressed in ovary tissues or young developing fruits, and 23 were transcriptionaly induced following fertilization. Thus, the self-probe subtraction did not select for genes weakly expressed in the target tissue while being highly expressed elsewhere in the plant. Of the receptor-like kinases (RLK) genes isolated, the leucine-rich repeat (LRR) family of RLK was by far the most represented with 25 members covering 11 LRR classes.

DNA, Complementary↗

Support vector machines for separation of mixed plant-pathogen EST collections based on codon usage.

MOTIVATION: Discovery of host and pathogen genes expressed at the plant-pathogen interface often requires the construction of mixed libraries that contain sequences from both genomes. Sequence identification requires high-throughput and reliable classification of genome origin. When using single-pass cDNA sequences difficulties arise from the short sequence length, the lack of sufficient taxonomically relevant sequence data in public databases and ambiguous sequence homology between plant and pathogen genes. RESULTS: A novel method is described, which is independent of the availability of homologous genes and relies on subtle differences in codon usage between plant and fungal genes. We used support vector machines (SVMs) to identify the probable origin of sequences. SVMs were compared to several other machine learning techniques and to a probabilistic algorithm (PF-IND) for expressed sequence tag (EST) classification also based on codon bias differences. Our software (Eclat) has achieved a classification accuracy of 93.1% on a test set of 3217 EST sequences from Hordeum vulgare and Blumeria graminis, which is a significant improvement compared to PF-IND (prediction accuracy of 81.2% on the same test set). EST sequences with at least 50 nt of coding sequence can be classified using Eclat with high confidence. Eclat allows training of classifiers for any host-pathogen combination for which there are sufficient classified training sequences. AVAILABILITY: Eclat is freely available on the Internet (http://mips.gsf.de/proj/est) or on request as a standalone version. CONTACT: friedel@informatik.uni-muenchen.de.

Algorithms↗

Identification of NdhL and Ssl1690 (NdhO) in NDH-1L and NDH-1M complexes of Synechocystis sp. PCC 6803.

The subunit compositions of two types of NAD(P)H dehydrogenase complexes of Synechocystis sp. PCC 6803, NDH-1L and NDH-1M, were studied by two-dimensional blue-native/SDS-PAGE followed by electrospray tandem mass spectrometry. Fifteen proteins were observed in NDH-1L including hydrophilic subunits (NdhH, -K, -I, -J, -M, and -N) and hydrophobic subunits (NdhA, -B, -E, -G, -D1, and -F1). In addition, NdhL and a novel subunit, Ssl1690 (designated NdhO), were shown to be components of this complex. All subunits mentioned above were present in the NDH-1M complex except NdhD1 and NdhF1. NdhL and Ssl1690 (NdhO) were homologous to hypothetical proteins encoded by genomic DNA in higher plants, suggesting that chloroplast NDH-1 complexes contain related subunits. Diagnostic sequence motifs were found for both NdhL and NdhO homologous proteins. Analysis of ndhL deletion mutant (M9) revealed the presence of assembled NDH-1L and NDH-1M complexes, but these complexes appear to be functionally impaired in the absence of NdhL. Both NDH-1 complexes were absent in the ndhB deletion mutant (M55).

Amino Acid Sequence↗

Characterization of the maize endosperm transcriptome and its comparison to the rice genome.

The cereal endosperm is a major organ of the seed and an important component of the world's food supply. To understand the development and physiology of the endosperm of cereal seeds, we focused on the identification of genes expressed at various times during maize endosperm development. We constructed several cDNA libraries to identify full-length clones and subjected them to a twofold enrichment. A total of 23,348 high-quality sequence-reads from 5'- and 3'-ends of cDNAs were generated and assembled into a unigene set representing 5326 genes with paired sequence-reads. Additional sequencing yielded a total of 3160 (59%) completely sequenced, full-length cDNAs. From 5326 unigenes, 4139 (78%) can be aligned with 5367 predicted rice genes and by taking only the "best hit" be mapped to 3108 positions on the rice genome. The 22% unigenes not present in rice indicate a rapid change of gene content between rice and maize in only 50 million years. Differences in rice and maize gene numbers also suggest that maize has lost a large number of duplicated genes following tetraploidization. The larger number of gene copies in rice suggests that as many as 30% of its genes arose from gene amplification, which would extrapolate to a significant proportion of the estimated 44,027 candidate genes of its entire genome. Functional classification of the maize endosperm unigene set indicated that more than a fourth of the novel functionally assignable genes found in this study are involved in carbohydrate metabolism, consistent with its role as a storage organ.

DNA, Complementary↗

Genome-wide in silico mapping of scaffold/matrix attachment regions in Arabidopsis suggests correlation of intragenic scaffold/matrix attachment regions with gene expression.

We carried out a genome-wide prediction of scaffold/matrix attachment regions (S/MARs) in Arabidopsis. Results indicate no uneven distribution on the chromosomal level but a clear underrepresentation of S/MARs inside genes. In cases where S/MARs were predicted within genes, these intragenic S/MARs were preferentially located within the 5'-half, most prominently within introns 1 and 2. Using Arabidopsis whole-genome expression data generated by the massively parallel signature sequencing methodology, we found a negative correlation between S/MAR-containing genes and transcriptional abundance. Expressed sequence tag data correlated the same way with S/MAR-containing genes. Thus, intragenic S/MARs show a negative correlation with transcription level. For various genes it has been shown experimentally that S/MARs can function as transcriptional regulators and that they have an implication in stabilizing expression levels within transgenic plants. On the basis of a genome-wide in silico S/MAR analysis, we found a significant correlation between the presence of intragenic S/MARs and transcriptional down-regulation.

Arabidopsis↗

Large-scale analysis of the barley transcriptome based on expressed sequence tags.

To provide resources for barley genomics, 110,981 expressed sequence tags (ESTs) were generated from 22 cDNA libraries representing tissues at various developmental stages. This EST collection corresponds to approximately one-third of the 380,000 publicly available barley ESTs. Clustering and assembly resulted in 14,151 tentative consensi (TCs) and 11 073 singletons, altogether representing 25 224 putatively unique sequences. Of these, 17.5% showed no significant similarity to other barley ESTs present in dbEST. More than 41% of all barley genes are supposed to belong to multigene families and approximately 4% of the barley genes undergo alternative splicing. Based on the functional annotation of the set of unique sequences, the functional category 'Energy' was further analysed to reveal tissue- and stage-specific differences in gene expression. Hierarchical clustering of 362 differentially expressed TCs resulted in the identification of seven major clusters. The clusters reflect biochemical pathways predominantly activated in specific tissues and at various developmental stages. During seed germination glycolysis could be identified as the most predominant biochemical pathway. Germination-specific glycolysis is characterized by the coordinated expression of phosphoenolpyruvate carboxylase and phosphoenolpyruvate carboxykinase, whose antagonistic actions possibly regulate the flux of amino acids into protein biosynthesis and gluconeogenesis respectively. The expression of defence-related and antioxidant genes during germination might be controlled by the ethylene-signalling pathway as concluded from the coordinated expression of those genes and the transcription factors (TF) EIN3 and EREBPG. Moreover, because of their predominant expression in germinating seeds, TF of the AP2 and MYB type are presumably major regulators of germination.

Expressed Sequence Tags↗

The genome sequence of the filamentous fungus Neurospora crassa.

Neurospora crassa is a central organism in the history of twentieth-century genetics, biochemistry and molecular biology. Here, we report a high-quality draft sequence of the N. crassa genome. The approximately 40-megabase genome encodes about 10,000 protein-coding genes--more than twice as many as in the fission yeast Schizosaccharomyces pombe and only about 25% fewer than in the fruitfly Drosophila melanogaster. Analysis of the gene set yields insights into unexpected aspects of Neurospora biology including the identification of genes potentially associated with red light photobiology, genes implicated in secondary metabolism, and important differences in Ca2+ signalling as compared with plants and animals. Neurospora possesses the widest array of genome defence mechanisms known for any eukaryotic organism, including a process unique to fungi called repeat-induced point mutation (RIP). Genome analysis suggests that RIP has had a profound impact on genome evolution, greatly slowing the creation of new genes through genomic duplication and resulting in a genome with an unusually low proportion of closely related genes.

Calcium Signaling↗