PubMed Health⌕ Search

Biomedical subjects

Bernard De Baets

Publications and source records attributed to Bernard De Baets.

6 recordsLinked to original sources

Genome analysis of the smallest free-living eukaryote Ostreococcus tauri unveils many unique features.

The green lineage is reportedly 1,500 million years old, evolving shortly after the endosymbiosis event that gave rise to early photosynthetic eukaryotes. In this study, we unveil the complete genome sequence of an ancient member of this lineage, the unicellular green alga Ostreococcus tauri (Prasinophyceae). This cosmopolitan marine primary producer is the world's smallest free-living eukaryote known to date. Features likely reflecting optimization of environmentally relevant pathways, including resource acquisition, unusual photosynthesis apparatus, and genes potentially involved in C(4) photosynthesis, were observed, as was downsizing of many gene families. Overall, the 12.56-Mb nuclear genome has an extremely high gene density, in part because of extensive reduction of intergenic regions and other forms of compaction such as gene fusion. However, the genome is structurally complex. It exhibits previously unobserved levels of heterogeneity for a eukaryote. Two chromosomes differ structurally from the other eighteen. Both have a significantly biased G+C content, and, remarkably, they contain the majority of transposable elements. Many chromosome 2 genes also have unique codon usage and splicing, but phylogenetic analysis and composition do not support alien gene origin. In contrast, most chromosome 19 genes show no similarity to green lineage genes and a large number of them are specialized in cell surface processes. Taken together, the complete genome sequence, unusual features, and downsized gene families, make O. tauri an ideal model system for research on eukaryotic genome evolution, including chromosome specialization and green lineage ancestry.

Animals↗

Mining fatty acid databases for detection of novel compounds in aerobic bacteria.

This study examines how the discriminatory power of an automated bacterial whole-cell fatty acid identification system can be significantly enhanced by exploring the vast amounts of information accumulated during 15 years of routine gas chromatographic analysis of the fatty acid content of aerobic bacteria. Construction of a global peak occurrence histogram based upon a large fatty acid database is shown to serve as a highly informative tool for assessing the delineation of the naming windows used during the automatic recognition of fatty acid compounds. Along the lines of this data mining application, it is suggested that several naming windows of the Sherlock MIS TSBA50 peak naming method may need to be re-evaluated in order to fit more closely with the bulk of observed fatty acid profiles. At the same time, the global peak occurrence histogram has put forward the delineation of 32 new peak naming windows, accounting for a 26% increase in the total number of fatty acid features taken into account for bacterial identification. By scrutinizing the relationships between the newly delineated naming windows and the many taxonomic units covered within a proprietary fatty acid database, all new naming windows were proven to correspond with stable features of some specific groups of microorganisms. This latter analysis clearly underscores the impact of incorporating the new fatty acid compounds for improving the resolution of the bacterial identification system and endorses the applicability of knowledge discovery in databases within the field of microbiology.

Bacteria, Aerobic↗

SpliceMachine: predicting splice sites from high-dimensional local context representations.

MOTIVATION: In this age of complete genome sequencing, finding the location and structure of genes is crucial for further molecular research. The accurate prediction of intron boundaries largely facilitates the correct prediction of gene structure in nuclear genomes. Many tools for localizing these boundaries on DNA sequences have been developed and are available to researchers through the internet. Nevertheless, these tools still make many false positive predictions. RESULTS: This manuscript presents a novel publicly available splice site prediction tool named SpliceMachine that (i) shows state-of-the-art prediction performance on Arabidopsis thaliana and human sequences, (ii) performs a computationally fast annotation and (iii) can be trained by the user on its own data. AVAILABILITY: Results, figures and software are available at http://www.bioinformatics.psb.ugent.be/supplementary_data/ CONTACT: sven.degroeve@psb.ugent.be; yves.vandepeer@psb.ugent.be.

Algorithms↗

Fuzzy rule-based models for decision support in ecosystem management.

To facilitate decision support in the ecosystem management, ecological expertise and site-specific data need to be integrated. Fuzzy logic can deal with highly variable, linguistic, vague and uncertain data or knowledge and, therefore, has the ability to allow for a logical, reliable and transparent information stream from data collection down to data usage in decision-making. Several environmental applications already implicate the use of fuzzy logic. Most of these applications have been set up by trial and error and are mainly limited to the domain of environmental assessment. In this article, applications of fuzzy logic for decision support in ecosystem management are reviewed and assessed, with an emphasis on rule-based models. In particular, the identification, optimisation, validation, the interpretability and uncertainty aspects of fuzzy rule-based models for decision support in ecosystem management are discussed.

Conservation of Natural Resources↗

Feature subset selection for splice site prediction.

MOTIVATION: The large amount of available annotated Arabidopsis thaliana sequences allows the induction of splice site prediction models with supervised learning algorithms (see Haussler (1998) for a review and references). These algorithms need information sources or features from which the models can be computed. For splice site prediction, the features we consider in this study are the presence or absence of certain nucleotides in close proximity to the splice site. Since it is not known how many and which nucleotides are relevant for splice site prediction, the set of features is chosen large enough such that the probability that all relevant information sources are in the set is very high. Using only those features that are relevant for constructing a splice site prediction system might improve the system and might also provide us with useful biological knowledge. Using fewer features will of course also improve the prediction speed of the system. RESULTS: A wrapper-based feature subset selection algorithm using a support vector machine or a naive Bayes prediction method was evaluated against the traditional method for selecting features relevant for splice site prediction. Our results show that this wrapper approach selects features that improve the performance against the use of all features and against the use of the features selected by the traditional method. AVAILABILITY: The data and additional interactive graphs on the selected feature subsets are available at http://www.psb.rug.ac.be/gps

Arabidopsis↗