PubMed Health⌕ Search

Biomedical subjects

David Gilbert

Publications and source records attributed to David Gilbert.

12 recordsLinked to original sources

When kinases meet mathematics: the systems biology of MAPK signalling.

The mitogen activated protein kinase/extracellular signal regulated kinase pathway regulates fundamental cellular function such as cell proliferation, survival, differentiation and motility, raising the question how these diverse functions are specified and coordinated. They are encoded through the activation kinetics of the pathway, a multitude of feedback loops, scaffold proteins, subcellular compartmentalisation, and crosstalk with other pathways. These regulatory motifs alone or in combination can generate a multitude of complex behaviour. Systems biology tries to decode this complexity through mathematical modelling and prediction in order to gain a deeper insight into the inner works of signalling networks.

Animals↗

Fast similarity search for protein 3D structures using topological pattern matching based on spatial relations.

Similarity search for protein 3D structures become complex and computationally expensive due to the fact that the size of protein structure databases continues to grow tremendously. Recently, fast structural similarity search systems have been required to put them into practical use in protein structure classification whilst existing comparison systems do not provide comparison results on time. Our approach uses multi-step processing that composes of a preprocessing step to represent geometry of protein structures with spatial objects, a filter step to generate a small candidate set using approximate topological string matching, and a refinement step to compute a structural alignment. This paper describes the preprocessing and filtering for fast similarity search using the discovery of topological patterns of secondary structure elements based on spatial relations. Our system is fully implemented by using Oracle 8i spatial. We have previously shown that our approach has the advantage of speed of performance compared with other approach such as DALI. This work shows that the discovery of topological relations of secondary structure elements in protein structures by using spatial relations of spatial databases is practical for fast structural similarity search for proteins.

Algorithms↗

FrankSum: new feature selection method for protein function prediction.

In the study of in silico functional genomics, improving the performance of protein function prediction is the ultimate goal for identifying proteins associated with defined cellular functions. The classical prediction approach is to employ pairwise sequence alignments. However this method often faces difficulties when no statistically significant homologous sequences are identified. An alternative way is to predict protein function from sequence-derived features using machine learning. In this case the choice of possible features which can be derived from the sequence is of vital importance to ensure adequate discrimination to predict function. In this paper we have successfully selected biologically significant features for protein function prediction. This was performed using a new feature selection method (FrankSum) that avoids data distribution assumptions, uses a data independent measurement (p-value) within the feature, identifies redundancy between features and uses an appropriate ranking criterion for feature selection. We have shown that classifiers generated from features selected by FrankSum outperforms classifiers generated from full feature sets, randomly selected features and features selected from the Wrapper method. We have also shown the features are concordant across all species and top ranking features are biologically informative. We conclude that feature selection is vital for successful protein function prediction and FrankSum is one of the feature selection methods that can be applied successfully to such a domain.

Amino Acid Sequence↗

Isolation and expression of the reverse transcriptase component of the Canis familiaris telomerase ribonucleoprotein (dogTERT).

The enzyme telomerase plays a crucial role in cellular proliferation and tumorigenesis. Telomerase is an RNA-directed DNA polymerase composed minimally of an RNA subunit (TR) and a catalytic protein component (TERT). The protein component acts as a reverse transcriptase (RT) and catalyses the addition of telomeric repeats onto the ends of chromosomes using the RNA subunit as a template. While both the RNA and catalytic subunits are essential for telomerase activity, the TERT component of telomerase is thought to be the primary determinant for enzyme activity as expression of TERT is largely limited to cells with telomerase activity. We describe here the isolation and sequence characterization of the telomerase catalytic subunit from Canis familiaris (dog), dogTERT. The predicted protein consists of 1123-aa residues and contains all the signature motifs of the TERT family members. Sequence comparisons with previously identified mammalian TERT proteins demonstrate that dogTERT shows the highest level of sequence similarity to the human TERT protein, supporting the dog as a model system for telomerase-based studies. Further, we demonstrate that TERT mRNA expression is associated with telomerase activity in canine-cultured cells, similar to TERT expression in human cells. This data will allow for further investigation of telomerase in canine malignancies as well as the development of the dog as a model system for human telomerase investigations.

Amino Acid Sequence↗

Effects of quitting smoking on EEG activation and attention last for more than 31 days and are more severe with stress, dependence, DRD2 A1 allele, and depressive traits.

Changes in physiology and attentional performance associated with smoking abstinence were characterized in 67 female smokers during low-stress and high-stress conditions. Abstinence was associated with decreases in cognitive performance, heart rate, and electroencephalographic (EEG) activation but with no change in serum estradiol or progesterone. Effects of quitting showed no tendency to resolve across the 31 days of abstinence. EEG deactivation and heart rate slowing were greater during a math task (high stress) than during relaxation (low stress). Individuals high in trait depression or nicotine dependence or with at least one dopamine D(2) receptor A1 allele experienced greater EEG deactivation following abstinence, especially in the right hemisphere during the stressful task. Thus, findings support the situation x trait adaptive response model of abstinence effects and emphasize the value of multiple dependent measures when characterizing abstinence responses.

Adolescent↗

Recommendation for the assessment of tobacco craving and withdrawal in smoking cessation trials.

This paper addresses methodological issues in the assessment of nicotine withdrawal and craving in clinical trials of smoking cessation therapies. We define withdrawal as a syndrome of behavioral, affective, cognitive, and physiological symptoms, typically transient, emerging upon cessation or reduction of tobacco use and causing distress or impairment of behavioral function. Offset effects (effects related to removal of a direct nicotine effect) are sustained effects of cessation or reduction of tobacco use that cause distress or impairment. Withdrawal and craving are important as potential predictors of relapse, as mediators and markers of treatment effects, and as clinical phenomena in their own right. Symptoms recommended for assessment include craving, irritability, depression, restlessness, sleep disturbance, difficulty concentrating, increased appetite, and weight gain; anxiety deserves further study. We recommend reporting of data on each of these individual symptoms, and use of multiple-item assessments. Although some standardized measures of withdrawal have promising psychometric properties, no measure has yet fully established its reliability, validity, and broad applicability and, therefore, we do not currently favor universal adoption of any one measure. Assessment of objective indices of withdrawal (e.g., hormonal changes) is currently technically challenging and of unknown value. Although weekly assessment may suffice in some large trials, more intensive measurement can provide better sensitivity. Analyses of withdrawal should include baseline measures and be sensitive to potential instability in baseline. Analytic approaches should take into account potential bias when only abstinent subjects are examined. Conversely, heterogeneity should be considered when smoking subjects are included in intent-to-treat analyses. Withdrawal data from clinical trials focused on assessing abstinence rates may be biased because of progressive subject loss to dropout and relapse; different designs and approaches are needed to investigate the process and natural history of craving and withdrawal.

Clinical Trials as Topic↗

Optimal placement of syringe-exchange programs.

Syringe-exchange programs (SEPs) will likely play a major role in slowing the spread of acquired immunodeficiency syndrome (AIDS) among injecting drug users (IDUs), but the success of any single SEP will depend to a large extent on where it is located. We show how the optimal position for a new SEP can be chosen given accurate knowledge of where IDUs live and how far they are willing to travel to an SEP. This information is not normally available, and one of our major points is that SEPs will necessarily be placed in suboptimal locations and will serve fewer IDUs than they otherwise might until it becomes available. Our method for choosing the best SEP placement is illustrated with Manhattan as an idealized example.

City Planning↗

MSAT: a multiple sequence alignment tool based on TOPS.

This article describes the development of a new method for multiple sequence alignment based on fold-level protein structure alignments, which provides an improvement in accuracy compared with the most commonly used sequence-only-based techniques. This method integrates the widely used, progressive multiple sequence alignment approach ClustalW with the Topology of Protein Structure (TOPS) topology-based alignment algorithm. The TOPS approach produces a structural alignment for the input protein set by using a topology-based pattern discovery program, providing a set of matched sequence regions that can be used to guide a sequence alignment using ClustalW. The resulting alignments are more reliable than a sequence-only alignment, as determined by 20-fold cross-validation with a set of 106 protein examples from the CATH database, distributed in seven superfold families. The method is particularly effective for sets of proteins that have similar structures at the fold level but low sequence identity. The aim of this research is to contribute towards bridging the gap between protein sequence and structure analysis, in the hope that this can be used to assist the understanding of the relationship between sequence, structure and function. The tool is available at http://balabio.dcs.gla.ac.uk/msat/.

Algorithms↗

Domain discovery method for topological profile searches in protein structures.

We describe a method for automated domain discovery for topological profile searches in protein structures. The method is used in a system TOPStructure for fast prediction of CATH classification for protein structures (given as PDB files). It is important for profile searches in multi-domain proteins, for which the profile method by itself tends to perform poorly. We also present an O(C(n)k + nk(2)) time algorithm for this problem, compared to the O(C(n)k + (nk)(2)) time used by a trivial algorithm (where n is the length of the structure, k is the number of profiles and C(n) is the time needed to check for a presence of a given motif in a structure of length n). This method has been developed and is currently used for TOPS representations of protein structures and prediction of CATH classification, but may be applied to other graph-based representations of protein or RNA structures and/or other prediction problems. A protein structure prediction system incorporating the domain discovery method is available at http://bioinf.mii.lu.lv/tops/.

Algorithms↗

An overview of data models for the analysis of biochemical pathways.

Biochemical pathways such as metabolic, regulatory or signal transduction pathways can be viewed as interconnected processes forming an intricate network of functional and physical interactions between molecular species in the cell. The amount of information available on such pathways for different organisms is increasing very rapidly. This is offering the possibility of performing various analyses on the structure of the full network of pathways for one organism as well as across different organisms, and has therefore generated interest in developing databases for storing and managing this information. Analysing these networks remains far from straightforward owing to the nature of the databases, which are often heterogeneous, incomplete or inconsistent. Pathway analysis is hence a challenging problem in systems biology and in bioinformatics. Various forms of data models have been devised for the analysis of biochemical pathways. This paper presents an overview of the types of models used for this purpose, concentrating on those concerned with the structural aspects of biochemical networks. In particular, the different types of data models found in the literature are classified using a unified framework. In addition, how these models have been used in the analysis of biochemical networks is described. This enables us to underline the strengths and weaknesses of the different approaches, as well as to highlight relevant future research directions.

Cell Physiological Phenomena↗

Ensemble machine learning on gene expression data for cancer classification.

Whole genome RNA expression studies permit systematic approaches to understanding the correlation between gene expression profiles to disease states or different developmental stages of a cell. Microarray analysis provides quantitative information about the complete transcription profile of cells that facilitate drug and therapeutics development, disease diagnosis, and understanding in the basic cell biology. One of the challenges in microarray analysis, especially in cancerous gene expression profiles, is to identify genes or groups of genes that are highly expressed in tumour cells but not in normal cells and vice versa. Previously, we have shown that ensemble machine learning consistently performs well in classifying biological data. In this paper, we focus on three different supervised machine learning techniques in cancer classification, namely C4.5 decision tree, and bagged and boosted decision trees. We have performed classification tasks on seven publicly available cancerous microarray data and compared the classification/prediction performance of these methods. We have observed that ensemble learning (bagged and boosted decision trees) often performs better than single decision trees in this classification task.

Algorithms↗

Multi-class protein fold classification using a new ensemble machine learning approach.

Protein structure classification represents an important process in understanding the associations between sequence and structure as well as possible functional and evolutionary relationships. Recent structural genomics initiatives and other high-throughput experiments have populated the biological databases at a rapid pace. The amount of structural data has made traditional methods such as manual inspection of the protein structure become impossible. Machine learning has been widely applied to bioinformatics and has gained a lot of success in this research area. This work proposes a novel ensemble machine learning method that improves the coverage of the classifiers under the multi-class imbalanced sample sets by integrating knowledge induced from different base classifiers, and we illustrate this idea in classifying multi-class SCOP protein fold data. We have compared our approach with PART and show that our method improves the sensitivity of the classifier in protein fold classification. Furthermore, we have extended this method to learning over multiple data types, preserving the independence of their corresponding data sources, and show that our new approach performs at least as well as the traditional technique over a single joined data source. These experimental results are encouraging, and can be applied to other bioinformatics problems similarly characterised by multi-class imbalanced data sets held in multiple data sources.

Amino Acid Sequence↗