PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Feature selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Cholera, rotavirus and ETEC diarrhoea: some clinico-epidemiological features.

This paper analyses a few selected features from the history and clinical examination of 1258 patients with acute diarrhoea and a single laboratory diagnosis of either cholera, rotavirus, or enterotoxigenic (ETEC) Escherichia coli infection. Age distribution and seasonality in Bangladesh were also studied. The duration of illness before admission was not significantly different in the 3 groups. Cholera occurred especially in the spring and early winter. Most cholera patients were between 3 and 10 years of age. Over 37% of the patients developed severe dehydration. In about 90% of cholera cases, the stools were alkaline (pH greater than 7). ETEC infections were seen mostly in April-May and September-October. Infants were frequently affected but from age 25 onwards the age distribution closely followed that of cholera. Severe dehydration occurred in 8.3% of patients and was more frequent than in rotavirus cases. Stool pH was as frequently acidic as basic. Rotavirus cases were concentrated during the winter in patients under 2 years of age. They had marked vomiting, yet severe dehydration was almost absent. Cough was present in half of them. The stools were usually acidic. In spite of considerable overlap of signs and symptoms between the 3 aetiological groups, a presumptive diagnosis of cholera could be made in patients past infancy and early childhood who showed very severe dehydration. However, age-specific prevalence was strikingly different and seasonal variations considerable.

Adolescent↗

Computerized classification of malignant and benign microcalcifications on mammograms: texture analysis using an artificial neural network.

We investigated the feasibility of using texture features extracted from mammograms to predict whether the presence of microcalcifications is associated with malignant or benign pathology. Eighty-six mammograms from 54 cases (26 benign and 28 malignant) were used as case samples. All lesions had been recommended for surgical biopsy by specialists in breast imaging. A region of interest (ROI) containing the microcalcifications was first corrected for the low-frequency background density variation. Spatial grey level dependence (SGLD) matrices at ten different pixel distances in both the axial and diagonal directions were constructed from the background-corrected ROI. Thirteen texture measures were extracted from each SGLD matrix. Using a stepwise feature selection technique, which maximized the separation of the two class distributions, subsets of texture features were selected from the multi-dimensional feature space. A backpropagation artificial neural network (ANN) classifier was trained and tested with a leave-one-case-out method to recognize the malignant or benign microcalcification clusters. The performance of the ANN was analysed with receiver operating characteristic (ROC) methodology. It was found that a subset of six texture features provided the highest classification accuracy among the feature sets studied. The ANN classifier achieved an area under the ROC curve of 0.88. By setting an appropriate decision threshold, 11 of the 28 benign cases were correctly identified (39% specificity) without missing any malignant cases (100% sensitivity) for patients who had undergone biopsy. This preliminary result indicates that computerized texture analysis can extract mammographic information that is not apparent by visual inspection. The computer-extracted texture information may be used to assist in mammographic interpretation, with the potential to reduce biopsies of benign cases and improve the positive predictive value of mammography.

Breast Diseases↗

Transcriptome Analysis and Experimental Validation of Palmitoylation- Related Biomarkers in Atherosclerosis.

INTRODUCTION: Protein palmitoylation contributes to membrane localisation, signal transduction, and cell-fate regulation. It is closely associated with lipid metabolic dysfunction, immune inflammation, and vascular remodelling in atherosclerosis (AS). However, key palmitoylation-related transcriptomic markers and their potential causal associations with AS remain incompletely defined. METHODS: The Gene Expression Omnibus (GEO) dataset GSE100927 was used as the training cohort, and GSE43292 was used as an external validation cohort. Differentially expressed genes were identified using limma and intersected with palmitoylation-related genes to obtain palmitoylation-related differentially expressed genes (PRDEGs). Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were then performed using clusterProfiler. Two-sample Mendelian randomisation was used to evaluate potential causal relationships between characteristic genes and AS. Feature selection was conducted using random forest and support vector machine recursive feature elimination (SVM-RFE), and the overlapping genes selected by both methods were retained. Receiver operating characteristic (ROC) curves were used to assess diagnostic performance. A five-gene nomogram was constructed, and its clinical utility was evaluated using calibration curves and decision curve analysis (DCA). Gene set variation analysis (GSVA) was applied to compare pathway activity between high- and low-expression groups for each core gene. Single-cell analysis using Seurat and expression-based cell-cell communication analysis using CellChat were conducted with GSE159677, and upstream transcription factors were predicted using NetworkAnalyst. For in vivo validation, an AS model was established in ApoE⁸/⁸ mice fed a high-fat diet, and aortic gene and protein expression were assessed by RT-qPCR and western blotting. RESULTS: In GSE100927, 51 PRDEGs were identified. GO and KEGG enrichment analyses highlighted pathways associated with regulation of monoatomic ion transport, sarcomere and myofibril organisation, and immune inflammation. Mendelian randomisation suggested a potential protective causal association between SLC7A7 and AS. By integrating MR with random forest and SVM-RFE feature selection, we prioritised five core genes: PLCB2, GMIP, NEXN, PLN, and SLC7A7. These genes showed good diagnostic performance in GSE43292. The resulting nomogram was well calibrated and demonstrated stable net benefit in decision curve and clinical impact curve analyses. Single-gene GSVA identified consistently activated pathways across multiple genes, including innate and adaptive immune recognition, calcium signalling and myocardial contraction/cardiomyopathy, extracellular matrix-receptor interaction, cell junction pathways, autophagy-lysosome pathways, and several metabolic programmes. At the single-cell level, PLCB2 and GMIP were predominantly expressed in T cells and macrophages, NEXN and PLN were enriched in vascular smooth muscle cells, and SLC7A7 was mainly expressed in macrophages. CellChat analysis indicated increased signals for immune-related ligand-receptor interactions. In ApoE⁸/⁸ mice fed a high-fat diet, PLCB2, GMIP, and SLC7A7 were upregulated, whereas NEXN and PLN were downregulated; protein-level changes were concordant with the transcriptomic trends. DISCUSSION: These findings indicate that palmitoylation-related dysregulation in AS converges on immune inflammation, calcium signalling/contractile programmes, ECM remodelling, and autophagy-linked metabolism. The five-gene panel is supported by external validation, single-cell localisation to immune and vascular compartments, and concordant results in ApoE⁸/⁸ mice. CONCLUSION: This study identified and validated five palmitoylation-related genes associated with AS. SLC7A7 showed a potential protective causal signal in MR analysis. The enriched pathway patterns linked these genes to immune inflammation, calcium signalling-contraction coupling, ECM remodelling, cell adhesion, and autophagy- associated metabolic reprogramming. The five-gene nomogram showed potential utility for diagnostic classification and decision support, nominating candidate biomarkers and pathway targets for AS molecular subtyping, diagnosis, and mechanistic investigation.

Atherosclerosis (AS)↗

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics–based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics–based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans↗

Cross-feature spread of global attentional modulation in human area MT+.

Feature-based attention affects the processing of the selected feature throughout the visual field. Here, we show that such global attentional modulation is not restricted to the attended feature but spreads to task-irrelevant features that are bound to the attended one. Attention to a color in one of the visual hemifields affected the processing of task-irrelevant motion in the other hemifield when it was associated with a stimulus that shared the attended color. This cross-feature global attentional selection increased the duration of the motion aftereffect and the strength of functional magnetic resonance imaging responses in the motion-sensitive area MT+, evoked by the task-irrelevant motion. These findings imply that features belonging to the same object are bound and selected jointly even outside the focus of attention.

Attention↗

Differences in quantitative nuclear features between ductal carcinoma in situ (DCIS) with and without accompanying invasive carcinoma in the surrounding breast.

The aim of this study was to demonstrate differences in nuclear morphology between ductal carcinoma in situ (DCIS) without an invasive component and DCIS associated with invasive carcinoma in adjacent breast tissue. DCIS specimens of 60 non-comedo and 21 comedo cases were obtained from two groups of patients with or without invasive carcinoma of the breast. The analysis of DCIS nuclei was performed on formalin fixed deparaffinized thin sections stained with a stoichiometric stain following the Feulgen procedure. Nuclear features, related to nuclear size, shape and DNA distribution, were quantitatively characterized by high resolution image cytometry. Features associated with the presence of invasive carcinoma in the surrounding breast tissue were identified in DCIS nuclei (independent of nuclear grade). Features selected by the stepwise procedure of the discriminant function analysis were typically texture features describing the DNA distribution in the nucleus. A classification function based on the selected nuclear features predicted accurately the presence of invasive carcinoma in all comedo DCIS and in 80% of non-comedo DCIS cases. Our results indicate that quantitative nuclear features of DCIS nuclei are predictive of the accompanying invasion and may be helpful as a new tool in evaluation of DCIS patients.

Breast Neoplasms↗

Accuracy of the diagnosis of physical features of fetal alcohol syndrome by pediatricians after specialized training.

OBJECTIVES: Accurate and early diagnosis of the fetal alcohol syndrome is important for secondary prevention, intervention, and treatment, yet many pediatricians lack expertise in recognition of the characteristic features of this disorder. After a structured training program for pediatricians, we examined the ability to accurately diagnose fetal alcohol syndrome. METHODS: Two dysmorphologists conducted a 2-day training program in the diagnosis of the physical features of fetal alcohol syndrome for 4 pediatricians in Moscow. Dysmorphologists and pediatricians worked in teams to examine children, demonstrate techniques, and validate that pediatricians could identify physical features of this disorder under direct observation. Subsequently, pediatricians independently evaluated children in 41 boarding schools and orphanages. Those children diagnosed with fetal alcohol syndrome or deferred (possible fetal alcohol syndrome) by the pediatricians were then evaluated by the dysmorphologists. Accuracy of the diagnosis of fetal alcohol syndrome or deferred was assessed, as well as the interrater agreement for specific selected features of the disorder. RESULTS: A total of 110 children were examined by both the pediatricians and the dysmorphologists. Of these, 79 were identified with fetal alcohol syndrome by the pediatricians; in 66 (83.5%) of these children, the diagnosis was confirmed by the dysmorphologists. Among 31 children who were classified as deferred by the pediatricians, 21 (67.7%) were confirmed with either fetal alcohol syndrome or deferred by the dysmorphologists. With respect to selected structural features characteristic of fetal alcohol syndrome, good interrater agreement was noted for height and head circumference < or = 10th centile, whereas moderate-to-fair agreement was noted for smooth philtrum, long philtrum, presence of "hockey-stick" palmar crease, and palpebral fissure length < or = 10th centile. Poor agreement was noted for thin upper lip. CONCLUSIONS: After a relatively short training session, pediatricians were reasonably accurate in diagnosing fetal alcohol syndrome on the basis of physical features and in recognizing most of the selected specific features associated with the disorder.

Adolescent↗

Data mining tools for biological sequences.

We describe a methodology, as well as some related data mining tools, for analyzing sequence data. The methodology comprises three steps: (a) generating candidate features from the sequences, (b) selecting relevant features from the candidates, and (c) integrating the selected features to build a system to recognize specific properties in sequence data. We also give relevant techniques for each of these three steps. For generating candidate features, we present various types of features based on the idea of k-grams. For selecting relevant features, we discuss signal-to-noise, t-statistics, and entropy measures, as well as a correlation-based feature selection method. For integrating selected features, we use machine learning methods, including C4.5, SVM, and Naive Bayes. We illustrate this methodology on the problem of recognizing translation initiation sites. We discuss how to generate and select features that are useful for understanding the distinction between ATG sites that are translation initiation sites and those that are not. We also discuss how to use such features to build reliable systems for recognizing translation initiation sites in DNA sequences.

Artificial Intelligence↗

Distinguishing features of 16S rDNA gene for five dominating bacterial genus observed in bioremediation.

Defining a microbial community and identifying bacteria, at least at the genus level, is a first step in predicting the behavior of a microbial community in bioremediation. In biological treatment systems, the most dominating groups observed are Pseudomonas, Moraxella, Acinetobactor, Burkholderia, and Alcaligenes. Our interest lies in identifying the distinguishing features of these bacterial groups based on their 16S rDNA sequence data, which could be used further for generating genus-specific probes. Accordingly, 20 sequences representing different species from each genus above were retrieved, which constituted a training set. A 16-dimensional feature vector comprised of transition probabilities of nucleotides was considered and each sampled sequence was expressed in terms of these features. A stepwise feature selection method was used to identify features that are distinct across the species of these five groups. Wilk's lambda selection criterion was used and resulted in a subset with six distinguishing features. The discriminating efficacy of this subset was tested through multiple group discriminant analysis. Two linear composites, as a function of these features, could discriminate the test set of forty-five sequences from these groups with 95% accuracy, thereby ascertaining the relevance of the identified features. The geometric representation of feature correlation in the reduced discriminant space demonstrated the dominance of identified features in specific groups. These features independently or in combination could be used to generate genus-specific patterns to design probes, so as to develop a tracking tool for the selected group of bacteria.

Bacteria↗

Malignant round-cell tumours of bone: an analytical histological study from the Cancer Research Campaign's bone tumour panel.

A study of 40 cases of malignant round-cell tumour of one was made from the files of the Cancer Research Campaign's Bone Tumour Panel. Five pathologists made a careful study of observer error, involving repeated examination of routine paraffin sections, to determine whether the cases were a homogeneous group or a collection of differing sub-groups. Cell outline, nuclear staining, nuclear pleomorphism, conspicuous nucleoli, reticulin pattern and intracellular glycogen were the histological features selected for study. For each feature, the results were analysed to assess the importance of differences between tumours, between samples of tissue from the same tumour, and between observers. It is concluded that round-cell tumours of bone are a heterogeneous group, although completely distinct sub-groups could not be identified. Certain histological features tend to be associated, and it is reasonable to distinguish on histological grounds between Ewing's sarcoma and reticulum-cell sarcoma, although some tumours are not typical of either group.

Age Factors↗

Basal forebrain stimulation changes cortical sensitivities to complex sound.

Experience affects how brains respond to sound. Here, we examined how the sensitivity and selectivity of auditory cortical neuronal responses were affected in adult rats by the repeated presentation of a complex sound that was paired with basal forebrain stimulation. The auditory cortical region that was responsive to complex sound was 2-5 five times greater in area in paired-stimulation rats than in naive rats. Magnitudes of neuronal responses evoked by complex sounds were also greatly increased by associative pairing, as were the percentages of neurons that responded selectively to the specific spectrotemporal features that were paired with stimulation. These findings demonstrate that feature selectivity within the auditory cortex can be flexibly altered in adult mammals through appropriate intensive training.

Acoustic Stimulation↗

A review of caveats in statistical nuclear image analysis.

A large body of the published literature in nuclear image analysis do not evaluate their findings on an independent data set. Hence, if several features are evaluated on a limited data set over-optimistic results are easily achieved. In order to find features that separate different outcome classes of interest, statistical evaluation of the nuclear features must be performed. Furthermore, to classify an unknown sample using image analysis, a classification rule must be designed and evaluated. Unfortunately, statistical evaluation methods used in the literature of nuclear image analysis are often inappropriate. The present article discusses some of the difficulties in statistical evaluation of nuclear image analysis, and a study of cervical cancer is presented in order to illustrate the problems. In conclusion, some of the most severe errors in nuclear image analysis occur in analysis of a large feature set, including few patients, without confirming the results on an independent data set. To select features, Bonferroni correction for multiple test is recommended, together with a standard feature set selection method. Furthermore, we consider that the minimum requirement of performing statistical evaluation in nuclear image analysis is confirmation of the results on an independent data set. We suggest that a consensus of how to perform evaluation of diagnostic and prognostic features is necessary, in order to develop reliable tools for clinical use, based on nuclear image analysis.

Data Interpretation, Statistical↗

Multiple neural network classification scheme for detection of colonic polyps in CT colonography data sets.

RATIONALE AND OBJECTIVES: A new classification system for colonic polyp detection, designed to increase sensitivity and reduce the number of false-positive findings with computed tomographic colonography, was developed and tested in this study. MATERIALS AND METHODS: The system involves classification by a committee of neural networks (NNs), each using largely distinct subsets of features selected from a general set. Back-propagation NNs trained with the Levenberg-Marquardt algorithm were used as primary classifiers (committee members). The set of features included region density, Gaussian and mean curvature and sphericity, lesion size, colon wall thickness, and the means and standard deviations of all of these values. Subsets of variables were initially selected because of their effectiveness according to training and test sample misclassification rates. The final decision for each case is based on the majority vote across the networks and reflects the weighted votes of all networks. The authors also introduce a smoothed cross-validation method designed to improve estimation of the true misclassification rates by reducing bias and variance. RESULTS: This committee method reduced the false-positive rate by 36%, a clinically meaningful reduction, and improved sensitivity by an average of 6.9% compared with decisions made by any single NN. The overall sensitivity and specificity were 82.9% and 95.3%, respectively, when sensitivity was estimated by means of smoothed cross-validation. CONCLUSION: The proposed method of using multiple classifiers and majority voting is recommended for classification tasks with large sets of input features, particularly when selected feature subsets may not be equally effective and do not provide satisfactory true- and false-positive rates. This approach reduces variance in estimates of misclassification rates.

Algorithms↗

Selective attention to specific features within objects: behavioral and electrophysiological evidence.

Evidence regarding the ability of attention to bias neural processing at the level of single features has been gathering steadily, but most of the experiments to date used arrays with multiple objects and locations, making it difficult to rule out indirect influences from object or spatial attention. To investigate feature-specific selective attention, we have assessed the ability to select and ignore individual features within the same object. We used a negative-priming paradigm in which the color or the direction of internal motion of the object could determine the relevant response. Bidimensional (colored and moving) and unidimensional (colored and stationary, or gray and moving) stimuli appeared in unpredictable order. In successive blocks, participants were instructed that one feature dimension was dominant. During that block, participants responded according to the dominant dimension for bidimensional stimuli. For unidimensional stimuli, participants responded to the only dimension of the stimulus that afforded a response, regardless of the instruction for the block. The ability to inhibit irrelevant task information at the level of specific features (negative priming for features) was indexed by a decrease in performance to detect one particular feature value (e.g., red) if the same feature value (red) but not another color value (green) had been ignored in the previous bidimensional stimulus. Behavioral results confirmed the existence of inhibitory, negative-priming mechanisms at the single-feature level for both color and motion dimensions of stimuli. Event-related potentials recorded during task performance revealed the dynamics of neural modulation by feature attention. Comparisons were made using the identical physical stimuli under different conditions of attention to isolate purely attentional effects. Processing of identical bidimensional stimuli was compared as a function of the dimension of attention (color, motion). Processing of identical unidimensional stimuli that followed bidimensional stimuli was also compared to identify possible effects of feature-specific negative priming. The electrophysiological effects revealed that inhibition of irrelevant features leads to modulation of brain activity during early stages of perceptual analysis.

Adult↗

Is contrast just another feature for visual selective attention?

The biased-competition theory of attention [Annual Review of Neuroscience 18 (1995) 193] suggests that attention and stimulus contrast trade off, and implies that high-contrast stimuli should be easy to attend to and hard to ignore. To test this, observers searched displays for a target digit. Observers were well able to exclude high-contrast distractors when attempting to search only among low-contrast stimuli (Experiment 1). In Experiments 2 and 3, location determined which stimuli were relevant. When contrast of relevant and irrelevant stimuli was uncertain (due to contrast varying between trials, Experiment 2), increasing the contrast of distractors impaired performance. However, when contrast was certain (due to blocking of trials, Experiment 3) and targets were of low contrast, high contrast distractors produced less interference than low contrast distractors. The ability of subjects to attend selectively to low vs. high contrast items in Experiments 1 and 3 suggests that selectivity for stimulus contrast might be similar to other types of feature selectivity (e.g., color and location). Such findings are inconsistent with the biased competition theory regarding the interplay of contrast and attention. However, results from Experiment 2 suggest that, when target contrast varies, the default tendency is to attend to high-contrast items.

Adaptation, Physiological↗

Bayesian feature and model selection for Gaussian mixture models.

We present a Bayesian method for mixture model training that simultaneously treats the feature selection and the model selection problem. The method is based on the integration of a mixture model formulation that takes into account the saliency of the features and a Bayesian approach to mixture learning that can be used to estimate the number of mixture components. The proposed learning algorithm follows the variational framework and can simultaneously optimize over the number of components, the saliency of the features, and the parameters of the mixture model. Experimental results using high-dimensional artificial and real data illustrate the effectiveness of the method.

Algorithms↗

Systematic benchmarking of microarray data classification: assessing the role of non-linearity and dimensionality reduction.

MOTIVATION: Microarrays are capable of determining the expression levels of thousands of genes simultaneously. In combination with classification methods, this technology can be useful to support clinical management decisions for individual patients, e.g. in oncology. The aim of this paper is to systematically benchmark the role of non-linear versus linear techniques and dimensionality reduction methods. RESULTS: A systematic benchmarking study is performed by comparing linear versions of standard classification and dimensionality reduction techniques with their non-linear versions based on non-linear kernel functions with a radial basis function (RBF) kernel. A total of 9 binary cancer classification problems, derived from 7 publicly available microarray datasets, and 20 randomizations of each problem are examined. CONCLUSIONS: Three main conclusions can be formulated based on the performances on independent test sets. (1) When performing classification with least squares support vector machines (LS-SVMs) (without dimensionality reduction), RBF kernels can be used without risking too much overfitting. The results obtained with well-tuned RBF kernels are never worse and sometimes even statistically significantly better compared to results obtained with a linear kernel in terms of test set receiver operating characteristic and test set accuracy performances. (2) Even for classification with linear classifiers like LS-SVM with linear kernel, using regularization is very important. (3) When performing kernel principal component analysis (kernel PCA) before classification, using an RBF kernel for kernel PCA tends to result in overfitting, especially when using supervised feature selection. It has been observed that an optimal selection of a large number of features is often an indication for overfitting. Kernel PCA with linear kernel gives better results.

Algorithms↗

The RIN: an RNA integrity number for assigning integrity values to RNA measurements.

BACKGROUND: The integrity of RNA molecules is of paramount importance for experiments that try to reflect the snapshot of gene expression at the moment of RNA extraction. Until recently, there has been no reliable standard for estimating the integrity of RNA samples and the ratio of 28S:18S ribosomal RNA, the common measure for this purpose, has been shown to be inconsistent. The advent of microcapillary electrophoretic RNA separation provides the basis for an automated high-throughput approach, in order to estimate the integrity of RNA samples in an unambiguous way. METHODS: A method is introduced that automatically selects features from signal measurements and constructs regression models based on a Bayesian learning technique. Feature spaces of different dimensionality are compared in the Bayesian framework, which allows selecting a final feature combination corresponding to models with high posterior probability. RESULTS: This approach is applied to a large collection of electrophoretic RNA measurements recorded with an Agilent 2100 bioanalyzer to extract an algorithm that describes RNA integrity. The resulting algorithm is a user-independent, automated and reliable procedure for standardization of RNA quality control that allows the calculation of an RNA integrity number (RIN). CONCLUSION: Our results show the importance of taking characteristics of several regions of the recorded electropherogram into account in order to get a robust and reliable prediction of RNA integrity, especially if compared to traditional methods.

Algorithms↗