PubMed Health⌕ Search

Biomedical subjects

Fritz Albregtsen

Publications and source records attributed to Fritz Albregtsen.

4 recordsLinked to original sources

Many are called, but few are chosen. Feature selection and error estimation in high dimensional spaces.

We address the problems of feature selection and error estimation when the number of possible feature candidates is large and the number of training samples is limited. A Monte Carlo study has been performed to illustrate the problems when using stepwise feature selection and discriminant analysis. The simulations demonstrate that in order to find the correct features, the necessary ratio of number of training samples to feature candidates is not a constant. It depends on the number of feature candidates, training samples and the Mahalanobis distance between the classes. Moreover, the leave-one-out error estimate may be a highly biased error estimate when feature selection is performed on the same data as the error estimation. It may even indicate complete separation of the classes, while no real difference between the classes exists. However, if feature selection and leave-one-out error estimation are performed in one process, an unbiased error estimate is achieved, but with high variance. The holdout error estimate gives a reliable estimate with low variance, depending on the size of the test set.

Algorithms↗

Low dimensional adaptive texture feature vectors from class distance and class difference matrices.

In many popular texture analysis methods, second or higher order statistics on the relation between pixel gray level values are stored in matrices. A high dimensional vector of predefined, nonadaptive features is then extracted from these matrices. Identifying a few consistently valuable features is important, as it improves classification reliability and enhances our understanding of the phenomena that we are modeling. Whatever sophisticated selection algorithm we use, there is a risk of selecting purely coincidental "good" feature sets, especially if we have a large number of features to choose from and the available data set is limited. In a unified approach to statistical texture feature extraction, we have used class distance and class difference matrices to obtain low dimensional adaptive feature vectors for texture classification. We have applied this approach to four relevant texture analysis methods. The new adaptive features outperformed the classical features when applied to the most difficult set of 45 Brodatz texture pairs. Class distance and difference matrices also clearly illustrated the difference in texture between cell nucleus images from two different prognostic classes of early ovarian cancer. For each of the texture analysis methods, one adaptive feature contained most of the discriminatory power of the method.

Algorithms↗

Normalizing the background and removing the trend in one-dimensional DNA fingerprint images.

Maximizing an individual's genetic information from its DNA fingerprint image depends on the number of bands distinguished from the background. To approach this goal, the background should be normalized while the information is preserved. Morphological operators have been used by some authors to normalize the background for two-dimensional gel images. Methods such as mean, median and "maxpolygon" are presented in this work to normalize the background in DNA fingerprint images. Mean and median methods will lead to some deformations. Visual evaluation of the results show that the original shape of the column signals are better preserved by the maxpolygon.

DNA Fingerprinting↗

Adaptive weighted least squares method for the estimation of DNA fragment lengths from agarose gels.

The size of DNA fragments is most frequently estimated from their electrophoretic mobilities. Agarose gels are used to estimate the size of DNA fragments ranging from a few hundred nucleotides to more than 20 kbp. The common practice when estimating the unknown fragment sizes is to plot the log of the size of molecular weight standards against their mobility and read the values of unknowns from this graph. However, due to perturbations in the gel, such plots often show pronounced curvature which may introduce significant subjectivity into the interpolation process. We present a new method "adaptive weighted least squares (AWLS)" based on the significance test to choose the order of polynomial. We compare this with the method introduced by Schaffer based on the modification of the Southern method. The results obtained by AWLS are significantly better than the method introduced by Schaffer. Different lanes are tested for consistency.

DNA↗