PubMed Health⌕ Search

Biomedical subjects

Robert T Gray

Publications and source records attributed to Robert T Gray.

2 recordsLinked to original sources

Image transform bootstrapping and its applications to semantic scene classification.

The performance of an exemplar-based scene classification system depends largely on the size and quality of its set of training exemplars, which can be limited in practice. In addition, in nontrivial data sets, variations in scene content as well as distracting regions may exist in many testing images to prohibit good matches with the exemplars. Various boosting schemes have been proposed in machine learning, focusing on the feature space. We introduce the novel concept of image-transform bootstrapping using transforms in the image space to address such issues. In particular, three major schemes are described for exploiting this concept to augment training, testing, and both. We have successfully applied it to three applications of increasing difficulty: sunset detection, outdoor scene classification, and automatic image orientation detection. It is shown that appropriate transforms and meta-classification methods can be selected to boost performance according to the domain of the problem and the features/classifier used.

Algorithms↗

Psychophysical study of image orientation perception.

The experiment reported here investigates the perception of orientation of color photographic images. A collection of 1000 images (mix of professional photos and consumer snapshots) was used in this study. Each image was examined by at least five observers and shown at varying resolutions. At each resolution, observers were asked to indicate the image orientation, the level of confidence, and the cues they used to make the decision. The results show that for typical images, accuracy is close to 98% when using all available semantic cues from high-resolution images, and 84% when using only low-level vision features and coarse semantics from thumbnails. The accuracy by human observers suggests an upper bound for the performance of an automatic system. In addition, the use of a large, carefully chosen image set that spans the 'photo space' (in terms of occasions and subject matter) and extensive interaction with the human observers reveals cues used by humans at various image resolutions: sky and people are the most useful and reliable among a number of important semantic cues.

Adult↗