PubMed Health⌕ Search

Biomedical subjects

Mark Steyvers

Publications and source records attributed to Mark Steyvers.

4 recordsLinked to original sources

Finding scientific topics.

A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.

Databases, Factual↗

The effect of normative context variability on recognition memory.

According to some theories of recognition memory (e.g., S. Dennis & M. S. Humphreys, 2001), the number of different contexts in which words appear determines how memorable individual occurrences of words will be: A word that occurs in a small number of different contexts should be better recognized than a word that appears in a larger number of different contexts. To empirically test this prediction, a normative measure is developed, referred to here as context variability, that estimates the number of different contexts in which words appear in everyday life. These findings confirm the prediction that words low in context variability are better recognized (on average) than words that are high in context variability.

Attention↗

Conceptual interrelatedness and caricatures.

Concepts are interrelated to the extent that the characterization of each concept is influenced by the other concepts, and are isolated to the extent that the characterization of one concept is independent of other concepts. The relative categorization accuracy of the prototype and caricature of a concept can be used as a measure of concept interrelatedness. The prototype is the central tendency of a concept, whereas a caricature deviates from the concept's central tendency in the direction opposite the central tendency of other acquired concepts. The prototype is predicted to be relatively well categorized when a concept is relatively independent of other concepts, but the caricature is predicted to be relatively well categorized when a concept is highly related to other concepts. Support for these predictions comes from manipulations of the labels given to simultaneously acquired concepts (Experiment 1) and of the order of categories during learning (Experiment 2).

Caricatures as Topic↗

Feature frequency effects in recognition memory.

Rare words are usually better recognized than common words, a finding in recognition memory known as the word-frequency effect. Some theories predict the word-frequency effect because they assume that rare words consist of more distinctive features than do common words (e.g., Shiffrin & Steyvers's, 1997, REM theory). In this study, recognition memory was tested for words that vary in the commonness of their orthographic features, and we found that recognition was best for words made up of primarily rare letters. In addition, a mirror effect was observed: Words with rare letters had a higher hit rate and a lower false-alarm rate than did words with common letters. We also found that normative word frequency affects recognition independently of letter frequency. Therefore, the distinctiveness of a word's orthographic features is one, but not the only, factor necessary to explain the word-frequency effect.

Humans↗