PubMed Health⌕ Search

Biomedical subjects

Martin Kulldorff

Publications and source records attributed to Martin Kulldorff.

At least 19 recordsLinked to original sources

Issues in applied statistics for public health bioterrorism surveillance using multiple data streams: research needs.

The objective of this report is to provide a basis to inform decisions about priorities for developing statistical research initiatives in the field of public health surveillance for emerging threats. Rapid information system advances have created a vast opportunity of secondary data sources for information to enhance the situational and health status awareness of populations. While the field of medical informatics and initiatives to standardize healthcare-seeking encounter records continue accelerating, it is necessary to adapt analytic and statistical methodologies to mature in sync with sibling information science technologies. One major right-of-passage for statistical inference is to advance the optimal application of analytic methodologies for using multiple data streams in detecting and characterizing public health population events of importance. This report first describes the problem in general and the data context, then delineates more specifically the practical nature of the problem and the related issues. Approaches currently applied to data with time-series, statistical process control and traditional inference concepts are described with examples in the section on Statistics and the Role of the Analytic Surveillance Data Monitor. These are the techniques that are providing substance to surveillance professionals and enabling use of multiple data streams. The next section describes use of a more complex approach that takes temporal as well as spatial dimensions into consideration for detection and situational awareness regarding event distributions. The space-time statistic has successfully been used to detect and track public health events of interest. Important research questions which are summarized at the end of this report are described in more detail with respect to the methodological application in the respective sections. This was thought to help elucidate the research requirements as summarized later in the report. Following the description of the space-time scan statistical application; this report extends to a less traditional area of promise given what has been observed in recent application of analytic methods. Bayesian networks (BNs) represent a conceptual step with advantages of flexibility for the public health surveillance community. Progression from traditional to the more extending statistical concepts in the context of the dynamic status quo of responsibility and challenge, leads to a conclusion consisting of categorical research needs. The report is structured by design to inform judgment about how to build on practical systems to achieve better analytic outcomes for public health surveillance. There are references to research issues throughout the sections with a summarization at the end, which also includes items previously unmentioned in the report.

Algorithms↗

Multivariate scan statistics for disease surveillance.

In disease surveillance, there are often many different data sets or data groupings for which we wish to do surveillance. If each data set is analysed separately rather than combined, the statistical power to detect an outbreak that is present in all data sets may suffer due to low numbers in each. On the other hand, if the data sets are added by taking the sum of the counts, then a signal that is primarily present in one data set may be hidden due to random noise in the other data sets. In this paper, we present an extension of the spatial and space-time scan statistic that simultaneously incorporates multiple data sets into a single likelihood function, so that a signal is generated whether it occurs in only one or in multiple data sets. This is done by defining the combined log likelihood as the sum of the individual log likelihoods for those data sets for which the observed case count is more than the expected. We also present another extension, where the concept of combining likelihoods from different data sets is used to adjust for covariates. Using data from the National Bioterrorism Syndromic Surveillance Demonstration Project, we illustrate the new method using physician telephone calls, regular physician visits and urgent care visits by Harvard Pilgrim Health Care members cared for by Harvard Vanguard Medical Associates, a large multi-specialty group practice in Massachusetts. For upper and lower gastrointestinal (GI) illness, there were on average 20 telephone calls, nine urgent care visits and 22 regular physician visits per day. The strongest signal was generated by a single data set and due to a familial outbreak of pinworm disease. The second and third strongest signals were generated by the combined strength of two of the three data sets.

Boston↗

A spatial scan statistic for ordinal data.

Spatial scan statistics are widely used for count data to detect geographical disease clusters of high or low incidence, mortality or prevalence and to evaluate their statistical significance. Some data are ordinal or continuous in nature, however, so that it is necessary to dichotomize the data to use a traditional scan statistic for count data. There is then a loss of information and the choice of cut-off point is often arbitrary. In this paper, we propose a spatial scan statistic for ordinal data, which allows us to analyse such data incorporating the ordinal structure without making any further assumptions. The test statistic is based on a likelihood ratio test and evaluated using Monte Carlo hypothesis testing. The proposed method is illustrated using prostate cancer grade and stage data from the Maryland Cancer Registry. The statistical power, sensitivity and positive predicted value of the test are examined through a simulation study.

Cluster Analysis↗

An elliptic spatial scan statistic.

The spatial scan statistic is commonly used for geographical disease cluster detection, cluster evaluation and disease surveillance. The most commonly used shape of the scanning window is circular. In this paper we explore an elliptic version of the spatial scan statistic, using a scanning window of variable location, shape (eccentricity), angle and size, and with and without an eccentricity penalty. The method is applied to breast cancer mortality data from Northeastern United States and female oral cancer mortality in the United States. Power comparisons are made with the circular scan statistic.

Adolescent↗

Evaluating spatial surveillance: detection of known outbreaks in real data.

Since the anthrax attacks of October 2001 and the SARS outbreaks of recent years, there has been an increasing interest in developing surveillance systems to aid in the early detection of such illness. Systems have been established which do this is by monitoring primary health-care visits, pharmacy sales, absenteeism records, and other non-traditional sources of data. While many resources have been invested in establishing such systems, relatively little effort has as yet been expended in evaluating their performance. One way to evaluate a given surveillance system is to compare the signals it generates with known outbreaks identified in other systems. In public health practice, for example, public health departments investigate reports of illness and sometimes track hospital admissions. Comparison of new systems with extant systems cannot generate estimates of test characteristics such as sensitivity and specificity, since the actual number of positives and negatives cannot be known. However, the comparison can reveal whether a new or proposed system's signals match outbreaks detected by the existing system. This could help support or reject the new system as an alternative or complement to the extant system. We propose three methods to test the null hypothesis that the new system does not signal true outbreaks more often than would be expected by chance. The methods differ in the restrictiveness of the assumptions required. Each test may detect weaknesses in the new system, depending on the distribution of outbreaks and can be used to construct confidence limits on the agreement between the new system's signals and the outbreaks, given the distribution of the signals. They can be used to assess whether the new system works in that it detects the outbreaks better than chance would suggest and can also determine if the new systems' signals are generated earlier than an extant system.

Data Interpretation, Statistical↗

Likelihood based tests for spatial randomness.

Many different methods have been proposed to test the spatial randomness of a point pattern adjusting for an inhomogeneous background population. These tests can be classified into cluster detection tests, concerned with the detection and inference of local clusters, and global clustering tests, which collect evidence for clustering throughout the study region. This paper is mainly concerned about global clustering tests. Some tests for spatial randomness are based on likelihoods, which include the spatial and space-time scan statistics with variable window size and Gangnon and Clayton's weighted average likelihood ratio tests. Both of these tests perform well compared to other tests for cluster detection and global clustering, respectively. In this study, we develop other likelihood based tests for global clustering and we explore the use of different weight functions with these tests. The power of these tests is evaluated using simulated data set and compared with existing methods.

Cluster Analysis↗

Geographically based investigation of prostate cancer mortality in four U.S. Northern Plain states.

BACKGROUND: Historically, prostate cancer mortality rates have been elevated in the U.S. Northern Plains states. The purpose of this study was to investigate possible contributing factors, especially whether there was any association with crop patterns. METHODS: Prostate cancer mortality rates (1950-2000) in four northern plains states (MN, MT, ND, and SD) were compared to rates for 46 other U.S. states. Within the four states, county rates in urban, less urban, and rural areas also were compared. For additional analysis, urban counties and counties with <10% of county area in crops were excluded. The average percent of county area in total cropland 1930-1950 and 1954-1974 was estimated. Using Poisson regression, we investigated whether the average percentage of county area in total cropland, 1930-1950 and 1954-1974, was associated with prostate cancer mortality rates, 1975-2000, respectively. Poisson regression analyses were also used to evaluate associations between rates and major crops, which included spring and durum wheat, winter wheat, corn, and other crops. Population centroids of the Census 2000 block groups were used to estimate the percentage of males aged 35 and older residing in close proximity to small grains crops. RESULTS: Mortality rates were higher in rural compared to urban counties in 1950-2000 (rate ratio [RR]=1.032; 95% CI=1.001-1.063). Rates in 1950-1974 were significantly associated with production of corn and other crops in 1930-1950 (corn: RR per 10% increase=1.033, 95% CI=1.012-1.054; other crops: RR=1.042, 95% CI=1.021-1.063). Mortality rates in 1975-2000 were significantly associated with spring and durum wheat production in 1954-1974 (RR per 10% increase=1.042, 95% CI=1.017-1.067). Prostate cancer mortality rates increased as the percentage of population living within 500 m of small grains crops increased. CONCLUSIONS: Epidemiologic studies to evaluate agricultural practices are warranted to further evaluate the observed associations.

Adult↗

Cancer map patterns: are they random or not?

BACKGROUND: Maps depicting the geographic variation in cancer incidence, mortality or treatment can be useful tools for developing cancer control and prevention programs, as well as for generating etiologic hypotheses. An important question with every cancer map is whether the geographic pattern seen is due to random fluctuations, as by pure chance there are always some areas with more cases than expected, or whether the map reflects true underlying geographic variation in screening, treatment practices, or etiologic risk factors. METHODS: Nine different tests for spatial randomness are evaluated in very practical settings by applying them to cancer maps for different types of data at different scales of spatial resolution: breast, prostate, and thyroid cancer incidence; breast cancer treatment and prostate cancer stage in Connecticut; and nasopharynx and prostate cancer mortality in the U.S. RESULTS: Tango's MEET, Oden's Ipop, and the spatial scan statistic performed well across all the data sets. Besag-Newell's R, Cuzick-Edwards k-NN, and Turnbull's CEPP often perform well, but the results are highly dependent on the parameter chosen. Moran's I performs poorly for most data sets, whereas Swartz Entropy Test and Whittemore's Test perform well for some data sets but not for other. CONCLUSIONS: When publishing cancer maps we recommend evaluating the spatial patterns observed using Tango's MEET, a global clustering test, and the spatial scan statistic, a cluster detection test.

Adolescent↗

Missing stage and grade in Maryland prostate cancer surveillance data, 1992-1997.

BACKGROUND: Missing data in cancer surveillance records are common; however, little information exists on the types of cases most likely to have missing data, or how missing data influence research or policy. Two clinical elements often missing in surveillance data are histologic grade and stage of disease. Missing data are either not clinically ascertained or not successfully abstracted. METHODS: Prostate cancer cases (N=22,217) reported to the Maryland Cancer Registry during 1992-1997 were geocoded by residence and analyzed. Multi-level logistic regression was used to examine case attributes and area-level demographic, economic, and health services characteristics predictive of either missing stage or grade. A scanning statistic was used to explore geographic clustering of high and low rates of missing stage and grade within the state, before and after adjustment for significant variables from multi-level models. RESULTS: Older age, black race, missing grade, and higher county-level median income increased the likelihood of missing stage, whereas more recent year of diagnosis, higher blockgroup-level median income, and county-level rurality decreased the likelihood. Older age, missing or later stage, higher blockgroup-level median income, and more urologists per case in one's county of residence increased the likelihood of missing tumor grade, and more recent year of diagnosis, higher county-level median income, and rurality decreased the likelihood. Adjustment reduced statistically significant clusters of missing stage from six to two, and clusters of missing grade from three to zero. CONCLUSIONS: Results suggest systematic influences on missing stage and grade, which could be investigated with case-control follow-back studies.

Adolescent↗

Tango's maximized excess events test with different weights.

BACKGROUND: Tango's maximized excess events test (MEET) has been shown to have very good statistical power in detecting global disease clustering. A nice feature of this test is that it considers a range of spatial scale parameters, adjusting for the multiple testing. This means that it has good power to detect a wide range of clustering processes. The test depends on the functional form of a weight function, and it is unknown how sensitive the test is to the choice of this weight function and what function provides optimal power for different clustering processes. In this study, we evaluate the performance of the test for a wide range of weight functions. RESULTS: The power varies greatly with different choice of weight. Tango's original choice for the weight function works very well. There are also other weight functions that provide good power. CONCLUSION: We recommend the use of Tango's MEET to test global disease clustering, either with the original weight or one of the alternate weights that have good power.

Journal Article↗

Geographic prediction of human onset of West Nile virus using dead crow clusters: an evaluation of year 2002 data in New York State.

The risk of becoming a West Nile virus case in New York State, excluding New York City, was evaluated for persons whose town of residence was proximal to spatial clusters of dead American crows (Corvus brachyrhynchos). Weekly clusters were delineated for June-October 2002 by using both the binomial spatial scan statistic and kernel density smoothing. The relative risk of a human case was estimated for different spatial-temporal exposure definitions after adjusting for population density and age distribution using Poisson regression, adjusting for week and geographic region, and conducting Cox proportional hazards modeling, where the week that a human case was identified was treated as the failure time and baseline hazard was stratified by region. The risk of becoming a West Nile virus case was positively associated with living in towns proximal to dead crow clusters. The highest risk was consistently for towns associated with a cluster in the current or prior 1-2 weeks. Weaker, but positive associations were found for towns associated with a cluster in just the 1-2 prior weeks, indicating an ability to predict onset in a timely fashion.

Animals↗

Meat, meat cooking methods and preservation, and risk for colorectal adenoma.

Cooking meat at high temperatures produces heterocyclic amines (HCAs) and polycyclic aromatic hydrocarbons (PAHs). Processed meats contain N-nitroso compounds. Meat intake may increase cancer risk as HCAs, PAHs, and N-nitroso compounds are carcinogenic in animal models. We investigated meat, processed meat, HCAs, and the PAH benzo(a)pyrene and the risk of colorectal adenoma in 3,696 left-sided (descending and sigmoid colon and rectum) adenoma cases and 34,817 endoscopy-negative controls. Dietary intake was assessed using a 137-item food frequency questionnaire, with additional questions on meats and meat cooking practices. The questionnaire was linked to a previously developed database to determine exposure to HCAs and PAHs. Intake of red meat, with known doneness/cooking methods, was associated with an increased risk of adenoma in the descending and sigmoid colon [odds ratio (OR), 1.26; 95% confidence interval (95% CI), 1.05-1.50 comparing extreme quintiles of intake] but not rectal adenoma. Well-done red meat was associated with increased risk of colorectal adenoma (OR, 1.21; 95% CI, 1.06-1.37). Increased risks for adenoma of the descending colon and sigmoid colon were observed for the two HCAs: 2-amino-3,8-dimethylimidazo[4,5]quinoxaline and 2-amino-1-methyl-6-phenylimidazo[4,5]pyridine (OR, 1.18; 95% CI, 1.01-1.38 and OR, 1.17, 95% CI, 1.01-1.35, respectively) as well as benzo(a)pyrene (OR, 1.18; 95% CI, 1.02-1.35). Greater intake of bacon and sausage was associated with increased colorectal adenoma risk (OR, 1.14; 95% CI, 1.00-1.30); however, total intake of processed meat was not (OR, 1.04; 95% CI, 0.90-1.19). Our study of screening-detected colorectal adenomas shows that red meat and meat cooked at high temperatures are associated with an increased risk of colorectal adenoma.

Adenoma↗

Meat and meat-mutagen intake and risk of non-Hodgkin lymphoma: results from a NCI-SEER case-control study.

Non-Hodgkin Lymphoma (NHL) incidence has risen dramatically over past decades, but the reasons for most of this increase are not known. Meat cooked well-done using high-temperature cooking techniques produces heterocyclic amines (HCAs) and polycyclic aromatic hydrocarbons (PAHs) such as benzo[a]pyrene (B[a]P). This study was conducted as a population-based case-control study in Iowa, Detroit, Seattle and Los Angeles and was designed to determine whether meat, meat-cooking methods, HCAs or PAHs from meat were associated with NHL risk. This study consisted of 458 NHL cases, diagnosed between 1998 and 2000, and 383 controls. Participants completed a 117-item food frequency questionnaire (FFQ), with graphical aids to assess the meat-cooking method and doneness level, which was linked to a HCA and B[a]P database. Logistic regression, comparing the fourth to the first quartile, found no association between red meat or processed meat intake and risk for NHL [odds ratio (OR) and 95% confidence interval (CI): 1.10 (0.67-1.81) and 1.18 (0.74-1.89), respectively]. A marginally significant elevated risk for NHL was associated with broiled meat [OR and 95% CI: 1.32 (0.99-1.77); P trend = 0.09], comparing those who consumed broiled meat with those who did not. The degree to which meat was cooked was not associated with the risk for NHL, although one of the HCAs, DiMeIQx (2-amino-3,4,8-trimethylimidazo[4,5-f]quinoxaline), was associated with an inverse risk. Fat intake was associated with a significantly elevated risk for NHL [OR and 95% CI: 1.60 (1.05-2.45); P trend = 0.12]; in contrast, animal protein was inversely associated with risk for NHL [OR and 95% CI: 0.39 (0.22-0.70); P trend = 0.004]. Overall, our study suggests that consumption of meat, whether or not it is well-done, does not increase the risk of NHL. Furthermore, neither HCAs nor B[a]P from meat increase the risk of NHL.

Adult↗

A space-time cluster of adverse events associated with canine rabies vaccine.

Electronic medical records of a large veterinary practice were used for surveillance of potential space-time clustering of adverse events associated with rabies vaccination in dogs. The study population was 257,564 dogs vaccinated in 169 hospitals in 13 US metropolitan areas during a 24-month period. Using a scan statistic for population rate data, significant space-time clusters were identified involving the Atlanta and Tampa/St. Petersburg areas during a 4-month period. Separate spatial-temporal analyses of these cities using coordinates for individual address coordinates identified one significant patient cluster (P=0.002), associated with a 23.26 km-radius area in Atlanta (20 adverse events in 702 dogs; 2.85%) from November 2002 through February 2003. This percentage of adverse events was significantly increased after adjustment for host-related factors and the number of concurrent vaccinations.

Adverse Drug Reaction Reporting Systems↗

Lumping or splitting: seeking the preferred areal unit for health geography studies.

BACKGROUND: Findings are compared on geographic variation of incident and late-stage cancers across Connecticut using different areal units for analysis. RESULTS: Few differences in results were found for analyses across areal units. Global clustering of incident prostate and breast cancer cases was apparent regardless of the level of geography used. The test for local clustering found approximately the same locales, populations at risk and estimated effects. However, some discrepancies were uncovered. CONCLUSION: In the absence of conditions calling for surveillance of small area cancer clusters ('hot spots'), the rationale for accepting the burdens of preparing data at levels of geography finer than the census tract may not be compelling.

Journal Article↗

A space-time permutation scan statistic for disease outbreak detection.

BACKGROUND: The ability to detect disease outbreaks early is important in order to minimize morbidity and mortality through timely implementation of disease prevention and control measures. Many national, state, and local health departments are launching disease surveillance systems with daily analyses of hospital emergency department visits, ambulance dispatch calls, or pharmacy sales for which population-at-risk information is unavailable or irrelevant. METHODS AND FINDINGS: We propose a prospective space-time permutation scan statistic for the early detection of disease outbreaks that uses only case numbers, with no need for population-at-risk data. It makes minimal assumptions about the time, geographical location, or size of the outbreak, and it adjusts for natural purely spatial and purely temporal variation. The new method was evaluated using daily analyses of hospital emergency department visits in New York City. Four of the five strongest signals were likely local precursors to citywide outbreaks due to rotavirus, norovirus, and influenza. The number of false signals was at most modest. CONCLUSION: If such results hold up over longer study times and in other locations, the space-time permutation scan statistic will be an important tool for local and national health departments that are setting up early disease detection surveillance systems.

Data Interpretation, Statistical↗

Geographical clustering of prostate cancer grade and stage at diagnosis, before and after adjustment for risk factors.

BACKGROUND: Spatial variation in patterns of disease outcomes is often explored with techniques such as cluster detection analysis. In other types of investigations, geographically varying individual or community level characteristics are often used as independent predictors in statistical models which also attempt to explain variation in disease outcomes. However, there is a lack of research which combines geographically referenced exploratory analysis with multilevel models. We used a spatial scan statistic approach, in combination with predicted block group-level disease patterns from multilevel models, to examine geographic variation in prostate cancer grade and stage at diagnosis. RESULTS: We examined data from 20928 Maryland men with incident prostate cancer reported to the Maryland Cancer Registry during 1992-1997. Initial cluster detection analyses, prior to adjustment, indicated that there were four statistically significant clusters of high and low rates of each outcome (later stage at diagnosis and higher histologic grade of tumor) for prostate cancer cases in Maryland during 1992-1997. After adjustment for individual case attributes, including age, race, year of diagnosis, patterns of clusters changed for both outcomes. Additional adjustment for Census block group and county-level socioeconomic measures changed the cluster patterns further. CONCLUSIONS: These findings provide evidence that, in locations where adjustment changed patterns of clusters, the adjustment factors may be contributing causes of the original clusters. In addition, clusters identified after adjusting for individual and area-level predictors indicate area of unexplained variation, and merit further small-area investigations.

Journal Article↗