PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Likelihood Functions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Gene family evolution and homology: genomics meets phylogenetics.

With the advent of high-throughput DNA sequencing and whole-genome analysis, it has become clear that the coding portions of the genome are organized hierarchically in gene families and superfamilies. Because the hierarchy of genes, like that of living organisms, reflects an ancient and continuing process of gene duplication and divergence, many of the conceptual and analytical tools used in phylogenetic systematics can and should be used in comparative genomics. Phylogenetic principles and techniques for assessing homology, inferring relationships among genes, and reconstructing evolutionary events provide a powerful way to interpret the ever increasing body of sequence data. In this review, we outline the application of phylogenetic approaches to comparative genomics, beginning with the inference of phylogeny and the assessment of gene orthology and paralogy. We also show how the phylogenetic approach makes possible novel kinds of comparative analysis, including detection of domain shuffling and lateral gene transfer, reconstruction of the evolutionary diversification of gene families, tracing of evolutionary change in protein function at the amino acid level, and prediction of structure-function relationships. A marriage of the principles of phylogenetic systematics with the copious data generated by genomics promises unprecedented insights into the nature of biological organization and the historical processes that created it.

Animals↗

Detecting overlapping coding sequences in virus genomes.

BACKGROUND: Detecting new coding sequences (CDSs) in viral genomes can be difficult for several reasons. The typically compact genomes often contain a number of overlapping coding and non-coding functional elements, which can result in unusual patterns of codon usage; conservation between related sequences can be difficult to interpret--especially within overlapping genes; and viruses often employ non-canonical translational mechanisms--e.g. frameshifting, stop codon read-through, leaky-scanning and internal ribosome entry sites--which can conceal potentially coding open reading frames (ORFs). RESULTS: In a previous paper we introduced a new statistic--MLOGD (Maximum Likelihood Overlapping Gene Detector)--for detecting and analysing overlapping CDSs. Here we present (a) an improved MLOGD statistic, (b) a greatly extended suite of software using MLOGD, (c) a database of results for 640 virus sequence alignments, and (d) a web-interface to the software and database. Tests show that, from an alignment with just 20 mutations, MLOGD can discriminate non-overlapping CDSs from non-coding ORFs with a typical accuracy of up to 98%, and can detect CDSs overlapping known CDSs with a typical accuracy of 90%. In addition, the software produces a variety of statistics and graphics, useful for analysing an input multiple sequence alignment. CONCLUSION: MLOGD is an easy-to-use tool for virus genome annotation, detecting new CDSs--in particular overlapping or short CDSs--and for analysing overlapping CDSs following frameshift sites. The software, web-server, database and supplementary material are available at http://guinevere.otago.ac.nz/mlogd.html.

Algorithms↗

Evaluation of point-of-care tests for diagnosis of disseminated intravascular coagulation in dogs admitted to an intensive care unit.

OBJECTIVE: To evaluate the accuracy of point-of-care tests for the diagnosis of disseminated intravascular coagulation (DIC) in dogs and assess the correlation and agreement of results between point-of-care and laboratory tests in the evaluation of hemostatic function. DESIGN: Prospective case series. ANIMALS: 59 critically ill dogs (affected dogs) with clinical signs of diseases known to predispose to DIC and 52 clinically normal dogs. PROCEDURES: Accuracy of the point-of-care tests (activated clotting time [ACT], estimated platelet count and number of schizocytes from a blood smear, plasma total solids [TS] concentration, and the protamine sulfate test) was evaluated, using receiver operating characteristic curves and likelihood ratios. A strategy, using likelihood ratios to calculate a posttest probability of DIC, was tested with 65% used as a threshold for initiation of treatment. Results of laboratory tests (coagulogram and plasma antithrombin III activity) were used as the standard for comparison in each dog. RESULTS: ACT and estimated platelet count provided the best accuracy for detection of DIC. The plasma TS concentration, schizocyte number, and protamine sulfate test had poor accuracy. The strategy using post-test probability of DIC identified 12 of 16 affected dogs that had DIC. Estimated platelet count was correlated and had acceptable clinical agreement with automated platelet count (r = 0.70). The plasma TS (r = 0.28) concentration and serum albumin (r = 0.63) concentration were not accurate predictors of plasma antithrombin III activity. The ACT did not correlate with activated partial thromboplastin time (r = 0.28). CONCLUSIONS AND CLINICAL RELEVANCE: Strategic use of likelihood ratios from point-of-care tests can assist clinicians in making treatment decisions for dogs suspected to have DIC when immediate laboratory support is unavailable.

Animals↗

Fractional Gaussian noise, functional MRI and Alzheimer's disease.

Fractional Gaussian noise (fGn) provides a parsimonious model for stationary increments of a self-similar process parameterised by the Hurst exponent, H, and variance, sigma2. Fractional Gaussian noise with H < 0.5 demonstrates negatively autocorrelated or antipersistent behaviour; fGn with H > 0.5 demonstrates 1/f, long memory or persistent behaviour; and the special case of fGn with H = 0.5 corresponds to classical Gaussian white noise. We comparatively evaluate four possible estimators of fGn parameters, one method implemented in the time domain and three in the wavelet domain. We show that a wavelet-based maximum likelihood (ML) estimator yields the most efficient estimates of H and sigma2 in simulated fGn with 0 < H < 1. Applying this estimator to fMRI data acquired in the "resting" state from healthy young and older volunteers, we show empirically that fGn provides an accommodating model for diverse species of fMRI noise, assuming adequate preprocessing to correct effects of head movement, and that voxels with H > 0.5 tend to be concentrated in cortex whereas voxels with H < 0.5 are more frequently located in ventricles and sulcal CSF. The wavelet-ML estimator can be generalised to estimate the parameter vector beta for general linear modelling (GLM) of a physiological response to experimental stimulation and we demonstrate nominal type I error control in multiple testing of beta, divided by its standard error, in simulated and biological data under the null hypothesis beta = 0. We illustrate these methods principally by showing that there are significant differences between patients with early Alzheimer's disease (AD) and age-matched comparison subjects in the persistence of fGn in the medial and lateral temporal lobes, insula, dorsal cingulate/medial premotor cortex, and left pre- and postcentral gyrus: patients with AD had greater persistence of resting fMRI noise (larger H) in these regions. Comparable abnormalities in the AD patients were also identified by a permutation test of local differences in the first-order autoregression AR(1) coefficient, which was significantly more positive in patients. However, we found that the Hurst exponent provided a more sensitive metric than the AR(1) coefficient to detect these differences, perhaps because neurophysiological changes in early AD are naturally better described in terms of abnormal salience of long memory dynamics than a change in the strength of association between immediately consecutive time points. We conclude that parsimonious mapping of fMRI noise properties in terms of fGn parameters efficiently estimated in the wavelet domain is feasible and can enhance insight into the pathophysiology of Alzheimer's disease.

Aged↗

On studentising and blocklength selection for the bootstrap on time series.

For independent data, non-parametric bootstrap is realised by resampling the data with replacement. This approach fails for dependent data such as time series. If the data generating process is at least stationary and mixing, the blockwise bootstrap by drawing subsamples or blocks of the data saves the concept. For the blockwise bootstrap a blocklength has to be selected. We propose a method for selecting the optimal blocklength. To improve the finite size properties of the blockwise bootstrap, studentised statistics is considered. If the statistic can be represented as a smooth function model this studentisation can be approximated efficiently. The studentised blockwise bootstrap method is applied for testing hypotheses on medical time series.

Algorithms↗

Modeling error variance in job specification ratings: the influence of rater, job, and organization-level factors.

The authors modeled sources of error variance in job specification ratings collected from 3 levels of raters across 5 organizations (N=381). Variance components models were used to estimate the variance in ratings attributable to true score (variance between knowledge, skills, abilities, and other characteristics [KSAOs]) and error (KSAO-by-rater and residual variance). Subsequent models partitioned error variance into components related to the organization, position level, and demographic characteristics of the raters. Analyses revealed that the differential ordering of KSAOs by raters was not a function of these characteristics but rather was due to unexplained rating differences among the raters. The implications of these results for job specification and validity transportability are discussed.

Adult↗

Replication of linkage studies of complex traits: an examination of variation in location estimates.

In linkage studies, independent replication of positive findings is crucial in order to distinguish between true positives and false positives. Recently, the following question has arisen in linkage studies of complex traits: at what distance do we reject the hypothesis that two location estimates in a genomic region represent the same gene? Here we attempt to address this question. Sampling distributions for location estimates were constructed by computer simulation. The conditions for simulation were chosen to reflect features of "typical" complex traits, including incomplete penetrance, phenocopies, and genetic heterogeneity. Our findings, which bear on what is considered a replication in linkage studies of complex traits, suggest that, even with relatively large numbers of multiplex families, chance variation in the location estimate is substantial. In addition, we report evidence that, for the conditions studied here, the standard error of a location estimate is a function of the magnitude of the expected LOD score.

Chromosome Mapping↗

Gene conversion drives the evolution of HINTW, an ampliconic gene on the female-specific avian W chromosome.

The HINTW gene on the female-specific W chromosome of chicken and other birds is amplified and present in numerous copies. Moreover, as HINTW is distinctly different from its homolog on the Z chromosome (HINTZ), is a candidate gene in avian sex determination, and evolves rapidly under positive selection, it shows several common features to ampliconic and testis-specific genes on the mammalian Y chromosome. A phylogenetic analysis within galliform birds (chicken, turkey, quail, and pheasant) shows that individual HINTW copies within each species are more similar to each other than to gene copies of related species. Such convergent evolution is most easily explained by recurrent events of gene conversion, the rate of which we estimated at 10(-6)-10(-5) per site and generation. A significantly higher GC content of HINTW than of other W-linked genes is consistent with biased gene conversion increasing the fixation probability of mutations involving G and C nucleotides. Furthermore, and as a likely consequence, the neutral substitution rate is almost twice as high in HINTW as in other W-linked genes. The region on W encompassing the HINTW gene cluster is not covered in the initial assembly of the chicken genome, but analysis of raw sequence reads indicates that gene copy number is significantly higher than a previous estimate of 40. While sexual selection is one of several factors that potentially affect the evolution of ampliconic, male-specific genes on the mammalian Y chromosome, data from HINTW provide evidence that gene amplification followed by gene conversion can evolve in female-specific chromosomes in the absence of sexual selection. The presence of multiple and highly similar copies of HINTW may be related to protein function, but, more generally, amplification and conversion offers a means to the avoidance of accumulation of deleterious mutations in nonrecombining chromosomes.

Animals↗

Multilevel generalized linear models for modelling age-related gender difference in violent behaviour and associated factors in the general household population.

It is preferable to use longitudinal data when studying patterns of violence and antisocial behaviour over the lifespan together with the associated risk factors in the general population. From the statistical modelling perspective, random samples of cross-sectional data, representative of the population, can be a reliable alternative. Sampling, weighting, and possible geographical clustering of the behaviour must be considered in the analysis together with correct choice of model as a function of age, although cohort effects and age effects are not separated from the analysis. This paper demonstrates the use of multilevel generalized linear models in the British National Survey of Psychiatric Morbidity in 2000. A multilevel logistic model as a special case of a generalized linear model with individual weightings was adapted for a dichotomous measure of violence and extended to Poisson and negative binomial outcomes. Three types of age function, discrete age effects, continuous age effects, and piecewise polynomial function of age intervals were evaluated for goodness of fit, and for their practical advantages and disadvantages. Models were developed for possible risk factors in relation to specific age groups of interest.

Adolescent↗

Evolution of base-substitution gradients in primate mitochondrial genomes.

Inferences of phylogenies and dates of divergence rely on accurate modeling of evolutionary processes; they may be confounded by variation in substitution rates among sites and changes in evolutionary processes over time. In vertebrate mitochondrial genomes, substitution rates are affected by a gradient along the genome of the time spent being single-stranded during replication, and different types of substitutions respond differently to this gradient. The gradient is controlled by biological factors including the rate of replication and functionality of repair mechanisms; little is known, however, about the consistency of the gradient over evolutionary time, or about how evolution of this gradient might affect phylogenetic analysis. Here, we evaluate the evolution of response to this gradient in complete primate mitochondrial genomes, focusing particularly on A-->G substitutions, which increase linearly with the gradient. We developed a methodology to evaluate the posterior probability densities of the response parameter space, and used likelihood ratio tests and mixture models with different numbers of classes to determine whether groups of genomes have evolved in a similar fashion. Substitution gradients usually evolve slowly in primates, but there have been at least two large evolutionary jumps: on the lineage leading to the great apes, and a convergent change on the lineage leading to baboons (Papio). There have also been possible convergences at deeper taxonomic levels, and different types of substitutions appear to evolve independently. The placements of the tarsier and the tree shrew within and in relation to primates may be incorrect because of convergence in these factors.

Animals↗

Blind separation of positive sources by globally convergent gradient search.

The instantaneous noise-free linear mixing model in independent component analysis is largely a solved problem under the usual assumption of independent nongaussian sources and full column rank mixing matrix. However, with some prior information on the sources, like positivity, new analysis and perhaps simplified solution methods may yet become possible. In this letter, we consider the task of independent component analysis when the independent sources are known to be nonnegative and well grounded, which means that they have a nonzero pdf in the region of zero. It can be shown that in this case, the solution method is basically very simple: an orthogonal rotation of the whitened observation vector into nonnegative outputs will give a positive permutation of the original sources. We propose a cost function whose minimum coincides with nonnegativity and derive the gradient algorithm under the whitening constraint, under which the separating matrix is orthogonal. We further prove that in the Stiefel manifold of orthogonal matrices, the cost function is a Lyapunov function for the matrix gradient flow, implying global convergence. Thus, this algorithm is guaranteed to find the nonnegative well-grounded independent sources. The analysis is complemented by a numerical simulation, which illustrates the algorithm.

Algorithms↗

Can Zipf's law be adapted to normalize microarrays?

BACKGROUND: Normalization is the process of removing non-biological sources of variation between array experiments. Recent investigations of data in gene expression databases for varying organisms and tissues have shown that the majority of expressed genes exhibit a power-law distribution with an exponent close to -1 (i.e. obey Zipf's law). Based on the observation that our single channel and two channel microarray data sets also followed a power-law distribution, we were motivated to develop a normalization method based on this law, and examine how it compares with existing published techniques. A computationally simple and intuitively appealing technique based on this observation is presented. RESULTS: Using pairwise comparisons using MA plots (log ratio vs. log intensity), we compared this novel method to previously published normalization techniques, namely global normalization to the mean, the quantile method, and a variation on the loess normalization method designed specifically for boutique microarrays. Results indicated that, for single channel microarrays, the quantile method was superior with regard to eliminating intensity-dependent effects (banana curves), but Zipf's law normalization does minimize this effect by rotating the data distribution such that the maximal number of data points lie on the zero of the log ratio axis. For two channel boutique microarrays, the Zipf's law normalizations performed as well as, or better than existing techniques. CONCLUSION: Zipf's law normalization is a useful tool where the Quantile method cannot be applied, as is the case with microarrays containing functionally specific gene sets (boutique arrays).

Algorithms↗

The uncertainty associated with the predictive value of test results.

The uncertainty associated with predictive value of test results was taken into consideration, as concerns both sampling error (related to the size of the statistical reference samples) and analytical imprecision (unavoidably involved by the measurement itself). A software package, developed for the statistical calculations, was used for the treatment of the results obtained for serum free thyroxin in euthyroid and dysthyroid subjects, assumed as an experimental model. Examples are shown for the obtainable functions predictive value vs. estimate and the related uncertainty regions. These data could help in comparing test results and, in particular, in preparing fully informative laboratory reports. The difficulties involved by an extensive application of the procedure are discussed.

Humans↗

Explosive lineage-specific expansion of the orphan nuclear receptor HNF4 in nematodes.

The nuclear receptor superfamily expanded in at least two episodes: one early in metazoan evolution, the second within the vertebrate lineage. An exception to this pattern is the genome of the nematode Caenorhabditis elegans, which encodes more than 270 nuclear receptors, most of them highly divergent. We generated 128 cDNA sequences for 76 C. elegans nuclear receptors, confirming that these are active genes. Among these numerous receptors are 13 orthologues of nuclear receptors found in arthropods and/or vertebrates. We show that the supplementary nuclear receptors (supnrs) originated from an explosive burst of duplications of a unique orphan receptor, HNF4. This origin has specific implications for the role of ligand binding in the function and evolution of the nematode supplementary nuclear receptors. Moreover, the supplementary nuclear receptors include a group of very rapidly evolving genes found primarily on chromosome V. We propose a model of lineage-specific duplications from a chromosome on which duplication and substitution rates are highly increased. Our results provide a framework to study nuclear receptors in nematodes, as well as to consider the functional and evolutionary consequences of lineage-specific duplications.

Animals↗

Estimating multiple temporal mechanisms in human vision.

When studying human ability to perceive temporal changes in luminance it is customary to estimate either temporal impulse response shapes or temporal modulation transfer functions, the representation of the impulse response in the frequency domain. The advantages and limitations of previous methods are summarized. We then describe an approach based on use of an impulse response basis set that resolves some of those limitations. We next present psychophysical results for spatiotemporal signal detection in spatiotemporal noise, together with an economical model of performance. The model is based on accepted notions of psychophysical detection mechanisms and the filter basis set described in the first part of the paper. The best-fitting model requires only eight parameters, as opposed to the 198 parameters required to separately fit each psychometric function, and captures both qualitative and quantitative properties of the psychophysical data. Finally, the best-fitting model indicates that only two temporal filters are necessary to describe the performance of each of three subjects under the specific stimulus conditions employed here.

Humans↗

Multi-way clustering of microarray data using probabilistic sparse matrix factorization.

MOTIVATION: We address the problem of multi-way clustering of microarray data using a generative model. Our algorithm, probabilistic sparse matrix factorization (PSMF), is a probabilistic extension of a previous hard-decision algorithm for this problem. PSMF allows for varying levels of sensor noise in the data, uncertainty in the hidden prototypes used to explain the data and uncertainty as to the prototypes selected to explain each data vector. RESULTS: We present experimental results demonstrating that our method can better recover functionally-relevant clusterings in mRNA expression data than standard clustering techniques, including hierarchical agglomerative clustering, and we show that by computing probabilities instead of point estimates, our method avoids converging to poor solutions.

Algorithms↗

Using gene-history and expression analyses to assess the involvement of LGI genes in human disorders.

Mutations in the leucine-rich, glioma-inactivated 1 gene, LGI1, cause autosomal-dominant lateral temporal lobe epilepsy via unknown mechanisms. LGI1 belongs to a subfamily of leucine-rich repeat genes comprising four members (LGI1-LGI4) in mammals. In this study, both comparative developmental as well as molecular evolutionary methods were applied to investigate the evolution of the LGI gene family and, subsequently, of the functional importance of its different gene members. Our phylogenetic studies suggest that LGI genes evolved early in the vertebrate lineage. Genetic and expression analyses of all five zebrafish lgi genes revealed duplications of lgi1 and lgi2, each resulting in two paralogous gene copies with mostly nonoverlapping expression patterns. Furthermore, all vertebrate LGI1 orthologs experience high levels of purifying selection that argue for an essential role of this gene in neural development or function. The approach of combining expression and selection data used here exemplarily demonstrates that in poorly characterized gene families a framework of evolutionary and expression analyses can identify those genes that are functionally most important and are therefore prime candidates for human disorders.

Animals↗

Local phylogenetic divergence and global evolutionary convergence of skull function in reef fishes of the family Labridae.

The Labridae is one of the most structurally and functionally diversified fish families on coral and rocky reefs around the world, providing a compelling system for examination of evolutionary patterns of functional change. Labrid fishes have evolved a diverse array of skull forms for feeding on prey ranging from molluscs, crustaceans, plankton, detritus, algae, coral and other fishes. The species richness and diversity of feeding ecology in the Labridae make this group a marine analogue to the cichlid fishes. Despite the importance of labrids to coastal reef ecology, we lack evolutionary analysis of feeding biomechanics among labrids. Here, we combine a molecular phylogeny of the Labridae with the biomechanics of skull function to reveal a broad pattern of repeated convergence in labrid feeding systems. Mechanically fast jaw systems have evolved independently at least 14 times from ancestors with forceful jaws. A repeated phylogenetic pattern of functional divergence in local regions of the labrid tree produces an emergent family-wide pattern of global convergence in jaw function. Divergence of close relatives, convergence among higher clades and several unusual 'breakthroughs' in skull function characterize the evolution of functional complexity in one of the most diverse groups of reef fishes.

Animals↗