PubMed HealthSearch

SEARCH · PubMed Health

Results for “Gaussian mixture model”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

17 recordsLinked to original sources

Model-based multifacet clustering with high-dimensional omics applications.

High-dimensional omics data often contain intricate and multifaceted information, resulting in the coexistence of multiple plausible sample partitions based on different subsets of selected features. Conventional clustering methods typically yield only one clustering solution, limiting their capacity to fully capture all facets of cluster structures in high-dimensional data. To address this challenge, we propose a model-based multifacet clustering (MFClust) method based on a mixture of Gaussian mixture models, where the former mixture achieves facet assignment for gene features and the latter mixture determines cluster assignment of samples. We demonstrate superior facet and cluster assignment accuracy of MFClust through simulation studies. The proposed method is applied to three transcriptomic applications from postmortem brain and lung disease studies. The result captures multifacet clustering structures associated with critical clinical variables and provides intriguing biological insights for further hypothesis generation and discovery.

Humans

Deconvolution of evolutionary architecture unmasks a high-risk, subclonal-rich subtype in treatment-naive small cell lung cancer.

BACKGROUND: Intratumoral heterogeneity (ITH) drives therapeutic resistance in small cell lung cancer (SCLC). However, conventional single-sample analysis has limited horizontal, cross-patient comparisons, leaving the overarching evolutionary architecture in treatment-naive tumors poorly understood. This study aims to deconvolve these architectures to identify clinically relevant evolutionary subtypes. METHODS: We analyzed whole-exome sequencing data from 41 treatment-naive SCLC patients. To overcome the cross-patient comparability bottleneck, we developed a novel probabilistic framework using a refined Gaussian Mixture Model (GMM). This standardized subclonal structures into four hierarchical strata, enabling the identification of evolutionary subtypes via unsupervised clustering. To address the scarcity of SCLC public data, prognostic concordance was robustly explored in The Cancer Genome Atlas (TCGA) lung squamous cell carcinoma (LUSC) based on shared smoking etiology, with lung adenocarcinoma (LUAD) serving as a negative control. RESULTS: The cohort robustly segregated into "Clonal-dominant" (Group 1, n=28) and "Subclonal-rich" (Group 2, n=13) subtypes. Group 1 evolution was primarily driven by tobacco signatures (SBS4). Conversely, Group 2 exhibited late-stage acquisition of a DNA mismatch repair deficiency (MMRd) signature (SBS15), fueling trace subclonal diversification. Clinically, Group 2 demonstrated a significantly lower objective response rate (ORR) to platinum-based regimens (25.0% vs. 81.3%, P=0.02). Furthermore, the Subclonal-rich architecture independently predicted inferior overall survival (OS) [adjusted hazard ratio (adj. HR) =2.93, P=0.02], driven predominantly by limited-stage disease. Cross-cancer analysis validated this histology-dependent, high-heterogeneity adverse pattern in early-stage LUSC but not in LUAD. CONCLUSIONS: This hypothesis-generating study demonstrates that a "Subclonal-rich" architecture, driven by acquired MMRd, identifies high-risk, chemo-resistant SCLC. Our GMM approach suggests that pre-existing heterogeneity may serve as a potential, histology-dependent prognostic marker that warrants prospective validation for tailoring future therapeutic regimens.

Gaussian Mixture Model (GMM)

A genome-wide assessment of the population structure of thirteen admixed and pure Australian beef cattle breeds.

Knowledge of population structure is a key factor for successful multi-breed genomic prediction, especially in single-step analysis when metafounders are considered. In Australia, current assessments mostly focus on single breeds using a single-step genomic prediction method. However, the effective integration of pedigree, phenotypic, and genomic data in a multi-breed framework still requires further research, especially for combined analyses including admixed and multi-breed populations. This study began with 602,952 genotyped individuals with 8K SNPs in common from 13 beef cattle breeds (Alexandria, Angus, Brahman, Brangus, Charolais, Droughtmaster, Hereford, Kynuna, Limousin, Santa Gertrudis, Shorthorn, Speckle Park, and Wagyu). Due to different numbers of animals being genotyped in each breed, a representative subset of animals was chosen by employing a validated sampling strategy using Gaussian Mixture Models (GMM) complemented by Principal Component Analysis (PCA) within each breed. Subsequently, a specific number of animals in each cluster were randomly selected to capture the entire genetic diversity per breed, with a total of 260 animals from each breed. The first three principal components explained 59.89% of the total variation, with PC1 (33.54%) clearly separating Bos indicus from Bos taurus lineages. Admixture analysis identified stable ancestral components and defined the genetic makeup of both pure and composite populations. The results showed extensive genetic diversity in some breeds and highlighted distinct genetic differences between Bos indicus and Bos taurus breeds. In addition, six composite breeds' admixture levels confirmed their origin and breed history, revealing a directional shift in ancestry proportions by a longitudinal increase in Brahman ancestry within tropical composites over time. Thus, the findings pave the way for more effective utilization of genetic diversity both within and across populations and provide a framework for designing multi-breed genetic evaluations and breeding programs to improve productivity and profitability in Australian beef production.

Animals

Dissecting the shared genetic architecture of schizophrenia with ventricular subregion volumes.

Schizophrenia is characterized by cerebral ventricular enlargement as an early and consistent structural anomaly. While genetic factors significantly influence both schizophrenia and cerebral ventricular enlargement, the shared genetic etiology between them requires further investigation. Using summary statistics from recent large genome-wide association studies on schizophrenia and 9 ventricular subregion volumes phenotypes. Gaussian causal mixture modeling was applied to characterize the genetic architecture and overlap between schizophrenia and ventricular subregion volumes phenotypes. Local genetic correlation was investigated with Local Analysis of Variant Association. The conjunctional false discovery rate framework was used to identify the specific shared genetic loci, annotated with FUMA. Gaussian causal mixture modeling estimated schizophrenia to be more polygenic more polygenic (9574 trait-influencing variants) than ventricular subregion volumes phenotypes (157-1267 trait-influencing variants). Conjunctional false discovery rate analysis identified 42 shared genetic loci, 17 loci were identified as novel for both schizophrenia and the ventricular subregion volumes phenotypes. Local Analysis of Variant Association revealed that 11 distinct loci demonstrated significant differences, among which 4 loci were situated in the Major Histocompatibility Complex region. Annotated genes in shared loci were enriched in molecular signaling pathways involved in inflammation and the brain structure. The shared loci between them were annotated and enriched in Major Histocompatibility Complex and inflammation-related pathways, highlighting new opportunities for future investigation.

Schizophrenia

Proteomics-driven discovery of intervention windows and risk subtypes in osteoporosis: A prospective cohort study.

Given the limited feasibility of population-wide bone mineral density screening and the infrequency of long-term monitoring in healthy individuals, identifying the window for early intervention and the populations to be prioritized for screening is critical. This study aimed to identify intervention windows for osteoporosis and to determine potential high-risk subtypes within the healthy population. Based on proteomic data from 41,408 healthy adults, we conducted the DE-SWAN method to identify change peaks in plasma protein during the pre-diagnostic osteoporosis phase, and employed finite Gaussian mixture model-based clustering to delineate high-risk subtypes of osteoporosis. We identified 122 protein biomarkers significantly associated with osteoporosis risk throughout the follow-up period. Importantly, we identified two critical peaks occurring approximately 10 and 6 years before diagnosis, with the former enriched in immune-related pathways and the latter prominently involving responses to retinoic acid and glucocorticoids. Furthermore, one high-risk subtype for osteoporosis was identified in both males and females, termed the Frailty and Obesity Subtype. This subtype is characterized by a high degree of frailty and obesity, accompanied by a significantly elevated risk of both osteoporosis and fractures. Finally, we developed a predictive model comprising 10 proteins for identifying high-risk subtypes of osteoporosis, which demonstrated better performance than the traditional risk factor model (AUC: 0.743 vs. 0.680). Our findings demonstrate that proteomic profiling can reveal early molecular changes and identify high-risk subtypes years before clinical onset, providing a foundation for screening and precision prevention of osteoporosis.

Proteomics

Segregation of noisy Mendelian traits and the effect of age-dependence: a prolegomenon.

A discussion of the primordial confusion between the multiplicative Galtonian ("lognormal") trait and the imperfectly segregating Mendelian trait is laid out from a probabilistic standpoint. Several criteria used in comparing them (bimodality; bitangentiality; goodness of fit to the multinomialized form of the distribution; cumulants of the distributions) are reviewed and the inadequacy of their probabilistic properties discussed in some detail. The logical asymmetry of the normalized score ("Roberts" correction") and hence its invalidity as a criterion for distinguishing between the models is pointed out. The form of a mixture of two Gaussian distributions with fixed and equal variances but with differing age-dependent means ("the Platt model") is explored. The epistemological implications are exhibited. As a first step to restoring symmetry, a general model is proposed of which these and other models in wide use emerge as special cases. No attempt is made to deal here with the statistical aspects of the problem.

Age Factors

Expert opinion elicitation for assisting deep learning based Lyme disease classifier with patient data.

BACKGROUND: Diagnosing erythema migrans (EM) skin lesion, the most common early symptom of Lyme disease, using deep learning techniques can be effective to prevent long-term complications. Existing works on deep learning based EM recognition only utilizes lesion image due to the lack of a dataset of Lyme disease related images with associated patient data. Doctors rely on patient information about the background of the skin lesion to confirm their diagnosis. To assist deep learning model with a probability score calculated from patient data, this study elicited opinions from fifteen expert doctors. To the best of our knowledge, this is the first expert elicitation work to calculate Lyme disease probability from patient data. METHODS: For the elicitation process, a questionnaire with questions and possible answers related to EM was prepared. Doctors provided relative weights to different answers to the questions. We converted doctors' evaluations to probability scores using Gaussian mixture based density estimation. We exploited formal concept analysis and decision tree for elicited model validation and explanation. We also proposed an algorithm for combining independent probability estimates from multiple modalities, such as merging the EM probability score from a deep learning image classifier with the elicited score from patient data. RESULTS: We successfully elicited opinions from fifteen expert doctors to create a model for obtaining EM probability scores from patient data. CONCLUSIONS: The elicited probability score and the proposed algorithm can be utilized to make image based deep learning Lyme disease pre-scanners robust. The proposed elicitation and validation process is easy for doctors to follow and can help address related medical diagnosis problems where it is challenging to collect patient data.

Humans

Lack of a bimodal distribution of ventricular size in schizophrenia: a Gaussian mixture analysis of 1056 cases and controls.

The finding of clinical and laboratory differences between schizophrenic patients with large and small cerebral ventricles has led to the widespread assumption that large ventricles are a marker that characterizes a subgroup of patients with schizophrenia. We reviewed all published English language ventricle-to-brain ratio (VBR) studies in which individual data points were available (schizophrenics: n = 691, medical controls; n = 205, normal volunteers: n = 160). Using a univariate normal mixture model to examine the distribution of ventricular size in each group, we found no evidence of a mixture of Gaussian distributions (i.e., "bimodality") within any of the three groups. The same analysis was then performed on the combined sample of schizophrenic patients and normal and medical controls, respectively. In each case the improvement in fit of a mixture of normal distributions compared to a single component normal distribution was significant. The data do not support the notion that ventricular enlargement is a discontinuous marker of a subtype of schizophrenia.

Analysis of Variance

A quantitative and qualitative description of electromyographic linear envelopes for synergy analysis.

The muscular synergy patterns of human locomotion can be described by the phasic activity of electromyographic linear envelopes (LE) and the interphasic spatio-temporal relations. To represent the phasic activity, the LE is modeled as the summation of Gaussian pulses of various lengths. The parameters of interest are the temporal features: time, duration, and amplitude of the phases of activity. A maximum likelihood approach to the parameter estimation for a mixture of normal distributions is adopted for extracting the temporal features. Based on the derived temporal features, a set of relational descriptors can be defined to describe the spatio-temporal relations between the multichannel phasic activities. The strength of this approach is not only that the phasic activity of LE can be quantitatively represented accurately, but also that the resulting synergy patterns can be easily interpreted by observers.

Algorithms

Statistical evaluation of cell kinetic data from DNA flow cytometry (FCM) by the EM algorithm.

Flow cytometric DNA measurements yield the amount of DNA for each of a large number of cells. A DNA histogram normally consists of a mixture of one or more constellations of G0/G1-, S-, G2/M-phase cells, together with internal standards, debris, background noise, and one or more populations of clumped cells. We have modelled typical DNA histograms as a mixed distribution with Gaussian densities for the G0/G1 and G2/M phases, an S-phase density, assumed to be uniform between the G0/G1 and G2/M peaks, observed with a Gaussian error, and with Gaussian densities for standards of chicken and trout red blood cells. The debris is modelled as a truncated exponential distribution, and we also have included a uniform background noise distribution over the whole observation interval. We have explored a new approach for maximum-likelihood analyses of complex DNA histograms by the application of the EM algorithm. This algorithm was used for four observed DNA histograms of varying complexity. Our results show that the algorithm works very well, and it converges to reasonable values for all parameters. In simulations from the estimated models, we have investigated bias, variance, and correlations of the estimates.

Algorithms

Investigations of the simian ontogenic switch from fetal to adult hemoglobin at the progenitor cell level.

The ontogenic switch from fetal to adult hemoglobin could result from discontinuous events, such as replacement of fetal erythroid progenitor cells by adult ones, or gradual modulation of the hemoglobin program of a single progenitor cell pool. The former would result in progenitors at midswitch with skewed fractional beta-globin synthesis programs, the latter in a Gaussian distribution. For these studies, we obtained bone marrow from rhesus monkey fetuses at 141-153 d (midswitch). Mononuclear cells were cultured in methyl cellulose with erythropoietin, and single BFU-E-derived colonies were removed and incubated with [3H]leucine. Globin synthesis was examined by gel electrophoresis and fluorography. The beta-globin synthesis pattern of single fetal colonies was skewed, and did not fit a normal distribution. The fetal pattern resembled the pattern of an artificial mixture of fetal and adult progenitors, suggesting that the fetal progenitor pool could contain populations with different beta-globin programs. This non-Gaussian distribution in the progenitors of midswitch fetuses is consistent with a discontinuous model for hemoglobin switching during ontogeny.

Animals

Decay time distribution analysis of Yt-base in benzene-methanol mixtures.

Frequency-domain fluorometry was used to measure intensity decays of synthetic Yt-base in mixtures of benzene-methanol at 20 degrees C. Multiexponential analysis shows that the decay of Yt-base fluorescence in benzene and methanol can be well fitted to a single-exponential model with tau = 9.67 ns and 6.25 ns respectively. In mixtures of benzene-methanol the decays became heterogeneous, and the maximum of heterogeneity observed was in a mixture containing 6% methanol. Since we expected a distribution of Yt-base solvation states in the solvent mixtures, and because the decay times of Yt-base are sensitive to solvent, we analyzed the data in terms of decay time distributions. The goodness-of-fit for the unimodal distribution model which has two floating parameters was equivalent to that found using the double exponential model with three floating parameters. The Lorentzian distribution model appears to provide a slightly superior fit relative to the Gaussian distribution model. These results suggest that the intensity decays of solvent-sensitive fluorophores in mixed solvents are described by a distribution of decay times.

Benzene

The molecular mobility of alpha-actinin and actin in a reconstituted model of gelation.

Dictyostelium discoideum alpha-actinin (D.d. alpha-actinin) is a calcium and pH-regulated actin-binding protein that can cross-link F-actin into a gel at a submicromolar free calcium concentration and a pH less than 7 [Fechheimer, et al., 1982]. We examined mixtures of actin and D.d. alpha-actinin at four pH and calcium concentrations that exhibited various degrees of gelation or solation. The macroscopic viscosities of these mixtures were measured by falling ball viscometry (FBV) and compared to the translational diffusion coefficients measured by gaussian spot and periodic-pattern fluorescence photobleaching recovery (FPR) of both the actin filaments and D.d. alpha-actinin. A homogeneous, macroscopic gel was not composed of a static actin network. Instead, the filament diffusion coefficient decreased to approximately 65% of the control value. If the D.d. alpha-actinin concentration was increased, the solution became inhomogeneous, consisting of domains of higher actin concentration. These domains were often composed of a static actin network. The mobility of D.d. alpha-actinin consisted of a major fraction that freely diffused and a minor fraction that appeared immobile under the conditions employed. This suggested that D.d. alpha-actinin binding to the actin filaments was static over the time course of measurement (approximately 5 sec). Under solation conditions, there was no apparent interaction of actin with D.d. alpha-actinin. These results demonstrate that 1) actin filaments need not be cross-linked into an immobile, static array in order to have macroscopic properties of a gel; 2) interpretation of the rheological properties of actin:alpha-actinin gels are complicated by spatial heterogeneity of the filament concentration and mobility; and 3) a fraction of D.d. alpha-actinin binds statically to actin in undisturbed gels. The implications of these results are discussed in relation to cytoplasmic structure and contractility.

Actinin

Proton NMR bandshape studies of lamellar liquid crystals and gel phases containing lecithins and cholesterol.

Proton NMR spectra for gel and liquid crystalline samples, composed of dimyristoyl and/or dipalmitoyl lecithin, cholesterol and water, can be consistently interpreted in terms of mesophase symmetry and molecular diffusion according to a model proposed by Wennerstrom (Wennerstrom, H. (1973) Chem. Phys. Lett. 18, 41-44). It is shown by computer simulation that the characteristic "super-lorentzian" bandshape of the lamellar mesophase can be described by the superposition of three gaussian curves. The NMR signal of the gel phase can be simulated by the superposition of two gaussian curves with widths at half height of 2.5 kHz and 19 kHz. An upper limit of the lateral diffusion coefficient of the lecithin molecules in the gel phase is calculated to be about 5-10(-15) m-2/s. It is therefore concluded that the static intermolecular dipolar couplings average to zero in the lamellar mesophase. An estimation of the order parameter of the liquid crystalline phase is made from experimental data and a calculated "rigid lattice" linewidth. A two phase system is shown to exist in the temperature range 28-34 degrees C for a mesophase of a mixture of dimyristoyl and dipalmitoyl lecithin. The presence of cholesterol results in enhanced lateral diffusion of the lecithin molecules at temperatures below the Chapman transition point.

Binding Sites

Protein dynamics. Comparative investigation on heme-proteins with different physiological roles.

We report the low temperature carbon monoxide recombination kinetics after photolysis and the temperature dependence of the visible absorption spectra of the isolated alpha SH-CO and beta SH-CO subunits from human hemoglobin A in ethylene glycol/water and in glycerol/water mixtures. Kinetic measurements on sperm whale (Physeter catodon) myoglobin and previously published optical spectroscopy data on the latter protein and on human hemoglobin A, in both solvents, (Cordone, L., A. Cupane, M. Leone, E. Vitrano, and D. Bulone. 1988. J. Mol. Biol. 199:312-218) are taken as reference. Low temperature flash photolysis data are analyzed within the multiple substates model proposed by Frauenfelder and co-workers (Austin, R. H., K. W. Beeson, L. Eisenstein, H. Frauenfelder, and I. C. Gunsalus. 1975. Biochemistry. 14:5355-5373). Within this model a distribution of activation enthalpies for ligand binding accounts for the structural heterogeneity of the protein, while the preexponential factor, containing also the entropic contribution to the free energy of the process, is considered to be constant for all conformational substates. Optical spectra are deconvoluted in gaussian components and the temperature dependence of the moments of the resulting bands is analyzed, within the harmonic Frank-Condon approximation, to obtain information on the stereodynamic properties of the heme pocket. The kinetic and spectral parameters thus obtained are found to be protein dependent also with respect to their sensitivity to changes in the composition of the external medium. A close correlation between the kinetic and spectral features is observed for the proteins examined under all experimental conditions studied. The results reported are discussed in terms of differences in the heme pocket structure and in the conformational heterogeneity among the various proteins, as related to their different capability to accommodate constraints imposed by the external medium.

Animals

Fluorescence lifetime distributions of diphenylhexatriene-labeled phosphatidylcholine as a tool for the study of phospholipid-cholesterol interactions.

Fluorescence lifetimes of 1-palmitoyl-2-diphenylhexatrienylpro-pionyl-phosphatidylc hol ine in vesicles of palmitoyloleoyl phosphatidylcholine (POPC) (1:300, mol/mol) in the liquid crystalline state were determined by multifrequency phase fluorometry. On the basis of statistic criteria (chi 2red) the measured phase angles and demodulation factors were equally well fitted to unimodal Lorentzian, Gaussian, or uniform lifetime distributions. No improvement in chi 2red could be observed if the experimental data were fitted to bimodal Lorentzian distributions or a double exponential decay. The unimodal Lorentzian lifetime distribution was characterized by a lifetime center of 6.87 ns and a full width at half maximum of 0.57 ns. Increasing amounts of cholesterol in the phospholipid vesicles (0-50 mol% relative to POPC) led to a slight increase of the lifetime center (7.58 ns at 50 mol% sterol) and reduced significantly the distributional width (0.14 ns at 50 mol% sterol). Lifetime distributions of POPC-cholesterol mixtures containing greater than 20 mol% sterol were within the resolution limit and could not be distinguished from monoexponential decays on the basis of chi 2red. Cholesterol stabilizes and rigidifies phospholipid bilayers in the fluid state. Considering its effect on lifetime distributions of fluorescent phospholipids it may also act as a membrane homogenizer.

Cholesterol

Determination of fluid and gel domain sizes in two-component, two-phase lipid bilayers. An electron spin resonance spin label study.

The average sizes of fluid and gel domains in the two-component, two-phase system formed from mixtures of dimyristoyl phosphatidylcholine and distearoyl phosphatidylcholine were determined from an analysis of the electron spin resonance spectral lineshapes of a dimyristoyl phosphatidylcholine-nitroxide spin label as a function of spin label concentration. The ratio, R, of the intensities measured at two magnetic field strengths was found to be diagnostic of a statistical distribution of spin labels in disconnected domains. R is defined as V'/2Vpp, where Vpp is the maximum intensity and V' is the intensity at a position in the wings of a first derivative electron spin resonance line that is a constant multiple of the peak-to-peak linewidth. The intensity ratio for Gaussian or Voigt lineshapes is less than or equal to the value for a Lorentzian lineshape. The intensity ratio was found to be greater than the value for a Lorentzian line when spectra from disconnected domains containing a statistical distribution of spin labels undergoing spin-spin interactions were summed. The intensity ratio, R, calculated by spectral simulations as a function of the average number of labels per domain, N, was found to increase to a maximum with increasing N and then to decrease. The dependence on spin label concentration of the experimentally measured intensity ratios paralleled this predicted behavior. A method is presented to calculate the average number of lipids per fluid or gel domain based on a knowledge of R, and of the distribution of the spin label between the fluid and gel phases determined from the phase diagram. The results demonstrate that the number of lipids per domain increases linearly from a fixed number of nucleation sites, as the fraction of the phase that is disconnected increases. At any given mole fraction of the particular phase, the gel domains are bigger than the fluid domains because they have a lower nucleation density. The results also suggest that the disconnected domains are, in most cases, nonrandomly distributed in the plane of the bilayer.

Dimyristoylphosphatidylcholine