Determinants of socioeconomic success: regression and latent variables analysis in samples of twins.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
High-dimensional genomics studies are frequently confounded by unmeasured biological processes that obscure disease-specific signals. While existing workflows can estimate these latent confounders, they fail to quantify how robust a discovery is to varying levels of hypothetical confounding. We introduce sensGAN, a deep-learning adversarial framework that systematically explores the confounding spectrum by learning "worst-case" latent variables that nullify the most gene associations under novel predictive-gain constraints. By identifying the minimum confounding strength required to explain away an observed effect, our method shifts the paradigm toward a formal, quantitative sensitivity analysis. In diverse simulations, sensGAN accurately recovers latent structures and outperforms existing methods in identifying confounder-sensitive genes. Applied to human Alzheimer's disease microglia, our framework prioritizes robust disease pathways while successfully isolating signals driven by unmeasured co-occurring neurodegenerative pathologies. Our method is publicly available, deposited at the GitHub repository yifanlinz/ADsensitivityICML.
People who inject drugs (PWID) are disproportionately impacted by viral hepatitis C (HCV). Among PWID, homelessness, poverty, and material insecurity elevate infectious disease risk. We hypothesized a positive association between material hardship (difficulty accessing basic needs) and lifetime diagnosis of HCV, and a negative association between cumulative risk (material hardship and years since first injection) and HCV treatment. From 2021-22, we conducted a survey among community-recruited PWID that included items on HCV outcomes, material hardship (a sum score of usually (4) to never (1) having difficulty finding food, clothing, shelter, restrooms, and showers in the past 3 months), and years since injection drug use initiation. We developed a latent variable, cumulative risk, by combining individual scores of material hardship indicators and years since first injection drug use. Among our sample of PWID (n = 471), 246 (53%) participants reported lifetime HCV diagnosis and 27% of those who tested positive for HCV reported ever or current treatment (n = 67). In modified-Poisson regression, a one-unit increase in material hardship score was associated with a 3% (OR: 1.03, 95% CI: 1.01%, 1.05%) increase in odds of HCV diagnosis. In latent modeling, among PWID testing positive for HCV, a one-unit increase in cumulative risk score was associated with a 0.28 (OR: -0.28, 95% CI: -0.50, -0.06), decrease in the Z-score odds of receiving HCV treatment. Findings emphasize structural interventions to strengthen material security and co-delivering basic needs with health services to improve HCV-related outcomes among PWID.
MOTIVATION: Studying protein isoforms is an essential step in biomedical research; at present, the main approach for analyzing proteins is via bottom-up mass spectrometry proteomics, which return peptide identifications, that are indirectly used to infer the presence of protein isoforms. However, the detection and quantification processes are noisy; in particular, peptides may be erroneously detected, and most peptides, known as shared peptides, are associated to multiple protein isoforms. As a consequence, studying individual protein isoforms is challenging, and inferred protein results are often abstracted to the gene-level or to groups of protein isoforms. RESULTS: Here, we introduce IsoBayes, a novel statistical method to perform inference at the isoform level. Our method enhances the information available, by integrating mass spectrometry proteomics and transcriptomics data in a Bayesian probabilistic framework. To account for the uncertainty in the measurement process, we propose a two-layer latent variable approach: first, we sample if a peptide has been correctly detected (or, alternatively filter peptides); second, we allocate the abundance of such selected peptides across the protein(s) they are compatible with. This enables us, starting from peptide-level data, to recover protein-level data; in particular, we: (i) infer the presence/absence of each protein isoform (via a posterior probability), (ii) estimate its abundance (and credible interval), and (iii) target isoforms where transcript and protein relative abundances significantly differ. We benchmarked our approach in simulations, and in two multi-protease real datasets: our method displays good sensitivity and specificity when detecting protein isoforms, its estimated abundances highly correlate with the ground truth, and can detect changes between protein and transcript relative abundances. AVAILABILITY AND IMPLEMENTATION: IsoBayes is freely distributed as a Bioconductor R package, and is accompanied by an example usage vignette.
The growth characteristics of alpha3a bacteriophage on stationary phase Achromobacter strain 14 are described. Phage alpha3a growth on stationary phase cells is characterized by a long and variable latent period of 6 to 9 h and an increased burst size of 710 p.f.u./cell as compared with 153 p.f.u./cell in exponential wild type cells. During the latent period the infected cells are very sensitive to changes in growth conditions and in particular, dilution. Pre-conditioning of the bacterial cells by allowing them to stand for 24 h after shaking for 3 days is an important aspect of the stationary phase phage growth system. Cells which have been allowed to stand retain the ability to be infected and to support phage growth for at least 16 days. Shaking cultures gradually lose the ability to support phage growth but the phage can persist in the host cell for 10 days until removal from shaking when the lytic cycle can proceed after allowing the cultures to stand.
IMPORTANCE: Brain health may facilitate resilience to detrimental consequences from neurological diseases. Infarct volume is associated with poor functional outcome after acute ischemic stroke (AIS), but potential mediating effects through stroke-related brain health loss have not been investigated. OBJECTIVE: To determine whether stroke-related brain health loss, quantified by change in MRI derived effective Reserve (eR), mediates the effect of acute infarct volume on functional outcome after AIS. DESIGN: Observational multicenter cohort study. SETTING: We analyzed data from the GASROS (n=488) and MRI-GENIE (n=560) cohorts, collected 2003-2011. PARTICIPANTS: Adult patients consecutively diagnosed with AIS, with available admission MRI. EXPOSURE: At admission, white matter hyperintensity (WMH) and normal-appearing brain volumes were assessed on T2-FLAIR, and acute infarct volume on diffusion weighted imaging. WMH was normalized by brain volume, creating WMH load. We quantified brain health using eR, a latent variable incorporating age, WMH load, and normal-appearing brain volume. ΔeR reflected the change in eR when acute infarct volume was included, representing stroke-related brain health decline. Mediation analysis was used to determine if ΔeR mediates the effect of infarct volume on functional outcome (modified Rankin Scale [mRS] at 90 days). MAIN OUTCOME MEASURE: Proportion of mediating effect. RESULTS: We included 1,048 patients (median age 67y, 38% females). At baseline, median NIHSS score was 3 (IQR 1-7), median infarct volume 3.1mL (IQR 0.9-15.5). At 90 days, median mRS score was 1 (IQR 1-3) and 51 (5%) patients had died. In mediation analysis, ΔeR significantly mediated 36% (95% CI 16-56%) of the total effect of infarct volume on functional outcome (direct effect (ß=0.15 [95% CI 0.09-0.22], p<0.001; indirect effect mediated through ΔeR: ß=0.09 [95% CI 0.04 to 0.14], p=0.001). In subgroup-analyses, the mediative effect was apparent among female but not male, and among patients aged >67y but not ≤67y. CONCLUSIONS AND RELEVANCE: Stroke-related structural brain health loss mediates about one third of the effect of acute infarct volume on functional outcome after ischemic stroke, with important sex and age differences. Brain health significantly influences outcome and recovery potential, and may be considered a key biomarker when modeling outcome after AIS.
Physical neglect is associated with poorer cognitive functioning in patients with schizophrenia, including deficits in facial emotion recognition. Recent research by our group showed the relationship between physical neglect and facial emotion recognition is mediated by inflammation. While this mediation effect was observed at the level of behaviour, examining the relationship between physical neglect, inflammation and neural activation during facial emotion recognition would help confirm brain regions impacted and support behavioural findings, but these relationships are unclear. Two hundred twelve participants (52 patients with schizophrenia and 160 healthy controls) underwent functional magnetic resonance imaging while performing an established facial emotion recognition task and a subset completed the Childhood Trauma Questionnaire and provided blood samples outside of the scanner. Inflammation was measured using a latent variable that combined basal plasma levels of interleukin-6, tumour necrosis factor-alpha and C-reactive protein. The relationships between physical neglect, inflammation and neural activation were examined using multiple regression. Neither physical neglect nor inflammation predicted altered neural response during facial emotion recognition at p < 0.05, family-wise error corrected for multiple comparisons within an anterior cingulate region of interest, across the whole brain, in the whole sample or separately in patients or controls. Future research should examine relationships between physical neglect, inflammation and brain activation in larger samples, which may have sensitivity to detect smaller effects, and use tasks that require active recognition of emotions, in addition to passive viewing of faces, which might be associated with additional neural responses.
Linear models, including those used for differential abundance analyses, are frequently used in microbiome research to assess how experimental conditions (e.g., disease state or age) affect microbial abundance. Linear mixed-effects models (MEMs) extend linear models to accommodate complex designs, such as longitudinal sampling or hierarchical study structures. However, when applied to microbiome data, existing MEM approaches suffer from high false positive and false negative rates because sequence counts are compositional - they reflect relative rather than absolute abundances. Current methods attempt to overcome this limitation through normalization, but these approaches rely on strong, often unrealistic assumptions about the unmeasured biological scale (e.g., total microbial load). Here we introduce scale-reliant mixed-effects models (SR-MEM), which extend our earlier scale-reliant inference framework by explicitly modeling uncertainty in the unmeasured scale via user-defined probability distributions. By treating scale as a latent variable rather than fixing it through normalization, SR-MEM enables robust inference for complex experimental designs. SR-MEM can incorporate external scale measurements (e.g., flow cytometry, qPCR) or leverage scale information from independent studies to further improve inference. Across simulations and multiple real-world case studies, SR-MEM consistently controls the false discovery rate while maintaining comparable or higher power than standard approaches relying on normalization or bias correction. In reanalyses of published datasets, SR-MEM yields results that are more reproducible across studies and more consistent with known biological and pharmacological effects. SR-MEM provides a principled and practical framework for mixed-effects modeling of microbiome sequence count data in the presence of unmeasured biological scale. By avoiding normalization-based assumptions and instead propagating scale uncertainty through inference, SR-MEM improves error control and reproducibility in longitudinal and hierarchical studies. An accessible implementation is provided in the ALDEx3 R package.
Single-cell transcriptomics experiments provide gene expression snapshots of heterogeneous cell populations across cell states. These snapshots have been used to infer trajectories and dynamic information even without intensive, time-series data by ordering cells according to gene expression similarity. However, while single-cell snapshots sometimes offer valuable insights into dynamic processes, current methods for ordering cells are limited by descriptive notions of "pseudotime" that lack intrinsic physical meaning. Instead of pseudotime, we propose inference of "process time" via a principled modeling approach to formulating trajectories and inferring latent variables corresponding to timing of cells subject to a biophysical process. Our implementation of this approach, called Chronocell, provides a biophysical formulation of trajectories built on cell state transitions. The Chronocell model is identifiable, making parameter inference meaningful. Furthermore, Chronocell can interpolate between trajectory inference, when cell states lie on a continuum, and clustering, when cells cluster into discrete states. By using a variety of datasets ranging from cluster-like to continuous, we show that Chronocell enables us to assess the suitability of datasets and reveals distinct cellular distributions along process time that are consistent with biological process times. We also compare our parameter estimates of degradation rates to those derived from metabolic labeling datasets, thereby showcasing the biophysical utility of Chronocell. Nevertheless, based on performance characterization on simulations, we find that process time inference can be challenging, highlighting the importance of dataset quality and careful model assessment.
Their frequency is estimated with difficulty, although on autopsy pulmonary edema is found almost routinely. It is a major complication of overdoses (48 p. 100 of severe intoxications). Their formation can be suspected, when after the first phase of respiratory depressions, with coma, myosis, and a variable latent period, a second attack of respiratory insufficiency occurs with tachypnea, and cyanosis. The chest X-ray shows diffuse alveolar infiltration, sparing the apices. The heart being generally of normal size. Rapid disappearance of this infiltrate (24 to 48 hours) enables the elimination of two diagnoses: pneumonia due to inhalation of gastric fluid, an infectious pneumonia. Their pathogenesis remains very debatable: - in the majority of cases abrupt L.V.F. can be eliminated: -on the other hand it could be an allergic accident of the anaphylactic type, or local liberation of histamine, or a local toxic action on the pulmonary capillaries; - hypoxia, secondary to respiratory depression, could lead to pulmonary edema, by the same mechanism as at altitude; - finally, owing to the central neurological disorders a neurogenic theory can be put forward. Their treatment is essentially a combination of Nalorphine with oxygen therapy (by mask, or if necessary by assisted, controlled ventilation) with prevention of inhalation of gastric fluid (gastric emptying) or curative treatment of possible aspiration by antibiotics, and cortico-steroids. Diuretics can be useful, as well as cardiotonics.
This paper tackles the challenge of estimating correlations between higher-level biological variables (e.g. proteins and gene pathways) when only lower-level measurements are directly observed (e.g. peptides and individual genes). Existing methods typically aggregate lower-level data into higher-level variables and then estimate correlations based on the aggregated data. However, different data aggregation methods can yield varying correlation estimates as they target different higher-level quantities. Our solution is a latent factor model that directly estimates these higher-level correlations from lower-level data without the need for data aggregation. We further introduce a shrinkage estimator to ensure the positive definiteness and improve the accuracy of the estimated correlation matrix. Furthermore, we establish the asymptotic normality of our estimator, enabling efficient computation of P-values for the identification of significant correlations. The effectiveness of our approach is demonstrated through comprehensive simulations and the analysis of proteomics and gene expression datasets. We develop the R package highcor for implementing our method.
Harmonized phenotyping and diverse population-specific studies are crucial for advancing gene discovery in psychiatric genetics. We conducted a genome-wide association (GWAS) mega-analysis of DSM-defined lifetime major depressive disorder (MDD) in 64 941 participants (25.7% cases) from the Dutch BIObanks Netherlands Internet Collaboration (BIONIC) consortium. Liability-scale SNP-based heritability was 12.0% (SE = 1.4%) as estimated by LDSC (assuming a lifetime prevalence of 15%) and 26.6% (SE = 1.1%) when estimated by LDAK-REML on individual-level genotype data, indicating substantial common-variant signal in this clinically harmonized sample. The genetic correlation with the latest major depression GWAS from the Psychiatric Genomics Consortium (PGC-MD) was high (rG = 0.89, SE = 0.048). Polygenic scores (PGSs) based on BIONIC predicted depression in UK Biobank, and PGSs derived from PGC-MD predicted MDD in BIONIC, supporting transferability of depression polygenic signal across cohorts and phenotype definitions. Within-family PGS analyses in twins suggested that the observed prediction was not primarily driven by detectable family-level confounding, and twin concordance for MDD increased with polygenic burden. We identified one genome-wide significant locus, indexed by rs3818852 in PALMD, but this finding currently lacks independent replication and should be interpreted cautiously. Finally, genetic correlation and latent causal variable analyses identified multiple traits showing shared or directionally consistent genetic associations with MDD. Together, these findings underscore the value of clinically harmonized phenotyping in regional biobank collaborations for studying the genetic architecture of MDD.
MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.
In this paper a regression analysis is performed with data on spinal cord injuries in order to demonstrate the benefits of determining which, if any, multicollinearities are present in prediction data. Existing multicollinearities are shown to be useful both in determining characteristics of the sampled population as well as explaining possible erratic behavior of variable selection procedures. Latent root regression is performed on the data to illustrate one method of using biased regression techniques to incorporate knowledge of multicollinearities in developing prediction equations.
The genetic diversity that exists in natural populations of Arachis duranensis, the wild diploid donor of the A subgenome of cultivated tetraploid peanut, has the potential to improve crop adaptability, resilience to major pests and diseases, and drought tolerance. Despite its potential value for peanut improvement, limited research has been focused on the association between allelic variation, environmental factors, and response to early (ELS) and late leaf spot (LLS) diseases. The present study implemented a landscape genomics approach to gain a better understanding of the genetic variability of A. duranensis represented in the ex-situ peanut germplasm collection maintained at the U.S. Department of Agriculture, which spans the entire geographic range of the species in its center of origin in South America. A set of 2810 single nucleotide polymorphism (SNP) markers allowed a high-resolution genome-wide characterization of natural populations. The analysis of population structure showed a complex pattern of genetic diversity with five putative groups. The incorporation of bioclimatic variables for genotype-environment associations, using the latent factor mixed model (LFMM2) method, provided insights into the genomic signatures of environmental adaptation, and led to the identification of SNP loci whose allele frequencies were correlated with elevation, temperature, and precipitation-related variables (q < 0.05). The LFMM2 analysis for ELS and LLS detected candidate SNPs and genomic regions on chromosomes A02, A03, A04, A06, and A08. These findings highlight the importance of the application of landscape genomics in ex situ collections of peanut and other crop wild relatives to effectively identify favorable alleles and germplasm for incorporation into breeding programs. We report new sources of A. duranensis germplasm harboring adaptive allelic variation, which have the potential to be utilized in introgression breeding for a single or multiple environmental factors, as well as for resistance to leaf spot diseases.
The records of 74 patients--young and almost exclusively males--with retinal detachment due to ocular penetration were reviewed to study the characteristics and surgical results of this type of traumatic retinal detachment. There was a high incidence of industrial and domestic accidents. The incidence and degree of myopia were significantly lower than in nontraumatic cases of retinal detachment. The time interval between the injury and the detection of the retinal detachment was variable. Many of the eyes with longer latent intervals had classic signs of traumatic retinal detachment. The most common type of retinal break was a dialysis in the oral zone at the posterior vitreous base border. The most common cause of these retinal breaks was severe traction from contracting vitreous bands and membranes that followed the loss of vitreous gel from the eye. Surgical results were relatively poor; recurrence was common because of progressive vitreous pathology. Postoperative traction on the vitreous base by the continued shrinkage of vitreous membranes was noted in 50% of the eyes that contained such membranes preoperatively.
PURPOSE: "Abrazos, no balazos" (Hugs, not gunshots) is the approach to violence used by the Federal Mexican government from 2018 to 2024. This study identifies distinct conceptually useful latent class growth trajectories in the annual intensity of risk for exposure to events and fatalities from violence against civilians and battles for adolescents and young adults (AYAs) in México from 2018 to 2024. METHODS: Data are from the Armed Conflict Location and Event Data project. Latent class growth analyses were conducted with four longitudinal dependent variables assessed in seven sequential years (2018-2024). RESULTS: In most of the Mexican states, AYAs experienced low intensity-flat trajectories (p >.096) in the annual intensity of risk for exposure to events and fatalities from violence against civilians and battles. AYAs who lived in states with extremely, markedly, or moderately elevated 2018 intensities of risk for exposure to violence experienced decreasing trajectories (p =.000). One latent class with a mildly elevated-increasing trajectory (p =.000) in the annual intensity of risk for exposure to violence was identified for each of the four longitudinal dependent variables. DISCUSSION: During the time that the Federal Mexican government used the "Abrazos, no balazos" approach, AYAs in the majority of the states in México had intensity of risk for exposure to violence and armed conflict trajectories that were consistent with the goals of primary and secondary violence prevention. Violence prevention efforts should be augmented in states with unfavorable trajectories.
Genetic data has shown that antigenic variants or allotypes of variable and constant regions of the heavy and light chains of immunoglobulins represent products of allelic structural genes. New findings in several animal systems have demonstrated the non-allelic behavior of immunoglobulin allotypes. In the present study we report the results of our efforts to confirm the presence of latent allotypes in rabbit sera. Varying amounts of latent allotypes for groups a and b markers were detected in 95% of all normal and immune sera tested. Several rabbits expressed high levels of latent a and b allotypes (in excess of 200 micrograms/ml of serum). By gel diffusion and radioimmune inhibition analysis, the latent b4 marker from one b6b6 rabbit, B240, was identical to the nominal b4 allotype. Also, the L chains bearing the latent b4 marker were isolated from B240 IgG and analyzed chemically by taking advantage of an acid-labile aspartic acid-proline bond at positions 109 and 110, which is unique for b4 and b9 kappa L chains.