PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data fragmentation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

A decentralized future for the open-science databases.

The continuous and reliable open access to curated biological data repositories is indispensable for accelerating rigorous scientific inquiry and fostering reproducible research outcomes. However, the current paradigm, which relies heavily on centralized infrastructure for the storage and distribution of foundational biomedical datasets, inherently introduces significant vulnerabilities. This centralized model is susceptible to single points of failure, including cyberattacks, technical malfunctions, natural disasters, and even political or funding uncertainties. Such disruptions can lead to widespread data unavailability, data loss, integrity compromises, and substantial delays in critical research, ultimately impeding scientific progress. The downstream effect of such interruptions can be the widespread paralysis of diverse research activities, including computational, clinical, molecular, and climate studies. This scenario vividly illustrates the inherent dangers of consolidating essential scientific resources within a single geopolitical or institutional locus. As data generation is accelerating and the global landscape continues to fluctuate, the sustainability of centralized models must be critically re-evaluated. A shift toward federated and decentralized architectures may offer a robust and forward-looking approach to enhancing the resilience of scientific data infrastructures by reducing exposure to governance instability, infrastructural fragility, and funding volatility, while also promoting equity and global accessibility. Inspired by established models such as ELIXIR's federated infrastructure and the policy and funding frameworks developed by CODATA and the Global Biodata Coalition (GBC), emerging Decentralized Science (DeSci) initiatives can contribute to building more resilient, fair, and incentive-aligned data ecosystems. The future of open science depends on integrating these complementary approaches to establish a globally distributed, economically sustainable, and institutionally robust infrastructure that safeguards scientific data as a public good, further ensuring continued accessibility, interoperability, and preservation for generations to come. Here, we examine the structural limitations of centralized repositories, evaluate federated and decentralized models, and propose a hybrid framework for resilient, fair, and sustainable scientific data stewardship.

data accessibility↗

Induction of atypical hyperplasia, apoptosis, and type II estrogen-binding sites in the ventral prostates of Noble rats treated with testosterone and pharmacologic doses of estradiol-17 beta.

BACKGROUND: We have previously shown that combined administration of testosterone (T) and a low dose of estradiol 17 beta (T+LDE2) for 16 weeks induces an atypical proliferative lesion, termed dysplasia, in the dorsolateral prostates of intact Noble rats (1, 2). The lesion was accompanied by increases in the levels of a moderate affinity, high capacity, estrogen-binding site (type II sites) found exclusively in dorsolateral prostates of these animals (1, 3). In contrast, a proliferative response and type II sites were not observed in the ventral prostates (VP) of the same rats treated with this hormonal regimen. In the current study, rats were treated with a higher dose of E2 (4 x LDE2) but the same dose of T (T+HDE2) for 16 weeks. Our aims were to determine how the VP would respond to the T+HDE2 treatment. EXPERIMENTAL DESIGN: Intact Noble rats were treated with T+HDE2 for 16 weeks. Prostatic tissues were removed for histology, electronmicroscopy, and type II site measurements. Proliferating cells were identified by the histochemical detection of proliferating cell nuclear antigen and colcemid-arrested mitotic figures. Apoptotic cells were recognized by their characteristic histologic and ultrastructural features and by in situ detection of nuclear DNA fragmentation. Data were compared with results previously obtained from VP of rats treated with T+LDE2. RESULTS: The VP of T+HDE2-treated animals contained focal atypical hyperplasia and wide-spread apoptosis. Proliferating cell nuclear Ag-positive-stained epithelial cells and mitotic figures were only present in foci of atypical hyperplasia. Total DNA content of the VP was significantly increased, but the tissue wet weight was not augmented. Nuclear type II sites, never observed in untreated or T+LDE2-treated rats, were detected in the VP of the majority of T+HDE2-treated animals. CONCLUSIONS: The administration of a high dose of E2 with T produced a unique lesion in the VP, characterized by simultaneous occurrence of apoptosis and proliferation. The synergy between androgens and estrogens, via type II site induction, likely produces the proliferative response. On the other hand, inhibition of intracellular androgen activation pathways, leading to reduction in cell survival factors, may be the cause for the apoptotic development. Our model, thus, provides a unique opportunity to further study the balance/switch between cell proliferation and apoptosis that is often disturbed during cancer development.

Animals↗

Terminal restriction fragment length polymorphism data analysis for quantitative comparison of microbial communities.

Terminal restriction fragment length polymorphism (T-RFLP) is a culture-independent method of obtaining a genetic fingerprint of the composition of a microbial community. Comparisons of the utility of different methods of (i) including peaks, (ii) computing the difference (or distance) between profiles, and (iii) performing statistical analysis were made by using replicated profiles of eubacterial communities. These samples included soil collected from three regions of the United States, soil fractions derived from three agronomic field treatments, soil samples taken from within one meter of each other in an alfalfa field, and replicate laboratory bioreactors. Cluster analysis by Ward's method and by the unweighted-pair group method using arithmetic averages (UPGMA) were compared. Ward's method was more effective at differentiating major groups within sets of profiles; UPGMA had a slightly reduced error rate in clustering of replicate profiles and was more sensitive to outliers. Most replicate profiles were clustered together when relative peak height or Hellinger-transformed peak height was used, in contrast to raw peak height. Redundancy analysis was more effective than cluster analysis at detecting differences between similar samples. Redundancy analysis using Hellinger distance was more sensitive than that using Euclidean distance between relative peak height profiles. Analysis of Jaccard distance between profiles, which considers only the presence or absence of a terminal restriction fragment, was the most sensitive in redundancy analysis, and was equally sensitive in cluster analysis, if all profiles had cumulative peak heights greater than 10,000 fluorescence units. It is concluded that T-RFLP is a sensitive method of differentiating between microbial communities when the optimal statistical method is used for the situation at hand. It is recommended that hypothesis testing be performed by redundancy analysis of Hellinger-transformed data and that exploratory data analysis be performed by cluster analysis using Ward's method to find natural groups or by UPGMA to identify potential outliers. Analyses can also be based on Jaccard distance if all profiles have cumulative peak heights greater than 10,000 fluorescence units.

Bacteria↗

A Monte Carlo method for finding important ligand fragments from receptor data.

A simulated annealing method for finding important ligand fragments is described. At a given temperature, ligand fragments are randomly selected and randomly placed within the given receptor cavity, often replacing or forming bonds with existing ligand fragments. For each new ligand fragment combination, the bonded, nonbonded, polarization and solvation energies of the new ligand-receptor system are compared to the previous configuration. Acceptance or rejection of the new system is decided using the Boltzmann distribution e-E/kT, where E is the energy difference between the old and new systems, k is the Boltzmann constant and T is the temperature. Thus, energetically unfavorable fragment switches are sometimes accepted, sacrificing immediate energy gains in the interest of findings a system with minimum energy. By lowering the temperature, the rate of unfavorable switches decreases and energetically favorable combinations become more difficult to change. The process is terminated when the frequency of switches becomes too small. As a test, the method predicted positions and types of important ligand fragments for neuraminidase that were in accord with the known ligand, sialic acid.

Algorithms↗

Estimating nucleotide diversity from random amplified polymorphic DNA and amplified fragment length polymorphism data.

A way to estimate the index of nucleotide diversity (pi) from band match frequencies in random amplified polymorphic DNA and amplified fragment length polymorphism data is described. pi is shown to be a simple function of the proportion of mismatched bands between two individuals drawn at random from a population (phi) and the number of discriminating sites in the amplification system. The method is computationally and conceptually simple and avoids some of the assumptions inherent in other approaches: the relationship is independent of the base composition of the target DNA and avoids the bias inherent in estimations of allelic frequencies in dominant systems. Only two individuals from a population are needed to estimate pi. This economy of material suggests utility of this approach in conservation genetics or other fields where obtaining large samples is impractical or undesirable.

Alleles↗

Comparison of genetic diversity estimates within and among populations of maritime pine using chloroplast simple-sequence repeat and amplified fragment length polymorphism data.

We compared the genetic variation of Pinus pinaster populations using amplified fragment length polymorphism (AFLP) and chloroplast simple-sequence repeat (cpSSR) loci. Populations' levels of diversity within groups were found to be similar with AFLPs, but not with cpSSRs. The high interlocus variance associated with the AFLP loci could account for the lack of differences in the former. Although AFLPs revealed much lower genetic diversity than cpSSRs, the levels of among-population differentiation found with the two types of marker were similar, provided that loci showing fewer than four null-homozygotes, in any population, were pruned from the AFLP data. Moreover, the French and Portuguese populations were clearly differentiated from each other, with both markers. The Mantel test showed that the genetic distance matrix calculated using the AFLP data was correlated with the matrix derived from the cpSSRs. Because of the concordance found between markers we conclude that gene flow was indeed the predominant force shaping nuclear and chloroplastic genetic variation of the populations within regions, at the geographical scale studied.

Alleles↗

Genetic diversity of human pathogenic members of the Fusarium oxysporum complex inferred from multilocus DNA sequence data and amplified fragment length polymorphism analyses: evidence for the recent dispersion of a geographically widespread clonal lineage and nosocomial origin.

Fusarium oxysporum is a phylogenetically diverse monophyletic complex of filamentous ascomycetous fungi that are responsible for localized and disseminated life-threatening opportunistic infections in immunocompetent and severely neutropenic patients, respectively. Although members of this complex were isolated from patients during a pseudoepidemic in San Antonio, Tex., and from patients and the water system in a Houston, Tex., hospital during the 1990s, little is known about their genetic relatedness and population structure. This study was conducted to investigate the global genetic diversity and population biology of a comprehensive set of clinically important members of the F. oxysporum complex, focusing on the 33 isolates from patients at the San Antonio hospital and on strains isolated in the United States from the water systems of geographically distant hospitals in Texas, Maryland, and Washington, which were suspected as reservoirs of nosocomial fusariosis. In all, 18 environmental isolates and 88 isolates from patients spanning four continents were genotyped. The major finding of this study, based on concordant results from phylogenetic analyses of multilocus DNA sequence data and amplified fragment length polymorphisms, is that a recently dispersed, geographically widespread clonal lineage is responsible for over 70% of all clinical isolates investigated, including all of those associated with the pseudoepidemic in San Antonio. Moreover, strains of the clonal lineage recovered from patients were conclusively shown to genetically match those isolated from the hospital water systems of three U.S. hospitals, providing support for the hypothesis that hospitals may serve as a reservoir for nosocomial fusarial infections.

Animals↗

Have we seen all structures corresponding to short protein fragments in the Protein Data Bank? An update.

Assembling short fragments from known structures has been a widely used approach to construct novel protein structures. To what extent there exist structurally similar fragments in the database of known structures for short fragments of a novel protein is a question that is fundamental to this approach. This work addresses that question for seven-, nine- and 15-residue fragments. For each fragment size, two databases, a query database and a template database of fragments from high-quality protein structures in SCOP20 and SCOP90, respectively, were constructed. For each fragment in the query database, the template database was scanned to find the lowest r.m.s.d. fragment among non-homologous structures. For seven-residue fragments, there is a 99% probability that there exists such a fragment within 0.7 A r.m.s.d. for each loop fragment. For nine-residue fragments there is a 96% probability of a fragment within 1 A r.m.s.d., while for 15-residue fragments there is a 91% probability of a fragment within 2 A r.m.s.d. These results, which update previous studies, show that there exists sufficient coverage to model even a novel fold using fragments from the Protein Data Bank, as the current database of known structures has increased enormously in the last few years. We have also explored the use of a grid search method for loop homology modeling and make some observations about the use of a grid search compared with a database search for the loop modeling problem.

Databases, Factual↗

Structure Elucidator: a versatile expert system for molecular structure elucidation from 1D and 2D NMR data and molecular fragments.

StrucEluc is an expert system that allows the computer-assisted elucidation of chemical structures based on the inputs of a series of spectral data including 1D and 2D NMR and mass spectra. The system has been enabled to allow a chemist to utilize fragments stored in a fragment database as well as user-defined fragments submitted by the chemist in the structure elucidation process. The association of fragments in this way has been shown to dramatically speed up the process of structure generation from 2D NMR data and has helped to minimize or eliminate the need for user intervention thereby further enabling the vision of automated elucidation. The use of fragments has frequently transformed very difficult 2D NMR elucidation challenges into easily solvable tasks. A strategy to utilize molecular fragments has been developed and optimized based on specific challenging examples. This strategy will be described here using real world examples. Experience gained by solving more than 150 structure elucidation problems from a variety of literature sources is also reviewed in this work.

Journal Article↗

Sequence tag identification of intact proteins by matching tanden mass spectral data against sequence data bases.

Molecular and fragment ion data of intact 8- to 43-kDa proteins from electrospray Fourier-transform tandem mass spectrometry are matched against the corresponding data in sequence data bases. Extending the sequence tag concept of Mann and Wilm for matching peptides, a partial amino acid sequence in the unknown is first identified from the mass differences of a series of fragment ions, and the mass position of this sequence is defined from molecular weight and the fragment ion masses. For three studied proteins, a single sequence tag retrieved only the correct protein from the data base; a fourth protein required the input of two sequence tags. However, three of the data base proteins differed by having an extra methionine or by missing an acetyl or heme substitution. The positions of these modifications in the protein examined were greatly restricted by the mass differences of its molecular and fragment ions versus those of the data base. To characterize the primary structure of an unknown represented in the data base, this method is fast and specific and does not require prior enzymatic or chemical degradation.

Amino Acid Sequence↗

Investigations concerning the cultivation of myxo- and paramyxoviruses on chorioallantoic membrane fragments. Note I. Data on the multiplication of several myxo- and paramyxoviruses.

Influenza viruses A(H1N1) and A(H3N2) and parainfluenza viruses (Sendai mumps) were cultivated in chorioallantoic membrane (CAM) fragments maintained in media with different formulae, with or without daily medium changes, in roller or stationary tubes. Inoculation was performed either directly on CAM fragments in Petri dishes or by dilution of the virus-containing material in the medium. Infectant titers obtained in CAM fragments were similar to those recorded in embryonated eggs at 48 hours post inoculation (p.i.) in the case of influenza virus A(H1N1) and at 72 hours p.i. in that of Sendai and influenza A(H3N2) viruses; at 96 hours p.i. all the three viruses had titers superior to those found in the egg.

Animals↗

Pulmonary tuberculosis in Norwegian patients. The role of reactivation, re-infection and primary infection assessed by previous mass screening data and restriction fragment length polymorphism analysis.

SETTING: Norwegian patients with pulmonary tuberculosis notified to the National Tuberculosis Register in 1975, 1985 and 1995. OBJECTIVE: To assess the proportion of cases attributable to endogenous reactivation, exogenous re-infection and primary infection. DESIGN: We reviewed patients notified with sputum smear and/or culture confirmed pulmonary tuberculosis in 1975 (50% random sample, 95 cases), 1985 (133 cases) and 1995 (70 cases). Information on previous chest X-ray, tuberculin and BCG status was collected from mass screening data files. Strains from 54 patients in 1995 were analysed by IS6110 restriction fragment length polymorphism (RFLP) typing and compared with culture-positive patients notified between 1994 and 1997. RESULTS: Most patients had previously had tuberculosis (65% in 1975, 53% in 1985 and 61% in 1995), either notified with tuberculosis or with X-ray findings indicating previous tuberculosis. Another 10% had a prior infection, but normal X-rays. No previous tuberculosis infection or disease was found in 10% in 1975, 19% in 1985, and 16% in 1995. Of 54 patients with RFLP results, three were caused by laboratory contamination. Of the remaining 51, eight (16%) belonged to a cluster. Among 45 patients with results of both RFLP typing and mass screening, 37 (82.2%) were probably caused by reactivation, six (13.3%) by re-infection and two (4.4%) by primary infection. CONCLUSION: Pulmonary tuberculosis in Norwegian patients can mainly be attributed to reactivation, predominantly in persons with previous changes on chest X-ray.

Adolescent↗

Long-distance seed dispersal in a metapopulation of Banksia hookeriana inferred from a population allocation analysis of amplified fragment length polymorphism data.

There is currently a poor understanding of the nature and extent of long-distance seed dispersal, largely due to the inherent difficulty of detection. New statistical approaches and molecular markers offer the potential to accurately address this issue. A log-likelihood population allocation test (AFLPOP) was applied to a plant metapopulation to characterize interpopulation seed dispersal. Banksia hookeriana is a fire-killed shrub, restricted to sandy dune crests in fire-prone shrublands of the Eneabba sandplain, southwest Australia. Population genetic variation was assessed for 221 individuals sampled from 21 adjacent dune-crest populations of B. hookeriana using amplified fragment length polymorphism. Genetic diversity was high, with 175 of 183 (96%) amplified fragment length polymorphism markers polymorphic. Of the total genetic diversity, 8% was partitioned among populations by amova and FST. There was no relationship between genetic diversity within populations and population demographic parameters such as population size and sample size. A population allocation test on these data unambiguously assigned 177 of 221 (80.1%) individuals to a single population. Of these, 171 (77.4% of total) were assigned to the population from which they were sampled and 6 (2.7% of total) were assigned to a known population other than the one from which they were sampled. A further 9 (4.1% of total) were assigned to outside the sampled metapopulation area, and 35 individuals (15.8%) could not be assigned unambiguously to any particular population. These results suggest that both the extent [15 of 221 (6.8%) individuals originating from a population other than the one in which they occur] and distance (1.6 to > 2.5 km), of seed dispersal between dune-crest populations is greater than expected from previous studies. The extent of long-distance interpopulation seed dispersal observed provides a basis for explaining the survival of populations of the fire-killed B. hookeriana in a landscape experiencing frequent fire, where local extinctions and recolonizations may be a regular occurrence.

Demography↗

Confocal optical sectioning and three-dimensional reconstruction of carcinoma fragments in Pap smears using sophisticated image data processing.

Carcinoma fragments found in Pap smears contain important diagnostic information not available to the light microscopist because of their thickness and consequent blurring. Optical sectioning by the confocal microscope allows us to reclaim the mitotic figures, glandular architecture, and abnormal chromatin patterns in the restained original smears. The high spatial resolution of the confocal microscope can be further exploited by processing the digital images with the sophisticated Application Visualization System (AVS) on a CONVEX computer. Serial sections in which the fluorescent signals are color coded by this software package and three-dimensional reconstructions of the nuclei and mitotic figures expand our knowledge of these malignant epithelial fragments.

Adenocarcinoma↗

Dose- and time-dependent effects of a novel (-)-hydroxycitric acid extract on body weight, hepatic and testicular lipid peroxidation, DNA fragmentation and histopathological data over a period of 90 days.

(-)-Hydroxycitric acid (HCA), a natural extract from the dried fruit rind of Garcinia cambogia (family Guttiferae), is a popular supplement for weight management. The dried fruit rind has been used for centuries as a condiment in Southeastern Asia to make food more filling and satisfying. A significant number of studies highlight the efficacy of Super CitriMax (HCA-SX, a novel 60% calcium-potassium salt of HCA derived from Garcinia cambogia) in weight management. These studies also demonstrate that HCA-SX promotes fat oxidation, inhibits ATP-citrate lyase (a building block for fat synthesis), and lowers the level of leptin in obese subjects. Acute oral, acute dermal, primary dermal irritation and primary eye irritation toxicity studies have demonstrated the safety of HCA-SX. However, no long-term safety of HCA-SX or any other (-)-hydroxycitric acid extract has been previously assessed. In this study, we have evaluated the dose- and time-dependent effects of HCA-SX in Sprague-Dawley rats on body weight, hepatic and testicular lipid peroxidation, DNA fragmentation, liver and testis weight, expressed as such and as a % of body weight and brain weight, and histopathological changes over a period of 90 days. The animals were treated with 0, 0.2, 2.0 and 5.0% HCA-SX as feed intake and the animals were sacrificed on 30, 60 or 90 days of treatment. The feed and water intake were assessed and correlated with the reduction in body weight. HCA-SX supplementation demonstrated a reduction in body weight in both male and female rats over a period of 90 days as compared to the corresponding control animals. An advancing age-induced marginal increase in hepatic lipid peroxidation was observed in both male and female rats as compared to the corresponding control animals. However, no such difference in hepatic DNA fragmentation and testicular lipid peroxidation and DNA fragmentation was observed. Furthermore, liver and testis weight, expressed as such and as a percentage of body weight and brain weight, at 30, 60 and 90 days of treatment, exhibited no significant difference between the four groups. Taken together, these results indicate that treatment of HCA-SX over a period of 90 days results in a reduction in body weight, but did not cause any changes in hepatic and testicular lipid peroxidation, DNA fragmentation, or histopathological changes.

ATP Citrate (pro-S)-Lyase↗

The HUPO proteomics standards initiative--overcoming the fragmentation of proteomics data.

Proteomics is a key field of modern biomolecular research, with many small and large scale efforts producing a wealth of proteomics data. However, the vast majority of this data is never exploited to its full potential. Even in publicly funded projects, often the raw data generated in a specific context is analysed, conclusions are drawn and published, but little attention is paid to systematic documentation, archiving, and public access to the data supporting the scientific results. It is often difficult to validate the results stated in a particular publication, and even simple global questions like "In which cellular contexts has my protein of interest been observed?" can currently not be answered with realistic effort, due to a lack of standardised reporting and collection of proteomics data. The Proteomics Standards Initiative (PSI), a work group of the Human Proteome Organisation (HUPO), defines community standards for data representation in proteomics to facilitate systematic data capture, comparison, exchange and verification. In this article we provide an overview of PSI organisational structure, activities, and current results, as well as ways to get involved in the broad-based, open PSI process.

Databases, Protein↗

Data from amplified fragment length polymorphism (AFLP) markers show indication of size homoplasy and of a relationship between degree of homoplasy and fragment size.

We investigate the distribution of sizes of fragments obtained from the amplified fragment length polymorphism (AFLP) marker technique. We find that empirical distributions obtained in two plant species, Phaseolus lunatus and Lolium perenne, are consistent with the expected distributions obtained from analytical theory and from numerical simulations. Our results indicate that the size distribution is strongly asymmetrical, with a much higher proportion of small than large fragments, that it is not influenced by the number of selective nucleotides nor by genome size but that it may vary with genome-wide GC-content, with a higher proportion of small fragments in cases of lower GC-content when considering the standard AFLP protocol with the enzyme MseI. Results from population samples of the two plant species show that there is a negative relationship between AFLP fragment size and fragment population frequency. Monte Carlo simulations reveal that size homoplasy, arising from pulling together nonhomologous fragments of the same size, generates patterns similar to those observed in P. lunatus and L. perenne because of the asymmetry of the size distribution. We discuss the implications of these results in the context of estimating genetic diversity with AFLP markers.

Alleles↗