PubMed HealthSearch

SEARCH · PubMed Health

Results for “Data Compression”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Optimizing sparse and skew hashing: faster k-mer dictionaries.

MOTIVATION: Representing a set of k-mers-strings of length k-in small space under fast lookup queries is a fundamental requirement for several applications in Bioinformatics. A data structure based on sparse and skew hashing (SSHash) was recently proposed for this purpose (Pibiri 2022): it combines good space effectiveness with fast lookup and streaming queries. It is also order-preserving, i.e. consecutive k-mers (sharing a prefix-suffix overlap of length k-1) are assigned consecutive hash codes which helps compressing satellite data typically associated with k-mers, like abundances and color sets in colored De Bruijn graphs. RESULTS: We study the problem of accelerating queries under the sparse and skew hashing indexing paradigm, without compromising its space effectiveness. We propose a refined data structure with less complex lookups and fewer cache misses. We give a simpler and faster algorithm for streaming lookup queries. The refined architecture translates to substantial performance gains, outperforming the original version of SSHash in both index construction speed and query efficiency. Compared to indexes with similar capabilities and based on the Burrows-Wheeler transform, like SBWT and FMSI, SSHash is significantly faster to build and query. SSHash is competitive in space with the fast (and default) modality of SBWT when both k-mer strands are indexed. While larger than FMSI, it is also more than one order of magnitude faster to query. AVAILABILITY AND IMPLEMENTATION: The SSHash software is available at https://github.com/jermp/sshash, and also distributed via Bioconda. A benchmark of data structures for k-mer sets is available at https://github.com/jermp/kmer_sets_benchmark. The datasets used in this article are described and available at https://zenodo.org/records/17582116.

Algorithms

FFC: a scalable FASTA compressor.

SUMMARY: FASTA is a widely used text-based format for storing nucleotide and protein sequences. The existing FASTA compressors usually focus on (slightly) improving the compression ratio, not on practical performance. We present FFC, a scalable FASTA compressor that achieves average compression speeds 4.7× and 11.4× higher than two high-performance compressors, zstd and NAF, respectively, across a benchmark set of seven single genomes. It also delivers average decompression speeds 3.5× and 2.7× higher than zstd and NAF, respectively. Although a chunk-based zstd variant with parallel decompression, pzstd, almost matches FFC speed, its compression ratio is on average by 23% worse than FFC's. For the experiment, a 14-core workstation and a RAM disk (to reduce the impact of I/O) were used. AVAILABILITY AND IMPLEMENTATION: FFC is freely available at github.com/kowallus/ffc and also as a Zenodo repository at 10.5281/zenodo.18892353, and the used datasets at 10.5281/zenodo.18873744.

Data Compression

Evaluation of polymeric materials for maxillofacial prosthetics.

The methods for initial evaluation of a new prospective maxillofacial prosthetic material, MDX-4-4210, have been described and tensile and compressive test data have been obtained. Evaluation of the data indicates that this is a promising material with significant potential for use in maxillofacial prosthetics. Further studies of the material's biologic implant compatibility, its chemical and environmental stabilities, and its clinical performance are being done.

Elasticity

[Clinical electroencephalographic characteristics of the phasic nature of the course of severe cerebrocranial injury].

In order to early recognize intracranial hematomas and contusions, differential diagnostic criteria for analysis of neurological symptomatology and EEG signs are considered from the standpoint of a phasic clinical progress of traumatic pathological conditions of the brain. The EEG syndromes of initial, pronounced and gross synchronization correlated with the clinical phases of subcompensation, mild and deep decompensation. The study demonstrates a contrary trend in the dynamics of the clinical and EEG data in compression and contusion of the brain.

Brain Concussion

A comparison of ultraviolet-curing and self-curing polymers in preventive, restorative and orthodontic dentistry.

Both self-cured and UV-cured resin-base dental materials are used in preventive, restorative, and orthodontic dentistry. Polymerization is initiated in both systems by free radicals. Self-curing materials generate free radicals by means of chemical compounds included in their formulation. UV-curing systems rely upon externally-supplied, long wavelength, ultraviolet radiation to produce free radicals within the material. Therefore, although the major chemical components of both systems are similar in many respects, each system has particular advantages and disadvantages over the other, which must be recognized by the practitioner. Substantial differences exist, for example, in the color stability of these two types of materials, because of the fact that the UV-cured system cannot include UV absorbers, which protect the self-cured systems from discoloration after exposure to sunlight. UV-cured systems require a limitation on the maximum depth of filled restorative that can be cured at one time, since the filler particles attenuate UV radiation. The limit-layer is generally established as 1-1-5 mm maximum thickness. Therefore, UV-cured filled systems are more time-consuming in restorations of deeper cavities. This liability is also in evidence as it affects the degree of polymerization of UV-cured filled systems. The uncertainty of complete polymerization is apparently responsible for highly erratic compressive strength data found with UV-cured restoratives. Normally, the amount of unpolymerized monomer is much less predictable in UV-cured systems, over that which is obtained in self-cured materials. The presence of a larger fraction of unpolymerized monomer creates a greater potential for pulpal injury from UV-cured restorative materials. The catalyst used in several UV-cured systems is benzoin methyl ether, a compound of rather high toxicity (LD50:300 mg/kg). The safety of using UV radiation in the vicinity of oral mucosa has not been firmly established. The design of the UV lamp should provide for focusing all radiation onto hard tissue. However, UV-cured systems do offer several advantages over self-cured systems. They normally are one-component systems and therefore are more convenient to use in certain types of applications, e.g., fissure sealing. UV-cured systems also provide an unlimited working time, an important advantage for specific applications.

Acrylic Resins

STABIX: summary-statistic-based GWAS indexing and compression.

MOTIVATION: Genome-wide association studies (GWAS) are widely used to investigate the role of genetics in disease traits, but the resulting file sizes from these studies are large, posing barriers to efficient storage, sharing, and querying. This issue is especially important for biobanks like the UK Biobank that publish GWAS for thousands of traits, increasing the volume of data that must be effectively managed. Current compression and query methods reduce file sizes and allow for quick genomic position-based queries but do not provide utility for quickly finding loci based on their summary statistics. For example, finding all SNVs in a particular p-value range would require decompressing and scanning the whole file. We propose a new tool, STABIX, which introduces summary-statistic-based queries and improves upon the standard bgzip compression and Tabix query tool in both compression ratio and decompression speed. RESULTS: When applied to 10 GWAS files from PanUKBB, STABIX created smaller compressed data and indices than Tabix for all files, where bgzip and tbi files were an average of 1.2 times the size of STABIX compressed files and indexes. In the same 10 files, STABIX per gene decompression was, on average 7× faster than Tabix per gene decompression, and achieved faster per gene decompression times for over 99% of nearly 20,000 genes. AVAILABILITY AND IMPLEMENTATION: Software freely available for download at GitHub: https://github.com/kristen-schneider/stabix/.

Genome-Wide Association Study

[Biomathematical evaluation of morphometric data--presented illustrated on the example of the relationship between axon diameter and thickness of the myelin sheath. II. Nonlinear approximations].

After the process of measuring morphometrical characteristics performing empirical regressions as a first stage of data processing is treated in the foregoing part I of the paper. The morphometrical datas have to characterize a functional relationship between two variables X and Y, in the cases in question this is the connection between axon caliber and thickness of myelin sheat of nerve fibres. Based on the results of empirical regression the second data processing stage consists in performing nonlinear approximations of the measured courses by a suitable chosen mathematical function. For this purpose the generalized logistic function is chosen for a quantitative and qualitative description of the connection between axon caliber and myelin sheat thickness. The numerical procedure and the various possibilities of the ALGOL-program for performing the approximation task are sketched. The results of this second data processing stage are discussed under the aspects of information compressing, of automatic data processing, and of providing an objective basis for the further process of evaluation and interpretation.

Anthropometry

Interpretation of dissolution rate data from in vitro testing of compressed tablets.

To find if theoretically and experimentally a relation existed between the dissolution rate theory of Kitazawa, Johno & others (1975) and that of Wagner (1969), a study was undertaken with uncoated caffeine, aspirin and proxyphylline tablets using two dissolution methods. Although the original treatment for surface area of drug available for dissolution was quite different between the two dissolution theories, the dissolution rate constants obtained were in fair agreement. Hence it might not be always necessary to take into consideration changes in the surface area as a function of dissolution rate, and the 1n W infinity/(W infinity) versus time plot devised by Kitazawa & others might be a useful and simple means of obtaining the dissolution rate constant of an active ingredient from a dosage form such as compressed tablet.

Aspirin

[The E.E.G. in ruptured cerebral aneurysms before operation. Use of carotid compression and the controlled hypotension test (author's transl)].

In 45 cases of ruptured intra-cranial aneurysms the authors studied the standard E.E.G. data under carotid compression and controlled hypotension; the changes observed were analysed and their value assessed for judging the state of the patient's cerebral circulation. A normal resting E.E.G. carries a good prognosis. Carotid compression and controlled hypotension provide a useful assessment of the supply and elasticity of the cerebral vessels; the occurrence of slow waves during controlled hypotension is an alarm signal which often helps in anaesthetic procedure.

Carotid Sinus

Evaluation of sequencing reads at scale using rdeval.

MOTIVATION: Large sequencing datasets are being produced and deposited into public archives at unprecedented rates. The availability of tools that can reliably and efficiently generate and store sequencing read summary statistics has become critical. RESULTS: As part of the effort by the Vertebrate Genomes Project (VGP) to generate high-quality reference genomes at scale, we sought to address the community's need for efficient sequence data evaluation by developing rdeval, a standalone tool to quickly compute and interactively display sequencing read metrics. Rdeval can either run on the fly or store key sequence data metrics in tiny read 'snapshot' files. Statistics can then be efficiently recalled from snapshots for additional processing. Rdeval can convert fa*[.gz] files to and from other popular formats including BAM and CRAM for better compression. Overall, while CRAM achieves the best compression, the gain compared to BAM is marginal, and BAM achieves the best compromise between data compression and access speed. Rdeval also generates a detailed visual report with multiple data analytics that can be exported in various formats. We showcase rdeval's functionalities using long-read data from different sequencing platforms and species, including human. For PacBio long-read sequencing, our analysis shows dramatic improvements in both read length and quality over time, as well as the benefit of increased coverage for genome assembly, though the magnitude varies by taxa. AVAILABILITY AND IMPLEMENTATION: Rdeval is implemented in C++ for data processing and in R for data visualization. Precompiled releases (Linux, MacOS, Windows) and commented source code for rdeval are available under MIT license at https://github.com/vgl-hub/rdeval. Documentation is available on ReadTheDocs (https://rdeval-documentation.readthedocs.io). Rdeval is also available in Bioconda and in Galaxy (https://usegalaxy.org). An automated test workflow ensures the consistency of software updates.

Software

[Biostatical and biomathematical evaluation of morphometric data--illustrated on the example of the relationship between axon diameter and thickness of the myelin sheat. I. 1st data evaluation step: empiric regression].

Application of the biostatistical procedure of empirical regression on a first data processing stage is demonstrated by examples from morphometrical research concerning the connection between axon caliber and thickness of myelin sheat of nerve fibers. This mathematical kind of describing connections between continuous random variables offers advantages which are discussed under the aspects of information compressing, of automatic data processing, and of providing objective basis for the process of data handling and decision making of further evaluations.

Animals

[Acoustic investigations in orbital tumor diagnosis (author's transl)].

360 patients with unilateral exophthalmos were examined with echoophthalmograph to study the diagnostical possibilities of a method of echography in orbital tumors. The findings of para- and transbulbal methodics were checked in 140 patients. The tonometry of tumors and soft orbital tissues was performed. The methodics of acoustic orbitonometry based on the measurement of retrobulbal neoplasms at the moment of their step-by-step compression was worked out. The echographic symptoms of different orbital neoplasms is described as well as those of pseudotumor and tireotropic exophthalmos. The possibility of acoustic measurement of orbital neoplasms is shown. The dependence between the extent and progress of deformation in the exophthalmed eye on one hand-nature of exophthalmos, character and localization of the tumor on the other was revealed. The quantitative data about the compression of different orbital tumors were obtained. It was shown that the surveying echography, biometry and acoustic orbitonometry facilitate the task of differential diagnosis, allow the localization and size of the tumor to be detected and help to control the results of the treatment.

Anthropometry

Adaptive segmentation of EEG records: a new approach to automatic EEG analysis.

The first step in a procedure for automatic EEG analysis is to compress the incoming data into a manageable format while preserving the essential diagnostic information. In our approach we mimic the visual procedure of looking through the record for segments and events of particular interest. We assume that the EEG is composed of roughly stationary segments of variable length, possibly superposed by sharp transients. By using an autoregressive model we have developed a procedure to detect the segment boundaries and locate transients, and to represent the information in the segments in terms of a set of parameters specifying their power spectra. In this way, the time structure as well as the frequency content of the signal is preserved. Examples of segmentation and transient detection are shown for several EEG signals, and the quality of the representation is demonstrated by simulating the original signal from the parameters. Possible applications to practical EEG analysis are discussed.

Adult

RLBWT-based LCP computation in compressed space for terabase-scale pangenome analysis.

MOTIVATION: Lossless full text indexes are utilized in a myriad of applications in bioinformatics. The continuously decreasing cost of generating biological data has resulted in the need to build full text indexes on biological datasets of increasing size. Many compressed full text indexes have been developed to address this problem. In particular, run-length Burrows-Wheeler transform (RLBWT) based compressed full text indexes have seen wide development and adoption. However, the construction of these RLBWT-based compressed full text indexes is still computationally expensive, sometimes prohibitively so, even for current dataset sizes. RESULTS: Therefore, we present algorithms for the construction of RLBWT-based compressed full text indexes and their supporting data structures in compressed space. The algorithms have a space complexity of O(r) words and run in O(n) time for repetitive datasets, where r is the number of runs in the BWT, n is the length of the text, and repetitive datasets implies nr∈Ω(log n). We provide the first algorithm to compute LCP-related information for repetitive datasets in optimal time and O(r) space, greatly reducing memory requirements. The key idea behind this algorithm is the utilization of r samples of the inverse suffix array at regular intervals. For example, on the Human Pangenome Reference Consortium Release 2 dataset, this reduces peak memory from 2135 GiB to 170 GiB (12.6x reduction) compared to the previous best method (pfp-thresholds). AVAILABILITY AND IMPLEMENTATION: The implementation is available at https://github.com/ucfcbb/TeraTools.

Algorithms

Limitations of bandwidth compression hearing aids applied to the voiced portion of speech.

Numerous speech processing techniques have been applied to assist hearing-impaired subjects with extreme high-frequency hearing losses who can be helped only to a limited degree with conventional hearing aids. The results of providing this class of deaf subjects with a speech encoding hearing aid, which is able to reproduce intelligible speech for their particular needs, have generally been disappointing. There are at least four problems related to bandwidth compression applied to the voiced portion of speech: (1) the problem of pitch extraction in real time; (2) pitch extraction under realistic listening conditions, i.e. when competing speech and noise sources are present; (3) an insufficient data base for successful compression of voiced speech; and (4) the introduction of undesirable spectral energies in the bandwidth-compressed signal, due to the compression process itself. Experiments seem to indicate that voiced speech segments bandwidth limited to f = 1000 Hz, even at a loss of higher formant frequencies, is in most instances superior in intelligibility compared to bandwidth-compressed voiced speech segments of the same bandwidth, even if pitch can be extracted with no error. With the added complexity of real-time pitch extraction which has to function in actual listening conditions, it is doubtful that a speech encoding hearing aid, based on bandwidth compression on the voiced portion of speech, could be successfully implemented. However, if bandwidth compression is applied to the unvoiced portions of speech only, the above limitations can be overcome (1).

Auditory Perception

Time, rate, and temperature factors in the onset of high-pressure convulsions.

An interrupted compression profile technique was used to develop data to separate the effects of time and pressure factors governing increase of high-pressure neurological syndrome (HPNS) convulsion threshold pressures (the compression rate effect) during different compression profiles. A single differential equation fits all data available to date for compression rate effect on convulsion thresholds of CD-1 mice (three distinct types of compression profile; mean compression rates 12-1,000 atm/h). The process leading to increase in HPNS convulsion pressure is initiated at the very beginning of compression, proceeds at increasingly rapid rates as higher pressures are attained, and approaches a limiting upper convulsion pressure. The convulsion threshold pressure in any given experiment is independent of the compression rate prevailing during the time immediately preceding onset of the seizure. The magnitude of the compression rate effect in the CD-1 mouse is independent of chamber temperature over a range of 27-36 degrees C, and rectal temperatures of 29.2-37.5 degrees C. The bearing of these results on the design of optimal compression schedules and on the analysis of the neurological mechanisms underlying the HPNS is discussed.

Animals

RP-REP Ribosomal Profiling Reports: an open-source cloud-enabled framework for reproducible ribosomal profiling data processing, analysis, and result reporting.

Ribosomal profiling is an emerging experimental technology to measure protein synthesis by sequencing short mRNA fragments undergoing translation in ribosomes. Applied on the genome wide scale, this is a powerful tool to profile global protein synthesis within cell populations of interest. Such information can be utilized for biomarker discovery and detection of treatment-responsive genes. However, analysis of ribosomal profiling data requires careful preprocessing to reduce the impact of artifacts and dedicated statistical methods for visualizing and modeling the high-dimensional discrete read count data. Here we present Ribosomal Profiling Reports (RP-REP), a new open-source cloud-enabled software that allows users to execute start-to-end gene-level ribosomal profiling and RNA-Seq analysis on a pre-configured Amazon Virtual Machine Image (AMI) hosted on AWS or on the user's own Ubuntu Linux server. The software works with FASTQ files stored locally, on AWS S3, or at the Sequence Read Archive (SRA). RP-REP automatically executes a series of customizable steps including filtering of contaminant RNA, enrichment of true ribosomal footprints, reference alignment and gene translation quantification, gene body coverage, CRAM compression, reference alignment QC, data normalization, multivariate data visualization, identification of differentially translated genes, and generation of heatmaps, co-translated gene clusters, enriched pathways, and other custom visualizations. RP-REP provides functionality to contrast RNA-SEQ and ribosomal profiling results, and calculates translational efficiency per gene. The software outputs a PDF report and publication-ready table and figure files. As a use case, we provide RP-REP results for a dengue virus study that tested cytosol and endoplasmic reticulum cellular fractions of human Huh7 cells pre-infection and at 6 h, 12 h, 24 h, and 40 h post-infection. Case study results, Ubuntu installation scripts, and the most recent RP-REP source code are accessible at GitHub. The cloud-ready AMI is available at AWS (AMI ID: RPREP RSEQREP (Ribosome Profiling and RNA-Seq Reports) v2.1 (ami-00b92f52d763145d3)).

AMI

Intonation and intelligibility of time-compressed speech. Supplementary report: English vs. French.

Comparative data are reported for the intelligibility of English and of French time-compressed speech when heard spoken either in normal intonation or in intonation patterns conflicting with underlying syntactic structure. Within an overall decrement in intelligibility with increasing compression, both French and English show similar superiority functions for sentences heard in normal intonation. Results suggest a role of prosodic features in perceptual processing of French comparable to that previously reported for English.

Adult