PubMed Health⌕ Search

Biomedical subjects

H Herzel

Publications and source records attributed to H Herzel.

At least 19 recordsLinked to original sources

Statistical analysis of the DNA sequence of human chromosome 22.

We study statistical patterns in the DNA sequence of human chromosome 22, the first completely sequenced human chromosome. We find that (i). the 33.4 x 10(6) nucleotide long human chromosome exhibits long-range power-law correlations over more than four orders of magnitude, (ii). the entropies H(n) of the frequency distribution of oligonucleotides of length n (n-mers) grow sublinearly with increasing n, indicating the presence of higher-order correlations for all of the studied lengths 1<or=n<or=10, and (iii). the generalized entropies H(n)(q) of n-mers decrease monotonically with increasing q and the decay of H(n)(q) with q becomes steeper with increasing n<or=10, indicating that the frequency distribution of oligonucleotides becomes increasingly nonuniform as the length n increases. We investigate to what degree known biological features may explain the observed statistical patterns. We find that (iv). the presence of interspersed repeats may cause the sublinear increase of H(n) with n, and that (v). the presence of monomeric tandem repeats as well as the suppression of CG dinucleotides may cause the observed decay of H(n)(q) with q.

Algorithms↗

Combining frequency and positional information to predict transcription factor binding sites.

MOTIVATION: Even though a number of genome projects have been finished on the sequence level, still only a small proportion of DNA regulatory elements have been identified. Growing amounts of gene expression data provide the possibility of finding coregulated genes by clustering methods. By analysis of the promoter regions of those genes, rather weak signals of transcription factor binding sites may be detected. RESULTS: We introduce the new algorithm ITB, an Integrated Tool for Box finding, which combines frequency and positional information to predict transcription factor binding sites in upstream regions of coregulated genes. Motifs are extracted by exhaustive analysis of regular expression-like patterns and by estimating probabilities of positional clusters of motifs. ITB detects consensus sequences of experimentally verified transcription factor binding sites of the yeast Saccharomyces cerevisiae. Moreover, a number of new binding site candidates with significant scores are predicted. Besides applying ITB on yeast upstream regions, the program is run on human promoter sequences. AVAILABILITY: ITB is available upon request.

Algorithms↗

Is there a bias in proteome research?

Advances in technology have enabled us to take a fresh look at data acquired by traditional single experiments and to compare them with genomewide data. The differences can be tremendous, as we show here, in the field of proteomics. We have compared data sets of protein-protein interactions in Saccharomyces cerevisiae that were detected by an identical underlying technical method, the yeast two-hybrid system. We found that the individually identified protein-protein interactions are considerably different from those identified by two genomewide scans. Interacting proteins in the pooled database from single publications are much more closely related to each other with respect to transcription profiles when compared to genomewide data. This difference may have been introduced by two factors: by a selection process in individual publications and by false positives in the whole-genome scans. If we assume that the differences are a result of false positives in the whole-genome data, the scans would contain 47%, 44%, and 91% of false positives for the UETZ, ITO-core, and ITO-full data, respectively. If, however, the true fraction of false positives is considerably lower than estimated here, the data from hypothesis-driven experiments must have been subjected to a serious selection process.

False Positive Reactions↗

The harmonic-to-noise ratio applied to dog barks.

Dog barks are typically a mixture of regular components and irregular (noisy) components. The regular part of the signal is given by a series of harmonics and is most probably due to regular vibrations of the vocal folds, whereas noise refers to any nonharmonic (irregular) energy in the spectrum of the bark signal. The noise components might be due to chaotic vibrations of the vocal-fold tissue or due to turbulence of the air. The ratio of harmonic to nonharmonic energy in dog barks is quantified by applying the harmonics-to-noise ratio (HNR). Barks of a single dog breed were recorded in the same behavioral context. Two groups of dogs were considered: a group of ten healthy dogs (the normal sample), and a group of ten unhealthy dogs, i.e., dogs treated in a veterinary clinic (the clinic sample). Although the unhealthy dogs had no voice disease, differences in emotion or pain or impacts of surgery might have influenced their barks. The barks of the dogs were recorded for a period of 6 months. The HNR computation is based on the Fourier spectrum of a 50-ms section from the middle of the bark. A 10-point moving average curve of the spectrum on a logarithmic scale is considered as estimator of the noise level in the bark, and the maximum difference of the original spectrum and the moving average is defined as the HNR measure. It is shown that a reasonable ranking of the voices is achievable based on the measurement of the HNR. The HNR-based classification is found to be consistent with perceptual evaluation of the barks. In addition, a multiparametric approach confirms the classification based on the HNR. Hence, it may be concluded that the HNR might be useful as a novel parameter in bioacoustics for quantifying the noise within a signal.

Animal Communication↗

Spatio-temporal analysis of irregular vocal fold oscillations: biphonation due to desynchronization of spatial modes.

This report is on direct observation and modal analysis of irregular spatio-temporal vibration patterns of vocal fold pathologies in vivo. The observed oscillation patterns are described quantitatively with multiline kymograms, spectral analysis, and spatio-temporal plots. The complex spatio-temporal vibration patterns are decomposed by empirical orthogonal functions into independent vibratory modes. It is shown quantitatively that biphonation can be induced either by left-right asymmetry or by desynchronized anterior-posterior vibratory modes, and the term "AP (anterior-posterior) biphonation" is introduced. The presented phonation examples show that for normal phonation the first two modes sufficiently explain the glottal dynamics. The spatio-temporal oscillation pattern associated with biphonation due to left-right asymmetry can be explained by the first three modes. Higher-order modes are required to describe the pattern for biphonation induced by anterior-posterior vibrations. Spatial irregularity is quantified by an entropy measure, which is significantly higher for irregular phonation than for normal phonation. Two asymmetry measures are introduced: the left-right asymmetry and the anterior-posterior asymmetry, as the ratios of the fundamental frequencies of left and right vocal fold and of anterior-posterior modes, respectively. These quantities clearly differentiate between left-right biphonation and anterior-posterior biphonation. This paper proposes methods to analyze quantitatively irregular vocal fold contour patterns in vivo and complements previous findings of desynchronization of vibration modes in computer modes and in in vitro experiments.

Adult↗

Optimization of coding potentials using positional dependence of nucleotide frequencies.

We study the coding potential of human DNA sequences, using the positional asymmetry function (D(p)) and the positional information function (I(q)). Both D(p)and I(q)are based on the positional dependence of single nucleotide frequencies. We investigate the accuracy of D(p)and I(q)in distinguishing coding and non-coding DNA as a function of the parameters p and q, respectively, and explore at which parameters p(opt)and q(opt)both D(p)and I(q)distinguish coding and non-coding DNA most accurately. We compare our findings with classically used parameter values and find that optimized coding potentials yield comparable accuracies as classical frame-independent coding potentials trained on prior data. We find that p(opt)and q(opt)vary only slightly with the sequence length.

Codon↗

Information content of protein sequences.

The complexity of large sets of non-redundant protein sequences is measured. This is done by estimating the Shannon entropy as well as applying compression algorithms to estimate the algorithmic complexity. The estimators are also applied to randomly generated surrogates of the protein data. Our results show that proteins are fairly close to random sequences. The entropy reduction due to correlations is only about 1%. However, precise estimations of the entropy of the source are not possible due to finite sample effects. Compression algorithms also indicate that the redundancy is in the order of 1%. These results confirm the idea that protein sequences can be regarded as slightly edited random strings. We discuss secondary structure and low-complexity regions as causes of the redundancy observed. The findings are related to numerical and biochemical experiments with random polypeptides.

Algorithms↗

Normalization strategies for cDNA microarrays.

Multiple Arabidopsis thaliana clones from an experimental series of cDNA microarrays are evaluated in order to identify essential sources of noise in the spotting and hybridization process. Theoretical and experimental strategies for an improved quantitative evaluation of cDNA microarrays are proposed and tested on a series of differently diluted control clones. Several sources of noise are identified from the data. Systematic and stochastic fluctuations in the spotting process are reduced by control spots and statistical techniques. The reliability of slide to slide comparison is critically assessed within the statistical framework of pattern matching and classification.

Arabidopsis↗

Are noncoding sequences of Rickettsia prowazekii remnants of "neutralized" genes?

It has been hypothesized that a large fraction of 24% noncoding DNA in R. prowazekii consists of degraded genes. This hypothesis has been based on the relatively high G+C content of noncoding DNA. However, a comparison with other genomes also having a low overall G+C content shows that this argument would also apply to other bacteria. To test this hypothesis, we study the coding potential in sets of genes, pseudogenes, and intergenic regions. We find that the correlation function and the chi(2)-measure are clearly indicative of the coding function of genes and pseudogenes. However, both coding potentials make almost no indication of a preexisting reading frame in the remaining 23% of noncoding DNA. We simulate the degradation of genes due to single-nucleotide substitutions and insertions/deletions and quantify the number of mutations required to remove indications of the reading frame. We discuss a reduced selection pressure as another possible origin of this comparatively large fraction of noncoding sequences.

DNA, Intergenic↗

Species independence of mutual information in coding and noncoding DNA.

We explore if there exist universal statistical patterns that are different in coding and noncoding DNA and can be found in all living organisms, regardless of their phylogenetic origin. We find that (i) the mutual information function [symbol: see text] has a significantly different functional form in coding and noncoding DNA. We further find that (ii) the probability distributions of the average mutual information [symbol: see text] are significantly different in coding and noncoding DNA, while (iii) they are almost the same for organisms of all taxonomic classes. Surprisingly, we find that [symbol: see text] is capable of predicting coding regions as accurately as organism-specific coding measures.

DNA↗

Nonlinear phenomena in the natural howling of a dog-wolf mix.

It was reported to the first author that a female dog-wolf mix showed anomalously rough-sounding vocalization. Spectral analysis of recordings of the vocalization revealed frequency occurrences of subharmonics, biphonation (two independent pitches) and chaos. Since these nonlinear phenomena are currently widely discussed as integral to mammalian vocalization [Wilden et al., Bioacoustics 9, 171-196 (1988)] or as indicators of vocal pathologies [Herzel et al., J. Speech Hearing Res. 37, 1008-1019 (1994); Riede et al., Z. Sgtkde 62 Suppl: 198-203 (1997)], we sought to understand the production mechanism of the observed vocal instabilities. First the frequency of nonlinear phenomena in the calls was determined for the female and four additional individuals. It turned out that these phenomena appear, but much less frequently in the repertoire of the four other animals. The larynges of the female and two other individuals were dissected post mortem. There was no apparent asymmetry of the vocal folds but a slight asymmetry of the arytenoid cartilages. The most pronounced difference, however, was an upward extension of both vocal folds of the female. This feature is reminiscent of "vocal lips" (syn. "vocal membranes") in some primates and bats. Spectral analysis of the female's voice showed clear similarities with an intensively studied voice of a human who produces biphonation intentionally. Finally, the possible communicative relevance of nonlinear phenomena is discussed.

Animals↗

Irregular vocal-fold vibration--high-speed observation and modeling.

Direct observations of nonstationary asymmetric vocal-fold oscillations are reported. Complex time series of the left and the right vocal-fold vibrations are extracted from digital high-speed image sequences separately. The dynamics of the corresponding high-speed glottograms reveals transitions between low-dimensional attractors such as subharmonic and quasiperiodic oscillations. The spectral components of either oscillation are given by positive linear combinations of two fundamental frequencies. Their ratio is determined from the high-speed sequences and is used as a parameter of laryngeal asymmetry in model calculations. The parameters of a simplified asymmetric two-mass model of the larynx are preset by using experimental data. Its bifurcation structure is explored in order to fit simulations to the observed time series. Appropriate parameter settings allow the reproduction of time series and differentiated amplitude contours with quantitative agreement. In particular, several phase-locked episodes ranging from 4:5 to 2:3 rhythms are generated realistically with the model.

Adult↗

Average mutual information of coding and noncoding DNA.

One basic problem in the analysis of DNA sequences is the recognition of protein-coding genes. Computer algorithms to facilitate gene identification have become important as genome sequencing projects have turned from mapping to large-scale sequencing, resulting in an exponentially growing number of sequenced nucleotides that await their annotation. Many statistical patterns have been discovered that are different in coding and noncoding DNA, but most of them vary from species to species, and hence require prior training on organism-specific data sets. Here, we investigate if there exist species-independent statistical patterns that are different in coding and noncoding DNA. We introduce an information-theoretic quantity, the average mutual information (AMI), and we find that the probability distribution functions of the AMI are significantly different in coding and noncoding DNA, while they are almost identical for different species. This finding suggests that the AMI might be useful for the recognition of protein-coding regions in genomes for which training sets do not exist.

Algorithms↗

10-11 bp periodicities in complete genomes reflect protein structure and DNA folding.

MOTIVATION: Completely sequenced genomes allow for detection and analysis of the relatively weak periodicities of 10-11 basepairs (bp). Two sources contribute to such signals: correlations in the corresponding protein sequences due to the amphipatic character of alpha-helices and the folding of DNA (nucleosomal patterns, DNA supercoiling). Since the topological state of genomic DNA is of importance for its replication, recombination and transcription, there is an immediate interest to obtain information about the supercoiled state from sequence periodicities. RESULTS: We show that correlations within proteins affect mainly the oscillations at distances below 35 bp. The long-ranging correlations up to 100 bp reflect primarily DNA folding. For the yeast genome these oscillations are consistent in detail with the chromatin structure. For eubacteria and archaea the periods deviate significantly from the 10.55 bp value for free DNA. These deviations suggest that while a period of 11 bp in bacteria reflects negative supercoiling, the significantly different period of thermophilic archaea close to 10 bp corresponds to positive supercoiling of thermophilic archaeal genomes. AVAILABILITY: Protein sets and C programs for the calculation of correlation functions are available on request from the authors (see http://itb.biologie.hu-berlin.de).

Archaea↗

Modeling the role of nonhuman vocal membranes in phonation.

Although the mammalian larynx exhibits little structural variation compared to sound-producing organs in other taxa (birds or insects), there are some morphological features which could lead to significant differences in acoustic functioning, such as air sacs and vocal membranes. The vocal membrane (or "vocal lip") is a thin upward extension of the vocal fold that is present in many bat and primate species. The vocal membrane was modeled as an additional geometrical element in a two-mass model of the larynx. It was found that vocal membranes of an optimal angle and length can substantially lower the subglottal pressure at which phonation is supported, thus increasing vocal efficiency, and that this effect is most pronounced at high frequencies. The implications of this finding are discussed for animals such as bats and primates which are able to produce loud, high-pitched calls. Modeling efforts such as this provide guidance for future empirical investigations of vocal membrane structure and function, can provide insight into the mechanisms of animal communication, and could potentially lead to better understanding of human clinical disorders such as sulcus vocalis.

Humans↗

Correlations in protein sequences and property codes.

Correlation functions in large sets of non-homologous protein sequences are analysed. Finite size corrections are applied and fluctuations are estimated. As symbol sequences have to be mapped to sequences of numbers to calculate correlation functions, several property codes are tested as such mappings. We found hydrophobicity autocorrelation functions to be strongly oscillating. Another strong signal is the monotonously decaying alpha-helix propensity autocorrelation function. Furthermore, we detected signals corresponding to an alteration of positively and negatively charged residues at a distance of 3-4 amino acids. To look beyond the property codes gained by the methods of physical chemistry, mappings yielding a strong correlation signal are sought for using a Monte Carlo simulation. The mappings leading to strong signals are found to be related to hydrophobicity of alpha-helix propensity. A cluster analysis of the top scoring mappings leads to two novel property codes. These two property codes are gained from sequence data only. They turn out to be similar to known property codes for hydrophobicity or polarity.

Amino Acids↗

Sequence periodicity in complete genomes of archaea suggests positive supercoiling.

The topological state of genomic DNA is of importance for its replication, recombination and transcription. The wrapping of the DNA around nucleosomes is associated with sequence periodicities (Trifonov and Sussman, Proc. Natl. Acad. Sci. USA, 77, pp. 3816-20). Recently, also the negative supercoiling of eubacterial DNA was related to 11 base pair (bp) periodicity (Herzel et al. Physica A, 249, pp. 449-59). Archaeal plasmids and a virus-like particle from Sulfolobus are positively supercoiled, but the superhelical conformation of archaeal genomic DNA is still uncertain. The problem of superhelicity can now be addressed via a comparative statistical analysis of the available complete genomes. For this purpose one has to look for periodicities which are in phase with the helical repeat of 10-11 bp. Similar periodicities are induced, however, by the amphipatic character of alpha-helices of encoded proteins (Zhurkin, Nucl. Acids Res., 9, pp. 1963-71). We show that these protein-induced periodicities are extended over a few periods only. The periods of additional long-ranging oscillations deviate significantly from the value for free DNA. A period of 11 bp in Eubacteria reflects negative supercoiling, whereas the significantly different period of thermophilic Archaea close to 10 bp suggests positive supercoiling of archaeal genomes.

DNA, Archaeal↗