PubMed Health⌕ Search

Biomedical subjects

Héctor Romero

Publications and source records attributed to Héctor Romero.

13 recordsLinked to original sources

Genomic GC level, optimal growth temperature, and genome size in prokaryotes.

Two years ago, we showed that positive correlations between optimal growth temperature (T(opt)) and genome GC are observed in 15 out of the 20 families of prokaryotes we analyzed, thus indicating that "T(opt) is one of the factors that influence genomic GC in prokaryotes". Our results were disputed, but these criticisms were demonstrated to be mistaken and based on misconceptions. In a recent report, Wang et al. [H.C. Wang, E. Susko, A.J. Roger, On the correlation between genomic G+C content and optimal growth temperature in prokaryotes: data quality and confounding factors, Biochem. Biophys. Res. Commun. 342 (2006) 681-684] criticize our results by stating that "all previous simple correlation analyses of GC versus temperature have ignored the fact that genomic GC content is influenced by multiple factors including both intrinsic mutational bias and extrinsic environmental factors". This statement, besides being erroneous, is surprising because it applies in fact not to ours but to the authors' article. Here, we rebut the points raised by Wang et al. and review some issues that have been a matter of debate, regarding the influence of environmental factors upon GC content in prokaryotes. Furthermore, we demonstrate that the relationship that exists between genome size and GC level is valid for aerobic, facultative, and microaerophilic species, but not for anaerobic prokaryotes.

Base Composition↗

Inferring parameters shaping amino acid usage in prokaryotic genomes via Bayesian MCMC methods.

Molar content of guanine plus cytosine (G + C) and optimal growth temperature (OGT) are main factors characterizing the frequency distribution of amino acids in prokaryotes. Previous work, using multivariate exploratory methods, has emphasized ascertainment of biological factors underlying variability between genomes, but the strength of each identified factor on amino acid content has not been quantified. We combine the flexibility of the phylogenetic mixed model (PMM) with the power of Bayesian inference via Markov Chain Monte Carlo (MCMC) methods, to obtain a novel evolutionary picture of amino acid usage in prokaryotic genomes. We implement a Bayesian PMM which incorporates the feature that evolutionary history makes observed data interdependent. As in previous studies with PMM, we present a variance partition; however, attention is also given to the posterior distribution of "systematic effects" that may shed light about the relative importance of and relationships between evolutionary forces acting at the genomic level. In particular, we analyzed influences of G + C, OGT, and respiratory metabolism. Estimates of G + C effects were significant for amino acids coded by G + C or molar content of adenine plus thymine (A + T) in first and second bases. OGT had an important effect on 12 amino acids, probably reflecting complex patterns of protein modifications, to cope with varying environments. The effect of respiratory metabolism was less clear, probably due to the already reported association of G + C with aerobic metabolism. A "heritability" parameter was always high and significant, reinforcing the importance of accommodating phylogenetic relationships in these analyses. "Heritable" component correlations displayed a pattern that tended to cluster "pure" G + C (A + T) in first and second codon positions, suggesting an inherited departure from linear regression on G + C.

Amino Acids↗

Genomic GC content prediction in prokaryotes from a sample of genes.

GC level is a key feature in prokaryotic genomes. Widely employed in evolutionary studies, new insights appear however limited because of the relatively low number of characterized genomes. Since public databases mainly comprise several hundreds of prokaryotes with a low number of sequences per genome, a reliable prediction method based on available sequences may be useful for studies that need a trustworthy estimation of whole genomic GC. As the analysis of completely sequenced genomes shows a great variability in distributional shapes, it is of interest to compare different estimators. Our analysis shows that the mean of GC values of a random sample of genes is a reasonable estimator, based on simplicity of the calculation and overall performance. However, usually sequences come from a process that cannot be considered as random sampling. When we analyzed two introduced sources of bias (gene length and protein functional categories) we were able to detect an additional bias in the estimation for some cases, although the precision was not affected. We conclude that the mean genic GC level of a sample of 10 genes is a reliable estimator of genomic GC content, showing comparable accuracy with many widely employed experimental methods.

Base Composition↗

Evolution of selenium utilization traits.

BACKGROUND: The essential trace element selenium is used in a wide variety of biological processes. Selenocysteine (Sec), the 21st amino acid, is co-translationally incorporated into a restricted set of proteins. It is encoded by an UGA codon with the help of tRNASec (SelC), Sec-specific elongation factor (SelB) and a cis-acting mRNA structure (SECIS element). In addition, Sec synthase (SelA) and selenophosphate synthetase (SelD) are involved in the biosynthesis of Sec on the tRNASec. Selenium is also found in the form of 2-selenouridine, a modified base present in the wobble position of certain tRNAs, whose synthesis is catalyzed by YbbB using selenophosphate as a precursor. RESULTS: We analyzed completely sequenced genomes for occurrence of the selA, B, C, D and ybbB genes. We found that selB and selC are gene signatures for the Sec-decoding trait. However, selD is also present in organisms that do not utilize Sec, and shows association with either selA, B, C and/or ybbB. Thus, selD defines the overall selenium utilization. A global species map of Sec-decoding and 2-selenouridine synthesis traits is provided based on the presence/absence pattern of selenium-utilization genes. The phylogenies of these genes were inferred and compared to organismal phylogenies, which identified horizontal gene transfer (HGT) events involving both traits. CONCLUSION: These results provide evidence for the ancient origin of these traits, their independent maintenance, and a highly dynamic evolutionary process that can be explained as the result of speciation, differential gene loss and HGT. The latter demonstrated that the loss of these traits is not irreversible as previously thought.

Evolution, Molecular↗

Correspondence analysis of amino acid usage within the family Bacillaceae.

When the amino acid usage of all completely sequenced prokaryotes is studied by multivariate analysis (MVA), it is known that the genomic molar content of guanine plus cytosine (GC) and optimal growth temperature (Topt) have a dominant effect. Furthermore, these two factors are associated to the first two axes of different MVA, and thus, nearly independent among them. However, it was recently shown that for several Families of prokaryotes there are significant and positive correlations between GC and Topt. This trend is particularly clear within Bacillaceae, where there are species displaying a broad range of variations for these two factors. In this paper we report that (a) Topt and genomic GC are the main factors shaping amino acid usage but are not independent between them, (b) the usage of cysteine is the second source of variability, and finally (c) the global hydrophobicity of the encoded proteins of each species is the third main factor.

Amino Acids↗

Correlations between genomic GC levels and optimal growth temperatures in prokaryotes.

In prokaryotes, GC levels range from 25% to 75%, and Topt from approximately 0 degrees C to >100 degrees C. When all species are considered together, no correlation is found between the two variables. Correlations are found, however, when Families of prokaryotes are analysed. Indeed, when Families comprising at least 10 species were studied (a set of 20 Families), positive correlations are found for 15 of them. Furthermore, a comparative analysis by independent contrasts made within the Families in order to control for phylogenetic non-independence showed qualitatively equivalent results. We conclude that Topt is one of the factors that influences genomic GC in prokaryotes.

Bacteria↗

Evidence of intratypic recombination in natural populations of hepatitis C virus.

Hepatitis C virus (HCV) has high genomic variability and, since its discovery, at least six different types and an increasing number of subtypes have been reported. Genotype 1 is the most prevalent genotype found in South America. In the present study, three different genomic regions (5'UTR, core and NS5B) of four HCV strains isolated from Peruvian patients were sequenced in order to investigate the congruence of HCV genotyping for these three genomic regions. Phylogenetic analysis using 5'UTR-core sequences found strain PE22 to be related to subtype 1b. However, the same analysis using the NS5B region found it to be related to subtype 1a. To test the possibility of genetic recombination, phylogenetic studies were carried out, revealing that a crossover event had taken place in the NS5B protein. We discuss the consequences of this observation on HCV genotype classification, laboratory diagnosis and treatment of HCV infection.

5' Untranslated Regions↗

The strength of translational selection for codon usage varies in the three replicons of Sinorhizobium meliloti.

The genome of the nitrogen-fixing bacterium Sinorhizobium meliloti is composed of three replicons of 3.65 (chromosome), 1.35 (pSymA) and 1.68 Mb (pSymB), respectively. While the chromosome encodes for most of the housekeeping functions, the three elements may contribute to symbiosis, though pSymA is absolutely necessary for nodulation and nitrogen fixation, since it harbours all the characterized nodulation and symbiotic fixation genes. On the other hand, the majority of the sequences located in this megaplasmid are probably not expressed during the free-living stage of the organism. Since most of the sequences located in pSymA are transcribed only at the stage of bacteroids when most probably the fate of the bacterium is to die, the mutations occurring at this stage will not be fixed in the population. Therefore, if natural selection contributes to the codon usage pattern in this species, its effect will be much weaker for the genes placed in pSymA. A codon usage analysis of the genes comprising the three replicons is consistent with the conclusion that selection for translational speed shapes the codon usage of the two replicons which are important for competitive cell growth while the codon usage of the third replicon reflects primarily the mutational bias.

Base Sequence↗

The influence of translational selection on codon usage in fishes from the family Cyprinidae.

In this paper, the main factors shaping codon usage in three species of fishes that belong to the family Cyprinidae (namely Brachidanio rerio, Cyprinus carpio, and Carassius auratus) are reported. Correspondence analysis (COA), a commonly used multivariate statistical approach, was used to analyze codon usage bias. Our results show that the main trend is strongly correlated with the GC(3) content at silent sites of each sequence. On the other hand, the second axis discriminates between presumed highly and lowly expressed genes, a result that is confirmed by the distribution of matching expressed sequence tags (ESTs) along that axis. Translational selection appears, therefore, to influence synonymous codon usage in these fishes. The comparison of codon usages of the sequences displaying the extreme values on the second axis indicates that several codons are significantly incremented among the heavily expressed sequences. Interestingly, several of these triplets are not only shared by the three fishes but also by Xenopus laevis, another cold-blooded vertebrate in which translational selection influences codon choices. We postulate that natural selection was operative for codon usage in the last common ancestor of these fishes and Xenopus, and will probably be detected in cold-blooded vertebrates in general. Finally, we raise the possibility that the same phenomena will be found among warm-blooded vertebrates.

Amino Acids↗

Translational selection is operative for synonymous codon usage in Clostridium perfringens and Clostridium acetobutylicum.

Here, the codon usage patterns of two Clostridium species (Clostridium perfringens and Clostridium acetobutylicum) are reported. These prokaryotes are characterized by a strong mutational bias towards A+T, a striking excess of coding sequences and purine-rich leading strands of replication, strong GC-skews and a high frequency of genomic rearrangements. As expected, it was found that the mutational bias dominates codon usage but there is some variation of synonymous codon choices among genes in the two species. This variation was investigated using a multivariate statistical approach. In the two species, two major trends were detected. One was related to the location of the sequences in the leading or lagging strand of replication, and the other was associated with the preferential use of putatively translational optimal codons in heavily expressed genes. Analyses of the estimated number of synonymous and non-synonymous substitutions among orthologous genes permit us to postulate that optimal codons might be selected not only for speed but also for accuracy during translation.

Amino Acids↗

Trends in codon and amino acid usage in Thermotoga maritima.

The usage of synonymous codons and the frequencies of amino acids were investigated in the complete genome of the bacterium Thermotoga maritima using a multivariate statistical approach. The GC3 content of each gene was the most prominent source of variation of codon usage. Surprisingly the usage of UGU and UGC (synonymous triplets coding for Cys, the least frequent amino acid in this species) was detected as the second most prominent source of variation. However, this result is probably an artifact due to the very low frequency of Cys together with the nonbiased composition of this genome. The third trend was related to the preferential usage of a subset of codons among highly expressed genes, and these triplets are presumed to be translationally optimal. Concerning the amino acid usage, the hydropathy level of each protein (and therefore the frequency of charged residues) was the main trend, while the second factor was related to the frequency of usage of the smaller residues, suggesting that the cell economy strongly influences the architecture of the proteins. The third axis of the analysis discriminated the usage of Phe, Tyr, Trp (aromatic residues) plus Cys, Met, and His. These six residues have in common the property of being the preferential targets of reactive oxygen species, and therefore the anaerobic condition of T. maritima is an important factor for the amino acid frequencies. Finally, the Cys content of each protein was the fourth trend.

Amino Acids↗

Aerobiosis increases the genomic guanine plus cytosine content (GC%) in prokaryotes.

The huge variation in the genomic guanine plus cytosine content (GC%) among prokaryotes has been explained by two mutually exclusive hypotheses, namely, selectionist and neutralist. The former proposals have in common the assumption that this feature is a form of adaptation to some ecological or physiological condition. On the other hand, the neutralist interpretation states that the variations are due only to different mutational biases. Since all of the traits that have been proposed by the selectionists either appeared to be limited to certain genera or were invalidated by the availability of more data, they cannot be considered as a selective force influencing the genomic GC% across all prokaryotes. In this report we show that aerobic prokaryotes display a significant increment in genomic GC% in relation to anaerobic ones. This is the first time that a link between a metabolic character and GC% has been found, independently of phylogenetic relationships and with a statistically significant amount of data.

Aerobiosis↗

Molecular evolution of hepatitis A virus: a new classification based on the complete VP1 protein.

Hepatitis A virus (HAV) is a positive-stranded RNA virus in the genus Hepatovirus in the family Picornaviridae So far, analysis of the genetic variability of HAV has been based on two discrete regions, the VP1/2A junction and the VP1 N terminus. In this report, we determined the nucleotide and deduced amino acid sequences of the complete VP1 gene of 81 strains from France, Kosovo, Mexico, Argentina, Chile, and Uruguay and compared them with the sequences of seven strains of HAV isolated elsewhere. Overall strain variation in the complete VP1 gene was found to be as high as 23.7% at the nucleotide level and 10.5% at the amino acid level. Different phylogenetic methods revealed that HAV sequences form five distinct and well-supported genetic lineages. Within these lineages, HAV sequences clustered by geographical origin only for European strains. The analysis of the complete VP1 gene allowed insight into the mode of evolution of HAV and revealed the emergence of a novel variant with a 15-amino-acid deletion located on the VP1 region where neutralization escape mutations were found. This could be the first antigenic variant of HAV so far identified.

Amino Acid Sequence↗