PubMed Health⌕ Search

Biomedical subjects

L Brocchieri

Publications and source records attributed to L Brocchieri.

13 recordsLinked to original sources

Phylogenetic inferences from molecular sequences: review and critique.

Conflicting results often accompany phylogenetic analyses of RNA, DNA, or protein sequences across diverse species. Causes contributing to these conflicts relate to ambiguities in identifying homologous characters of alignments, sensitivity of tree-making methods to unequal evolutionary rates, biases in species sampling, unrecognized paralogy, functional differentiation, loss of phylogenetic informational content due to long branches or fast evolution, and difficulties with the assumptions and approximations used to infer phylogenetic relationships. Attempts to surmount these conflicts by averaging over many proteins are problematic due to inherent biases of selected families, lack of signal in others, and events of lateral transfer, fusion, and/or chimerism. The process of assessing reliability of the results using the bootstrap method is strewn with obstacles because of lack of independence and inhomogeneity in the molecular data. Problems inherent to the three major procedures for developing phylogenetic trees--parsimony, likelihood, distance--are reviewed. Special attention is given to the problem of inferring evolutionary distances from patterns of similarity among sequences. The difficulties encountered by methods of phylogenetic reconstructions based on the analysis of divergent sequence families make new methods based on the analysis of complete genomes reasonable alternatives. Several of these are considered, including the signature sequences of Gupta and associates, the study of genome profiles, and the genomic signature set forth by Karlin and colleagues.

Animals↗

Heat shock protein 60 sequence comparisons: duplications, lateral transfer, and mitochondrial evolution.

Heat shock proteins 60 (GroEL) are highly expressed essential proteins in eubacterial genomes and in eukaryotic organelles. These chaperone proteins have been advanced as propitious marker sequences for tracing the evolution of mitochondrial (Mt) genomes. Similarities among HSP60 sequences based on significant segment pair alignment calculations are used to deduce associations of sequences taking into account GroEL functional/structural domain differences and to relate HSP60 duplications pervasive in alpha-proteobacterial lineages to the dynamics of lateral transfer and plasmid integration. Multiple alignments with consensuses are determined for 10 natural groups. The group consensuses sharpen the similarity contrasts among individual sequences. In particular, the Mt group matches best with the classical alpha-proteobacteria and closely with Rickettsia but significantly worse with the rickettsial groups Ehrlichia and Orientia. However, across broad protein sequence comparisons, there appears to be no consistent prokaryote whose protein sequences align best with animal Mt genomes. There are plausible scenarios indicating that the nuclear-encoded HSP60 (and HSP70) sequences functioning in Mt are results of lateral transfer and are probably derived from an alpha-proteobacterium. This hypothesis relates to the plethora of duplicated HSP60 sequences among the classical alpha-proteobacteria contrasted with no duplications of HSP60 among other clades of proteobacterial genomes. Evolutionary relations are confounded by differential selection pressures, convergence, variable mutational rates, site variability, and lateral gene transfer.

Bacteria↗

Conservation among HSP60 sequences in relation to structure, function, and evolution.

The chaperonin HSP60 (GroEL) proteins are essential in eubacterial genomes and in eukaryotic organelles. Functional regions inferred from mutation studies and the Escherichia coli GroEL 3D crystal complexes are evaluated in a multiple alignment across 43 diverse HSP60 sequences, centering on ATP/ADP and Mg2+ binding sites, on residues interacting with substrate, on GroES contact positions, on interface regions between monomers and domains, and on residues important in allosteric conformational changes. The most evolutionary conserved residues relate to the ATP/ADP and Mg2+ binding sites. Hydrophobic residues that contribute in substrate binding are also significantly conserved. A large number of charged residues line the central cavity of the GroEL-GroES complex in the substrate-releasing conformation. These span statistically significant intra- and inter-monomer three-dimensional (3D) charge clusters that are highly conserved among sequences and presumably play an important role interacting with the substrate. Unaligned short segments between blocks of alignment are generally exposed at the outside wall of the Anfinsen cage complex. The multiple alignment reveals regions of divergence common to specific evolutionary groups. For example, rickettsial sequences diverge in the ATP/ADP binding domain and gram-positive sequences diverge in the allosteric transition domain. The evolutionary information of the multiple alignment proffers attractive sites for mutational studies.

Adenosine Triphosphate↗

A chimeric prokaryotic ancestry of mitochondria and primitive eukaryotes.

We provide data and analysis to support the hypothesis that the ancestor of animal mitochondria (Mt) and many primitive amitochondrial (a-Mt) eukaryotes was a fusion microbe composed of a Clostridium-like eubacterium and a Sulfolobus-like archaebacterium. The analysis is based on several observations: (i) The genome signatures (dinucleotide relative abundance values) of Clostridium and Sulfolobus are compatible (sufficiently similar) and each has significantly more similarity in genome signatures with animal Mt sequences than do all other available prokaryotes. That stable fusions may require compatibility in genome signatures is suggested by the compatibility of plasmids and hosts. (ii) The expanded energy metabolism of the fusion organism was strongly selective for cementing such a fusion. (iii) The molecular apparatus of endospore formation in Clostridium serves as raw material for the development of the nucleus and cytoplasm of the eukaryotic cell.

Amino Acid Sequence↗

A symmetric-iterated multiple alignment of protein sequences.

A new symmetric-iterative method for multiple alignment of protein sequences is presented. The method can be described as a combination of motif finding and dynamic programming procedures. It uses each sequence as a standard to which all sequences are aligned based on the significant segment pair alignment (SSPA) protocol. Sequences are further matched using a reduced scoring threshold to provide fillers and extensions between highly significant segment pair matches. The method produces alignment blocks that accommodate indels and are separated by variable-length unaligned segments. Construction of consensus sequences is iterative, assigning greater weights to more distantly related sequences. A consensus sequence and various measures of conservation at each aligned position can be used for comparisons between protein families, for data base searches, and for analysis of functional and evolutionary features. The method is illustrated on the extended family of prokaryotic and eukaryotic RecA-like sequences. The RecA-like sequences reveal extended alignments among eubacterial RecA and separately among eukaryotic/archaebacterial Rad51/RadA. Eleven conserved blocks are common to both groups, two of them encompassing the ATP-binding A and B-sites. Among the most conserved positions are glycine residues. For example, they occur twice as doublets putatively serving as hinge connections that provide opportunity for alternative structural conformations. Also several charged/polar residues are highly conserved, probably consequent upon the extensive intermonomer interactions in RecA/Rad51 filament formation and possibly relevant protein-protein and protein-nucleic acid interactions.

Algorithms↗

Heat shock protein 70 family: multiple sequence comparisons, function, and evolution.

The heat shock protein 70 kDa sequences (HSP70) are of great importance as molecular chaperones in protein folding and transport. They are abundant under conditions of cellular stress. They are highly conserved in all domains of life: Archaea, eubacteria, eukaryotes, and organelles (mitochondria, chloroplasts). A multiple alignment of a large collection of these sequences was obtained employing our symmetric-iterative ITERALIGN program (Brocchieri and Karlin 1998). Assessments of conservation are interpreted in evolutionary terms and with respect to functional implications. Many archaeal sequences (methanogens and halophiles) tend to align best with the Gram-positive sequences. These two groups also miss a signature segment [about 25 amino acids (aa) long] present in all other HSP70 species (Gupta and Golding 1993). We observed a second signature sequence of about 4 aa absent from all eukaryotic homologues, significantly aligned in all prokaryotic sequences. Consensus sequences were developed for eight groups [Archaea, Gram-positive, proteobacterial Gram-negative, singular bacteria, mitochondria, plastids, eukaryotic endoplasmic reticulum (ER) isoforms, eukaryotic cytoplasmic isoforms]. All group consensus comparisons tend to summarize better the alignments than do the individual sequence comparisons. The global individual consensus "matches" 87% with the consensus of consensuses sequence. A functional analysis of the global consensus identifies a (new) highly significant mixed charge cluster proximal to the carboxyl terminus of the sequence highlighting the hypercharge run EEDKKRRER (one-letter aa code used). The individual Archaea and Gram-positive sequences contain a corresponding significant mixed charge cluster in the location of the charge cluster of the consensus sequence. In contrast, the four Gram-negative proteobacterial sequences of the alignment do not have a charge cluster (even at the 5% significance level). All eukaryotic HSP70 sequences have the analogous charge cluster. Strikingly, several of the eukaryotic isoforms show multiple mixed charged clusters. These clusters were interpreted with supporting data related to HSP70 activity in facilitating chaperone, transport, and secretion function. We observed that the consensus contains only a single tryptophan residue and a single conserved cysteine. This is interpreted with respect to the target rule for disaggregating misfolded proteins. The mitochondrial HSP70 connections to bacterial HSP70 are analyzed, suggesting a polyphyletic split of Trypanosoma and Leishmania protist mitochondrial (Mt) homologues separated from Mt-animal/fungal/plant homologues. Moreover, the HSP70 sequences from the amitochondrial Entamoeba histolytica and Trichomonas vaginalis species were analyzed. The E. histolytica HSP70 is most similar to the higher eukaryotic cytoplasmic sequences, with significantly weaker alignments to ER sequences and much diminished matching to all eubacterial, mitochondrial, and chloroplast sequences. This appears to be at variance with the hypothesis that E. histolytica rather recently lost its mitochondrial organelle. T. vaginalis contains two HSP70 sequences, one Mt-like and the second similar to eukaryotic cytoplasmic sequences suggesting two diverse origins.

Amino Acid Sequence↗

Evolutionary comparisons of RecA-like proteins across all major kingdoms of living organisms.

Protein sequences with similarities to Escherichia coli RecA were compared across the major kingdoms of eubacteria, archaebacteria, and eukaryotes. The archaeal sequences branch monophyletically and are most closely related to the eukaryotic paralogous Rad51 and Dmc1 groups. A multiple alignment of the sequences suggests a modular structure of RecA-like proteins consisting of distinct segments, some of which are conserved only within subgroups of sequences. The eukaryotic and archaeal sequences share an N-terminal domain which may play a role in interactions with other factors and nucleic acids. Several positions in the alignment blocks are highly conserved within the eubacteria as one group and within the eukaryotes and archaebacteria as a second group, but compared between the groups these positions display nonconservative amino acid substitutions. Conservation within the RecA-like core domain identifies possible key residues involved in ATP-induced conformational changes. We propose that RecA-like proteins derive evolutionarily from an assortment of independent domains and that the functional homologs of RecA in noneubacteria comprise an array of RecA-like proteins acting in series or cooperatively.

Amino Acid Sequence↗

Evolutionary conservation of RecA genes in relation to protein structure and function.

Functional and structural regions inferred from the Escherichia coli R ecA protein crystal structure and mutation studies are evaluated in terms of evolutionary conservation across 63 RecA eubacterial sequences. Two paramount segments invariant in specific amino acids correspond to the ATP-binding A site and the functionally unassigned segment from residues 145 to 149 immediately carboxyl to the ATP hydrolysis B site. Not only are residues 145 to 149 conserved individually, but also all three-dimensional structural neighbors of these residues are invariant, strongly attesting to the functional or structural importance of this segment. The conservation of charged residues at the monomer-monomer interface, emphasizing basic residues on one surface and acidic residues on the other, suggests that RecA monomer polymerization is substantially mediated by electrostatic interactions. Different patterns of conservation also allow determination of regions proposed to interact with DNA, of LexA binding sites, and of filament-filament contact regions. Amino acid conservation is also compared with activities and properties of certain RecA protein mutants. Arginine 243 and its strongly cationic structural environment are proposed as the major site of competition for DNA and LexA binding to RecA. The conserved acidic and glycine residues of the disordered loop L1 and its proximity to the RecA acidic monomer interface suggest its involvement in monomer-monomer interactions rather than DNA binding. The conservation of various RecA positions and regions suggests a model for RecA-double-stranded DNA interaction and other functional and structural assignments.

Amino Acid Sequence↗

Pbx modulation of Hox homeodomain amino-terminal arms establishes different DNA-binding specificities across the Hox locus.

Pbx cofactors are implicated to play important roles in modulating the DNA-binding properties of heterologous homeodomain proteins, including class I Hox proteins. To assess how Pbx proteins influence Hox DNA-binding specificity, we used a binding-site selection approach to determine high-affinity target sites recognized by various Pbx-Hox homeoprotein complexes. Pbx-Hox heterodimers preferred to bind a bipartite sequence 5'-ATGATTNATNN-3' consisting of two adjacent half sites in which the Pbx component of the heterodimer contacted the 5' half (ATGAT) and the Hox component contacted the more variable 3' half (TNATNN). Binding sites matching the consensus were also obtained for Pbx1 complexed with HoxA10, which lacks a hexapeptide but requires a conserved tryptophan-containing motif for cooperativity with Pbx. Interactions with Pbx were found to play an essential role in modulating Hox homeodomain amino-terminal arm contact with DNA in the core of the Hox half site such that heterodimers of different compositions could distinguish single nucleotide alterations in the Hox half site both in vitro and in cellular assays measuring transactivation. When complexed with Pbx, Hox proteins B1 through B9 and A10 showed stepwise differences in their preferences for nucleotides in the Hox half site core (TTAT to TGAT, 5' to 3') that correlated with the locations of their respective genes in the Hox cluster. These observations demonstrate previously undetected DNA-binding specificity for the amino-terminal arm of the Hox homeodomain and suggest that different binding activities of Pbx-Hox complexes are at least part of the position-specific activities of the Hox genes.

Animals↗

How are close residues of protein structures distributed in primary sequence?

Structurally neighboring residues are categorized according to their separation in the primary sequence as proximal (1-4 positions apart) and otherwise distal, which in turn is divided into near (5-20 positions), far (21-50 positions), very far ( > 50 positions), and interchain (from different chains of the same structure). These categories describe the linear distance histogram (LDH) for three-dimensional neighboring residue types. Among the main results are the following: (i) nearest-neighbor hydrophobic residues tend to be increasingly distally separated in the linear sequence, thus most often connecting distinct secondary structure units. (ii) The LDHs of oppositely charged nearest-neighbors emphasize proximal positions with a subsidiary maximum for very far positions. (iii) Cysteine-cysteine structural interactions rarely involve proximal positions. (iv) The greatest numbers of interchain specific nearest-neighbors in protein structures are composed of oppositely charged residues. (v) The largest fraction of side-chain neighboring residues from beta-strands involves near positions, emphasizing associations between consecutive strands. (vi) Exposed residue pairs are predominantly located in proximal linear positions, while buried residue pairs principally correspond to far or very far distal positions. The results are principally invariant to protein sizes, amino acid usages, linear distance normalizations, and over- and underrepresentations among nearest-neighbor types. Interpretations and hypotheses concerning the LDHs, particularly those of hydrophobic and charged pairings, are discussed with respect to protein stability and functionality. The pronounced occurrence of oppositely charged interchain contacts is consistent with many observations on protein complexes where multichain stabilization is facilitated by electrostatic interactions.

Amino Acid Sequence↗

Geometry of interplanar residue contacts in protein structures.

The relative spatial disposition of interacting side-chain planar groups (aromatic, guanidinium, amide, carboxyl, imidazole) is analyzed for 186 non-homologous well-resolved protein structures. The dihedral angle of amide or carboxyl planar groups with other planar groups accords with a random distribution of planes. By contrast, the dihedral angle of the planes between close aromatic rings or of the histidine ring interacting with aromatic residues is significantly nonrandom, showing an approximately uniform distribution. Our results indicate that edge-to-edge and edge-to-center spatial dispositions of residue planar sections are prevalent, while complete stacking configurations are uncommon. The hypothesis that electrostatic forces are a major determinant of the geometry of interactions between side-chain planar groups is discussed.

Mathematics↗

Measuring residue associations in protein structures. Possible implications for protein folding.

We propose a number of distance measures between residues in protein structures based on average, minimum and maximum distances of all atom (backbone and side-chain) coordinates or with respect to side-chain atom coordinates only. The d1-distance (D1-distance) refers to the average distance between side-chain (backbone and side-chain) atoms of a residue pair in a given structure. The dm-distance (Dm-distance) refers to the minimum distance between side-chain atoms (non-trivial minimum distance between all atoms of a residue pair). For each distance measure, averaging and normalizing over representative protein structures, association values and closeness orderings for all amino acid types are determined. The expected associations of side-chain interactions between oppositely charged residues, among hydrophobic residues and of cysteine with cysteine are confirmed. Several surprising associations are observed relative to (1) the aromatic residues tyrosine and tryptophan, but not phenylalanine; (2) multiple histidine residues; (3) asymmetries of arginine versus lysine, aspartate versus glutamate, alanine versus glycine, and asparagine versus glutamine; (4) absence of correlations of alpha-carbon distances with side-chain distances. The all atoms D1-distance attractions are dominated by steric relationships, with glycine and alanine significantly close to all amino acids, whereas large residues are under-associated with all residue types. In contrast, for the closeness ordering corresponding to the minimum side-chain dm-distance, glycine and alanine are among the least associated. However, in the d1-distance alanine is significantly close to all hydrophobic residues with the exception of tryptophan. The dm-distance preferences display a pervasive attraction for tyrosine by almost all residue types, the prominence of tyrosine and tryptophan in cation-aromatic interactions, and the versatility of histidine in functionality. The principal findings suggest a new perspective on the early and intermediate stages of protein folding.

Amino Acid Sequence↗