PubMed Health⌕ Search

Biomedical subjects

M J Gibbs

Publications and source records attributed to M J Gibbs.

At least 19 recordsLinked to original sources

The variable codons of H3 influenza A virus haemagglutinin genes.

We have analyzed several sets of well-studied haemagglutinin (HA) gene sequences of H3 subtype influenza A viruses to identify codons that are unusually variable, using a simple pairwise sliding window method, DnDscanning. For two of the sets there were results of detailed phylogenetic modeling studies of selection already published. A third set had been the subject of an antigen mapping study, the results of which provide a completely independent benchmark of selected changes in H3 HA genes. Our analyses show that the codons with greatest DnDscan scores (i.e. the most variable) were mostly those reported in the published studies as being positively selected; indeed the DnDscan results matched the antigenic mapping results more closely than did those of the phylogenetic modeling methods. These results suggest that codons under selection can be found even when, as with some sets of virus sequences, a phylogeny is uncertain or cannot be obtained because, for example, the sequences are recombinants, or when selection is not necessarily linked with phylogeny, as in host-switching events. The program DnDscan is available at (biojanus.anu.edu.au).

Algorithms↗

A broader definition of 'the virus species'.

We propose that the formal definition of a virus species by the International Committee on Taxonomy of Viruses (ICTV) should be broadened by removing the restrictive word "polythetic" from the current definition, so that any characters can be used to define species. This change will bring the definition of virus species into line with the species definitions of cellular organisms and broaden the range of characters available for describing virus species.

Terminology as Topic↗

The phylogeny of SARS coronavirus.

Different tree-building methods consistently place the SARS corona-virus (SARS-CoV) as a basal Group 2 coronavirus rather than as an ungrouped species as concluded by others. Detailed comparisons of the SARS-CoV genomic sequence with those of six other coronaviruses failed to find evidence of recombination or genomic rearrangement using computational methods designed for that purpose.

Animals↗

A type of nucleotide motif that distinguishes tobamovirus species more efficiently than nucleotide signatures.

The complete genomic sequences of forty-eight tobamoviruses were classified and found to form at least twelve species clusters. Individual species were not conveniently defined by 'nucleotide signatures' (i.e. strings of one or more nucleotides unique to a taxon) as these were scattered sparsely throughout the genomes and were mostly single nucleotides. By contrast all the species were concisely and uniquely distinguished by short nucleotide motifs consisting of conserved genus-specific sites intercalated with variable sites that provided species-specific combinations of nucleotides (nucleotide combination motifs; NC-motifs). We describe the procedure for finding NC-motifs in a convenient and phylogenetically conserved region of the tobamovirus RNA polymerase gene, the '4404-50 motif'. NC-motifs have been found in other sets of homologous sequences, and are convenient for use in published taxonomic descriptions.

Conserved Sequence↗

The haemagglutinin gene, but not the neuraminidase gene, of 'Spanish flu' was a recombinant.

Published analyses of the sequences of three genes from the 1918 Spanish influenza virus have cast doubt on the theory that it came from birds immediately before the pandemic. They showed that the virus was of the H1N1 subtype lineage but more closely related to mammal-infecting strains than any known bird-infecting strain. They provided no evidence that the virus originated by gene reassortment nor that the virus was the direct ancestor of the two lineages of H1N1 viruses currently found in mammals; one that mostly infects human beings, the other pigs. The unusual virulence of the virus and why it produced a pandemic have remained unsolved. We have reanalysed the sequences of the three 1918 genes and found conflicting patterns of relatedness in all three. Various tests showed that the patterns in its haemagglutinin (HA) gene were produced by true recombination between two different parental HA H1 subtype genes, but that the conflicting patterns in its neuraminidase and non-structural-nuclear export proteins genes resulted from selection. The recombination event that produced the 1918 HA gene probably coincided with the start of the pandemic, and may have triggered it.

Animals↗

Recombination in the hemagglutinin gene of the 1918 "Spanish flu".

When gene sequences from the influenza virus that caused the 1918 pandemic were first compared with those of related viruses, they yielded few clues about its origins and virulence. Our reanalysis indicates that the hemagglutinin gene, a key virulence determinant, originated by recombination. The "globular domain" of the 1918 hemagglutinin protein was encoded by a part of a gene derived from a swine-lineage influenza, whereas the "stalk" was encoded by parts derived from a human-lineage influenza. Phylogenetic analyses showed that this recombination, which probably changed the virulence of the virus, occurred at the start of, or immediately before, the pandemic and thus may have triggered it.

Animals↗

Sister-scanning: a Monte Carlo procedure for assessing signals in recombinant sequences.

MOTIVATION: To devise a method that, unlike available methods, directly measures variations in phylogenetic signals in gene sequences that result from recombination, tests the significance of the signal variations and distinguishes misleading signals. RESULTS: We have developed a method, that we call 'sister-scanning', for assessing phylogenetic and compositional signals in the various patterns of identity that occur between four nucleotide sequences. A Monte Carlo randomization is done for all columns (positions) within a window and Z-scores are obtained for four real sequences or three real sequences with an outlier that is also randomized. The usefulness of the approach is demonstrated using tobamovirus and luteovirus sequences. Contradictory phylogenetic signals were distinguished in both datasets, as were regions of sequence that contained no clear signal or potentially misleading signals related to compositional similarities. In the tobamovirus dataset, contradictory phylogenetic signals were separated by coding sequences up to a kilobase long that contained no clear signal. Our re-analysis of this dataset using sister-scanning also yielded the first evidence known to us of an inter-species recombination site within a viral RNA-dependent RNA polymerase gene together with evidence of an unusual pattern of conservation in the three codon positions.

Computer Simulation↗

Phylogenetic analysis of some large double-stranded RNA replicons from plants suggests they evolved from a defective single-stranded RNA virus.

Sequences were recently obtained from four double-stranded (ds) RNAs from different plant species. These dsRNAs are not associated with particles and as they appeared not to be horizontally transmitted, they were thought to be a kind of RNA plasmid. Here we report that the RNA-dependent RNA polymerase (RdRp) and helicase domains encoded by these dsRNAs are related to those of viruses of the alpha-like virus supergroup. Recent work on the RdRp sequences of alpha-like viruses raised doubts about their relatedness, but our analyses confirm that almost all the viruses previously assigned to the supergroup are related. Alpha-like viruses have single-stranded (ss) RNA genomes and produce particles, and they are much more diverse than the dsRNAs. This difference in diversity suggests the ssRNA alpha-like virus form is older, and we speculate that the transformation to a dsRNA form began when an ancestral ssRNA virus lost its virion protein gene. The phylogeny of the dsRNAs indicates this transformation was not recent and features of the dsRNA genome structure and translation strategy suggest it is now irreversible. Our analyses also show some dsRNAs from distantly related plants are closely related, indicating they have not strictly co-speciated with their hosts. In view of the affinities of the dsRNAs, we believe they should be classified as viruses and we suggest they be recognized as members of a new virus genus (Endornavirus) and family (Endoviridae).

Defective Viruses↗

Sugarcane yellow leaf virus: a novel member of the Luteoviridae that probably arose by inter-species recombination.

The 5895 nucleotide long single-stranded RNA genome of Sugarcane yellow leaf virus Florida isolate (SCYLV-F) includes six major ORFs. All but the first of these are homologous to genes of known function encoded by viruses of the three newly defined genera in the LUTEOVIRIDAE: ('luteovirids'), i.e. poleroviruses, luccccteoviruses and the enamoviruses. SCYLV-F ORFs 1 and 2 are most closely related to their polerovirus counterparts, whereas SCYLV-F ORFs 3 and 4 are most closely related to counterparts in luteovirus genomes, and SCYLV-F ORF5 is most closely related to the read-through protein gene of the only known enamovirus. These differences in affinity result from inter-species recombination. Two recombination sites in the genome of SCYLV-F map to the same genomic locations as previously described recombinations involving other luteovirids. A fourth type of luteovirid, Soybean dwarf virus, has already been described. Our analyses indicate that SCYLV-F represents a distinct fifth type.

Base Sequence↗

Evidence that a plant virus switched hosts to infect a vertebrate and then recombined with a vertebrate-infecting virus.

There are several similarities between the small, circular, single-stranded-DNA genomes of circoviruses that infect vertebrates and the nanoviruses that infect plants. We analyzed circovirus and nanovirus replication initiator protein (Rep) sequences and confirmed that an N-terminal region in circovirus Reps is similar to an equivalent region in nanovirus Reps. However, we found that the remaining C-terminal region is related to an RNA-binding protein (protein 2C), encoded by picorna-like viruses, and we concluded that the sequence encoding this region of Rep was acquired from one of these single-stranded RNA viruses, probably a calicivirus, by recombination. This is clear evidence that a DNA virus has incorporated a gene from an RNA virus, and the fact that none of these viruses code for a reverse transcriptase suggests that another agent with this capacity was involved. Circoviruses were thought to be a sister-group of nanoviruses, but our phylogenetic analyses, which take account of the recombination, indicate that circoviruses evolved from a nanovirus. A nanovirus DNA was transferred from a plant to a vertebrate. This transferred DNA included the viral origin of replication; the sequence conservation clearly indicates that it maintained the ability to replicate. In view of these properties, we conclude that the transferred DNA was a kind of virus and the transfer was a host-switch. We speculate that this host-switch occurred when a vertebrate was exposed to sap from an infected plant. All characterized caliciviruses infect vertebrates, suggesting that the host-switch happened first and that the recombination took place in a vertebrate.

Amino Acid Sequence↗

Australian isolates of ryegrass mosaic rymovirus and their relationships.

The sequences of the 3'-terminal 1.8 kb of the genomes of three Australian and three Welsh isolates of ryegrass mosaic rymovirus (RGMV) were determined, as too were the virion protein genes of two New Zealand isolates of RGMV. They were compared with each other and with the published sequences of a Danish and a South African isolate by distance and maximum likelihood methods, and found to be very closely related (mean nucleotide difference 5.5%). All three Australian isolates and one from North Island of New Zealand formed one consistent cluster, and the Danish and South African isolates formed another. However the relationships between these two clusters and the other isolates were not consistent; they depended on the method of comparison used, and on the protein, gene or codon position compared. Nonetheless the European (Welsh and Danish) sequences were 2-4 times more different from one another than those from the Antipodes, suggesting that the European RGMV population may be older than the Antipodean. The Danish isolate has 39 nucleotides of the 5'-terminal region of its virion protein gene frameshifted -1 relative to the 'common' sequence. Interestingly the South African isolate has a similar frameshift, but sequence comparisons indicate that this frameshift must have occurred independently; a possible example of 'convergent frameshifting'.

Amino Acid Sequence↗

The genome organization and affinities of an Australian isolate of carrot mottle umbravirus.

The genomic sequence of an Australian isolate of carrot mottle umbravirus (CMoV-A) was determined from cDNA generated from dsRNA. This provides the first data on the genome organization and phylogeny of an umbravirus. The 4201-nucleotide genome contains four major open reading frames (ORFs). Analysis suggests that ORF2 encodes an RNA-dependent RNA polymerase, that ORF4 encodes a movement protein, and that the virus has no coat protein gene. The functions of ORFs 1 and 3 remain unknown. ORF2 is probably translated following ribosomal frameshifting. ORFs 3 and 4 are probably translated from a subgenomic mRNA. Sequence comparisons showed CMoV-A to be closely related to pea enation mosaic RNA2 (PEMV-RNA2), but also to have affinities with the Bromoviridae. These findings shed light on the relationships between the luteoviruses, PEMV, and the umbraviruses and on the relationships between the carmo-like viruses and the Bromoviridae.

Australia↗

A reevaluation of the higher taxonomy of viruses based on RNA polymerases.

In order to assess the validity of classifications of RNA viruses, published alignments and phylogenies of RNA-dependent RNA and DNA polymerase sequences were reevaluated by a Monte Carlo randomization procedure, bootstrap resampling, and phylogenetic signal analysis. Although clear relationships between some viral taxa were identified, overall the sequence similarities and phylogenetic signals were insufficient to support many of the proposed evolutionary groupings of RNA viruses. Likewise, no support for the common ancestry of RNA-dependent RNA polymerases and reverse transcriptases was found.

Animals↗

A recombinational event in the history of luteoviruses probably induced by base-pairing between the genomes of two distinct viruses.

Alignments of luteovirus readthrough protein amino acid sequences show they consist of two distinct regions, here named the N domain and the C domain. N domain sequences were classified, and comparison of this gene phylogeny to phylogenies of other luteovirus genes revealed an anomaly in the relationships between beet western yellows luteovirus, cucurbit aphidborne yellows luteovirus (CABYV), and pea enation mosaic RNA1 (PEMV1). Together with alignments of virion protein and readthrough protein amino acid sequences, these gene phylogenies indicate the anomaly to be the result of two recombinational events, probably between ancestors of CABYV and PEMV1 and leading to the transfer of RNA coding for the N domain to an ancestor of CABYV. Two likely recombination sites were identified from the alignments, one at the 5' end of the readthrough protein gene and the other at the 5' end of the sequence coding for the C domain. Alignments of the nucleotide sequences encompassing the probable recombination sites suggest that base-pairing between the genomes of the two ancestral luteoviruses, resulting from local sequence similarity at the 5' end of the readthrough protein gene, probably induced one of the interspecies recombinational events.

Amino Acid Sequence↗

Partial nucleotide sequence of poplar mosaic virus RNA confirms its classification as a carlavirus.

The nucleotide sequence of the 3'-proximal 1328 nucleotides of poplar mosaic virus (PMV) was determined and shown to contain two large open reading frames (ORFs). The ORF nearer to the 3' terminus of the RNA is capable of encoding a polypeptide of 14K with a 'zinc-finger' motif, and is homologous to sequences in corresponding positions in five other carlaviruses. The other ORF encodes a protein of 36K which includes two sequences of amino acids identified in tryptic digests as virion capsid protein, and has amino acid sequences in common with both carlaviruses and potexviruses.

Amino Acid Sequence↗

Comparisons of rbcL genes for the large subunit of ribulose-bisphosphate carboxylase from closely related C3 and C4 plant species.

Ribulosebisphosphate carboxylase/oxygenase from C4 plants exhibits higher turnover rates and lower affinities for CO2 than the enzyme from C3 plants or C3-C4 intermediate species. This property is shown to be inherited maternally in reciprocal interspecific crosses between two Flaveria species, and thus must be specified by the chloroplast-encoded large subunits. To investigate the amino acid changes responsible, the chloroplast rbcL genes from three pairs of C3 and C4 species from three genera (Flaveria, Atriplex, and Neurachne) were cloned and sequenced. Comparisons of the predicted amino acid sequences from species of the same genus revealed a limited number of changes within each pair, ranging from three to six, of which only one (309Met (C3) to Ile (C4] was consistently observed. This residue occurs in the loop connecting the carboxyl end of beta strand 5 with the amino end of alpha helix 5 in the alpha/beta barrel of the large subunit, and is close to the active site in a region which makes interdomain and intersubunit contacts. However, it is unlikely that a change of this residue alone is responsible for the alteration of kinetic properties. Nucleotide sequence comparisons of the rbcL genes showed no significant or consistent changes in the promoter and transcribed but nontranslated regions to suggest why rbcL is not expressed in C4 leaf mesophyll cells. It is concluded that mutations in rbcL have led to an alteration of the kinetics but not the expression of ribulose-bisphosphate carboxylase.

Amino Acid Sequence↗