PubMed Health⌕ Search

Biomedical subjects

Aaron J Mackey

Publications and source records attributed to Aaron J Mackey.

12 recordsLinked to original sources

The genome of the sea urchin Strongylocentrotus purpuratus.

We report the sequence and analysis of the 814-megabase genome of the sea urchin Strongylocentrotus purpuratus, a model for developmental and systems biology. The sequencing strategy combined whole-genome shotgun and bacterial artificial chromosome (BAC) sequences. This use of BAC clones, aided by a pooling strategy, overcame difficulties associated with high heterozygosity of the genome. The genome encodes about 23,300 genes, including many previously thought to be vertebrate innovations or known only outside the deuterostomes. This echinoderm genome provides an evolutionary outgroup for the chordates and yields insights into the evolution of deuterostomes.

Animals↗

Common inheritance of chromosome Ia associated with clonal expansion of Toxoplasma gondii.

Toxoplasma gondii is a globally distributed protozoan parasite that can infect virtually all warm-blooded animals and humans. Despite the existence of a sexual phase in the life cycle, T. gondii has an unusual population structure dominated by three clonal lineages that predominate in North America and Europe, (Types I, II, and III). These lineages were founded by common ancestors approximately10,000 yr ago. The recent origin and widespread distribution of the clonal lineages is attributed to the circumvention of the sexual cycle by a new mode of transmission-asexual transmission between intermediate hosts. Asexual transmission appears to be multigenic and although the specific genes mediating this trait are unknown, it is predicted that all members of the clonal lineages should share the same alleles. Genetic mapping studies suggested that chromosome Ia was unusually monomorphic compared with the rest of the genome. To investigate this further, we sequenced chromosome Ia and chromosome Ib in the Type I strain, RH, and the Type II strain, ME49. Comparative genome analyses of the two chromosomal sequences revealed that the same copy of chromosome Ia was inherited in each lineage, whereas chromosome Ib maintained the same high frequency of between-strain polymorphism as the rest of the genome. Sampling of chromosome Ia sequence in seven additional representative strains from the three clonal lineages supports a monomorphic inheritance, which is unique within the genome. Taken together, our observations implicate a specific combination of alleles on chromosome Ia in the recent origin and widespread success of the clonal lineages of T. gondii.

Animals↗

SynView: a GBrowse-compatible approach to visualizing comparative genome data.

UNLABELLED: We present SynView, a simple and generic approach to dynamically visualize multi-species comparative genome data. It is a light-weight application based on the popular and configurable web-based GBrowse framework. It can be used with a variety of databases and provides the user with a high degree of interactivity. The tool is written in Perl and runs on top of the GBrowse framework. It is in use in the PlasmoDB (http://www.PlasmoDB.org) and the CryptoDB (http://www.CryptoDB.org) projects and can be easily integrated into other cross-species comparative genome projects. AVAILABILITY: The program and instructions are freely available at http://www.ApiDB.org/apps/SynView/ CONTACT: jkissing@uga.edu.

Algorithms↗

OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.

The OrthoMCL database (http://orthomcl.cbil.upenn.edu) houses ortholog group predictions for 55 species, including 16 bacterial and 4 archaeal genomes representing phylogenetically diverse lineages, and most currently available complete eukaryotic genomes: 24 unikonts (12 animals, 9 fungi, microsporidium, Dictyostelium, Entamoeba), 4 plants/algae and 7 apicomplexan parasites. OrthoMCL software was used to cluster proteins based on sequence similarity, using an all-against-all BLAST search of each species' proteome, followed by normalization of inter-species differences, and Markov clustering. A total of 511,797 proteins (81.6% of the total dataset) were clustered into 70,388 ortholog groups. The ortholog database may be queried based on protein or group accession numbers, keyword descriptions or BLAST similarity. Ortholog groups exhibiting specific phyletic patterns may also be identified, using either a graphical interface or a text-based Phyletic Pattern Expression grammar. Information for ortholog groups includes the phyletic profile, the list of member proteins and a multiple sequence alignment, a statistical summary and graphical view of similarities, and a graphical representation of domain architecture. OrthoMCL software, the entire FASTA dataset employed and clustering results are available for download. OrthoMCL-DB provides a centralized warehouse for orthology prediction among multiple species, and will be updated and expanded as additional genome sequence data become available.

Animals↗

The transcriptome of Toxoplasma gondii.

BACKGROUND: Toxoplasma gondii gives rise to toxoplasmosis, among the most prevalent parasitic diseases of animals and man. Transformation of the tachzyoite stage into the latent bradyzoite-cyst form underlies chronic disease and leads to a lifetime risk of recrudescence in individuals whose immune system becomes compromised. Given the importance of tissue cyst formation, there has been intensive focus on the development of methods to study bradyzoite differentiation, although the molecular basis for the developmental switch is still largely unknown. RESULTS: We have used serial analysis of gene expression (SAGE) to define the Toxoplasma gondii transcriptome of the intermediate-host life cycle that leads to the formation of the bradyzoite/tissue cyst. A broad view of gene expression is provided by >4-fold coverage from nine distinct libraries (approximately 300,000 SAGE tags) representing key developmental transitions in primary parasite populations and in laboratory strains representing the three canonical genotypes. SAGE tags, and their corresponding mRNAs, were analyzed with respect to abundance, uniqueness, and antisense/sense polarity and chromosome distribution and developmental specificity. CONCLUSION: This study demonstrates that phenotypic transitions during parasite development were marked by unique stage-specific mRNAs that accounted for 18% of the total SAGE tags and varied from 1-5% of the tags in each developmental stage. We have also found that Toxoplasma mRNA pools have a unique parasite-specific composition with 1 in 5 transcripts encoding Apicomplexa-specific genes functioning in parasite invasion and transmission. Developmentally co-regulated genes were dispersed across all Toxoplasma chromosomes, as were tags representing each abundance class, and a variety of biochemical pathways indicating that trans-acting mechanisms likely control gene expression in this parasite. We observed distinct similarities in the specificity and expression levels of mRNAs in primary populations (Day-6 post-sporozoite infection) that occur prior to the onset of bradyzoite development that were uniquely shared with the virulent Type I-RH laboratory strain suggesting that development of RH may be arrested. By contrast, strains from Type II-Me49B7 and Type III-VEGmsj contain SAGE tags corresponding to bradyzoite genes, which suggests that priming of developmental expression likely plays a role in the greater capacity of these strains to complete bradyzoite development.

Animals↗

Composite genome map and recombination parameters derived from three archetypal lineages of Toxoplasma gondii.

Toxoplasma gondii is a highly successful protozoan parasite in the phylum Apicomplexa, which contains numerous animal and human pathogens. T.gondii is amenable to cellular, biochemical, molecular and genetic studies, making it a model for the biology of this important group of parasites. To facilitate forward genetic analysis, we have developed a high-resolution genetic linkage map for T.gondii. The genetic map was used to assemble the scaffolds from a 10X shotgun whole genome sequence, thus defining 14 chromosomes with markers spaced at approximately 300 kb intervals across the genome. Fourteen chromosomes were identified comprising a total genetic size of approximately 592 cM and an average map unit of approximately 104 kb/cM. Analysis of the genetic parameters in T.gondii revealed a high frequency of closely adjacent, apparent double crossover events that may represent gene conversions. In addition, we detected large regions of genetic homogeneity among the archetypal clonal lineages, reflecting the relatively few genetic outbreeding events that have occurred since their recent origin. Despite these unusual features, linkage analysis proved to be effective in mapping the loci determining several drug resistances. The resulting genome map provides a framework for analysis of complex traits such as virulence and transmission, and for comparative population genetic studies.

Animals↗

The genome of Cryptosporidium hominis.

Cryptosporidium species cause acute gastroenteritis and diarrhoea worldwide. They are members of the Apicomplexa--protozoan pathogens that invade host cells by using a specialized apical complex and are usually transmitted by an invertebrate vector or intermediate host. In contrast to other Apicomplexans, Cryptosporidium is transmitted by ingestion of oocysts and completes its life cycle in a single host. No therapy is available, and control focuses on eliminating oocysts in water supplies. Two species, C. hominis and C. parvum, which differ in host range, genotype and pathogenicity, are most relevant to humans. C. hominis is restricted to humans, whereas C. parvum also infects other mammals. Here we describe the eight-chromosome approximately 9.2-million-base genome of C. hominis. The complement of C. hominis protein-coding genes shows a striking concordance with the requirements imposed by the environmental niches the parasite inhabits. Energy metabolism is largely from glycolysis. Both aerobic and anaerobic metabolisms are available, the former requiring an alternative electron transport system in a simplified mitochondrion. Biosynthesis capabilities are limited, explaining an extensive array of transporters. Evidence of an apicoplast is absent, but genes associated with apical complex organelles are present. C. hominis and C. parvum exhibit very similar gene complements, and phenotypic differences between these parasites must be due to subtle sequence divergence.

Animals↗

CRP: Cleavage of Radiolabeled Phosphoproteins.

The CRP (Cleavage of Radiolabeled Phosphoproteins) program guides the design and interpretation of experiments to identify protein phosphorylation sites by Edman sequencing of unseparated peptides. Traditionally, phosphorylation sites are determined by cleaving the phosphoprotein and separating the peptides for Edman 32P-phosphate release sequencing. CRP analysis of a phosphoprotein's sequence accelerates this process by omitting the separation step: given a protein sequence of interest, the CRP program performs an in silico proteolytic cleavage of the sequence and reports the predicted Edman cycles in which radioactivity would be observed if a given serine, threonine or tyrosine were phosphorylated. Experimentally observed cycles containing 32P can be compared with CRP predictions to confirm candidate sites and/or explore the ability of additional cleavage experiments to resolve remaining ambiguities. To reduce ambiguity, the phosphorylated residue (P-Tyr, P-Ser or P-Thr) can be determined experimentally, and CRP will ignore sites with alternative residues. CRP also provides simple predictions of likely phosphorylation sites using known kinase recognition motifs. The CRP interface is available at http://fasta.bioch.virginia.edu/crp.

Humans↗

Identification of residues in glutathione transferase capable of driving functional diversification in evolution. A novel approach to protein redesign.

Evolution of protein function can be driven by positive selection of advantageous nonsynonymous codon mutations that arise following gene duplication. By observing the presence and degree of site-specific positive selection for change between divergent paralogs, residue positions responsible for functional changes can be identified. We applied this analysis to genes encoding Mu class glutathione transferases, which differ widely in substrate specificities. Approximately 3% of the amino acid residue positions, both near to and distant from the active site, are under statistically significant positive selection for change. Relevant human glutathione transferase (GST) M1-1 and GST M2-2 codons were mutated. A chemically conservative threonine to serine mutation in GST M2-2 elicited a 1,000-fold increase in specific activity with the GST M1-1-specific substrate trans-stilbene oxide and a 30-fold increase with the alternative epoxide substrates styrene oxide and nitrophenyl glycidol. The reverse mutation in GST M1-1 resulted in reciprocal decreases in activity. Thus, identification of hypervariable codon positions can be a powerful aid in the redesign of protein function, lessening the requirement for extensive mutagenesis or structural knowledge and sometimes suggesting mutations that would otherwise be considered functionally conservative.

Evolution, Molecular↗

Getting more from less: algorithms for rapid protein identification with multiple short peptide sequences.

We describe two novel sequence similarity search algorithms, FASTS and FASTF, that use multiple short peptide sequences to identify homologous sequences in protein or DNA databases. FASTS searches with peptide sequences of unknown order, as obtained by mass spectrometry-based sequencing, evaluating all possible arrangements of the peptides. FASTF searches with mixed peptide sequences, as generated by Edman sequencing of unseparated mixtures of peptides. FASTF deconvolutes the mixture, using a greedy heuristic that allows rapid identification of high scoring alignments while reducing the total number of explored alternatives. Both algorithms use the heuristic FASTA comparison strategy to accelerate the search but use alignment probability, rather than similarity score, as the criterion for alignment optimality. Statistical estimates are calculated using an empirical correction to a theoretical probability. These calculated estimates were accurate within a factor of 10 for FASTS and 1000 for FASTF on our test dataset. FASTS requires only 15-20 total residues in three or four peptides to robustly identify homologues sharing 50% or greater protein sequence identity. FASTF requires about 25% more sequence data than FASTS for equivalent sensitivity, but additional sequence data are usually available from mixed Edman experiments. Thus, both algorithms can identify homologues that diverged 100 to 500 million years ago, allowing proteomic identification from organisms whose genomes have not been sequenced.

Algorithms↗

A strategy for the rapid identification of phosphorylation sites in the phosphoproteome.

Edman phosphate ((32)P) release sequencing provides a high sensitivity means of identifying phosphorylation sites in proteins that complements mass spectrometry techniques. We have developed a bioinformatic assessment tool, the cleavage of radiolabeled protein (CRP) program, which enables experimental identification of phosphorylation sites via (32)P labeling and Edman degradation of cleaved proteins obtained at femtomole levels. By observing the Edman cycle(s) in which radioactivity is found, candidate phosphorylation sites are identified by determining which residues occur at the observed number of cycles downstream from a peptide cleavage site. In cases where more than one residue could be responsible for the observed radioactivity, additional experiments with cleavage reagents having alternative specificities may resolve the ambiguity. Given a protein sequence and a cleavage site, CRP performs these experiments in silico, identifying resolved sites based on user-supplied experimental data, as well as suggesting combinations of reagents for additional analyses. Analysis of the PhosphoBase protein sequence database suggests that CRP data from two cleavage experiments can be used to identify unambiguously 60% of known phosphorylation sites. Data from additional cleavage experiments may increase the overall coverage to 70% of known sites. By comparing theoretical data obtained from the CRP program with (32)P release data obtained from an Edman sequencer, a known phosphorylation site was identified unambiguously and correctly. In addition, our results show that in vivo phosphorylation sites can be determined routinely by differential proteolysis analysis and Edman cycling with less than 1 fmol of protein and 1000 cpm.

Amino Acid Sequence↗

Entamoeba histolytica: sequence conservation of the Gal/GalNAc lectin from clinical isolates.

The Gal/GalNAc lectin gene of Entamoeba histolytica is a major amebic virulence protein responsible for interaction with host tissues. We investigated sequence differences in the Gal/GalNAc lectin heavy subunit in three isolates from Bangladesh and one isolate from Georgia, each of which was determined to be genetically distinct by SREHP AluI digestion. Interestingly, we observed only slight genetic diversity in the lectin gene as compared with the HM1:IMSS laboratory strain, originally a clinical isolate from Mexico. Genetic conservation of the Gal/GalNAc lectin between isolates may reflect that the lectin is under strong functional selection or possibly, that E. histolytica is a clonal population. Sequence conservation of the lectin indicates that immune responses against it should be cross-protective.

Amino Acid Sequence↗