PubMed Health⌕ Search

Biomedical subjects

Brian J Haas

Publications and source records attributed to Brian J Haas.

At least 19 recordsLinked to original sources

Serum, cell-free, HPV-human DNA junction detection and HPV typing for predicting and monitoring cervical cancer recurrence.

Almost all cervical cancers are caused by human papillomaviruses (HPVs). In most cases, HPV DNA is integrated into the human genome. We found that tumor-specific, HPV-human DNA junctions are detectable in serum cell-free DNA of a fraction of cervical cancer patients at the time of initial treatment and/or at 6 months following treatment. Retrospective analysis revealed these junctions were more frequently detectable in women in whom the cancer later recurred. We also found that cervical cancers caused by HPV types outside of phylogenetic clade α9 had a higher recurrence frequency than those caused by α9 types in both our study and The Cancer Genome Atlas cervical cancer database, despite the higher prevalence ofα9 types, including HPV16, in cervical cancer. Thus, HPV-human DNA junction detection in serum cell-free DNA and HPV type determination in tumor tissue may help predict recurrence risk. Screening serum cell-free DNA for junctions may also offer an unambiguous non-invasive means to monitor absence of recurrence following treatment.

Humans↗

Serum, Cell-Free, HPV-Human DNA Junction Detection and HPV Typing for Predicting and Monitoring Cervical Cancer Recurrence.

Almost all cervical cancers are caused by human papillomaviruses (HPVs). In most cases, HPV DNA is integrated into the human genome. We found that tumor-specific, HPV-human DNA junctions are detectable in serum cell-free DNA of a fraction of cervical cancer patients at the time of initial treatment and/or at six months following treatment. Retrospective analysis revealed these junctions were more frequently detectable in women in whom the cancer later recurred. We also found that cervical cancers caused by HPV types outside of phylogenetic clade α9 had a higher recurrence frequency than those caused by α9 types in both our study and The Cancer Genome Atlas cervical cancer database, despite the higher prevalence of α9 types including HPV16 in cervical cancer. Thus, HPV-human DNA junction detection in serum cell-free DNA and HPV type determination in tumor tissue may help predict recurrence risk. Screening serum cell-free DNA for junctions may also offer an unambiguous, non-invasive means to monitor absence of recurrence following treatment.

DNA integration↗

Sequencing Medicago truncatula expressed sequenced tags using 454 Life Sciences technology.

BACKGROUND: In this study, we addressed whether a single 454 Life Science GS20 sequencing run provides new gene discovery from a normalized cDNA library, and whether the short reads produced via this technology are of value in gene structure annotation. RESULTS: A single 454 GS20 sequencing run on adapter-ligated cDNA, from a normalized cDNA library, generated 292,465 reads that were reduced to 252,384 reads with an average read length of 92 nucleotides after cleaning. After clustering and assembly, a total of 184,599 unique sequences were generated containing over 400 SSRs. The 454 sequences generated hits to more genes than a comparable amount of sequence from MtGI. Although short, the 454 reads are of sufficient length to map to a unique genome location as effectively as longer ESTs produced by conventional sequencing. Functional interpretation of the sequences was carried out by Gene Ontology assignments from matches to Arabidopsis and was shown to cover a broad range of GO categories. 53,796 assemblies and singletons (29%) had no match in the existing MtGI. Within the previously unobserved Medicago transcripts, thousands had matches in a comprehensive protein database and one or more of the TIGR Plant Gene Indices. Approximately 20% of these novel sequences could be found in the Medicago genome sequence. A total of 70,026 reads generated by the 454 technology were mapped to 785 Medicago finished BACs using PASA and over 1,000 gene models required modification. In parallel to 454 sequencing, 4,445 5'-prime reads were generated by conventional sequencing using the same library and from the assembled sequences it was shown to contain about 52% full length cDNAs encoding proteins from 50 to over 500 amino acids in length. CONCLUSION: Due to the large number of reads afforded by the 454 DNA sequencing technology, it is effective in revealing the expression of transcripts from a broad range of GO categories and contains many rare transcripts in normalized cDNA libraries, although only a limited portion of their sequence is uncovered. As with longer ESTs, 454 reads can be mapped uniquely onto genomic sequence to provide support for, and modifications of, gene predictions.

Base Sequence↗

Comparative genomics of Brassica oleracea and Arabidopsis thaliana reveal gene loss, fragmentation, and dispersal after polyploidy.

We sequenced 2.2 Mb representing triplicated genome segments of Brassica oleracea, which are each paralogous with one another and homologous with a segmentally duplicated region of the Arabidopsis thaliana genome. Sequence annotation identified 177 conserved collinear genes in the B. oleracea genome segments. Analysis of synonymous base substitution rates indicated that the triplicated Brassica genome segments diverged from a common ancestor soon after divergence of the Arabidopsis and Brassica lineages. This conclusion was corroborated by phylogenetic analysis of protein families. Using A. thaliana as an outgroup, 35% of the genes inferred to be present when genome triplication occurred in the Brassica lineage have been lost, most likely via a deletion mechanism, in an interspersed pattern. Genes encoding proteins involved in signal transduction or transcription were not found to be significantly more extensively retained than those encoding proteins classified with other functions, but putative proteins predicted in the A. thaliana genome were underrepresented in B. oleracea. We identified one example of gene loss from the Arabidopsis lineage. We found evidence for the frequent insertion of gene fragments of nuclear genomic origin and identified four apparently intact genes in noncollinear positions in the B. oleracea and A. thaliana genomes.

Arabidopsis↗

Macronuclear genome sequence of the ciliate Tetrahymena thermophila, a model eukaryote.

The ciliate Tetrahymena thermophila is a model organism for molecular and cellular biology. Like other ciliates, this species has separate germline and soma functions that are embodied by distinct nuclei within a single cell. The germline-like micronucleus (MIC) has its genome held in reserve for sexual reproduction. The soma-like macronucleus (MAC), which possesses a genome processed from that of the MIC, is the center of gene expression and does not directly contribute DNA to sexual progeny. We report here the shotgun sequencing, assembly, and analysis of the MAC genome of T. thermophila, which is approximately 104 Mb in length and composed of approximately 225 chromosomes. Overall, the gene set is robust, with more than 27,000 predicted protein-coding genes, 15,000 of which have strong matches to genes in other organisms. The functional diversity encoded by these genes is substantial and reflects the complexity of processes required for a free-living, predatory, single-celled organism. This is highlighted by the abundance of lineage-specific duplications of genes with predicted roles in sensing and responding to environmental conditions (e.g., kinases), using diverse resources (e.g., proteases and transporters), and generating structural complexity (e.g., kinesins and dyneins). In contrast to the other lineages of alveolates (apicomplexans and dinoflagellates), no compelling evidence could be found for plastid-derived genes in the genome. UGA, the only T. thermophila stop codon, is used in some genes to encode selenocysteine, thus making this organism the first known with the potential to translate all 64 codons in nuclear genes into amino acids. We present genomic evidence supporting the hypothesis that the excision of DNA from the MIC to generate the MAC specifically targets foreign DNA as a form of genome self-defense. The combination of the genome sequence, the functional diversity encoded therein, and the presence of some pathways missing from other model organisms makes T. thermophila an ideal model for functional genomic studies to address biological, biomedical, and biotechnological questions of fundamental importance.

Animals↗

Analysis of the cDNAs of hypothetical genes on Arabidopsis chromosome 2 reveals numerous transcript variants.

In the fully sequenced Arabidopsis (Arabidopsis thaliana) genome, many gene models are annotated as "hypothetical protein," whose gene structures are predicted solely by computer algorithms with no support from either expressed sequence matches from Arabidopsis, or nucleic acid or protein homologs from other species. In order to confirm their existence and predicted gene structures, a high-throughput method of rapid amplification of cDNA ends (RACE) was used to obtain their cDNA sequences from 11 cDNA populations. Primers from all of the 797 hypothetical genes on chromosome 2 were designed, and, through 5' and 3' RACE, clones from 506 genes were sequenced and cDNA sequences from 399 target genes were recovered. The cDNA sequences were obtained by assembling their 5' and 3' RACE polymerase chain reaction products. These sequences revealed that (1) the structures of 151 hypothetical genes were different from their predictions; (2) 116 hypothetical genes had alternatively spliced transcripts and 187 genes displayed polyadenylation sites; and (3) there were transcripts arising from both strands, from the strand opposite to that of the prediction and possible dicistronic transcripts. Promoters from five randomly chosen hypothetical genes (At2g02540, At2g31270, At2g33640, At2g35550, and At2g36340) were cloned into report constructs, and their expressions are tissue or development stage specific. Our results indicate at least 50% of hypothetical genes on chromosome 2 are expressed in the cDNA populations with about 38% of the gene structures differing from their predictions. Thus, by using this targeted approach, high-throughput RACE, we revealed numerous transcripts including many uncharacterized variants from these hypothetical genes.

Alternative Splicing↗

Comparative genomics of trypanosomatid parasitic protozoa.

A comparison of gene content and genome architecture of Trypanosoma brucei, Trypanosoma cruzi, and Leishmania major, three related pathogens with different life cycles and disease pathology, revealed a conserved core proteome of about 6200 genes in large syntenic polycistronic gene clusters. Many species-specific genes, especially large surface antigen families, occur at nonsyntenic chromosome-internal and subtelomeric regions. Retroelements, structural RNAs, and gene family expansion are often associated with syntenic discontinuities that-along with gene divergence, acquisition and loss, and rearrangement within the syntenic regions-have shaped the genomes of each parasite. Contrary to recent reports, our analyses reveal no evidence that these species are descended from an ancestor that contained a photosynthetic endosymbiont.

Animals↗

Compilation of mRNA polyadenylation signals in Arabidopsis revealed a new signal element and potential secondary structures.

Using a novel program, SignalSleuth, and a database containing authenticated polyadenylation [poly(A)] sites, we analyzed the composition of mRNA poly(A) signals in Arabidopsis (Arabidopsis thaliana), and reevaluated previously described cis-elements within the 3'-untranslated (UTR) regions, including near upstream elements and far upstream elements. As predicted, there are absences of high-consensus signal patterns. The AAUAAA signal topped the near upstream elements patterns and was found within the predicted location to only approximately 10% of 3'-UTRs. More importantly, we identified a new set, named cleavage elements, of poly(A) signals flanking both sides of the cleavage site. These cis-elements were not previously revealed by conventional mutagenesis and are contemplated as a cluster of signals for cleavage site recognition. Moreover, a single-nucleotide profile scan on the 3'-UTR regions unveiled a distinct arrangement of alternate stretches of U and A nucleotides, which led to a prediction of the formation of secondary structures. Using an RNA secondary structure prediction program, mFold, we identified three main types of secondary structures on the sequences analyzed. Surprisingly, these observed secondary structures were all interrupted in previously constructed mutations in these regions. These results will enable us to revise the current model of plant poly(A) signals and to develop tools to predict 3'-ends for gene annotation.

3' Untranslated Regions↗

Complete reannotation of the Arabidopsis genome: methods, tools, protocols and the final release.

BACKGROUND: Since the initial publication of its complete genome sequence, Arabidopsis thaliana has become more important than ever as a model for plant research. However, the initial genome annotation was submitted by multiple centers using inconsistent methods, making the data difficult to use for many applications. RESULTS: Over the course of three years, TIGR has completed its effort to standardize the structural and functional annotation of the Arabidopsis genome. Using both manual and automated methods, Arabidopsis gene structures were refined and gene products were renamed and assigned to Gene Ontology categories. We present an overview of the methods employed, tools developed, and protocols followed, summarizing the contents of each data release with special emphasis on our final annotation release (version 5). CONCLUSION: Over the entire period, several thousand new genes and pseudogenes were added to the annotation. Approximately one third of the originally annotated gene models were significantly refined yielding improved gene structure annotations, and every protein-coding gene was manually inspected and classified using Gene Ontology terms.

Alternative Splicing↗

The genome of the basidiomycetous yeast and human pathogen Cryptococcus neoformans.

Cryptococcus neoformans is a basidiomycetous yeast ubiquitous in the environment, a model for fungal pathogenesis, and an opportunistic human pathogen of global importance. We have sequenced its approximately 20-megabase genome, which contains approximately 6500 intron-rich gene structures and encodes a transcriptome abundant in alternatively spliced and antisense messages. The genome is rich in transposons, many of which cluster at candidate centromeric regions. The presence of these transposons may drive karyotype instability and phenotypic variation. C. neoformans encodes unique genes that may contribute to its unusual virulence properties, and comparison of two phenotypically distinct strains reveals variation in gene content in addition to sequence polymorphisms between the genomes.

Alternative Splicing↗

Whole genome shotgun sequencing of Brassica oleracea and its application to gene discovery and annotation in Arabidopsis.

Through comparative studies of the model organism Arabidopsis thaliana and its close relative Brassica oleracea, we have identified conserved regions that represent potentially functional sequences overlooked by previous Arabidopsis genome annotation methods. A total of 454,274 whole genome shotgun sequences covering 283 Mb (0.44 x) of the estimated 650 Mb Brassica genome were searched against the Arabidopsis genome, and conserved Arabidopsis genome sequences (CAGSs) were identified. Of these 229,735 conserved regions, 167,357 fell within or intersected existing gene models, while 60,378 were located in previously unannotated regions. After removal of sequences matching known proteins, CAGSs that were close to one another were chained together as potentially comprising portions of the same functional unit. This resulted in 27,347 chains of which 15,686 were sufficiently distant from existing gene annotations to be considered a novel conserved unit. Of 192 conserved regions examined, 58 were found to be expressed in our cDNA populations. Rapid amplification of cDNA ends (RACE) was used to obtain potentially full-length transcripts from these 58 regions. The resulting sequences led to the creation of 21 gene models at 17 new Arabidopsis loci and the addition of splice variants or updates to another 19 gene structures. In addition, CAGSs overlapping already annotated genes in Arabidopsis can provide guidance for manual improvement of existing gene models. Published genome-wide expression data based on whole genome tiling arrays and massively parallel signature sequencing were overlaid on the Brassica-Arabidopsis conserved sequences, and 1399 regions of intersection were identified. Collectively our results and these data sets suggest that several thousand new Arabidopsis genes remain to be identified and annotated.

Arabidopsis↗

Transcriptional divergence of the duplicated oxidative stress-responsive genes in the Arabidopsis genome.

Previous studies have indicated that Arabidopsis thaliana experienced a genome-wide duplication event shortly before its divergence from Brassica followed by extensive chromosomal rearrangements and deletions. While a large number of the duplicated genes have significantly diverged or lost their sister genes, we found 4222 pairs that are still highly conserved, and as a result had similar functional assignments during the annotation of the genome sequence. Using whole-genome DNA microarrays, we identified 906 duplicated gene pairs in which at least one member exhibited a significant response to oxidative stress. Among these, only 117 pairs were up- or down-regulated in both pairs and many of these exhibited dissimilar patterns of expression. Examination of the expression patterns of PAL1 and PAL2, ACD1 and ACD2, genes coding for two Hsp20s, various P450s, and electron transfer flavoproteins suggests Arabidopsis evolved a number of distinct oxidative stress response mechanisms using similar gene sets following the duplication of its genome.

Arabidopsis↗

DAGchainer: a tool for mining segmental genome duplications and synteny.

SUMMARY: Given the positions of protein-coding genes along genomic sequence and probability values for protein alignments between genes, DAGchainer identifies chains of gene pairs sharing conserved order between genomic regions, by identifying paths through a directed acyclic graph (DAG). These chains of collinear gene pairs can represent segmentally duplicated regions and genes within a single genome or syntenic regions between related genomes. Automated mining of the Arabidopsis genome for segmental duplications illustrates the use of DAGchainer.

Algorithms↗

Development and evaluation of an Arabidopsis whole genome Affymetrix probe array.

We describe the development of a high-density Arabidopsis'whole genome' oligonucleotide probe array for expression analysis (the Affymetrix ATH1 GeneChip probe array) that contains approximately 22 750 probe sets. Precedence on the array was given to genes for which either expression evidence or a credible database match existed. The remaining space was filled with 'hypothetical' genes. The new ATH1 array represents approximately 23 750 genes of which 60% were detected in RNA from cultured seedlings. Sensitivity of the array, determined using spiking controls, was approximately one transcript per cell. The array demonstrated high technical reproducibility and concordance with real-time PCR results. Indole-3 acetic acid (IAA)-induced changes in gene expression were used for biological validation of the array. A total of 222 genes were significantly upregulated and 103 significantly downregulated by exposure to IAA. Of the genes whose products could be functionally classified, the largest specific classes of upregulated genes were transcriptional regulators and protein kinases, many fewer of which were represented among the downregulated genes. Over one-third of the auxin-regulated genes have no known function, although many belong to gene families with members that have previously been shown to be auxin regulated. For the 6714 genes represented both on this and the earlier Arabidopsis Genome (AG) array, both signal intensities and gene expression ratios were very similar. Mapping of the oligonucleotides on the ATH1 array to the latest (version 4.0) annotation showed that over 95% of the probe sets (based on version 2.0 annotation) still fully represented their original target genes.

Arabidopsis↗

Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.

The spliced alignment of expressed sequence data to genomic sequence has proven a key tool in the comprehensive annotation of genes in eukaryotic genomes. A novel algorithm was developed to assemble clusters of overlapping transcript alignments (ESTs and full-length cDNAs) into maximal alignment assemblies, thereby comprehensively incorporating all available transcript data and capturing subtle splicing variations. Complete and partial gene structures identified by this method were used to improve The Institute for Genomic Research Arabidopsis genome annotation (TIGR release v.4.0). The alignment assemblies permitted the automated modeling of several novel genes and >1000 alternative splicing variations as well as updates (including UTR annotations) to nearly half of the approximately 27 000 annotated protein coding genes. The algorithm of the Program to Assemble Spliced Alignments (PASA) tool is described, as well as the results of automated updates to Arabidopsis gene annotations.

Algorithms↗

Kinetics and binding of the thymine-DNA mismatch glycosylase, Mig-Mth, with mismatch-containing DNA substrates.

We have examined the removal of thymine residues from T-G mismatches in DNA by the thymine-DNA mismatch glycosylase from Methanobacterium thermoautrophicum (Mig-Mth), within the context of the base excision repair (BER) pathway, to investigate why this glycosylase has such low activity in vitro. Using single-turnover kinetics and steady-state kinetics, we calculated the catalytic and product dissociation rate constants for Mig-Mth, and determined that Mig-Mth is inhibited by product apyrimidinic (AP) sites in DNA. Electrophoretic mobility shift assays (EMSA) provide evidence that the specificity of product binding is dependent upon the base opposite the AP site. The binding of Mig-Mth to DNA containing the non-cleavable substrate analogue difluorotoluene (F) was also analyzed to determine the effect of the opposite base on Mig-Mth binding specificity for substrate-like duplex DNA. The results of these experiments support the idea that opposite strand interactions play roles in determining substrate specificity. Endonuclease IV, which cleaves AP sites in the next step of the BER pathway, was used to analyze the effect of product removal on the overall rate of thymine hydrolysis by Mig-Mth. Our results support the hypothesis that endonuclease IV increases the apparent activity of Mig-Mth significantly under steady-state conditions by preventing reassociation of enzyme to product.

DNA↗

Computational discovery of internal micro-exons.

Very short exons, also known as micro-exons, occur in large numbers in some eukaryotic genomes. Existing annotation tools have a limited ability to recognize these short sequences, which range in length up to 25 bp. Here, we describe a computational method for the identification of micro-exons using near-perfect alignments between cDNA and genomic DNA sequences. Using this method, we detected 319 micro-exons in 4 complete genomes, of which 224 were previously unknown, human (170), the nematode Caenorhabditis elegans (4), the fruit fly Drosophila melanogaster (14), and the mustard plant Arabidopsis thaliana (36). Comparison of our computational method with popular cDNA alignment programs shows that the new algorithm is both efficient and accurate. The algorithm also aids in the discovery of micro-exon-skipping events and cross-species micro-exon conservation.

Alternative Splicing↗