PubMed Health⌕ Search

Biomedical subjects

Victoria Zismann

Publications and source records attributed to Victoria Zismann.

5 recordsLinked to original sources

SNiPer-HD: improved genotype calling accuracy by an expectation-maximization algorithm for high-density SNP arrays.

MOTIVATION: The technology to genotype single nucleotide polymorphisms (SNPs) at extremely high densities provides for hypothesis-free genome-wide scans for common polymorphisms associated with complex disease. However, we find that some errors introduced by commonly employed genotyping algorithms may lead to inflation of false associations between markers and phenotype. RESULTS: We have developed a novel SNP genotype calling program, SNiPer-High Density (SNiPer-HD), for highly accurate genotype calling across hundreds of thousands of SNPs. The program employs an expectation-maximization (EM) algorithm with parameters based on a training sample set. The algorithm choice allows for highly accurate genotyping for most SNPs. Also, we introduce a quality control metric for each assayed SNP, such that poor-behaving SNPs can be filtered using a metric correlating to genotype class separation in the calling algorithm. SNiPer-HD is superior to the standard dynamic modeling algorithm and is complementary and non-redundant to other algorithms, such as BRLMM. Implementing multiple algorithms together may provide highly accurate genotyping calls, without inflation of false positives due to systematically miss-called SNPs. A reliable and accurate set of SNP genotypes for increasingly dense panels will eliminate some false association signals and false negative signals, allowing for rapid identification of disease susceptibility loci for complex traits. AVAILABILITY: SNiPer-HD is available at TGen's website: http://www.tgen.org/neurogenomics/data.

Algorithms↗

Expressed sequence tags from loblolly pine embryos reveal similarities with angiosperm embryogenesis.

The process of embryogenesis in gymnosperms differs in significant ways from the more widely studied process in angiosperms. To further our understanding of embryogenesis in gymnosperms, we have generated Expressed Sequence Tags (ESTs) from four cDNA libraries constructed from un-normalized, normalized, and subtracted RNA populations of zygotic and somatic embryos of loblolly pine (Pinus taeda L.). A total of 68,721 ESTs were generated from 68,131 cDNA clones. Following clustering and assembly, these sequences collapsed into 5,274 contigs and 6,880 singleton sequences for a total of 12,154 non-redundant sequences. Searches of a non-identical amino acid database revealed a putative homolog for 9,189 sequences, leaving 2,965 sequences with no known function. More extensive searches of additional plant sequence data sets revealed a putative homolog for all but 1,388 (11.4%) of the sequences. Using gene ontologies, a known function could be assigned for 5,495 of the 12,154 total non-redundant sequences with 13,633 associations in total assigned. When compared to approximately 72,000 sequences in a collated P. taeda transcript assembly derived from >245,000 ESTs derived from root, xylem, stem, needles, pollen cone, and shoot ESTs, 3,458 (28.5%) of the non-redundant embryo sequences were unique and thereby provide a valuable addition to development of a complete loblolly pine transcriptome. To assess similarities between angiosperm and gymnosperm embryo development, we examined our EST collection for putative homologs of angiosperm genes implicated in embryogenesis. Out of 108 angiosperm embryogenesis-related genes, homologs were present for 83 of these genes suggesting that pine contains similar genes for embryogenesis and that our RNA sampling methods were successful. We also identified sequences from the pine embryo transcriptome that have no known function and may contribute to the programming of gene expression and embryo development.

Amino Acid Sequence↗

Mitochondrial genome sequences and molecular evolution of the Irish potato famine pathogen, Phytophthora infestans.

The mitochondrial genomes of haplotypes of the Irish potato famine pathogen, Phytophthora infestans, were sequenced. The genome sizes were 37,922, 39,870 and 39,840 bp for the type Ia, IIa and IIb mitochondrial DNA (mtDNA) haplotypes, respectively. The mitochondrial genome size for the type Ib haplotype, previously sequenced by others, was 37,957 bp. More than 90% of the genome contained coding regions. The GC content was 22.3%. A total of 18 genes involved in electron transport, 2 RNA-encoding genes, 16 ribosomal protein genes and 25 transfer RNA genes were coded on both strands with a conserved arrangement among the haplotypes. The type I haplotypes contained six unique open reading frames (ORFs) of unknown function while the type II haplotypes contained 13 ORFs of unknown function. Polymorphisms were observed in both coding and non-coding regions although the highest variation was in non-coding regions. The type I haplotypes (Ia and Ib) differed by only 14 polymorphic sites, whereas the type II haplotypes (IIa and IIb) differed by 50 polymorphic sites. The largest number (152) of polymorphic sites was found between the type IIb and Ia haplotypes. A large spacer flanked by the genes coding for tRNA-Tyr (trnY) and the small subunit RNA (rns) contained the largest number of polymorphic sites and corresponds to the region where a large indel that differentiates type II from type I haplotypes is located. The size of this region was 785, 2,666 and 2,670 bp in type Ia, IIa and IIb haplotypes, respectively. Among the four haplotypes, 81 mutations were identified. Phylogenetic and coalescent analysis revealed that although the type I and II haplotypes shared a common ancestor, they clearly formed two independent lineages that evolved independently. The type II haplotypes diverged earlier than the type I haplotypes. Thus our data do not support the previous hypothesis that the type II lineages evolved from the type I lineages. The type I haplotypes diverged more recently and the mutations associated with the evolution of the Ia and Ib types were identified.

Evolution, Molecular↗

Sequence, annotation, and analysis of synteny between rice chromosome 3 and diverged grass species.

Rice (Oryza sativa L.) chromosome 3 is evolutionarily conserved across the cultivated cereals and shares large blocks of synteny with maize and sorghum, which diverged from rice more than 50 million years ago. To begin to completely understand this chromosome, we sequenced, finished, and annotated 36.1 Mb ( approximately 97%) from O. sativa subsp. japonica cv Nipponbare. Annotation features of the chromosome include 5915 genes, of which 913 are related to transposable elements. A putative function could be assigned to 3064 genes, with another 757 genes annotated as expressed, leaving 2094 that encode hypothetical proteins. Similarity searches against the proteome of Arabidopsis thaliana revealed putative homologs for 67% of the chromosome 3 proteins. Further searches of a nonredundant amino acid database, the Pfam domain database, plant Expressed Sequence Tags, and genomic assemblies from sorghum and maize revealed only 853 nontransposable element related proteins from chromosome 3 that lacked similarity to other known sequences. Interestingly, 426 of these have a paralog within the rice genome. A comparative physical map of the wild progenitor species, Oryza nivara, with japonica chromosome 3 revealed a high degree of sequence identity and synteny between these two species, which diverged approximately 10,000 years ago. Although no major rearrangements were detected, the deduced size of the O. nivara chromosome 3 was 21% smaller than that of japonica. Synteny between rice and other cereals using an integrated maize physical map and wheat genetic map was strikingly high, further supporting the use of rice and, in particular, chromosome 3, as a model for comparative studies among the cereals.

Arabidopsis↗

Analyzing the potato abiotic stress transcriptome using expressed sequence tags.

To further increase our understanding of responses in potato to abiotic stress and the potato transcriptome in general, we generated 20 756 expressed sequence tags (ESTs) from a cDNA library constructed by pooling mRNA from heat-, cold-, salt-, and drought-stressed potato leaves and roots. These ESTs were clustered and assembled into a collection of 5240 unique sequences with 3344 contigs and 1896 singleton ESTs. Assignment of gene ontology terms (GOSlim/Plant) to the sequences revealed that 8101 assignments could be made with a total of 3863 molecular function assignments. Alignment to a set of 78 825 ESTs from other potato cDNA libraries derived from root, leaf, stolon, tuber, germinating eye, and callus tissues revealed 1476 sequences unique to abiotic stressed potato leaf and root tissue. Sequences present within the 5240 sequence set had similarity to genes known to be involved in abiotic stress responses in other plant species such as transcription factors, stress response genes, and signal transduction processes. In addition, we identified a number of genes unique to the abiotic stress library with unknown function, providing new candidate genes for investigation of abiotic stress responses in potato.

Arabidopsis↗