PubMed Health⌕ Search

Biomedical subjects

Lincoln Stein

Publications and source records attributed to Lincoln Stein.

At least 19 recordsLinked to original sources

In silico generation of synthetic cancer genomes using generative AI.

Understanding how genomic alterations drive cancer is key to advancing precision oncology. To detect these alterations, accurate algorithms are used; however, due to privacy concerns, few deeply sequenced cancer genomes can be shared, limiting benchmarking and representing a major obstacle to the improvement of analytic tools. To address this, we developed OncoGAN, a generative AI model combining adversarial networks and variational autoencoders to create realistic synthetic cancer genomes. Trained on large-scale genomic datasets, OncoGAN accurately reproduces somatic mutations, copy number alterations, and structural variants across cancer types while preserving donors' privacy. The synthetic genomes reflect tumor-specific mutational signatures and positional mutation patterns. Using DeepTumour, we validated the synthetic data's fidelity, showing high concordance between generated and predicted tumors. Moreover, augmenting the training data with synthetic genomes improved DeepTumour's accuracy, underscoring OncoGAN's potential to generate shareable datasets with known ground truths for benchmarking and enhancement of cancer genome analysis tools.

Humans↗

Whole-plant growth stage ontology for angiosperms and its application in plant biology.

Plant growth stages are identified as distinct morphological landmarks in a continuous developmental process. The terms describing these developmental stages record the morphological appearance of the plant at a specific point in its life cycle. The widely differing morphology of plant species consequently gave rise to heterogeneous vocabularies describing growth and development. Each species or family specific community developed distinct terminologies for describing whole-plant growth stages. This semantic heterogeneity made it impossible to use growth stage description contained within plant biology databases to make meaningful computational comparisons. The Plant Ontology Consortium (http://www.plantontology.org) was founded to develop standard ontologies describing plant anatomical as well as growth and developmental stages that can be used for annotation of gene expression patterns and phenotypes of all flowering plants. In this article, we describe the development of a generic whole-plant growth stage ontology that describes the spatiotemporal stages of plant growth as a set of landmark events that progress from germination to senescence. This ontology represents a synthesis and integration of terms and concepts from a variety of species-specific vocabularies previously used for describing phenotypes and genomic information. It provides a common platform for annotating gene function and gene expression in relation to the developmental trajectory of a plant described at the organismal level. As proof of concept the Plant Ontology Consortium used the plant ontology growth stage ontology to annotate genes and phenotypes in plants with initial emphasis on those represented in The Arabidopsis Information Resource, Gramene database, and MaizeGDB.

Arabidopsis↗

Evolution of Arabidopsis microRNA families through duplication events.

Recently there has been a great interest in the identification of microRNAs and their targets as well as understanding the spatial and temporal regulation of microRNA genes. To understand how microRNA genes evolve, we looked at several rapidly evolving families in Arabidopsis thaliana, and found that they arose from a process of genome-wide duplication, tandem duplication, and segmental duplication followed by dispersal and diversification, similar to the processes that drive the evolution of protein gene families. Using multiple expression data sets to examine the transcription patterns of different members of the microRNA families, we find the sequence diversification of duplicated microRNA genes to be accompanied by a change in spatial and temporal expression patterns, suggesting that duplicated copies acquire new functionality as they evolve.

Arabidopsis↗

Look-Align: an interactive web-based multiple sequence alignment viewer with polymorphism analysis support.

UNLABELLED: We have developed Look-Align, an interactive web-based viewer to display pre-computed multiple sequence alignments. Although initially developed to support the visualization needs of the maize diversity website Panzea (http://www.panzea.org), the viewer is a generic stand-alone tool that can be easily integrated into other websites. AVAILABILITY: Look-Align is written in Perl using open-source components and is available under an open-source license. Live installation and download information can be found at the Panzea website (http://www.panzea.org/software/alignment_viewer.html). CONTACT: ware@cshl.edu SUPPLEMENTARY INFORMATION: The Supplementary information includes sample lists of multiple sequence alignment software and sample screenshots of the viewer.

Databases, Genetic↗

Panzea: a database and resource for molecular and functional diversity in the maize genome.

Serving as a community resource, Panzea (http://www.panzea.org) is the bioinformatics arm of the Molecular and Functional Diversity in the Maize Genome project. Maize, a classical model for genetic studies, is an important crop species and also the most diverse crop species known. On average, two randomly chosen maize lines have one single-nucleotide polymorphism every approximately 100 bp; this divergence is roughly equivalent to the differences between humans and chimpanzees. This exceptional genotypic diversity underlies the phenotypic diversity maize needs to be cultivated in a wide range of environments. The Molecular and Functional Diversity in the Maize Genome project aims to understand how selection has shaped molecular diversity in maize and then relate molecular diversity to functional phenotypic variation. The project will screen 4000 loci for the signature of selection and create a wide range of maize and maize-teosinte mapping populations. These populations will be genotyped and phenotyped, permitting high-power and high-resolution dissection of the traits and relating the molecular diversity to functional variation. Panzea provides access to the genotype, phenotype and polymorphism data produced by the project through user-friendly web-based database searches and data retrieval/visualization tools, as well as a wide variety of information and services related to maize diversity.

Chromosome Mapping↗

Gramene: a bird's eye view of cereal genomes.

Rice, maize, sorghum, wheat, barley and the other major crop grasses from the family Poaceae (Gramineae) are mankind's most important source of calories and contribute tens of billions of dollars annually to the world economy (FAO 1999, http://www.fao.org; USDA 1997, http://www.usda.gov). Continued improvement of Poaceae crops is necessary in order to continue to feed an ever-growing world population. However, of the major crop grasses, only rice (Oryza sativa), with a compact genome of approximately 400 Mbp, has been sequenced and annotated. The Gramene database (http://www.gramene.org) takes advantage of the known genetic colinearity (synteny) between rice and the major crop plant genomes to provide maize, sorghum, millet, wheat, oat and barley researchers with the benefits of an annotated genome years before their own species are sequenced. Gramene is a one stop portal for finding curated literature, genetic and genomic datasets related to maps, markers, genes, genomes and quantitative trait loci. The addition of several new tools to Gramene has greatly facilitated the potential for comparative analysis among the grasses and contributes to our understanding of the anatomy, development, environmental responses and the factors influencing agronomic performance of cereal crops. Since the last publication on Gramene database by D. H. Ware, P. Jaiswal, J. Ni, I. V. Yap, X. Pan, K. Y. Clark, L. Teytelman, S. C. Schmidt, W. Zhao, K. Chang et al. [(2002), Plant Physiol., 130, 1606-1613], the database has undergone extensive changes that are described in this publication.

Arabidopsis↗

Chromosome evolution in eukaryotes: a multi-kingdom perspective.

In eukaryotes, chromosomal rearrangements, such as inversions, translocations and duplications, are common and range from part of a gene to hundreds of genes. Lineage-specific patterns are also seen: translocations are rare in dipteran flies, and angiosperm genomes seem prone to polyploidization. In most eukaryotes, there is a strong association between rearrangement breakpoints and repeat sequences. Current data suggest that some repeats promoted rearrangements via non-allelic homologous recombination, for others the association might not be causal but reflects the instability of particular genomic regions. Rearrangement polymorphisms in eukaryotes are correlated with phenotypic differences, so are thought to confer varying fitness in different habitats. Some seem to be under positive selection because they either trap favorable allele combinations together or alter the expression of nearby genes. There is little evidence that chromosomal rearrangements cause speciation, but they probably intensify reproductive isolation between species that have formed by another route.

Animals↗

Distinct regulatory elements mediate similar expression patterns in the excretory cell of Caenorhabditis elegans.

Identification of cis-regulatory elements and their binding proteins constitutes an important part of understanding gene function and regulation. It is well accepted that co-expressed genes tend to share transcriptional elements. However, recent findings indicate that co-expression data show poor correlation with co-regulation data even in unicellular yeast. This motivates us to experimentally explore whether it is possible that co-expressed genes are subject to differential regulatory control using the excretory cell of Caenorhabditis elegans as an example. Excretory cell is a functional equivalent of human kidney. Transcriptional regulation of gene expression in the cell is largely unknown. We isolated a 10-bp excretory cell-specific cis-element, Ex-1, from a pgp-12 promoter. The significance of the element has been demonstrated by its capacity of converting an intestine-specific promoter into an excretory cell-specific one. We also isolated a cDNA encoding an Ex-1 binding transcription factor, DCP-66, using a yeast one-hybrid screen. Role of the factor in regulation of pgp-12 expression has been demonstrated both in vitro and in vivo. Search for occurrence of Ex-1 reveals that only a small portion of excretory cell-specific promoters contain Ex-1. Two other distinct cis-elements isolated from two different promoters can also dictate the excretory cell-specific expression but are independent of regulation by DCP-66. The results indicate that distinct regulatory elements are able to mediate the similar expression patterns.

ATP Binding Cassette Transporter, Subfamily B↗

SynBrowse: a synteny browser for comparative sequence analysis.

MOTIVATION: The recent efforts of various sequence projects to sequence deeply into various phylogenies provide great resources for comparative sequence analysis. A generic and portable tool is essential for scientists to visualize and analyze sequence comparisons. RESULTS: We have developed SynBrowse, a synteny browser for visualizing and analyzing genome alignments both within and between species. It is intended to help scientists study macrosynteny, microsynteny and homologous genes between sequences. It can also aid with the identification of uncharacterized genes, putative regulatory elements and novel structural features of a species. SynBrowse is a GBrowse (the Generic Genome Browser) family software tool that runs on top of the open source BioPerl modules. It consists of two components: a web-based front end and a set of relational database back ends. Each database stores pre-computed alignments from a focus sequence to reference sequences in addition to the genome annotations of the focus sequence. The user interface lets end users select a key comparative alignment type and search for syntenic blocks between two sequences and zoom in to view the relationships among the corresponding genome annotations in detail. SynBrowse is portable with simple installation, flexible configuration, convenient data input and easy integration with other components of a model organism system. AVAILABILITY: The software is available at http://www.gmod.org CONTACT: vbrendel@iastate.edu

Algorithms↗

The Sequence Ontology: a tool for the unification of genome annotations.

The Sequence Ontology (SO) is a structured controlled vocabulary for the parts of a genomic annotation. SO provides a common set of terms and definitions that will facilitate the exchange, analysis and management of genomic data. Because SO treats part-whole relationships rigorously, data described with it can become substrates for automated reasoning, and instances of sequence features described by the SO can be subjected to a group of logical operations termed extensional mereology operators.

Alternative Splicing↗

The oryza map alignment project: the golden path to unlocking the genetic potential of wild rice species.

The wild species of the genus Oryza offer enormous potential to make a significant impact on agricultural productivity of the cultivated rice species Oryza sativa and Oryza glaberrima. To unlock the genetic potential of wild rice we have initiated a project entitled the 'Oryza Map Alignment Project' (OMAP) with the ultimate goal of constructing and aligning BAC/STC based physical maps of 11 wild and one cultivated rice species to the International Rice Genome Sequencing Project's finished reference genome--O. sativa ssp. japonica c. v. Nipponbare. The 11 wild rice species comprise nine different genome types and include six diploid genomes (AA, BB, CC, EE, FF and GG) and four tetrapliod genomes (BBCC, CCDD, HHKK and HHJJ) with broad geographical distribution and ecological adaptation. In this paper we describe our strategy to construct robust physical maps of all 12 rice species with an emphasis on the AA diploid O. nivara--thought to be the progenitor of modern cultivated rice.

Chromosome Mapping↗

Site preferences of insertional mutagenesis agents in Arabidopsis.

We have performed a comparative analysis of the insertion sites of engineered Arabidopsis (Arabidopsis thaliana) insertional mutagenesis vectors that are based on the maize (Zea mays) transposable elements and Agrobacterium T-DNA. The transposon-based agents show marked preference for high GC content, whereas the T-DNA-based agents show preference for low GC content regions. The transposon-based agents show a bias toward insertions near the translation start codons of genes, while the T-DNAs show a predilection for the putative transcriptional regulatory regions of genes. The transposon-based agents also have higher insertion site densities in exons than do the T-DNA insertions. These observations show that the transposon-based and T-DNA-based mutagenesis techniques could complement one another well, and neither alone is sufficient to achieve the goal of saturation mutagenesis in Arabidopsis. These results also suggest that transposon-based mutagenesis techniques may prove the most effective for obtaining gene disruptions and for generating gene traps, while T-DNA-based agents may be more effective for activation tagging and enhancer trapping. From the patterns of insertion site distributions, we have identified a set of nucleotide sequence motifs that are overrepresented at the transposon insertion sites. These motifs may play a role in the transposon insertion site preferences. These results could help biologists to study the mechanisms of insertions of the insertional mutagenesis agents and to design better strategies for genome-wide insertional mutagenesis.

Arabidopsis↗

Maize-targeted mutagenesis: A knockout resource for maize.

We describe an efficient system for site-selected transposon mutagenesis in maize. A total of 43,776 F1 plants were generated by using Robertson's Mutator (Mu) pollen parents and self-pollinated to establish a library of transposon-mutagenized seed. The frequency of new seed mutants was between 10-4 and 10-5 per F1 plant. As a service to the maize community, maize-targeted mutagenesis selects insertions in genes of interest from this library by using the PCR. Pedigree, knockout, sequence, phenotype, and other information is stored in a powerful interactive database (maize-targeted mutagenesis database) that enables analysis of the entire population and the handling of knockout requests. By inhibiting Mu activity in most F1 plants, we sought to reduce somatic insertions that may cause false positives selected from pooled tissue. By monitoring the remaining Mu activity in the F2, however, we demonstrate that seed phenotypes depend on it, and false positives occur in lines that appear to lack it. We conclude that more than half of all mutations arising in this population are suppressed on losing Mu activity. These results have implications for epigenetic models of inbreeding and for functional genomics.

Base Sequence↗

A 3.9-centimorgan-resolution human single-nucleotide polymorphism linkage map and screening set.

Recent advances in technologies for high-throughout single-nucleotide polymorphism (SNP)-based genotyping have improved efficiency and cost so that it is now becoming reasonable to consider the use of SNPs for genomewide linkage analysis. However, a suitable screening set of SNPs and a corresponding linkage map have yet to be described. The SNP maps described here fill this void and provide a resource for fast genome scanning for disease genes. We have evaluated 6,297 SNPs in a diversity panel composed of European Americans, African Americans, and Asians. The markers were assessed for assay robustness, suitable allele frequencies, and informativeness of multi-SNP clusters. Individuals from 56 Centre d'Etude du Polymorphisme Humain pedigrees, with >770 potentially informative meioses altogether, were genotyped with a subset of 2,988 SNPs, for map construction. Extensive genotyping-error analysis was performed, and the resulting SNP linkage map has an average map resolution of 3.9 cM, with map positions containing either a single SNP or several tightly linked SNPs. The order of markers on this map compares favorably with several other linkage and physical maps. We compared map distances between the SNP linkage map and the interpolated SNP linkage map constructed by the deCode Genetics group. We also evaluated cM/Mb distance ratios in females and males, along each chromosome, showing broadly defined regions of increased and decreased rates of recombination. Evaluations indicate that this SNP screening set is more informative than the Marshfield Clinic's commonly used microsatellite-based screening set.

Alleles↗

ATIDB: Arabidopsis thaliana insertion database.

Insertional mutagenesis techniques, including transposon- and T-DNA-mediated mutagenesis, are key resources for systematic identification of gene function in the model plant species Arabidopsis thaliana. We have developed a database (http://atidb.cshl.org/) for archiving, searching and analyzing insertional mutagenesis lines. Flanking sequences from approximately 10 500 insertion lines (including transposon and T-DNA insertions) from several tagging programs in Arabidopsis were mapped to the genome sequence through our annotation system before being entered into the database. The database front end provides World Wide Web searching and analyzing interfaces for genome researchers and other biologists. Users can search the database to identify insertions in a particular gene or perform genome-wide analysis to study the distribution and preference of insertions. Tools integrated with the database include a graphical genome browser, a protein search function, a graphical representation of the insertion distribution and a Blast search function. The database is based on open source components and is available under an open source license.

Amino Acid Sequence↗

Comparison of genes among cereals.

Comparison of partially sequenced cereal genomes suggests a mosaic structure consisting of recombinationally active gene-rich islands that are separated by blocks of high-copy DNA. Annotation of the whole rice genome suggests that most, but not all, cereal genes are present within the rice genome and that the high number of reported genes in this genome is probably due to duplications. Within the cereals, macrocolinearity is conserved but, at the level of individual genes, microcolinearity is frequently disrupted. Preliminary evidence from limited comparative analysis of sequenced orthologous genomic segments suggests that local gene amplification and translocation within a plant genome may be linked in some cases.

Gene Amplification↗

Development and mapping of 2240 new SSR markers for rice (Oryza sativa L.).

A total of 2414 new di-, tri- and tetra-nucleotide non-redundant SSR primer pairs, representing 2240 unique marker loci, have been developed and experimentally validated for rice (Oryza sativa L.). Duplicate primer pairs are reported for 7% (174) of the loci. The majority (92%) of primer pairs were developed in regions flanking perfect repeats > or = 24 bp in length. Using electronic PCR (e-PCR) to align primer pairs against 3284 publicly sequenced rice BAC and PAC clones (representing about 83% of the total rice genome), 65% of the SSR markers hit a BAC or PAC clone containing at least one genetically mapped marker and could be mapped by proxy. Additional information based on genetic mapping and "nearest marker" information provided the basis for locating a total of 1825 (81%) of the newly designed markers along rice chromosomes. Fifty-six SSR markers (2.8%) hit BAC clones on two or more different chromosomes and appeared to be multiple copy. The largest proportion of SSRs in this data set correspond to poly(GA) motifs (36%), followed by poly(AT) (15%) and poly(CCG) (8%) motifs. AT-rich microsatellites had the longest average repeat tracts, while GC-rich motifs were the shortest. In combination with the pool of 500 previously mapped SSR markers, this release makes available a total of 2740 experimentally confirmed SSR markers for rice, or approximately one SSR every 157 kb.

Chromosome Mapping↗