PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 865 records · Page 48Linked to original sources

Genome-wide non-mendelian inheritance of extra-genomic information in Arabidopsis.

A fundamental tenet of classical mendelian genetics is that allelic information is stably inherited from one generation to the next, resulting in predictable segregation patterns of differing alleles. Although several exceptions to this principle are known, all represent specialized cases that are mechanistically restricted to either a limited set of specific genes (for example mating type conversion in yeast) or specific types of alleles (for example alleles containing transposons or repeated sequences). Here we show that Arabidopsis plants homozygous for recessive mutant alleles of the organ fusion gene HOTHEAD (HTH) can inherit allele-specific DNA sequence information that was not present in the chromosomal genome of their parents but was present in previous generations. This previously undescribed process is shown to occur at all DNA sequence polymorphisms examined and therefore seems to be a general mechanism for extra-genomic inheritance of DNA sequence information. We postulate that these genetic restoration events are the result of a template-directed process that makes use of an ancestral RNA-sequence cache.

Alleles↗

Genome assembly comparison identifies structural variants in the human genome.

Numerous types of DNA variation exist, ranging from SNPs to larger structural alterations such as copy number variants (CNVs) and inversions. Alignment of DNA sequence from different sources has been used to identify SNPs and intermediate-sized variants (ISVs). However, only a small proportion of total heterogeneity is characterized, and little is known of the characteristics of most smaller-sized (<50 kb) variants. Here we show that genome assembly comparison is a robust approach for identification of all classes of genetic variation. Through comparison of two human assemblies (Celera's R27c compilation and the Build 35 reference sequence), we identified megabases of sequence (in the form of 13,534 putative non-SNP events) that were absent, inverted or polymorphic in one assembly. Database comparison and laboratory experimentation further demonstrated overlap or validation for 240 variable regions and confirmed >1.5 million SNPs. Some differences were simple insertions and deletions, but in regions containing CNVs, segmental duplication and repetitive DNA, they were more complex. Our results uncover substantial undescribed variation in humans, highlighting the need for comprehensive annotation strategies to fully interpret genome scanning and personalized sequencing projects.

Base Sequence↗

Whole-genome phenotype prediction with machine learning: open problems in bacterial genomics.

MOTIVATION: How can we identify causal genetic mechanisms governing bacterial traits? Initial efforts entrusting machine learning models to handle the task of predicting phenotype from genotype yield high accuracy scores. However, attempts to extract meaningful interpretations from the predictive models are found to be corrupted by falsely identified 'causal' features. Relying solely on pattern recognition and correlations is unreliable, significantly so in bacterial genomics settings where high-dimensionality and spurious associations are the norm. Though it is not yet clear whether we can overcome this hurdle, significant efforts are being made towards discovering potential high-risk bacterial genetic variants. In view of this, we set up open problems surrounding phenotype prediction from bacterial whole-genome datasets and extending those approaches to learning causal effects, and discuss challenges that impact the reliability of a machine's decision-making when faced with datasets of this nature. RESULTS: We identify major sources of non-injectivity in the formulation of the genotype-to-phenotype mapping function-linkage-disequilibrium, limited sampling, information loss in representations, unmeasured confounders and observational noise-and analyse their implications for machine learning applications. Using a collection of 4,140 Staphylococcus aureus isolates, we illustrate challenges surrounding the defined open problems. AVAILABILITY AND IMPLEMENTATION: Raw sequencing data are available from the European Nucleotide Archive (ENA) under project accessions ERP001012, PRJEB3174, PRJEB2655, PRJEB2756, and PRJEB2944. Assemblies and annotations were generated with the Sanger bacterial pipeline (https://github.com/sanger-pathogens/vr-codebase) and unitigs extracted using DBGWAS (https://gitlab.com/leoisl/dbgwas).

Machine Learning↗

Structural analysis of a Lotus japonicus genome. V. Sequence features and mapping of sixty-four TAC clones which cover the 6.4 mb regions of the genome.

We determined the nucleotide sequences of 64 TAC (transformation-competent artificial chromosome) clones selected from genomic libraries of Lotus japonicus accession Miyakojima MG-20 based on the sequence information of expressed sequence tags (ESTs), cDNAs, genes and DNA markers from L. japonicus and other legumes. The length of the DNA regions sequenced in this study was 6,370,255 bp, and the total length of the L. japonicus genome sequenced so far is 32,537,698 bp together with the nucleotide sequences of 256 TAC clones previously reported. Five hundred forty-eight potential protein-encoding genes with known or predicted functions, 127 gene segments and 224 pseudogenes were assigned to the newly sequenced regions by computer prediction and similarity searches against the sequences in protein and EST databases. Based on the nucleotide sequences of the clones, simple sequence repeat length polymorphism (SSLP) or derived cleaved amplified polymorphic sequence (dCAPS) markers were generated, and each clone was genetically localized onto the linkage map of two accessions of L. japonicus, MG-20 and Gifu B-129. The sequence data, gene information and mapping information are available through the World Wide Web at http://www.kazusa.or.jp/lotus/.

Base Sequence↗

Genomic and post-genomic approaches to polycystic ovary syndrome--progress so far: Mini Review.

Genomic studies in polycystic ovary syndrome (PCOS) have focused on ovarian tissues and gene expression changes related to the gynaecological manifestations of PCOS. These studies have revealed a variety of altered genes that fall into many functional categories. Of these, the genes involved in steroidogenesis, including genes related to retinoic acid biosynthesis and LH-stimulated gene pathways, are generally up-regulated in PCOS samples. Genes involved in the Wnt signalling pathway appear down-regulated. Immune response genes and those involved in apoptosis are altered, but the net effect of these alterations is unclear at present. However, these altered gene expression patterns are yet to produce a defined aetiological basis or diagnostic biomarker for PCOS. The use of proteomic technologies for the study of the PCOS proteome is in its infancy; however, a few pilot studies have been published and the data are reviewed. Proteomics looks directly at the functional units within a cell, the proteins. This approach should thus serve to validate some of the gene expression changes identified and then build on the genomic results collected to date.

DNA↗

The chloroplast genome of Nymphaea alba: whole-genome analyses and the problem of identifying the most basal angiosperm.

Angiosperms (flowering plants) dominate contemporary terrestrial flora with roughly 250,000 species, but their origin and early evolution are still poorly understood. In recent years, molecular evidence has accumulated suggesting a dicotyledonous origin of monocots. Phylogenetic reconstructions have suggested that several dicotyledonous groups that include taxa such as Amborella, Austrobaileya, and Nymphaea branch off as the most basal among angiosperms. This has led to the concept of monocots, "eudicots," "basal dicots," and "ANITA" groupings. Here, we present the sequence and phylogenetic analyses of the chloroplast DNA of Nymphaea alba. Phylogenetic analyses of our 14-species data set, consisting of 29,991 aligned nucleotide positions per chloroplast genome, revealed consistent support for Nymphaea being a divergent member of a monophyletic dicot assemblage. Three distinct angiosperm lineages were supported in the majority of our phylogenetic analyses-eudicots, Magnoliopsida, and monocots. However, the monocot lineage leading to the grasses was the deepest branching. Although analyses of only one individual gene alignment (out of 61) is consistent with some recently proposed hypotheses for the paraphyly of dicots, we also report observations that nine genes do not support paraphyly of dicots. Instead, they support the basal monocot-dicot split. Consistent with this finding, we also report observations suggesting that the monocot lineage leading to the grasses has the strongest phylogenetic affinity to gymnosperms. Our findings have general implications for studies of substitution model specification and analyses of concatenated genome data.

Base Sequence↗

The Genome Sequence DataBase: towards an integrated functional genomics resource.

During 1998 the primary focus of the Genome Sequence DataBase (GSDB; http://www.ncgr.org/gsdb ) located at the National Center for Genome Resources (NCGR) has been to improve data quality, improve data collections, and provide new methods and tools to access and analyze data. Data quality has been improved by extensive curation of certain data fields necessary for maintaining data collections and for using certain tools. Data quality has also been increased by improvements to the suite of programs that import data from the International Nucleotide Sequence Database Collaboration (IC). The Sequence Tag Alignment and Consensus Knowledgebase (STACK), a database of human expressed gene sequences developed by the South African National Bioinformatics Institute (SANBI), became available within the last year, allowing public access to this valuable resource of expressed sequences. Data access was improved by the addition of the Sequence Viewer, a platform-independent graphical viewer for GSDB sequence data. This tool has also been integrated with other searching and data retrieval tools. A BLAST homology search service was also made available, allowing researchers to search all of the data, including the unique data, that are available from GSDB. These improvements are designed to make GSDB more accessible to users, extend the rich searching capability already present in GSDB, and to facilitate the transition to an integrated system containing many different types of biological data.

Animals↗

3D-GENOMICS: a database to compare structural and functional annotations of proteins between sequenced genomes.

The 3D-GENOMICS database (http://www.sbg.bio. ic.ac.uk/3dgenomics/) provides structural annotations for proteins from sequenced genomes. In August 2003 the database included data for 93 proteomes. The annotations stored in the database include homologous sequences from various sequence databases, domains from SCOP and Pfam, patterns from Prosite and other predicted sequence features such as transmembrane regions and coiled coils. In addition to annotations at the sequence level, several precomputed cross- proteome comparative analyses are available based on SCOP domain superfamily composition. Annotations are available to the user via a web interface to the database. Multiple points of entry are available so that a user is able to: (i) directly access annotations for a single protein sequence via keywords or accession codes, (ii) examine a sequence of interest chosen from a summary of annotations for a particular proteome, or (iii) access precomputed frequency-based cross-proteome comparative analyses.

Amino Acid Sequence↗

miRNAMap: genomic maps of microRNA genes and their target genes in mammalian genomes.

Recent work has demonstrated that microRNAs (miRNAs) are involved in critical biological processes by suppressing the translation of coding genes. This work develops an integrated database, miRNAMap, to store the known miRNA genes, the putative miRNA genes, the known miRNA targets and the putative miRNA targets. The known miRNA genes in four mammalian genomes such as human, mouse, rat and dog are obtained from miRBase, and experimentally validated miRNA targets are identified in a survey of the literature. Putative miRNA precursors were identified by RNAz, which is a non-coding RNA prediction tool based on comparative sequence analysis. The mature miRNA of the putative miRNA genes is accurately determined using a machine learning approach, mmiRNA. Then, miRanda was applied to predict the miRNA targets within the conserved regions in 3'-UTR of the genes in the four mammalian genomes. The miRNAMap also provides the expression profiles of the known miRNAs, cross-species comparisons, gene annotations and cross-links to other biological databases. Both textual and graphical web interface are provided to facilitate the retrieval of data from the miRNAMap. The database is freely available at http://mirnamap.mbc.nctu.edu.tw/.

Animals↗

PrimerStation: a highly specific multiplex genomic PCR primer design server for the human genome.

PrimerStation (http://ps.cb.k.u-tokyo.ac.jp) is a web service that calculates primer sets guaranteeing high specificity against the entire human genome. To achieve high accuracy, we used the hybridization ratio of primers in liquid solution. Calculating the status of sequence hybridization in terms of the stringent hybridization ratio is computationally costly, and no web service checks the entire human genome and returns a highly specific primer set calculated using a precise physicochemical model. To shorten the response time, we precomputed candidates for specific primers using a massively parallel computer with 100 CPUs (SunFire 15 K) about 3 months in advance. This enables PrimerStation to search and output qualified primers interactively. PrimerStation can select highly specific primers suitable for multiplex PCR by seeking a wider temperature range that minimizes the possibility of cross-reaction. It also allows users to add heuristic rules to the primer design, e.g. the exclusion of single nucleotide polymorphisms (SNPs) in primers, the avoidance of poly(A) and CA-repeats in the PCR products, and the elimination of defective primers using the secondary structure prediction. We performed several tests to verify the PCR amplification of randomly selected primers for ChrX, and we confirmed that the primers amplify specific PCR products perfectly.

DNA Primers↗

Putting the Leishmania genome to work: functional genomics by transposon trapping and expression profiling.

Leishmania are important protozoan pathogens of humans in temperate and tropical regions. The study of gene expression during the infectious cycle, in mutants or after environmental or chemical stimuli, is a powerful approach towards understanding parasite virulence and the development of control measures. Like other trypanosomatids, Leishmania gene expression is mediated by a polycistronic transcriptional process that places increased emphasis on post-transcriptional regulatory mechanisms including RNA processing and protein translation. With the impending completion of the Leishmania genome, global approaches surveying mRNA and protein expression are now feasible. Our laboratory has developed the Drosophila transposon mariner as a tool for trapping Leishmania genes and studying their regulation in the form of protein fusions; a classic approach in other microbes that can be termed 'proteogenomics'. Similarly, we have developed reagents and approaches for the creation of DNA microarrays, which permit the measurement of RNA abundance across the parasite genome. Progress in these areas promises to greatly increase our understanding of global mechanisms of gene regulation at both mRNA and protein levels, and to lead to the identification of many candidate genes involved in virulence.

Animals↗

DNA replication. Genomic views of genome duplication.

DNA replication is initiated at numerous origins of replication (oris) within the chromosomes. In a pair of ambitious studies, two groups have used different techniques to pinpoint the locations of all of the oris throughout the yeast genome at different times during S phase (Raghuraman et al., Wyrick et al.). Stillman, in his Perspective, compares and contrasts the different methods and their findings, and speculates on the value of combining these techniques to look at oris in the human genome.

Binding Sites↗

Genomics. Public-private project to deliver mouse genome in 6 months.

Research on the mouse genome lurched into the fast lane last week, as private donors joined the U.S. government to step on the gas. A public-private consortium announced on 6 October that it's kicking $58 million into a new fund that will pay to sequence the DNA of the "black six" (C57BL/6J) strain of laboratory mouse. The consortium aims to produce a draft version of the genome by the end of February.

Animals↗

Integration of chicken genomic resources to enable whole-genome sequencing.

Different genomic resources in chicken were integrated through the Wageningen chicken BAC library. First, a BAC anchor map was created by screening this library with two sets of markers: microsatellite markers from the consensus linkage map and markers created from BAC end sequencing in chromosome walking experiments. Second, HINdIII digestion fingerprints were created for all BACs of the Wageningen chicken BAC library. Third, cytogenetic positions of BACs were assigned by FISH. These integrated resources will facilitate further chromosome-walking experiments and whole-genome sequencing.

Animals↗

Characterization of the genome of molluscum contagiosum virus type 1 between the genome coordinates 0.045 and 0.075 by DNA nucleotide sequence analysis of a 5.6-kb HindIII/MluI DNA fragment.

The complete DNA nucleotide sequence of a HindIII/MluI genomic DNA fragment (0.045-0.075 viral map units) from molluscum contagiosum virus type 1 (MCV-1) was determined. The HindIII/MluI DNA fragment comprises 5,646 bp with a base composition of 64.4% G + C and 35.6% A + T. The DNA sequence contains many perfect direct repeats. A cluster of three repetitive DNA elements R1, R2 and R3, with a complex structural arrangement was detected between nucleotide positions 1802 and 2107. The unit length (box) of the repetitive DNA sequences was found to be 6 bp (15 boxes) and 9 bp (24 boxes) for R1 and R2, respectively. The repetitive DNA element R3 is organized in fifteen boxes (15 bp) in which a unit length of R1 is combined with a unit length of R2. The arrangement of the repetition R3 within the DNA sequences of this particular region of the MCV-1 genome was found to be (5 x R3) + (2 x R2) + (1 x R3) + (6 x R2) + (1 x R3) + (1 x R2) + (8 x R3). Twenty-three open reading frames (ORFs) of 60-1,175 amino acid (AA) residues were detected. The largest ORF (number 17) comprises 1,175 AA with a predicted molecular weight of 126 kD. This ORF harbors a promoter signal which is located 21 nucleotides upstream from the start codon and is very similar to the early promoter signals known for vaccinia virus. This putative protein contains glutamine-enriched regions between AA residues 427 and 682 which show homologies to the corresponding glutamine-enriched regions of a variety of cellular genes like human transcriptional initiation factor (TFIID: TATA box factor).

Amino Acid Sequence↗

Global genomic and antimicrobial resistance profiling of Neisseria gonorrhoeae: Insights from whole genome sequencing and minimum inhibitory concentration analysis.

BACKGROUND: The rising antimicrobial resistance (AMR) of Neisseria gonorrhoeae is a major global health concern that limits treatment options and complicates disease management. Efflux pump systems and resistance genes are key to bacteria's ability to evade antibiotics. This study examined the genetic and phenotypic resistance landscape using a large dataset of whole-genome sequences to identify key resistance mechanisms, assess efflux pump gene prevalence, and analyze regional variations in Minimum Inhibitory Concentration (MIC) values to inform treatment strategies and public health interventions. METHODS: A total of 38,585 whole-genome sequences of N. gonorrhoeae were analyzed to identify AMR determinants. This study focused on the presence and distribution of efflux pump genes (mtrC, farB, norM, and mtrA) and specific resistance genes, including tet(C) (tetracycline resistance) and aph(3')-Ia (aminoglycoside resistance). The MIC values were assessed for multiple antibiotics to evaluate resistance trends and regional variations, including penicillin, spectinomycin, zoliflodacin, gentamicin, and fluoroquinolones. RESULTS: This analysis revealed widespread resistance to multiple antibiotics. Efflux pump genes (mtrC, farB, norM, and mtrA) were found in nearly all isolates, highlighting their essential roles in resistance and adaptation. The presence of tet(C) and aph (3')-Ia varied across different Gene Presence Patterns, suggesting that regional or therapeutic factors may influence tetracycline and aminoglycoside resistance. High MIC values for penicillin were observed, likely because of blaTEM, a beta-lactamase gene responsible for beta-lactam resistance. Resistance to spectinomycin is also widespread, raising concerns about the diminishing efficacy of this antibiotic. In contrast, zoliflodacin, gentamicin, and fluoroquinolones exhibited relatively low MIC values, indicating their sustained effectiveness against N. gonorrhoeae. DISCUSSION: Efflux pump systems are key to N. gonorrhoeae resistance and adaptability. Regional MIC variations indicate that local antibiotic use shapes resistance patterns. The high resistance to penicillin and spectinomycin highlights the need for alternative treatments, whereas zoliflodacin and fluoroquinolones remain effective but require monitoring. This study emphasizes global AMR surveillance, novel therapies, and targeted antimicrobial stewardship to address multidrug-resistant infections.

Neisseria gonorrhoeae↗