PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 469 records · Page 26Linked to original sources

Discovery of previously unidentified genomic disorders from the duplication architecture of the human genome.

Genomic disorders are characterized by the presence of flanking segmental duplications that predispose these regions to recurrent rearrangement. Based on the duplication architecture of the genome, we investigated 130 regions that we hypothesized as candidates for previously undescribed genomic disorders. We tested 290 individuals with mental retardation by BAC array comparative genomic hybridization and identified 16 pathogenic rearrangements, including de novo microdeletions of 17q21.31 found in four individuals. Using oligonucleotide arrays, we refined the breakpoints of this microdeletion, defining a 478-kb critical region containing six genes that were deleted in all four individuals. We mapped the breakpoints of this deletion and of four other pathogenic rearrangements in 1q21.1, 15q13, 15q24 and 17q12 to flanking segmental duplications, suggesting that these are also sites of recurrent rearrangement. In common with the 17q21.31 deletion, these breakpoint regions are sites of copy number polymorphism in controls, indicating that these may be inherently unstable genomic regions.

Chromosome Breakage↗

Comparative genomics: genome-wide analysis in metazoan eukaryotes.

The increasing number of complete and nearly complete metazoan genome sequences provides a significant amount of material for large-scale comparative genomic analysis. Finding new effective methods to analyse such enormous datasets has been the object of intense research. Three main areas in comparative genomics have recently shown important developments: whole-genome alignment, gene prediction and regulatory-region prediction. Each of these areas improves the methods of deciphering long genomic sequences and uncovering what lies hidden in them.

Animals↗

Simultaneous identification of A, B, D and R genomes by genomic in situ hybridization in wheat-rye derivatives.

Multicolour genomic in situ hybridization was carried out in wheat-rye hybrids and in a wheat-rye translocation line. Different hybridization conditions and mixture compositions were used, and A, B and D genomes of hexaploid wheat as well as the R genome of rye were distinguished simultaneously in somatic cells. Combination of genomic and rDNA probes in multicolour in situ hybridization was also performed to identify chromosomes within a specific genome.

DNA, Plant↗

Identification of novel genomic markers related to progression to glioblastoma through genomic profiling of 25 primary glioma cell lines.

Identification of genetic copy number changes in glial tumors is of importance in the context of improved/refined diagnostic, prognostic procedures and therapeutic decision-making. In order to detect recurrent genomic copy number changes that might play a role in glioma pathogenesis and/or progression, we characterized 25 primary glioma cell lines including 15 non glioblastoma (non GBM) (I-III WHO grade) and 10 GBM (IV WHO grade), by array comparative genomic hybridization, using a DNA microarray comprising approx. 3500 BACs covering the entire genome with a 1 Mb resolution and additional 800 BACs covering chromosome 19 at tiling path resolution. Combined evaluation by single clone and whole chromosome analysis plus 'moving average (MA) approach' enabled us to confirm most of the genetic abnormalities previously identified to be associated with glioma progression, including +1q32, +7, -10, -22q, PTEN and p16 loss, and to disclose new small genomic regions, some correlating with grade malignancy. Grade I-III gliomas exclusively showed losses at 3p26 (53%), 4q13-21 (33%) and 7p15-p21 (26%), whereas only GBMs exhibited 4p16.1 losses (40%). Other recurrent imbalances, such as losses at 4p15, 5q22-q23, 6p23-25, 12p13 and gains at 11p11-q13, were shared by different glioma grades. Three intervals with peak of loss could be further refined for chromosome 10 by our MA approach. Data analysis of full-coverage chromosome 19 highlighted two main regions of copy number gain, never described before in gliomas, at 19p13.11 and 19q13.13-13.2. The well-known 19q13.3 loss of heterozygosity area in gliomas was not frequently affected in our cell lines. Genomic hotspot detection facilitated the identification of small intervals resulting in positional candidate genes such as PRDM2 (1p36.21), LRP1B (2q22.3), ADARB2 (10p15.3), BCCIP (10q26.2) and ING1 (13q34) for losses and ECT2 (3q26.3), MDK, DDB2, IG20 (11p11.2) for gains. These data increase our current knowledge about cryptic genetic changes in gliomas and may facilitate the further identification of novel genetic elements, which may provide us with molecular tools for the improved diagnostics and therapeutic decision-making in these tumors.

Cell Line, Tumor↗

Automated array-based genomic profiling in chronic lymphocytic leukemia: development of a clinical tool and discovery of recurrent genomic alterations.

B cell chronic lymphocytic leukemia (B-CLL) is characterized by a highly variable clinical course. Recurrent chromosomal imbalances provide significant prognostic markers. Risk-adapted therapy based on genomic alterations has become an option that is currently being tested in clinical trials. To supply a robust tool for such large scale studies, we developed a comprehensive DNA microarray dedicated to the automated analysis of recurrent genomic imbalances in B-CLL by array-based comparative genomic hybridization (matrix-CGH). Validation of this chip in a series of 106 B-CLL cases revealed a high specificity and sensitivity that fulfils the criteria for application in clinical oncology. This chip is immediately applicable within clinical B-CLL treatment trials that evaluate whether B-CLL cases with distinct chromosomal abnormalities should be treated with chemotherapy of different intensities and/or stem cell transplantation. Through the control set of DNA fragments equally distributed over the genome, recurrent genomic imbalances were discovered: trisomy of chromosome 19 and gain of the MYCN oncogene correlating with an elevation of MYCN mRNA expression.

Automation↗

Comparative genomic hybridization using oligonucleotide microarrays and total genomic DNA.

Array-based comparative genomic hybridization (CGH) measures copy-number variations at multiple loci simultaneously, providing an important tool for studying cancer and developmental disorders and for developing diagnostic and therapeutic targets. Arrays for CGH based on PCR products representing assemblies of BAC or cDNA clones typically require maintenance, propagation, replication, and verification of large clone sets. Furthermore, it is difficult to control the specificity of the hybridization to the complex sequences that are present in each feature of such arrays. To develop a more robust and flexible platform, we created probe-design methods and assay protocols that make oligonucleotide microarrays synthesized in situ by inkjet technology compatible with array-based comparative genomic hybridization applications employing samples of total genomic DNA. Hybridization of a series of cell lines with variable numbers of X chromosomes to arrays designed for CGH measurements gave median ratios for X-chromosome probes within 6% of the theoretical values (0.5 for XY/XX, 1.0 for XX/XX, 1.4 for XXX/XX, 2.1 for XXXX/XX, and 2.6 for XXXXX/XX). Furthermore, these arrays detected and mapped regions of single-copy losses, homozygous deletions, and amplicons of various sizes in different model systems, including diploid cells with a chromosomal breakpoint that has been mapped and sequenced to a precise nucleotide and tumor cell lines with highly variable regions of gains and losses. Our results demonstrate that oligonucleotide arrays designed for CGH provide a robust and precise platform for detecting chromosomal alterations throughout a genome with high sensitivity even when using full-complexity genomic samples.

Cell Line↗

Complete genome sequence and comparative genomic analysis of an emerging human pathogen, serotype V Streptococcus agalactiae.

The 2,160,267 bp genome sequence of Streptococcus agalactiae, the leading cause of bacterial sepsis, pneumonia, and meningitis in neonates in the U.S. and Europe, is predicted to encode 2,175 genes. Genome comparisons among S. agalactiae, Streptococcus pneumoniae, Streptococcus pyogenes, and the other completely sequenced genomes identified genes specific to the streptococci and to S. agalactiae. These in silico analyses, combined with comparative genome hybridization experiments between the sequenced serotype V strain 2603 V/R and 19 S. agalactiae strains from several serotypes using whole-genome microarrays, revealed the genetic heterogeneity among S. agalactiae strains, even of the same serotype, and provided insights into the evolution of virulence mechanisms.

Amino Acid Sequence↗

Impact of genomic imprinting on genomic instability and radiation-induced mutation.

PURPOSE: The purpose of this review is to assess the effect of radiation-induced mutation on genes subject to genomic imprinting, and the consequences of this on the understanding of genomic instability. Genomic imprinting is the phenomenon in which one of the two alleles of a gene is expressed or suppressed depending on the gamete from which it was inherited, thus effectively rendering a cell hemizygous for the expression of certain key genes. The consequence of this is that such loci are potentially more likely targets for mutagenesis since one allele is normally inactive. This is not only important in the recognition of a subgroup of target genes for radiation-induced damage, but also raises the possibility of mutations affecting the epigenotype of key tumour suppressor or tumour promoting genes. Such mutations may in principle affect the stability of imprinting and may fall into a novel class of 'epimutation', where the DNA sequence is not affected, but post-transcriptional mechanisms of epigenotype maintenance are stably altered. These novel mechanisms are discussed in relation with radiation-induced genomic instability and the heritability of tumour predisposition from radiation-exposed parents. CONCLUSIONS: As yet there is only circumstantial evidence that the targets for radiation-induced DNA and epigenetic damage are imprinted genes, or genes involved in the maintenance of the epigenotype. However, the potential consequences of such genes being important targets for the generation of genomic instability or other forms of damage are serious and could affect the interpretation of the risks of low dose radiation exposure and of epidemiological data.

Alleles↗

Using linkage genome scans to improve power of association in genome scans.

Scanning the genome for association between markers and complex diseases typically requires testing hundreds of thousands of genetic polymorphisms. Testing such a large number of hypotheses exacerbates the trade-off between power to detect meaningful associations and the chance of making false discoveries. Even before the full genome is scanned, investigators often favor certain regions on the basis of the results of prior investigations, such as previous linkage scans. The remaining regions of the genome are investigated simultaneously because genotyping is relatively inexpensive compared with the cost of recruiting participants for a genetic study and because prior evidence is rarely sufficient to rule out these regions as harboring genes with variation of conferring liability (liability genes). However, the multiple testing inherent in broad genomic searches diminishes power to detect association, even for genes falling in regions of the genome favored a priori. Multiple testing problems of this nature are well suited for application of the false-discovery rate (FDR) principle, which can improve power. To enhance power further, a new FDR approach is proposed that involves weighting the hypotheses on the basis of prior data. We present a method for using linkage data to weight the association P values. Our investigations reveal that if the linkage study is informative, the procedure improves power considerably. Remarkably, the loss in power is small, even when the linkage study is uninformative. For a class of genetic models, we calculate the sample size required to obtain useful prior information from a linkage study. This inquiry reveals that, among genetic models that are seemingly equal in genetic information, some are much more promising than others for this mode of analysis.

Genetic Linkage↗

Triplet repeats in human genome: distribution and their association with genes and other genomic regions.

MOTIVATION: Simple sequence repeats (SSRs) or microsatellite repeats are found abundantly in many prokaryotic and eukaryotic genomes. Among SSRs, triplet repeats are of special significance because some of them have been linked to various genetic disorders. The objective of the study is to analyze the triplet repeats of complete human genome and to identify the genes that contain the triplet repeats in their coding region. The analysis will help us to identify the candidate genes that have potential for repeat expansion. RESULTS: We have analyzed triplet repeats in the complete human genome from the publicly available sequences. Our analysis revealed that AGC and CCG repeat were predominantly present in the coding regions of the genome while UTRs and the upstream sequences contained CCG repeats in relative abundance. Analysis of density of triplet repeats (bp/Mb) revealed that AAT and AAC were the abundant repeats whereas ACT and ACG were the rare repeats found in human genome. We could identify about 2135 known or predicted genes that were associated with at least one of the triplet repeat types. A large proportion of putative transcripts that were identified by gene finding programs were found to be associated with triplet repeats. These transcripts will be the candidate genes for analysis of triplet repeat expansion and a possible association with disease phenotypes. Identification of 171 genes which contain a minimum of ten repeat units will be of particular interest in future in correlating their association with any disease phenotype due to the expansion potential of repeats present in them. The list of genes and other details of analysis are given in the online supplementary data (http://www.ingenovis.com/tripletrepeats).

Databases, Nucleic Acid↗

Genomic island identification in Vibrio vulnificus reveals significant genome plasticity in this human pathogen.

UNLABELLED: Genomic islands (GIs) are large chromosomal regions present in a subset of bacterial strains that increase the fitness of the organism under specific conditions. We compared the complete genome sequences of two Vibrio vulnificus strains YJ016 and CMCP6 and identified 14 regions (ranging in size from 14 to 117 kb), which had the characteristics of GIs. Bioinformatic analysis of these 14 GI regions identified the presence of phage-like integrase genes, aberrant GC content and genome signature (dinucleotide frequency) within each GI compared with the core genome indicating that these regions were acquired from an anomalous source. We examined the distribution of the nine GIs from strain YJ016 among 27 V. vulnificus isolates and found that most GIs were absent from the majority of these isolates. The chromosomal insertion sites of three GIs were adjacent to tRNA sites, which contained novel horizontally acquired DNA in all six available sequenced Vibrionaceae genomes. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

Adaptation, Physiological↗

Genomic Footprints of Historical Introgression Between Ancient Lineages of Wild Oryza AA-Genome Species With Widely Separated Contemporary Distributions.

Phylogenetic incongruence is increasingly recognized as pervasive, yet the extent to which reticulate evolution occurs between groups separated by substantial geographical distances and deep phylogenetic divergence remains poorly characterized. In the Oryza AA-genome group-a model for plant speciation and domestication-the traditional bifurcation model posits that Australian Oryza meridionalis and African Oryza longistaminata occupy basal branches, distinct from the more recently diversified monophyletic clade comprising Asian and other African lineages, including major cultivars. However, recent evidence from endogenous viral sequences has hinted at unexpected genetic relatedness between African O. longistaminata and Asian Oryza sativa, which are geographically and phylogenetically distant. Here, we conducted a genome-wide survey across 11 Oryza species to systematically identify genomic regions exhibiting phylogenetic incongruence. Widespread phylogenetic discordance was observed, notably involving genomic segments in which O. longistaminata showed phylogenetic proximity to Asian species, contradicting their established deep divergence. To distinguish between introgression and incomplete lineage sorting, we performed four-taxon ABBA-BABA tests, which provided statistical support for introgression. Furthermore, divergence time estimates for these incongruent regions were younger than the species divergence times, suggesting historical introgression between the ancestors of lineages that are currently separated by vast geographical distances. Systematic assessments indicated that potential analytical artifacts, such as compositional bias and substitution saturation, were unlikely to explain the observations. These convergent lines of evidence suggest that ancient introgression had occurred between currently geographically separated and evolutionarily divergent Oryza lineages, leaving detectable footprints across their modern genomes.

Oryza↗

Genomic divergence between human and chimpanzee estimated from large-scale alignments of genomic sequences.

To study the genomic divergence between human and chimpanzee, large-scale genomic sequence alignments were performed. The genomic sequences of human and chimpanzee were first masked with the RepeatMasker and the repeats were excluded before alignments. The repeats were then reinserted into the alignments of nonrepetitive segments and entire sequences were aligned again. A total of 2.3 million base pairs (Mb) of genomic sequences, including repeats, were aligned and the average nucleotide divergence was estimated to be 1.22%. The Jukes-Cantor (JC) distances (nucleotide divergences) in nonrepetitive (1.44 Mb) and repetitive sequences (0.86 Mb) are 1.14% and 1.34%, respectively, suggesting a slightly higher average rate in repetitive sequences. Annotated coding and noncoding regions of homologous chimpanzee genes were also retrieved from GenBank and compared. The average synonymous and nonsynonymous divergences in 88 coding genes are 1.48% and 0.55%, respectively. The JC distances in intron, 5' flanking, 3' flanking, promoter, and pseudogene regions are 1.47%, 1.41%, 1.68%, 0.75%, and 1.39%, respectively. It is not clear why the genetic distances in most of these regions are somewhat higher than those in genomic sequences. One possible explanation is that some of the genes may be located in regions with higher mutation rates.

Animals↗

Reorganization of adjacent gene relationships in yeast genomes by whole-genome duplication and gene deletion.

In Saccharomyces, an ancient whole-genome duplication (WGD) and widespread duplicate gene deletion resulted in extensive reorganization of adjacent gene relationships. We have studied the evolution of adjacent gene pairs' identity, orientation, and spacing following whole-genome duplication and deletion (WGD-D) using comparative genomic analyses and simulations. Surveying adjacent gene organization across the Saccharomyces species complex, we find a genome-wide bias toward divergently and convergently transcribed gene pairs in all species but a reduction in this bias in the species that underwent WGD-D. Among neutral models of WGD-D, only single-gene deletion can produce the appropriate reduction in orientation bias and recapitulate the pattern of short, highly dispersed deletions we observe in Saccharomyces cerevisiae. To characterize the dynamics of WGD-D, we trace the conservation and creation of adjacent gene pairs along the S. cerevisiae lineage. We find that newly created adjacencies have a tandem orientation bias, while adjacencies conserved from prior to WGD-D have the same divergent-convergent bias as found in the species that diverged before WGD. We also find that adjacent gene pairs produced by WGD-D gained greater intergenic spacing but that this is reduced in the older adjacencies. Given this, and the preponderance of short deleted blocks, we argue that the deletion phase of WGD-D occurred primarily by small inactivating mutations followed by numerous small deletions. Newly created adjacent gene pairs also have an initial increase in mean log2 expression ratios and maximal expression levels, suggesting that increased intergenic spacing caused a genome-wide reduction in transcriptional interference.

Evolution, Molecular↗

Complete genome sequence of the alkaliphilic bacterium Bacillus halodurans and genomic sequence comparison with Bacillus subtilis.

The 4 202 353 bp genome of the alkaliphilic bacterium Bacillus halodurans C-125 contains 4066 predicted protein coding sequences (CDSs), 2141 (52.7%) of which have functional assignments, 1182 (29%) of which are conserved CDSs with unknown function and 743 (18. 3%) of which have no match to any protein database. Among the total CDSs, 8.8% match sequences of proteins found only in Bacillus subtilis and 66.7% are widely conserved in comparison with the proteins of various organisms, including B.subtilis. The B. halodurans genome contains 112 transposase genes, indicating that transposases have played an important evolutionary role in horizontal gene transfer and also in internal genetic rearrangement in the genome. Strain C-125 lacks some of the necessary genes for competence, such as comS, srfA and rapC, supporting the fact that competence has not been demonstrated experimentally in C-125. There is no paralog of tupA, encoding teichuronopeptide, which contributes to alkaliphily, in the C-125 genome and an ortholog of tupA cannot be found in the B.subtilis genome. Out of 11 sigma factors which belong to the extracytoplasmic function family, 10 are unique to B. halodurans, suggesting that they may have a role in the special mechanism of adaptation to an alkaline environment.

ATP-Binding Cassette Transporters↗

Genomic repeats, genome plasticity and the dynamics of Mycoplasma evolution.

Mycoplasmas evolved by a drastic reduction in genome size, but their genomes contain numerous repeated sequences with important roles in their evolution. We have established a bioinformatic strategy to detect the major recombination hot-spots in the genomes of Mycoplasma pneumoniae, Mycoplasma genitalium, Ureaplasma urealyticum and Mycoplasma pulmonis. This allowed the identification of large numbers of potentially variable regions, as well as a comparison of the relative recombination potentials of different genomic regions. Different trends are perceptible among mycoplasmas, probably due to different functional and structural constraints. The largest potential for illegitimate recombination in M.pulmonis is found at the vsa locus and its comparison in two different strains reveals numerous changes since divergence. On the other hand, the main M.pneumoniae and M.genitalium adhesins rely on large distant repeats and, hence, homologous recombination for variation. However, the relation between the existence of repeats and antigenic variation is not necessarily straightforward, since repeats of P1 adhesin were found to be anti-correlated with epitopes recognized by patient antibodies. These different strategies have important consequences for the structures of genomes, since large distant repeats correlate well with the major chromosomal rearrangements. Probably to avoid such events, mycoplasmas strongly avoid inverse repeats, in comparison to co-oriented repeats.

Adhesins, Bacterial↗

The complete nucleotide sequence and RNA editing content of the mitochondrial genome of rapeseed (Brassica napus L.): comparative analysis of the mitochondrial genomes of rapeseed and Arabidopsis thaliana.

The entire mitochondrial genome of rapeseed (Brassica napus L.) was sequenced and compared with that of Arabidopsis thaliana. The 221 853 bp genome contains 34 protein-coding genes, three rRNA genes and 17 tRNA genes. This gene content is almost identical to that of Arabidopsis: However the rps14 gene, which is a pseudo-gene in Arabidopsis, is intact in rapeseed. On the other hand, five tRNA genes are missing in rapeseed compared to Arabidopsis, although the set of mitochondrially encoded tRNA species is identical in the two Cruciferae. RNA editing events were systematically investigated on the basis of the sequence of the rapeseed mitochondrial genome. A total of 427 C to U conversions were identified in ORFs, which is nearly identical to the number in Arabidopsis (441 sites). The gene sequences and intron structures are mostly conserved (more than 99% similarity for protein-coding regions); however, only 358 editing sites (83% of total editings) are shared by rapeseed and Arabidopsis: Non-coding regions are mostly divergent between the two plants. One-third (about 78.7 kb) and two-thirds (about 223.8 kb) of the rapeseed and Arabidopsis mitochondrial genomes, respectively, cannot be aligned with each other and most of these regions do not show any homology to sequences registered in the DNA databases. The results of the comparative analysis between the rapeseed and Arabidopsis mitochondrial genomes suggest that higher plant mitochondria are extremely conservative with respect to coding sequences and somewhat conservative with respect to RNA editing, but that non-coding parts of plant mitochondrial DNA are extraordinarily dynamic with respect to structural changes, sequence acquisition and/or sequence loss.

Amino Acid Sequence↗

TMBETA-GENOME: database for annotated beta-barrel membrane proteins in genomic sequences.

We have developed the database, TMBETA-GENOME, for annotated beta-barrel membrane proteins in genomic sequences using statistical methods and machine learning algorithms. The statistical methods are based on amino acid composition, reside pair preference and motifs. In machine learning techniques, the combination of amino acid and dipeptide compositions has been used as main attributes. In addition, annotations have been made using the criterion based on the identification of beta-barrel membrane proteins and exclusion of globular and transmembrane helical proteins. A web interface has been developed for identifying the annotated beta-barrel membrane proteins in all known genomes. The users have the feasibility of selecting the genome from the three kingdoms of life, archaea, bacteria and eukaryote, and five different methods. Further, the statistics for all genomes have been provided along with the links to different algorithms and related databases. It is freely available at http://tmbeta-genome.cbrc.jp/annotation/.

Algorithms↗