PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 343 records · Page 19Linked to original sources

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence↗

The detection of nucleotide sequences with strong similarity to hormone responsive elements in the genome of eubacteria and archaebacteria and their possible relation to similar sequences present in the mitochondrial genome.

To account for the presence of nucleotide sequences in mitochondria with similarity to the Hormone Response Elements (HREs) of the nuclear genomes of man, rat and mouse, the genomes of several procaryotes have been screened for the presence of the sequences AGAACA NNN TGTTCT and GGTACA NNN TGTTCT, which represent perfect palindromic and consensus class I HREs, respectively, and for the sequence AGGTCA NNN TGACCT, which represents class II HRE. In many of the examined procaryotes, eubacteria and archaebacteria, almost perfect palindromic class I HREs and perfect or almost perfect class II half palindromic HREs have been detected in various genes, some of which encode proteins involved in energy metabolism, in replication and in transcription control. These findings support the hypothesis that the similar sequences found in mitochondria, potentially involved in hormonal regulation of respiratory enzyme biosynthesis, were introduced into eucaryotic cell by the procaryotic endosymbionts.

Animals↗

Analysis of six DNA components of the faba bean necrotic yellows virus genome and their structural affinity to related plant virus genomes.

Faba bean necrotic yellows virus (FBNYV) has a multicomponent circular ssDNA genome. In addition to a previously described genome component (C1) coding for a replicase-associated protein (Rep), five further components (C2 to C6) have now been identified. Each of the six components is about 1 kb in size, contains one major open reading frame (ORF) in the virion sense with a TATA box and polyadenylation signal, and has a noncoding region containing a highly conserved sequence possibly forming a stem-loop structure. Similar to C1, C2 encodes another putative Rep of 33.1 kDa, which is closely related to the Rep of banana bunchy top virus (BBTV). Based on bacterial expression and immunoblot analysis, the ORF of C5 encodes the capsid protein (CP) with a deduced molecular mass of 19 kDa. The FBNYV CP shares the highest amino acid (aa) identity (56.2%) with that of subterranean clover stunt virus (SCSV). The ORF of C4 potentially codes for a hydrophobic protein which appears to be structurally and functionally similar to the BBTV-C4 and SCSV-C1 proteins. No protein sequence similarities were found in databases for the C3 and C6 ORFs of FBNYV. FBNYV is clearly distinct from any known virus but is taxonomically related to BBTV and SCSV.

Amino Acid Sequence↗

Phylogenetic relationship of the complete Rauscher murine leukemia virus genome with other murine leukemia virus genomes.

We report the complete nucleotide sequence of the genome of Rauscher murine leukemia virus (R-MuLV), the replication-competent helper virus present in the Rauscher virus complex, and its phylogenetic relationship with other murine leukemia virus genomes. An overall sequence identity of 97.6% was found between R-MuLV and the Friend helper virus (F-MuLV), and the two viruses were closely related on the phylogenetic trees constructed from either gag, pol, or env sequences. Moloney murine leukemia virus (Mo-MuLV) was the next closest relative to R-MuLV and F-MuLV on all trees, followed by Akv and radiation leukemia virus (RadLV). The most distantly related helper virus was Hortulanus murine leukemia virus (Ho-MuLV). Interestingly, Cas-Br-E branched with Mo-MuLV on the gag and pol trees, whereas on the env tree, it revealed the highest degree of relatedness to Ho-MuLV, possibly due to an ancient recombination with an Ho-MuLV ancestor. In summary, a phylogenetic analysis involving various MuLVs has been performed, in which the postulated close relationship between R-MuLV and F-MuLV has been confirmed, consistent with the pathobiology of the two viruses.

Algorithms↗

The genome nucleotide sequence of a contemporary wild strain of measles virus and its comparison with the classical Edmonston strain genome.

The only complete genome nucleotide sequences of measles virus (MeV) reported to date have been for the Edmonston (Ed) strain and derivatives, which were isolated decades ago, passaged extensively under laboratory conditions, and appeared to be nonpathogenic. Partial sequencing of many other strains has identified >/=15 genotypes. Most recent isolates, including those typically pathogenic, belong to genotypes distinct from the Edmonston type. Therefore, the sequence of Ed and related strains may not be representative of those of pathological measles circulating at that or any time in human populations. Taking into account these issues as well as the fact that so many studies have been based upon Ed-related strains, we have sequenced the entire genome of a recently isolated pathogenic strain, 9301B. Between this recent isolate and the classical Ed strain, there were 465 nucleotide differences (2.93%) and 114 amino acid differences (2.19%). Computation of nonsynonymous and synonymous substitutions in open reading frames as well as direct comparisons of noncoding regions of each gene and extracistronic regulatory regions clearly revealed the regions where changes have been permissible and nonpermissible. Notably, considerable nonsynonymous substitutions appeared to be permissible for the P frame to maintain a high degree of sequence conservation for the overlapping C frame. However, the cause and the effect were largely unclear for any substitution, indicating that there is a considerable gap between the two strains that cannot be filled. The sequence reported here would be useful as a reference of contemporary wild-type MeV.

3' Untranslated Regions↗

Characteristics of nucleotide substitution in the hepatitis C virus genome: constraints on sequence change in coding regions at both ends of the genome.

Comparison of complete genome sequences for different variants of hepatitis C virus (HCV) reveals several different constraints on sequence change. Synonymous changes are suppressed in coding regions at both 5' and 3' ends of the genome. No evidence was found for the existence of alternative reading frames or for a lower mutation frequency in these regions. Instead, suppression may be due to constraints imposed by RNA secondary structures identified within the core and NS5b genes. Nonsynonymous substitutions are less frequent than synonymous ones except in the hypervariable region of E2 and, to a lesser extent, in E1, NS2, and NS5b. Transitions are more frequent than transversions, particularly at the third position of codons where the bias is 16:1. In addition, nucleotide substitutions may not occur symmetrically since there is a bias toward G or C at the third position of codons, while T left and right arrow C transitions were twice as frequent as A left and right arrow G transitions. These different biases do not affect the phylogenetic analysis of HCV variants but need to be taken into account in interpreting sequence change in longitudinal studies.

Base Sequence↗

Complete nucleotide sequence and genome organization of sweet potato feathery mottle virus (S strain) genomic RNA: the large coding region of the P1 gene.

The complete nucleotide sequence of a sweet potato feathery mottle virus severe strain (SPFMV-S) genomic RNA was determined from overlapping cDNA clones and by directly sequencing viral RNA. The viral RNA genome is 10,820 nucleotides long, excluding the poly(A) tail and contains one open reading frame (ORF) starting at nucleotide 118 and ending at 10,599, potentially encoding a polyprotein of 3,493 amino acids (Mr 393,800). The ORF was followed by a 3' untranslated region of 221 nucleotides. The deduced polyprotein includes P1 (74K), HC-Pro (52K), P3 (46K), 6K1, CI (72K), 6K2, NIa-VPg (22K), NIa-Pro (28K), NIb (60K) and coat (35K) proteins, after an analysis of protein cleavage sites analogous to other potyvirus polyproteins. The polyprotein had a high level of amino acid identity with those of other potyviruses, except in the regions of P1 and P3. The P1 of SPFMV-S RNA has 664 amino acid residues, and is the largest and least similar to those of other potyviruses. HC-Pro and CI show high identity with those of other potyviruses. P3 has relatively low identity, however, the length of P3 was within the range of variability among other potyviruses. The 6K1 protein between P3 and C1 is also highly similar to those of other potyviruses. This is the first report on the complete nucleotide sequence of the sweet potato-infecting virus.

Genome, Viral↗

Identification of four genomic loci highly related to casein-kinase-2-alpha cDNA and characterization of a casein kinase-2-alpha pseudogene within the mouse genome.

Using the coding region of the human CK-2 alpha cDNA as a probe for screening a genomic mouse library, positive clones representing four different genomic loci were isolated. Partial DNA sequences of these loci encompassing the first 120 nucleotides of the putative coding region are reported. One positive clone was further analyzed by sequencing a 3.1 kb XbaI fragment. This clone displays the characteristics of a pseudogene, i.e. lack of introns and several nucleotide insertions and deletions. In its 3' region it contains a 91 bp large CT-rich stretch which consists of (CCTT) and (CT) repeats; in the 5' region three (CCCCCT) repeats.

Animals↗

Cereal genome evolution: pastoral pursuits with 'Lego' genomes.

The rapid progress in comparative analysis of cereal genomes reveals that they are composed of similar genomic building blocks. It seems that by simply rearranging these blocks and amplifying some of the repetitive sequences contained within them, it is possible to reconstitute the 56 different chromosomes found in wheat, rice, maize, sorghum, millet and sugarcane. Comparison of the orders of blocks in these reconstituted chromosomes reveals that the cleavage of a single chromosome formed from the blocks could give rise to all the combinations found in the chromosomes of the above species. A framework is now in place for collating all the information which has been generated from studying the individual cereals.

Biological Evolution↗

Genome partitioning and whole-genome analysis.

Standard DNA marker-based approaches to mapping genes that influence complex traits typically consider a limited number of hypotheses. Most of these hypotheses concentrate on the effect of a single individual locus (or relatively few loci) on the trait of interest. Although of tremendous importance scientifically, such hypotheses do not accommodate the full range of genetic phenomena that may contribute to phenotypic expression. We present novel approaches to complex trait analysis that make as complete use of marker information as is possible. The proposed methodologies can be used to entertain a wide variety of hypotheses, including those that engage, for example, the contribution of a particular chromosome, genome-wide heterozygosity, and multiple genomic regions, to phenotypic expression. We consider a number of possible extensions of the proposed methods as well as their limitations. Although we discuss many methodological details in the context of quantitative trait locus mapping involving sampling units such as human pedigrees and hybrids resulting from crosses between inbred strains of model organisms, our procedures can be easily adapted to standard sibpair and other sampling unit-based designs. Ultimately, the proposed approaches not only have the potential to increase power to identify individual loci that harbor trait-influencing genes, but also present a framework for testing a number of hypotheses about the nature of the genetic determinants of phenotypes in general.

Alleles↗

Genome sequences: genome sequence of a model prokaryote.

The complete Escherichia coli genome sequence is now known; it should greatly facilitate the analysis of other genomes, but a lot remains to be learnt about E. coli itself. About half the genes were previously uncharacterized, but expanding databases and improving analysis methods will help predict their functions.

Bacterial Proteins↗

Whole genome analysis: experimental access to all genome sequenced segments through larger-scale efficient oligonucleotide synthesis and PCR.

The recent ability to sequence whole genomes allows ready access to all genetic material. The approaches outlined here allow automated analysis of sequence for the synthesis of optimal primers in an automated multiplex oligonucleotide synthesizer (AMOS). The efficiency is such that all ORFs for an organism can be amplified by PCR. The resulting amplicons can be used directly in the construction of DNA arrays or can be cloned for a large variety of functional analyses. These tools allow a replacement of single-gene analysis with a highly efficient whole-genome analysis.

Animals↗

Chromosomal losses and gains in meningiomas: comparative genomic hybridization (CGH) study of the whole genome.

We investigated chromosomal aberrations in meningiomas using newly developed comparative genomic hybridization (CGH) technique and compared the results with the proliferating potential of the tumors. This technique permits the entire genome to be surveyed in one session of experiments. Our results revealed chromosomal aberrations in 5 out of 10 (50%) of the tumor samples studied. Losses of the distal parts of chromosome 1p (5 out of 10) and 22q (3 out of 10) were the two most frequent chromosomal aberrations. Losses and/or gains in other regions were only sporadic. The MIB-1 staining indices (MIB-SI, %) were 1.9 +/- 0.9% (mean +/- SD) in benign (n = 8), 4.5% in atypical (n = 1), and 11.7% in anaplastic (n = 1) meningiomas. The comparison of MIB-SI between the tumors with (2.3 +/- 0.6%) and without (1.6 +/- 0.3%) chromosomal aberrations demonstrated a trend towards an increased MIB-SI in meningiomas with chromosomal aberrations (p < 0.07) by unpaired Student's t-test. This study suggests that alterations in chromosomes 1p and 22q could be a primary focus of further detailed assessment of tumorigenesis and in understanding the biological behavior of meningiomas.

Adolescent↗

Genomes OnLine Database (GOLD 1.0): a monitor of complete and ongoing genome projects world-wide.

UNLABELLED: GOLD (Genomes On Line Database) is a World Wide Web resource for comprehensive access to information regarding complete and ongoing genome projects around the world. AVAILABILITY: GOLD is based at the University of Illinois at Urbana-Champaign and is available at http://geta.life.uiuc.edu/ approximately nikos/genomes. html. It is also mirrored at the European Bioinformatics Institute at http://www.ebi.ac.uk/research/cgg/genomes.html. CONTACT: genomes@ebi.ac.uk

Databases, Factual↗

Viral Genome DataBase: storing and analyzing genes and proteins from complete viral genomes.

SUMMARY: The Viral Genome DataBase (VGDB) contains detailed information of the genes and predicted protein sequences from 15 completely sequenced genomes of large (&100 kb) viruses (2847 genes). The data that is stored includes DNA sequence, protein sequence, GenBank and user-entered notes, molecular weight (MW), isoelectric point (pI), amino acid content, A + T%, nucleotide frequency, dinucleotide frequency and codon use. The VGDB is a mySQL database with a user-friendly JAVA GUI. Results of queries can be easily sorted by any of the individual parameters. AVAILABILITY: The software and additional figures and information are available at http://athena.bioc.uvic.ca/genomes/index.html .

Databases, Factual↗

Sequence analysis of the genome of the unicellular cyanobacterium Synechocystis sp. strain PCC6803. I. Sequence features in the 1 Mb region from map positions 64% to 92% of the genome.

The contiguous sequence of 1,003,450 bp spanning map positions 64% to 92% of the genome of Synechocystis sp. strain PCC6803 has been deduced. Computer analysis of the sequence predicts that this region contains at least 818 potential ORFs, in which 255 (31%) were either genes that had already been identified or their homologues, 84 (10%) were homologues to registered hypothetical genes, and 149 (18%) showed weak similarities to reported genes. The remaining 330 ORFs showed no apparent similarity to any reported genes or carried no significant protein motifs. The potential ORFs as a whole occupied 86% of the sequenced region, implying compact arrangement of genes in the genome. As to the structural RNA genes, one rRNA operon consisting of 5,028 bp and at least 11 species of tRNA genes were identified. It is noteworthy that 10 out of the 11 tRNA species showed significant sequence similarities to tRNAs reported in plant chloroplasts. As other notable unique sequences, three classes of IS-like elements each with characteristics typical of IS elements were identified, and a typical unit of WD(Trp-Asp)-repeats which have only been detected in the regulatory proteins of eukaryotes was identified within the large 5,079-bp ORF located at map position 69%.

Base Sequence↗

Complete genome sequence of enterohemorrhagic Escherichia coli O157:H7 and genomic comparison with a laboratory strain K-12.

Escherichia coli O157:H7 is a major food-borne infectious pathogen that causes diarrhea, hemorrhagic colitis, and hemolytic uremic syndrome. Here we report the complete chromosome sequence of an O157:H7 strain isolated from the Sakai outbreak, and the results of genomic comparison with a benign laboratory strain, K-12 MG1655. The chromosome is 5.5 Mb in size, 859 Kb larger than that of K-12. We identified a 4.1-Mb sequence highly conserved between the two strains, which may represent the fundamental backbone of the E. coli chromosome. The remaining 1.4-Mb sequence comprises of O157:H7-specific sequences, most of which are horizontally transferred foreign DNAs. The predominant roles of bacteriophages in the emergence of O157:H7 is evident by the presence of 24 prophages and prophage-like elements that occupy more than half of the O157:H7-specific sequences. The O157:H7 chromosome encodes 1632 proteins and 20 tRNAs that are not present in K-12. Among these, at least 131 proteins are assumed to have virulence-related functions. Genome-wide codon usage analysis suggested that the O157:H7-specific tRNAs are involved in the efficient expression of the strain-specific genes. A complete set of the genes specific to O157:H7 presented here sheds new insight into the pathogenicity and the physiology of O157:H7, and will open a way to fully understand the molecular mechanisms underlying the O157:H7 infection.

Bacterial Proteins↗

A genome-wide association study identified 10 novel genomic loci associated with intrinsic capacity.

BACKGROUND: Intrinsic capacity (IC) is a multidimensional concept within the World Health Organization framework for healthy aging. It refers to the composite of an individual's physical and mental capacities that enable them to maintain well-being, functional ability, and engagement in valued activities throughout life. While substantial evidence supports the biological basis of IC and its subdomains, the extent to which genetic factors influence IC remains largely unexplored, with no studies currently available. METHODS: Using datasets from the UK Biobank (UKB; N&#x2009;=&#x2009;44 631) and the Canadian Longitudinal Study on Aging (CLSA; N&#x2009;=&#x2009;13 085), we implemented the restricted maximum likelihood method to estimate SNP-based heritability (h2snp), followed by a Genome-Wide Association Study (GWAS) to identify genetic variants associated with IC, and post-GWAS analyses to pinpoint biological implications. RESULTS: The h2snp for IC was estimated at 25.2% in UKB and 19.5% in CLSA. Our GWAS identified 38 independent SNPs for IC across 10 genomic loci and 4289 candidate SNPs, mapped to 197 genes. Post-GWAS analysis revealed the role of these genes in cellular processes such as cell proliferation, immune function, metabolism, and neurodegeneration, with high expression in muscle, heart, brain, adipose, and nerve tissues. Of the 52 traits tested, 23 showed significant genetic correlations with IC, and a higher genetic loading for IC was associated with higher IC scores. CONCLUSIONS: Overall, this study provides comprehensive evidence on the genetic architecture of IC, identifying novel genetic variants and biological pathways, advancing our current knowledge and laying the foundation for ongoing and future research on healthy aging.

Adult↗