PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 883 records · Page 49Linked to original sources

Genome-wide single-nucleotide polymorphism arrays demonstrate high fidelity of multiple displacement-based whole-genome amplification.

Whole-genome DNA amplification by multiple displacement (MD-WGA) is a promising tool to obtain sufficient DNA amounts from samples of limited quantity. Using Affymetrix' GeneChip Human Mapping 10K Arrays, we investigated the accuracy and allele amplification bias in DNA samples subjected to MD-WGA. We observed an excellent concordance (99.95%) between single-nucleotide polymorphisms (SNPs) called both in the nonamplified and the corresponding amplified DNA. This concordance was only 0.01% lower than the intra-assay reproducibility of the genotyping technique used. However, MD-WGA failed to amplify an estimated 7% of polymorphic loci. Due to the algorithm used to call genotypes, this was detected only for heterozygous loci. We achieved a 4.3-fold reduction of noncalled SNPs by combining the results from two independent MD-WGA reactions. This indicated that inter-reaction variations rather than specific chromosomal loci reduced the efficiency of MD-WGA. Consistently, we detected no regions of reduced amplification, with the exception of several SNPs located near chromosomal ends. Altogether, despite a substantial loss of polymorphic sites, MD-WGA appears to be the current method of choice to amplify genomic DNA for array-based SNP analyses. The number of nonamplified loci can be substantially reduced by amplifying each DNA sample in duplicate.

Genome, Human↗

The challenge of documenting mutation across the genome: the human genome variation society approach.

New methods for the detection of mutations and the completion of the human genome sequencing project have contributed to an exponential rise in variation information that must be collected, quality controlled, documented, and stored safely to ensure future availability to health care professionals, researchers, and others. There may be anywhere from one to more than 1,000 mutations in any given gene. To date, this information has been collected by general databases such as Online Mendelian Inheritance in Man (OMIM) or the Human Gene Mutation Database (HGMD), which collect only published mutations and, in the case of OMIM, selected published mutations. Unpublished mutations have made their way into Locus Specific Databases (LSDBs), and these can often contain as many unpublished mutations as published ones, in addition to other more detailed gene-specific information. LSDBs, however, do not exist for all genes at this time. Through their interactions, a number of members of the Human Genome Variation Society (HGVS) have developed nomenclature, standard software to curate mutations in gene specific databases, a WayStation to collect and review new mutations from research and diagnostic laboratories, and central databases to store and display these mutations and their associated phenotypes. Nomenclature is now well defined for the commonest types of mutation, with work continuing on systematically naming the more complex types. Other projects, such as dedicated specialized software for LSDBs, are in the early stages of development.

Computational Biology↗

De Novo Whole Genome Assemblies of Unusual Case-Making Caddisflies (Trichoptera) Highlight Genomic Convergence in the Composition of the Major Silk Gene (h-fibroin).

Trichoptera (caddisflies) is one of the most species-rich orders of aquatic insects. Species of caddisflies cover a broad ecological diversity as exemplified by various uses of underwater silk secretions. Diversity of silk use generally aligns with the evolution of major caddisfly lineages, specifically at the subordinal level: Annulipalpia (retreat makers) and Integripalpia (cocoon and tube-case makers). However, silk use within suborders differs for a few exceptional species in these clades. In this study, we provide the first whole genome assemblies and annotations for two unusual Integripalpia species: Limnocentropus insolitus, whose hard tube-case is anchored to boulders by a rigid, elongated silken stalk, and Phryganopsyche brunnea which builds a "floppy" cylindrical case that lacks the typical robustness of tube-cases. Its texture rather resembles that of the flexible retreats built by Annulipalpia. Using the two high-quality genome assemblies, we identified and annotated the major silk gene, h-fibroin, and compared its amino acid composition across various groups, including retreat, cocoon, and tube-case makers. Our phylogenetic analysis confirmed the phylogenetic position of the two species in the tube-case-making clade. The major silk gene of L. insolitus shows a similar amino acid composition to other tube-case-making species. In contrast, the amino acid composition of P. brunnea resembles that of retreat-making species, in particular with regard to the high content of proline. This is consistent with the hypothesis that proline could be linked to enhanced extensibility of silk fibers. Taken together, our results underscore the role of silk genes in shaping the evolutionary ecology of retreat- and tube-case-making in caddisflies.

Animals↗

Genomic insights into end-use grain quality and nutritional traits of an ancient Indian dwarf wheat ( Triticum sphaerococcum Percival) population using a multi-locus genome-wide association study.

BACKGROUND: Triticum sphaerococcum, an ancient hexaploid wheat species, is renowned for its stress resilience and superior nutritional quality. A panel of 116 T. sphaerococcum accessions (the largest known collection at a single site globally), with six bread wheat released varieties, was evaluated for its potential for genetic quality improvement. Field experiments were conducted under standard, heat and moisture-deficit conditions across two cropping seasons for ten grain end-use quality and nutritional traits. RESULTS: Genotypes showed highly significant differences (P ≤ 0.001) for measured traits, with high broad-sense heritability resulting from substantial genotypic variance contributions. Triticum sphaerococcum consistently outperformed T. aestivum across environments, with moisture-deficit stress proving more detrimental to quality parameters than heat stress, while micronutrient content increased under stressed conditions. Trait correlations revealed that the gluten index (GI) correlated negatively with the grain hardness index (GHI), wet gluten (WG), and water-binding capacity (WB), while positively correlating with dry gluten (DG) and protein content (PRO), whereas grain iron (GFE), zinc (GZN), and protein showed consistent positive interrelationships. Two superior accessions, PAUTS10 (WG 35.13%, DG 13.71%, PRO 16.42%, GZN 50.89 ppm) and Sonamoti (WG 33.33%, DG 12.92%, PRO 16.27%, GZN 56.03 ppm), were identified, surpassing the best check variety HD3226 for quality and nutritional parameters. Multi-locus genome-wide association studies identified 30 stable quantitative trait nucleotides across environments, with candidate gene analysis revealing genes involved in transcription regulation, biosynthetic processes, metal ion homeostasis, and transport. CONCLUSIONS: Triticum sphaerococcum demonstrated superior grain quality and micronutrient potential compared with modern wheat, highlighting its value as a genetic resource for biofortification. The identification of elite accessions and stable quantitative trait nucleotides (QTNs) provides useful targets for breeding programs aimed at improving protein and micronutrient content. Integrating ancient germplasm with modern genomic tools can accelerate the development of nutritionally enhanced wheat varieties. © 2026 Society of Chemical Industry.

Triticum↗

Complete nucleotide sequences of the domestic cat (Felis catus) mitochondrial genome and a transposed mtDNA tandem repeat (Numt) in the nuclear genome.

The complete 17,009-bp mitochondrial genome of the domestic cat, Felis catus, has been sequenced and conforms largely to the typical organization of previously characterized mammalian mtDNAs. Codon usage and base composition also followed canonical vertebrate patterns, except for an unusual ATC (non-AUG) codon initiating the NADH dehydrogenase subunit 2 (ND2) gene. Two distinct repetitive motifs at opposite ends of the control region contribute to the relatively large size (1559 bp) of this carnivore mtDNA. Alignment of the feline mtDNA genome to a homologous 7946-bp nuclear mtDNA tandem repeat DNA sequence in the cat, Numt, indicates simple repeat motifs associated with insertion/deletion mutations. Overall DNA sequence divergence between Numt and cytoplasmic mtDNA sequence was only 5.1%. Substitutions predominate at the third codon position of homologous feline protein genes. Phylogenetic analysis of mitochondrial gene sequences confirms the recent transfer of the cytoplasmic mtDNA sequences to the domestic cat nucleus and recapitulates evolutionary relationships between mammal species.

Amino Acid Sequence↗

The detection of nucleotide sequences with strong similarity to hormone responsive elements in the genome of eubacteria and archaebacteria and their possible relation to similar sequences present in the mitochondrial genome.

To account for the presence of nucleotide sequences in mitochondria with similarity to the Hormone Response Elements (HREs) of the nuclear genomes of man, rat and mouse, the genomes of several procaryotes have been screened for the presence of the sequences AGAACA NNN TGTTCT and GGTACA NNN TGTTCT, which represent perfect palindromic and consensus class I HREs, respectively, and for the sequence AGGTCA NNN TGACCT, which represents class II HRE. In many of the examined procaryotes, eubacteria and archaebacteria, almost perfect palindromic class I HREs and perfect or almost perfect class II half palindromic HREs have been detected in various genes, some of which encode proteins involved in energy metabolism, in replication and in transcription control. These findings support the hypothesis that the similar sequences found in mitochondria, potentially involved in hormonal regulation of respiratory enzyme biosynthesis, were introduced into eucaryotic cell by the procaryotic endosymbionts.

Animals↗

Analysis of six DNA components of the faba bean necrotic yellows virus genome and their structural affinity to related plant virus genomes.

Faba bean necrotic yellows virus (FBNYV) has a multicomponent circular ssDNA genome. In addition to a previously described genome component (C1) coding for a replicase-associated protein (Rep), five further components (C2 to C6) have now been identified. Each of the six components is about 1 kb in size, contains one major open reading frame (ORF) in the virion sense with a TATA box and polyadenylation signal, and has a noncoding region containing a highly conserved sequence possibly forming a stem-loop structure. Similar to C1, C2 encodes another putative Rep of 33.1 kDa, which is closely related to the Rep of banana bunchy top virus (BBTV). Based on bacterial expression and immunoblot analysis, the ORF of C5 encodes the capsid protein (CP) with a deduced molecular mass of 19 kDa. The FBNYV CP shares the highest amino acid (aa) identity (56.2%) with that of subterranean clover stunt virus (SCSV). The ORF of C4 potentially codes for a hydrophobic protein which appears to be structurally and functionally similar to the BBTV-C4 and SCSV-C1 proteins. No protein sequence similarities were found in databases for the C3 and C6 ORFs of FBNYV. FBNYV is clearly distinct from any known virus but is taxonomically related to BBTV and SCSV.

Amino Acid Sequence↗

Phylogenetic relationship of the complete Rauscher murine leukemia virus genome with other murine leukemia virus genomes.

We report the complete nucleotide sequence of the genome of Rauscher murine leukemia virus (R-MuLV), the replication-competent helper virus present in the Rauscher virus complex, and its phylogenetic relationship with other murine leukemia virus genomes. An overall sequence identity of 97.6% was found between R-MuLV and the Friend helper virus (F-MuLV), and the two viruses were closely related on the phylogenetic trees constructed from either gag, pol, or env sequences. Moloney murine leukemia virus (Mo-MuLV) was the next closest relative to R-MuLV and F-MuLV on all trees, followed by Akv and radiation leukemia virus (RadLV). The most distantly related helper virus was Hortulanus murine leukemia virus (Ho-MuLV). Interestingly, Cas-Br-E branched with Mo-MuLV on the gag and pol trees, whereas on the env tree, it revealed the highest degree of relatedness to Ho-MuLV, possibly due to an ancient recombination with an Ho-MuLV ancestor. In summary, a phylogenetic analysis involving various MuLVs has been performed, in which the postulated close relationship between R-MuLV and F-MuLV has been confirmed, consistent with the pathobiology of the two viruses.

Algorithms↗

The genome nucleotide sequence of a contemporary wild strain of measles virus and its comparison with the classical Edmonston strain genome.

The only complete genome nucleotide sequences of measles virus (MeV) reported to date have been for the Edmonston (Ed) strain and derivatives, which were isolated decades ago, passaged extensively under laboratory conditions, and appeared to be nonpathogenic. Partial sequencing of many other strains has identified >/=15 genotypes. Most recent isolates, including those typically pathogenic, belong to genotypes distinct from the Edmonston type. Therefore, the sequence of Ed and related strains may not be representative of those of pathological measles circulating at that or any time in human populations. Taking into account these issues as well as the fact that so many studies have been based upon Ed-related strains, we have sequenced the entire genome of a recently isolated pathogenic strain, 9301B. Between this recent isolate and the classical Ed strain, there were 465 nucleotide differences (2.93%) and 114 amino acid differences (2.19%). Computation of nonsynonymous and synonymous substitutions in open reading frames as well as direct comparisons of noncoding regions of each gene and extracistronic regulatory regions clearly revealed the regions where changes have been permissible and nonpermissible. Notably, considerable nonsynonymous substitutions appeared to be permissible for the P frame to maintain a high degree of sequence conservation for the overlapping C frame. However, the cause and the effect were largely unclear for any substitution, indicating that there is a considerable gap between the two strains that cannot be filled. The sequence reported here would be useful as a reference of contemporary wild-type MeV.

3' Untranslated Regions↗

Characteristics of nucleotide substitution in the hepatitis C virus genome: constraints on sequence change in coding regions at both ends of the genome.

Comparison of complete genome sequences for different variants of hepatitis C virus (HCV) reveals several different constraints on sequence change. Synonymous changes are suppressed in coding regions at both 5' and 3' ends of the genome. No evidence was found for the existence of alternative reading frames or for a lower mutation frequency in these regions. Instead, suppression may be due to constraints imposed by RNA secondary structures identified within the core and NS5b genes. Nonsynonymous substitutions are less frequent than synonymous ones except in the hypervariable region of E2 and, to a lesser extent, in E1, NS2, and NS5b. Transitions are more frequent than transversions, particularly at the third position of codons where the bias is 16:1. In addition, nucleotide substitutions may not occur symmetrically since there is a bias toward G or C at the third position of codons, while T left and right arrow C transitions were twice as frequent as A left and right arrow G transitions. These different biases do not affect the phylogenetic analysis of HCV variants but need to be taken into account in interpreting sequence change in longitudinal studies.

Base Sequence↗

The complete nucleotide sequence and multipartite organization of the tobacco mitochondrial genome: comparative analysis of mitochondrial genomes in higher plants.

Tobacco is a valuable model system for investigating the origin of mitochondrial DNA (mtDNA) in amphidiploid plants and studying the genetic interaction between mitochondria and chloroplasts in the various functions of the plant cell. As a first step, we have determined the complete mtDNA sequence of Nicotiana tabacum. The mtDNA of N. tabacum can be assumed to be a master circle (MC) of 430,597 bp. Sequence comparison of a large number of clones revealed that there are four classes of boundaries derived from homologous recombination, which leads to a multipartite organization with two MCs and six subgenomic circles. The mtDNA of N. tabacum contains 36 protein-coding genes, three ribosomal RNA genes and 21 tRNA genes. Among the first class, we identified the genes rps1 and psirps14, which had previously been thought to be absent in tobacco mtDNA on the basis of Southern analysis. Tobacco mtDNA was compared with those of Arabidopsis thaliana, Beta vulgaris, Oryza sativa and Brassica napus. Since repeated sequences show no homology to each other among the five angiosperms, it can be supposed that these were independently acquired by each species during the evolution of angiosperms. The gene order and the sequences of intergenic spacers in mtDNA also differ widely among the five angiosperms, indicating multiple reorganizations of genome structure during the evolution of higher plants. Among the conserved genes, the same potential conserved nonanucleotide-motif-type promoter could only be postulated for rrn18-rrn5 in four of the dicotyledonous plants, suggesting that a coding sequence does not necessarily move with the promoter upon reorganization of the mitochondrial genome.

Base Sequence↗

Seventh international meeting on single nucleotide polymorphism and complex genome analysis: 'ever bigger scans and an increasingly variable genome'.

In September 2005, the seventh international meeting on single nucleotide polymorphism (SNP) and complex genome analysis was held in Hinckley, near Leicester, UK and the meeting was organised by Anthony Brookes, Stephen Chanock, Ivo Gut, Alec Jeffreys and Pui-Yan Kwok. Similar to prior meetings, the 3-day meeting focused on new trends and methods in the analysis of SNPs and complex human disease. A substantial portion of the meeting was devoted to preliminary analyses of data emerging from the International HapMap Consortium and addressed key issues in patterns of recombination, linkage disequilibrium and population genetics. Of great interest were the sessions that addressed SNP analysis in other species and the emerging field of copy number variation. Overall, there have been a number of recent advances in genomics that promise to accelerate the pace of dissecting the genetic basis of many complex diseases in humans-and perhaps in other species.

Evolution, Molecular↗

Completion of the full-length genome sequence of Menangle virus: characterisation of the polymerase gene and genomic 5' trailer region.

Menangle virus (MenV), isolated in 1997 from stillborn piglets during an outbreak of reproductive disease at a large commercial piggery, is the only new paramyxovirus to be identified in Australia since Hendra virus in 1994. Following partial characterisation of the MenV genome, we previously showed that MenV is a novel member of the genus Rubulavirus. Here we report the characterisation of the large (L) polymerase gene and the adjacent 5' trailer region of MenV, which completes the full-length genome sequence of this novel paramyxovirus (15,516 nucleotides), and thereby confirm its taxonomic position within the family Paramyxoviridae.

Amino Acid Sequence↗

Complete nucleotide sequence and genome organization of sweet potato feathery mottle virus (S strain) genomic RNA: the large coding region of the P1 gene.

The complete nucleotide sequence of a sweet potato feathery mottle virus severe strain (SPFMV-S) genomic RNA was determined from overlapping cDNA clones and by directly sequencing viral RNA. The viral RNA genome is 10,820 nucleotides long, excluding the poly(A) tail and contains one open reading frame (ORF) starting at nucleotide 118 and ending at 10,599, potentially encoding a polyprotein of 3,493 amino acids (Mr 393,800). The ORF was followed by a 3' untranslated region of 221 nucleotides. The deduced polyprotein includes P1 (74K), HC-Pro (52K), P3 (46K), 6K1, CI (72K), 6K2, NIa-VPg (22K), NIa-Pro (28K), NIb (60K) and coat (35K) proteins, after an analysis of protein cleavage sites analogous to other potyvirus polyproteins. The polyprotein had a high level of amino acid identity with those of other potyviruses, except in the regions of P1 and P3. The P1 of SPFMV-S RNA has 664 amino acid residues, and is the largest and least similar to those of other potyviruses. HC-Pro and CI show high identity with those of other potyviruses. P3 has relatively low identity, however, the length of P3 was within the range of variability among other potyviruses. The 6K1 protein between P3 and C1 is also highly similar to those of other potyviruses. This is the first report on the complete nucleotide sequence of the sweet potato-infecting virus.

Genome, Viral↗

Genome histories clarify evolution of the expansin superfamily: new insights from the poplar genome and pine ESTs.

Expansins comprise a superfamily of plant cell wall-loosening proteins that has been divided into four distinct families, EXPA, EXPB, EXLA and EXLB. In a recent analysis of Arabidopsis thaliana and Oryza sativa expansins, we proposed a further subdivision of the families into 17 clades, representing independent lineages in the last common ancestor of monocots and eudicots. This division was based on both traditional sequence-based phylogenetic trees and on position-based trees, in which genomic locations and dated segmental duplications were used to reconstruct gene phylogeny. In this article we review recent work concerning the patterns of expansin evolution in angiosperms and include additional insights gained from the genome of a second eudicot species, Populus trichocarpa, which includes at least 36 expansin genes. All of the previously proposed monocot-eudicot orthologous groups, but no additional ones, are represented in this species. The results also confirm that all of these clades are truly independent lineages. Furthermore, we have used position-based phylogeny to clarify the history of clades EXPA-II and EXPA-IV. Most of the growth of the expansin superfamily in the poplar lineage is likely due to a recent polyploidy event. Finally, some monocot-eudicot clades are shown to have diverged before the separation of the angiosperm and gymnosperm lineages.

Arabidopsis↗

From array to array: confirmation of genomic gains and losses discovered by array-based comparative genomic hybridization utilizing fluorescence in situ hybridization on tissue microarrays.

The combination of array-based comparative genomic hybridization (CGH) with fluorescence in situ hybridization utilizing custom-designed bacterial artificial chromosome (BAC) probes applied to tissue microarrays represents a powerful compendium of techniques-greatly enhancing the throughput of genomic analysis and subsequent target validation. Such approach can be automated at various levels and allows managing large volume of targets and samples in a few experiments. As such, this approach facilitates discovery, validation and implementation of findings in the process of identification of new diagnostic, prognostic and potentially therapeutic molecular markers.

DNA↗

Identification of four genomic loci highly related to casein-kinase-2-alpha cDNA and characterization of a casein kinase-2-alpha pseudogene within the mouse genome.

Using the coding region of the human CK-2 alpha cDNA as a probe for screening a genomic mouse library, positive clones representing four different genomic loci were isolated. Partial DNA sequences of these loci encompassing the first 120 nucleotides of the putative coding region are reported. One positive clone was further analyzed by sequencing a 3.1 kb XbaI fragment. This clone displays the characteristics of a pseudogene, i.e. lack of introns and several nucleotide insertions and deletions. In its 3' region it contains a 91 bp large CT-rich stretch which consists of (CCTT) and (CT) repeats; in the 5' region three (CCCCCT) repeats.

Animals↗

Cereal genome evolution: pastoral pursuits with 'Lego' genomes.

The rapid progress in comparative analysis of cereal genomes reveals that they are composed of similar genomic building blocks. It seems that by simply rearranging these blocks and amplifying some of the repetitive sequences contained within them, it is possible to reconstitute the 56 different chromosomes found in wheat, rice, maize, sorghum, millet and sugarcane. Comparison of the orders of blocks in these reconstituted chromosomes reveals that the cleavage of a single chromosome formed from the blocks could give rise to all the combinations found in the chromosomes of the above species. A framework is now in place for collating all the information which has been generated from studying the individual cereals.

Biological Evolution↗