PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Huntington disease: a case study describing the complexities and nuances of predictive testing of monozygotic twins.

When a candidate for predictive testing for the Huntington disease gene is a monozygotic twin, confidentiality of the co-twin's diagnosis and autonomy of participation are among the critical genetic counseling issues. Predictive testing can proceed when twins voluntarily and simultaneously request counseling and evaluation in an HD testing program. This case describes a young man referred for predictive testing to an HD testing site on the East Coast of the United States. Family history revealed a twin brother of unknown zygosity who resided on the West Coast of the United States. The genetic counselors on opposite coasts collaborated to provide genetic counseling and evaluation for voluntary, informed predictive testing of the twins, protecting their rights while observing national protocol guidelines

Adult↗

SWORDS: a statistical tool for analysing large DNA sequences.

In this article, we present some simple yet effective statistical techniques for analysing and comparing large DNA sequences. These techniques are based on frequency distributions of DNA words in a large sequence, and have been packaged into a software called SWORDS. Using sequences available in public domain databases housed in the Internet, we demonstrate how SWORDS can be conveniently used by molecular biologists and geneticists to unmask biologically important features hidden in large sequences and assess their statistical significance.

Animals↗

Ab initio gene identification: prokaryote genome annotation with GeneScan and GLIMMER.

We compare the annotation of three complete genomes using the ab initio methods of gene identification GeneScan and GLIMMER. The annotation given in GenBank, the standard against which these are compared, has been made using GeneMark. We find a number of novel genes which are predicted by both methods used here, as well as a number of genes that are predicted by GeneMark, but are not identified by either of the nonconsensus methods that we have used. The three organisms studied here are all prokaryotic species with fairly compact genomes. The Fourier measure forms the basis for an efficient non-consensus method for gene prediction, and the algorithm GeneScan exploits this measure. We have bench-marked this program as well as GLIMMER using 3 complete prokaryotic genomes. An effort has also been made to study the limitations of these techniques for complete genome analysis. GeneScan and GLIMMER are of comparable accuracy insofar as gene-identification is concerned, with sensitivities and specificities typically greater than 0.9. The number of false predictions (both positive and negative) is higher for GeneScan as compared to GLIMMER, but in a significant number of cases, similar results are provided by the two techniques. This suggests that there could be some as-yet unidentified additional genes in these three genomes, and also that some of the putative identifications made hitherto might require re-evaluation. All these cases are discussed in detail.

Algorithms↗

From gene function to improved health: genome research in the United Kingdom.

The United Kingdom has a prestigious track record in genetic and genomic research, with the structure of DNA, and delivery of a third of the human genome being significant landmarks. UK genomic research benefits from a variety of funding sources in both the public and private sector, and efforts to co-ordinate research strategies. Here we describe how this investment has impacted on the UK research capacity and highlight examples of progress in translating genetic information into improved healthcare.

Cardiovascular Diseases↗

A genome-wide scan suggests a locus on chromosome 1q21-q23 contributes to normal variation in plasma cholesterol concentration.

To identify genes that influence plasma cholesterol, triglyceride, and high-density and low-density lipoproteins concentrations we conducted a genome-wide scan using 354 polymorphic markers spaced at 10-cM intervals in 75 obese but otherwise normal human families. The results of the genome scan using sibling pair analysis of quantitative phenotypes suggested that 1q21-q23 contains a locus that influences plasma cholesterol concentration. Chromosome 12 gave evidence of linkage to plasma triglyceride concentration (D12SPAH) and chromosomes 3, 6, 7, 10, 11, 17, and 20 yielded additional evidence of linkage for lipid phenotypes at lower levels of statistical significance. Allele sharing for markers near prominent candidate genes was either very weakly related or unrelated to sibling similarity for lipid concentrations. Together these results suggest that genes with important roles in regulating normal cholesterol and triglyceride concentrations do not coincide with the location of previously known candidate genes.

Adult↗

Characterization and molecular genetic mapping of microsatellite loci in pepper.

Microsatellites or simple sequence repeats are highly variable DNA sequences that can be used as informative markers for the genetic analysis of plants and animals. For the development of microsatellite markers in Capsicum, microsatellites were isolated from two small-insert genomic libraries and the GenBank database. Using five types of oligonucleotides, (AT)(15), (GA)(15), (GT)(15), (ATT)(10) and (TTG)(10), as probes, positive clones were isolated from the genomic libraries, and sequenced. Out of 130 positive clones, 77 clones showed microsatellite motifs, out of which 40 reliable microsatellite markers were developed. (GA)(n) and (GT)(n) sequences were found to occur most frequently in the pepper genome, followed by (TTG)(n) and (AT)(n). Additional 36 microsatellite primers were also developed from GenBank and other published data. To measure the information content of these markers, the polymorphism information contents (PICs) were calculated. Capsicum microsatellite markers from the genomic libraries have shown a high level of PIC value, 0.76, twice the value for markers from GenBank data. Forty six microsatellite loci were placed on the SNU-RFLP linkage map, which had been derived from the interspecific cross between Capsicum annuum "TF68" and Capsicum chinense "Habanero". The current "SNU2" pepper map with 333 markers in 15 linkage groups contains 46 SSR and 287 RFLP markers covering 1,761.5 cM with an average distance of 5.3 cM between markers.

Capsicum↗

Coffee and tomato share common gene repertoires as revealed by deep sequencing of seed and cherry transcripts.

An EST database has been generated for coffee based on sequences from approximately 47,000 cDNA clones derived from five different stages/tissues, with a special focus on developing seeds. When computationally assembled, these sequences correspond to 13,175 unigenes, which were analyzed with respect to functional annotation, expression profile and evolution. Compared with Arabidopsis, the coffee unigenes encode a higher proportion of proteins related to protein modification/turnover and metabolism-an observation that may explain the high diversity of metabolites found in coffee and related species. Several gene families were found to be either expanded or unique to coffee when compared with Arabidopsis. A high proportion of these families encode proteins assigned to functions related to disease resistance. Such families may have expanded and evolved rapidly under the intense pathogen pressure experienced by a tropical, perennial species like coffee. Finally, the coffee gene repertoire was compared with that of Arabidopsis and Solanaceous species (e.g. tomato). Unlike Arabidopsis, tomato has a nearly perfect gene-for-gene match with coffee. These results are consistent with the facts that coffee and tomato have a similar genome size, chromosome karyotype (tomato, n=12; coffee n=11) and chromosome architecture. Moreover, both belong to the Asterid I clade of dicot plant families. Thus, the biology of coffee (family Rubiacaeae) and tomato (family Solanaceae) may be united into one common network of shared discoveries, resources and information.

Arabidopsis↗

An expression profile of human pancreatic islet mRNAs by Serial Analysis of Gene Expression (SAGE).

AIMS/HYPOTHESIS: The Human Genome Project seeks to identify all genes with the ultimate goal of evaluation of relative expression levels in physiology and in disease states. The purpose of the current study was the identification of the most abundant transcripts in human pancreatic islets and their relative expression levels using Serial Analysis of Gene Expression. METHODS: By cutting cDNAs into small uniform fragments (tags) and concatemerizing them into larger clones, the identity and relative abundance of genes can be estimated for a cDNA library. Approximately 49,000 SAGE tags were obtained from three human libraries: (i) ficoll gradient-purified islets (ii) islets further individually isolated by hand-picking, and (iii) pancreatic exocrine tissue. RESULTS: The relative abundance of each of the genes identified was approximated by the frequency of the tags. Gene ontology functions showed that all three libraries contained transcripts mostly encoding secreted factors. Comparison of the two islet libraries showed various degrees of contamination from the surrounding exocrine tissue (11 vs 25%). After removal of exocrine transcripts, the relative abundance of 2180 islet transcripts was determined. In addition to the most common genes (e.g. insulin, transthyretin, glucagon), a number of other abundant genes with ill-defined functions such as proSAAS or secretagogin, were also observed. CONCLUSION/INTERPRETATION: This information could serve as a resource for gene discovery, for comparison of transcript abundance between tissues, and for monitoring gene expression in the study of beta-cell dysfunction of diabetes. Since the chromosomal location of the identified genes is known, this SAGE expression data can be used in setting priorities for candidate genes that map to linkage peaks in families affected with diabetes.

Chromosomes, Human, Pair 1↗

Multiplex PCR design strategy used for the simultaneous amplification of 10 Y chromosome short tandem repeat (STR) loci.

The simultaneous amplification of multiple regions of a DNA template is routinely performed using the polymerase chain reaction (PCR) in a process termed multiplex PCR. A useful strategy involving the design, testing, and optimization of multiplex PCR primer mixtures will be presented. Other multiplex design protocols have focused on the testing and optimization of primers, or the use of chimeric primers. The design of primers, through the close examination of predicted DNA oligomer melting temperatures ( T(m)) and primer-dimer interactions, can reduce the amount of testing and optimization required to obtain a well-balanced set of amplicons. The testing and optimization of the multiplex PCR primer mixture constructed here revolves around varying the primer concentrations rather than testing multiple primer combinations. By solely adjusting primer concentrations, a well-balanced set of amplicons should result if the primers were designed properly. As a model system to illustrate this multiplex design protocol, a 10-loci multiplex (10plex) Y chromosome short tandem repeat (STR) assay is used.

Base Sequence↗

Mutation exposed: a neutral explanation for extreme base composition of an endosymbiont genome.

The influence of neutral mutation pressure versus selection on base composition evolution is a subject of considerable controversy. Yet the present study represents the first explicit population genetic analysis of this issue in prokaryotes, the group in which base composition variation is most dramatic. Here, we explore the impact of mutation and selection on the dynamics of synonymous changes in Buchnera aphidicola, the AT-rich bacterial endosymbiont of aphids. Specifically, we evaluated three forms of evidence. (i) We compared the frequencies of directional base changes (AT-->GC vs. GC-->AT) at synonymous sites within and between Buchnera species, to test for selective preference versus effective neutrality of these mutational categories. Reconstructed mutational changes across a robust intraspecific phylogeny showed a nearly 1:1 AT-->GC:GC-->AT ratio. Likewise, stationarity of base composition among Buchnera species indicated equal rates of AT-->GC and GC-->AT substitutions. The similarity of these patterns within and between species supported the neutral model. (ii) We observed an equivalence of relative per-site AT mutation rate and current AT content at synonymous sites, indicating that base composition is at mutational equilibrium. (iii) We demonstrated statistically greater equality in the frequency of mutational categories in Buchnera than in parallel mammalian studies that documented selection on synonymous sites. Our results indicate that effectively neutral mutational pressure, rather than selection, represents the major force driving base composition evolution in Buchnera. Thus they further corroborate recent evidence for the critical role of reduced N(e) in the molecular evolution of bacterial endosymbionts.

Animals↗

Bilaterian phylogeny based on analyses of a region of the sodium-potassium ATPase beta-subunit gene.

Molecular investigations of deep-level relationships within and among the animal phyla have been hampered by a lack of slowly evolving genes that are amenable to study by molecular systematists. To provide new data for use in deep-level metazoan phylogenetic studies, primers were developed to amplify a 1.3-kb region of the alpha subunit of the nuclear-encoded sodium-potassium ATPase gene from 31 bilaterians representing several phyla. Maximum parsimony, maximum likelihood, and Bayesian analyses of these sequences (combined with ATPase sequences for 23 taxa downloaded from GenBank) yield congruent trees that corroborate recent findings based on analyses of other data sets (e.g., the 18S ribosomal RNA gene). The ATPase-based trees support monophyly for several clades (including Lophotrochozoa, a form of Ecdysozoa, Vertebrata, Mollusca, Bivalvia, Gastropoda, Arachnida, Hexapoda, Coleoptera, and Diptera) but do not support monophyly for Deuterostomia, Arthropoda, or Nemertea. Parametric bootstrapping tests reject monophyly for Arthropoda and Nemertea but are unable to reject deuterostome monophyly. Overall, the sodium-potassium ATPase alpha-subunit gene appears to be useful for deep-level studies of metazoan phylogeny.

Animals↗

Detecting Traces of Prehistoric Human Migrations by Geographic Synthetic Maps of Polyomavirus JC.

The polyomavirus JC (JCV) is a double-stranded DNA virus that is ubiquitous in human populations and is excreted in urine by a large percentage of individuals (20-70%). The strong genetic stability, combined with a mechanism of transmission mainly within the family, makes JCV a good marker of human migrations. In this study, the coevolution of JCV with its human host is investigated by using over a thousand nucleotide sequences deposited in the EMBL database; they correspond to the IG region, which is the genomic region with the highest rate of variation. The pattern of genetic diversity in JCV is evaluated by the principal coordinates analysis and the construction of synthetic maps. The first principal coordinate supports the existence of two distinct virus lineages, both arising from the ancestral African type. The first synthetic map suggests a two-migration model of the human dispersal out of Africa, thus implying a more complex picture than that known from human genes. The second principal coordinate points out the distinctiveness of strains coming from Asian/Amerind populations. The picture yielded by the second synthetic map appears to be more consistent with that known from human genes. In fact, it provides evidence of a deep split of the Asian lineage of JCV into two main branches: one diffusing in Japan and Americas, the other in Southeast Asia. The view that JCV, with its peculiar feature of a dual early emergence from Africa, can provide new information about the evolutionary history of our ancestors is discussed.

Databases, Nucleic Acid↗

Independent origins of subgroup Bl + B2 and subgroup B3 metallo-beta-lactamases.

The metallo-beta-lactamases constitute Class B in the Ambler classification of beta-lactamases and are divided into three subclasses: Bl, B2, and B3. Bayesian phylogenies of the Subclass B1 + B2 and Subclass B3 metallo-beta-lactamases and their homologs show that the beta-lactam-hydrolyzing function evolved independently within each group. In Subclass B1+B2 that function evolved about 1 billion years ago, and in Subclass B3 it evolved before the divergence of the Gram-positive and Gram-negative eubacteria, about 2 billion years ago. These results lend additional support to the proposal that the metallo-beta-lactamases should be divided into two distinct classes.

Archaea↗

Why are young and old repetitive elements distributed differently in the human genome?

Alu elements are not distributed homogeneously throughout the human genome: old elements are preferentially found in the GC-rich parts of the genome, while young Alus are more often found in the GC-poor parts of the genome. The process giving rise to this differential distribution remains poorly understood. Here we investigate whether this pattern could be due to a preferential degradation of Alu elements integrated in GC-poor regions by small indel mutations. We aligned 5.1 Mb of human and chimpanzee sequences and examined whether the rate of insertion and deletion inside Alu elements differed according to the base composition surrounding them. We found that Alu elements are not preferentially degraded in GC-poor regions by indel events. We also looked at whether very young L1 elements show the same change in distribution compared to older ones. This analysis indicated that L1 elements also show a shift in their distribution, although we could not assess it as precisely as for Alu elements. We propose that the differential distribution of Alu elements is likely to be due to a change in their pattern of insertion or their probability of fixation through evolutionary time.

Alu Elements↗

Can codon usage bias explain intron phase distributions and exon symmetry?

More introns exist between codons (phase 0) than between the first and the second bases (phase 1) or between the second and the third base (phase 2) within the codon. Many explanations have been suggested for this excess of phase 0. It has, for example, been argued to reflect an ancient utility for introns in separating exons that code for separate protein modules. There may, however, be a simple, alternative explanation. Introns typically require, for correct splicing, particular nucleotides immediately 5' in exons (typically a G) and immediately 3' in the following exon (also often a G). Introns therefore tend to be found between particular nucleotide pairs (e.g., G|G pairs) in the coding sequence. If, owing to bias in usage of different codons, these pairs are especially common at phase 0, then intron phase biases may have a trivial explanation. Here we take codon usage frequencies for a variety of eukaryotes and use these to generate random sequences. We then ask about the phase of putative intron insertion sites. Importantly, in all simulated data sets intron phase distribution is biased in favor of phase 0. In many cases the bias is of the magnitude observed in real data and can be attributed to codon usage bias. It is also known that exons may carry either the same phase (symmetric) or different phases (asymmetric) at the opposite ends. We simulated a distribution of different types of exons using frequencies of introns observed in real genes assuming random combination of intron phases at the opposite sides of exons. Surprisingly the simulated pattern was quite similar to that observed. In the simulants we typically observe a prevalence of symmetric exons carrying phase 0 at both ends, which is common for eukaryotic genes. However, at least in some species, the extent of the bias in favor of symmetric (0,0) exons is not as great in simulants as in real genes. These results emphasize the need to construct a biologically relevant null model of successful intron insertion.

Animals↗

Large subunit mitochondrial rRNA secondary structures and site-specific rate variation in two lizard lineages.

A phylogenetic-comparative approach was used to assess and refine existing secondary structure models for a frequently studied region of the mitochondrial encoded large subunit (16S) rRNA in two large lizard lineages within the Scincomorpha, namely the Scincidae and the Lacertidae. Potential pairings and mutual information were analyzed to identify site interactions present within each lineage and provide consensus secondary structures. Many of the interactions proposed by previous models were supported, but several refinements were possible. The consensus structures allowed a detailed analysis of rRNA sequence evolution. Phylogenetic trees were inferred from Bayesian analyses of all sites, and the topologies used for maximum likelihood estimation of sequence evolution parameters. Assigning gamma-distributed relative rate categories to all interacting sites that were homologous between lineages revealed substantial differences between helices. In both lineages, sites within helix G2 were mostly conserved, while those within helix E18 evolved rapidly. Clear evidence of substantial site-specific rate variation (covarion-like evolution) was also detected, although this was not strongly associated with specific helices. This study, in conjunction with comparable findings on different, higher-level taxa, supports the ubiquitous nature of site-specific rate variation in this gene and justifies the incorporation of covarion models in phylogenetic inference.

Animals↗