PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “noncoding genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Evidence for turnover of functional noncoding DNA in mammalian genome evolution.

The vast majority of the mammalian genome does not code for proteins, and a fundamental question in genomics is: What proportion of the noncoding mammalian genome is functional? Most attempts to address this issue use sequence comparisons between highly diverged mammals such as human and mouse to identify conservation due to negative selection. But such comparisons will underestimate the true proportion of functional noncoding DNA if there is turnover, if patterns of negative selection change over time. Here we test whether the inferred level of negative selection differs between different pairwise species comparisons. Using a multiple alignment of more than a megabase of contiguous sequence from eight mammalian species, we find a strong negative relationship between inferred levels of negative selection and pairwise divergence using 21 pairwise comparisons. This result suggests that there is a high rate of turnover of functional noncoding elements in the mammalian genome, so measures of functional constraint based on human-mouse comparisons may seriously underestimate the true value.

Animals↗

Genetic organization and diversity of the 3' noncoding region of the hepatitis C virus genome.

The 3' noncoding region (3' NCR) of the hepatitis C virus (HCV) genome contained in viral particles was analyzed by an RNA linker ligation followed by reverse transcription-polymerase chain reaction. Sequence analysis of the amplified fragment from four strains, including different genotypes 1b, 2b, 3a, and 3b indicated that the 3' NCR is composed of between 200 and 235 nts. The sequence of the 3' NCR consists of a type-specific region (immediately following the termination codon), a poly(U) stretch, a C(U)n-repeat, and highly conserved region termed the core element. The poly(U) stretch and C(U)n-repeat regions varied in length and in sequence among different genotypes. Core elements having putative secondary structure consisted of 98 or 100 nts and were highly conserved in all genotypes. Most of the nt changes found in different genotypes did not affect the secondary structure of the core elements, suggesting that this region may play an important role in replication, stabilization of the HCV RNA, and/or packaging of the genome. Most of the HCV-1b strains carried two U residues at the 3' end of the core element, while the minor HCV-1b strains had no U residues, demonstrating that there are two variants in type 1b strains. Amplification of the core element using linker-primed cDNA was comparable with that using the 3' proximal core element-primed cDNA, indicating that the 3' end of HCV genome was terminated by an OH group.

Base Sequence↗

Selective constraint on noncoding regions of hominid genomes.

An important challenge for human evolutionary biology is to understand the genetic basis of human-chimpanzee differences. One influential idea holds that such differences depend, to a large extent, on adaptive changes in gene expression. An important step in assessing this hypothesis involves gaining a better understanding of selective constraint on noncoding regions of hominid genomes. In noncoding sequence, functional elements are frequently small and can be separated by large nonfunctional regions. For this reason, constraint in hominid genomes is likely to be patchy. Here we use conservation in more distantly related mammals and amniotes as a way of identifying small sequence windows that are likely to be functional. We find that putatively functional noncoding elements defined in this manner are subject to significant selective constraint in hominids.

Animals↗

The influence of specific neighboring bases on substitution bias in noncoding regions of the plant chloroplast genome.

Substitutions occurring in noncoding sequences of the plant chloroplast genome violate the independence of sites that is assumed by substitution models in molecular evolution. The probability that a substitution at a site is a transversion, as opposed to a transition, increases significantly with increasing A + T content of the two adjacent nucleotides. In the present study, this dependency of substitutions on local context is examined further in a number of noncoding regions from the chloroplast genome of members of the grass family (Poaceae). Two features were examined; the influence of specific neighboring bases, as opposed to the general A + T content, on transversion proportion and an influence on substitutions by nucleotides other than the two immediately adjacent to the site of substitution. In both cases, a significant effect was found. In the case of specific nucleotides, transversion proportion is significantly higher at sites with a pyrimidine immediately 5' on either strand. Substitutions at sites of the type YNR, where N is the site of substitution, have the highest rate of transversion. This specific effect is secondary to the A + T content effect such that, in terms of proportion of substitutions that are transversions, the nucleotides are ranked T > A > C > G as to their effect when they are immediately 5' to the site of substitution. In the case of nucleotides other than the immediate neighbors, a significant influence on substitution dynamics is observed in the case where the two neighboring bases are both A and/or T. Thus, substitutions are primarily, but not exclusively, influenced by the composition of the two nucleotides that are immediately adjacent. These results indicate that the pattern of molecular evolution of the plant chloroplast genome is extremely complex as a result of a variety of inter-site dependencies.

Chloroplasts↗

Meiotic recombination, noncoding DNA and genomic organization in Caenorhabditis elegans.

The genetic map of each Caenorhabditis elegans chromosome has a central gene cluster (less pronounced on the X chromosome) that contains most of the mutationally defined genes. Many linkage group termini also have clusters, though involving fewer loci. We examine the factors shaping the genetic map by analyzing the rate of recombination and gene density across the genome using the positions of cloned genes and random cDNA clones from the physical map. Each chromosome has a central gene-dense region (more diffuse on the X) with discrete boundaries, flanked by gene-poor regions. Only autosomes have reduced rates of recombination in these gene-dense regions. Cluster boundaries appear discrete also by recombination rate, and the boundaries defined by recombination rate and gene density mostly, but not always, coincide. Terminal clusters have greater gene densities than the adjoining arm but similar recombination rates. Thus, unlike in other species, most exchange in C. elegans occurs in gene-poor regions. The recombination rate across each cluster is constant and similar; and cluster size and gene number per chromosome are independent of the physical size of chromosomes. We propose a model of how this genome organization arose.

Animals↗

Neutral substitutions occur at a faster rate in exons than in noncoding DNA in primate genomes.

Point mutation rates in exons (synonymous sites) and noncoding (introns and intergenic) regions are generally assumed to be the same. However, comparative sequence analyses of synonymous substitutions in exons (81 genes) and that of long intergenic fragments (141.3 kbp) of human and chimpanzee genomes reveal a 30%-60% higher mutation rate in exons than in noncoding DNA. We propose a differential CpG content hypothesis to explain this fundamental, and seemingly unintuitive, pattern. We find that the increased exonic rate is the result of the relative overabundance of synonymous sites involved in CpG dinucleotides, as the evolutionary divergence in non-CpG sites is similar in noncoding DNA and synonymous sites of exons. Expectations and predictions of our hypothesis are confirmed in comparisons involving more distantly related species, including human-orangutan, human-baboon, and human-macaque. Our results suggest an underlying mechanism for higher mutation rate in GC-rich genomic regions, predict nonlinear accumulation of mutations in pseudogenes over time, and provide a possible explanation for the observed higher diversity of single nucleotide polymorphisms (SNPs) in the synonymous sites of exons compared to the noncoding regions.

Animals↗

Mapping of conserved RNA secondary structures predicts thousands of functional noncoding RNAs in the human genome.

In contrast to the fairly reliable and complete annotation of the protein coding genes in the human genome, comparable information is lacking for noncoding RNAs (ncRNAs). We present a comparative screen of vertebrate genomes for structural noncoding RNAs, which evaluates conserved genomic DNA sequences for signatures of structural conservation of base-pairing patterns and exceptional thermodynamic stability. We predict more than 30,000 structured RNA elements in the human genome, almost 1,000 of which are conserved across all vertebrates. Roughly a third are found in introns of known genes, a sixth are potential regulatory elements in untranslated regions of protein-coding mRNAs and about half are located far away from any known gene. Only a small fraction of these sequences has been described previously. A comparison with recent tiling array data shows that more than 40% of the predicted structured RNAs overlap with experimentally detected sites of transcription. The widespread conservation of secondary structure points to a large number of functional ncRNAs and cis-acting mRNA structures in the human genome.

Animals↗

Comparative genomic analysis as a tool for biological discovery.

The recent completion of the human genome sequence has enabled the identification of a large fraction of our gene catalogue and their physical chromosomal position. However, current efforts lag at defining the cis-regulatory sequences that control the spatial and temporal patterns of each gene's expression. This task remains difficult due to our lack of knowledge of the vocabulary controlling gene regulation and the vast genomic search space, with greater than 95% of our genome being noncoding. Recent comparative genomic-based strategies are beginning to aid in the identification of functional sequences based on their high levels of evolutionary conservation. This has proven successful for comparisons between closely related species such as human-primate or human-mouse, but also holds true for distant evolutionary comparisons, such as human-fish or human-bird. In this review we provide support for the utility of cross-species sequence comparisons by illustrating several applications of this strategy, including the identification of new genes and functional non-coding sequences. We also discuss emerging concepts as this field matures, such as how to properly select which species for comparison, which may differ significantly between independent studies.

Animals↗

Computational identification of noncoding RNAs in E. coli by comparative genomics.

Some genes produce noncoding transcripts that function directly as structural, regulatory, or even catalytic RNAs [1, 2]. Unlike protein-coding genes, which can be detected as open reading frames with distinctive statistical biases, noncoding RNA (ncRNA) gene sequences have no obvious inherent statistical biases [3]. Thus, genome sequence analyses reveal novel protein-coding genes, but any novel ncRNA genes remain invisible. Here, we describe a computational comparative genomic screen for ncRNA genes. The key idea is to distinguish conserved RNA secondary structures from a background of other conserved sequences using probabilistic models of expected mutational patterns in pairwise sequence alignments. We report the first whole-genome screen for ncRNA genes done with this method, in which we applied it to the "intergenic" spacers of Escherichia coli using comparative sequence data from four related bacteria. Starting from >23,000 conserved interspecies pairwise alignments, the screen predicted 275 candidate structural RNA loci. A sample of 49 candidate loci was assayed experimentally. At least 11 loci expressed small, apparently noncoding RNA transcripts of unknown function. Our computational approach may be used to discover structural ncRNA genes in any genome for which appropriate comparative genome sequence data are available.

Animals↗

Secondary structure of the 3'-noncoding region of flavivirus genomes: comparative analysis of base pairing probabilities.

The prediction of the complete matrix of base pairing probabilities was applied to the 3' noncoding region (NCR) of flavivirus genomes. This approach identifies not only well-defined secondary structure elements, but also regions of high structural flexibility. Flaviviruses, many of which are important human pathogens, have a common genomic organization, but exhibit a significant degree of RNA sequence diversity in the functionally important 3'-NCR. We demonstrate the presence of secondary structures shared by all flaviviruses, as well as structural features that are characteristic for groups of viruses within the genus reflecting the established classification scheme. The significance of most of the predicted structures is corroborated by compensatory mutations. The availability of infectious clones for several flaviviruses will allow the assessment of these structural elements in processes of the viral life cycle, such as replication and assembly.

Algorithms↗

A poliovirus temperature-sensitive RNA synthesis mutant located in a noncoding region of the genome.

We have constructed an 8-base-pair insertion mutation in the 3' noncoding region of an infectious poliovirus cDNA clone that gives rise to a temperature-sensitive RNA synthesis mutant upon transfection into mammalian cells. The mutated cDNA was used to establish a cell line that releases the mutant poliovirus in a temperature-dependent fashion, representing a unique persistent viral infection. A poliovirus mutant mapping in the noncapsid region of the viral genome can be complemented in this cell line, implying that the cell line expresses viral proteins at the nonpermissive temperature.

Base Sequence↗

Restricted variability of a 17 nucleotide stretch within the 5'-noncoding region of poliovirus genome.

The outbreak of poliomyelitis in Finland in 1984 was caused by a wild strain of poliovirus 3 with uncommon molecular and antigenic properties. We prepared a synthetic oligonucleotide probe complementary to nucleotides 494-510 in the 5'-noncoding part of the genome of a representative strain of the outbreak. This short nucleotide stretch was found to be relatively well conserved within the outbreak and uncommon among 82 independent poliovirus isolates. It may thus be a useful marker for screening isolates to identify those requiring more detailed genetic comparison. The sequences of the corresponding region of the genome are known for 32 separate poliovirus strains and 3 coxsackie B virus strains and show 6 fully conserved nucleotides that could assume a constant hairpin-loop position in a hypothetical secondary structure of the RNA. This could explain the persistence of a particular 17 nucleotide sequence for 40 years in nature in this highly variable region of the poliovirus genome.

Animals↗

Construction of less neurovirulent polioviruses by introducing deletions into the 5' noncoding sequence of the genome.

Viral attenuation may be due to lowered efficiency of certain steps essential for viral multiplication. For the construction of less neurovirulent strains of poliovirus in vitro, we introduced deletions into the 5' noncoding sequence (742 nucleotides long) of the genomes of the Mahoney and Sabin 1 strains of poliovirus type 1 by using infectious cDNA clones of the virus strains. Plaque sizes shown by deletion mutants were used as a marker for rate of viral proliferation. Deletion mutants of both the strains thus constructed lacked a genome region of nucleotide positions 564 to 726. The sizes of plaques displayed by these deletion mutants were smaller than those by the respective parental viruses, although a phenotype referring to reproductive capacity at different temperatures (rct) of viruses was not affected by introduction of the deletion. Monkey neurovirulence tests were performed on the deletion mutants. The results clearly indicated that the deletion mutants had much less neurovirulence than with the corresponding parent viruses. Production of infectious particles and virus-specific protein synthesis in cells infected with the deletion mutants started later than in those infected with the parental viruses. The rate at which cytopathic effect progressed was also slower in cells infected with the mutants. Phenotypic stability of the deletion mutant for small-plaque phenotype and temperature sensitivity was investigated after passaging the mutant at an elevated temperature of 37.5 degrees C. Our data strongly suggested that the less neurovirulent phenotype introduced by the deletion is very stable during passaging of the virus.

Animals↗

Genomic selective constraints in murid noncoding DNA.

Recent work has suggested that there are many more selectively constrained, functional noncoding than coding sites in mammalian genomes. However, little is known about how selective constraint varies amongst different classes of noncoding DNA. We estimated the magnitude of selective constraint on a large dataset of mouse-rat gene orthologs and their surrounding noncoding DNA. Our analysis indicates that there are more than three times as many selectively constrained, nonrepetitive sites within noncoding DNA as in coding DNA in murids. The majority of these constrained noncoding sites appear to be located within intergenic regions, at distances greater than 5 kilobases from known genes. Our study also shows that in murids, intron length and mean intronic selective constraint are negatively correlated with intron ordinal number. Our results therefore suggest that functional intronic sites tend to accumulate toward the 5' end of murid genes. Our analysis also reveals that mean number of selectively constrained noncoding sites varies substantially with the function of the adjacent gene. We find that, among others, developmental and neuronal genes are associated with the greatest numbers of putatively functional noncoding sites compared with genes involved in electron transport and a variety of metabolic processes. Combining our estimates of the total number of constrained coding and noncoding bases we calculate that over twice as many deleterious mutations have occurred in intergenic regions as in known genic sequence and that the total genomic deleterious point mutation rate is 0.91 per diploid genome, per generation. This estimated rate is over twice as large as a previous estimate in murids.

Animals↗

Pyrimidine-rich region mutations compensate for a stem-loop V lesion in the 5' noncoding region of poliovirus genomic RNA.

Five revertants of a linker-scanning mutation adjacent to the stem-loop V attenuation determinant (X472) in the 5' noncoding region of poliovirus RNA were independently isolated from neuroblastoma cells and contained RNAs with seven nucleotide changes in the pyrimidine-rich region. Generation of the identical rare second-site mutations suggests the existence of a replicase-dependent mutagenesis mechanism during poliovirus replication. Enzymatic structure probing of the mutated pyrimidine-rich domain identified secondary structure changes between stem-loops V and VI. A consensus secondary structure model is presented for wild-type stem-loops V and VI and the pyrimidine-rich region located in the 5' noncoding region of poliovirus RNA. A pyrimidine-rich region mutant (X472-R4N) produced large plaques in neuroblastoma cells and small plaques in HeLa cells, but the plaque size differences were not due to cell-type differences in viral translation or RNA replication. Release of X472-R4N from HeLa cells was 10-fold lower than release from neuroblastoma cells, which may explain the small plaque phenotype of X472-R4N in HeLa cells. Wild-type poliovirus was also released more efficiently from neuroblastoma cells (approximately 4-fold increase compared with release from HeLa cells), indicating that poliovirus neurotropism may be influenced by the cell-type efficiency of virus release. Thermal treatment increased the levels of infectious X472-R4N virions but not wild-type virus particles; thus RNA sequence and structural changes in the mutated 5' noncoding region of X472-R4N may have altered RNA-protein interactions necessary for virus infectivity.

5' Untranslated Regions↗

Sequence comparison and secondary structure analysis of the 3' noncoding region of flavivirus genomes reveals multiple pseudoknots.

Sequences of 191 flavivirus RNAs belonging to four sero-groups were used to predict the secondary structure of the 3' noncoding region (3' NCR) directly upstream of the conserved terminal hairpin. In mosquito-borne flavivirus RNAs (n = 164) a characteristic structure element was identified that includes a phylogenetically well-supported pseudoknot. This element is repeated in the dengue and Japanese encephalitis RNAs and centers around the conserved sequences CS2 and RCS2. In yellow fever virus RNAs that contain one CS2 motif, only one copy of this pseudoknotted structure was found. The conserved pseudoknotted element is absent from the 3' NCR of tick-borne virus RNAs, which altogether adopt a secondary structure that is very different from that of mosquito-borne virus RNAs. The strong conservation of the pseudoknot in mosquito-borne flavivirus RNAs implies a stronger relationship between these viruses than concluded from previous secondary structure analyses. The role of the (tandem) pseudoknots in flavivirus replication is discussed.

Base Sequence↗

Ancient noncoding elements conserved in the human genome.

Cartilaginous fishes represent the living group of jawed vertebrates that diverged from the common ancestor of human and teleost fish lineages about 530 million years ago. We generated approximately 1.4x genome sequence coverage for a cartilaginous fish, the elephant shark (Callorhinchus milii), and compared this genome with the human genome to identify conserved noncoding elements (CNEs). The elephant shark sequence revealed twice as many CNEs as were identified by whole-genome comparisons between teleost fishes and human. The ancient vertebrate-specific CNEs in the elephant shark and human genomes are likely to play key regulatory roles in vertebrate gene expression.

Animals↗

Intracellular modifications induced by poliovirus reduce the requirement for structural motifs in the 5' noncoding region of the genome involved in internal initiation of protein synthesis.

A series of genetic deletions based partly on two RNA secondary structure models (M. A. Skinner, V. R. Racaniello, G. Dunn, J. Cooper, P. D. Minor, and J. W. Almond, J. Mol. Biol. 207:379-392, 1989; E. V. Pilipenko, V. M. Blinov, L. I. Romanova, A. N. Sinyakov, S. V. Maslova, and V. I. Agol, Virology 168:201-209, 1989) was made in the cDNA encoding the 5' noncoding region (5' NCR) of the poliovirus genome in order to study the sequences that direct the internal entry of ribosomes. The modified cDNAs were placed between two open reading frames in a single transcriptional unit and used to transfect cells in culture. Internal entry of ribosomes was detected by measuring translation from the second open reading frame in the bicistronic mRNA. When assayed alone, a large proportion of the poliovirus 5' NCR superstructure including several well-defined stem-loops was required for ribosome entry and efficient translation. However, in cells cotransfected with a complete infectious poliovirus cDNA, the requirement for the stem-loops in this large superstructure was reduced. The results suggest that virus infection modifies the cellular translational machinery, so that shortened forms of the 5' NCR are sufficient for cap-independent translation, and that the internal entry of ribosomes occurs by two distinct modes during the virus replication cycle.

DNA Mutational Analysis↗