PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence biases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Phylogenetic relationships among phrynosomatid lizards as inferred from mitochondrial ribosomal DNA sequences: substitutional bias and information content of transitions relative to transversions.

The phylogenetic relationships among 40 species, representing all genera, within the North American lizard family Phrynosomatidae were inferred from mitochondrial ribosomal RNA gene sequences. Cladistic analysis of the DNA sequence data (779 bp; 162 informative characters) supported the monophyly of the sand lizards (Callisaurus, Cophosaurus, Holbrookia, and Uma), Petrosaurus, Phrynosoma, Urosaurus, and Uta. All the species of Sceloporus, except S. variabilis and S. chrysostictus, formed a clade. Except for a sand lizard + Phrynosoma clade, the intergeneric relationships inferred from the mtDNA were largely incongruent with recent cladistic analyses based on morphology. Sceloporus group monophyly was not supported, with Petrosaurus being a member of a clade containing Sator, Sceloporus, and Urosaurus, to the exclusion of Uta. The phylogenetic placement of Uta was ambiguous. The substitutional bias in the phrynosomatid mitochondrial rDNA sequences was examined, as well as the phylogenetic information content of transitions relative to transversions. There appeared to be a lower transition bias than observed in other vertebrate sequences, with some classes of transversions occurring as frequently as G <-> A transitions. Transitions were no less informative for phylogeny reconstruction than transversions. Therefore, transitions should not be down-weighted in phylogenetic analysis, as is often done.

Animals↗

Substrate selection rules for the hairpin ribozyme determined by in vitro selection, mutation, and analysis of mismatched substrates.

Substrate recognition by the hairpin ribozyme has been proposed to involve two short intermolecular helices, termed helix 1 and helix 2. We have used a combination of three methods (cleavage of mismatched substrates, in vitro selection, and site-specific mutational analysis) to systematically determine the substrate recognition rules for this RNA enzyme. Assays measuring substrate cleavage in trans under multiple turnover conditions were conducted using the wild-type ribozyme and substrates containing mismatches in all sites potentially recognized by the ribozyme. Molecules containing single- and multiple-base mismatches in helix 2 at sites distant from the cleavage site (g-4c, u-5a, g-4c: u-5a) were cleaved with reduced efficiency, whereas those with mismatches proximal to the cleavage site (c-2a, a-3c, c-2a: a-3c) were not cut. Analogous results were obtained for helix 1, where mismatches distal from the cleavage site (u+7a, u+8a, u+9a, u+7a: u+8a: u+9a) were used much more efficiently than those proximal to the cleavage site (c+4a, u-5a, g+6c, c+4a: u+5a: g+6c). In vitro selection experiments were carried out to identify active variants from populations of molecules in which either helix 1 or helix 2 was randomized. Results constitute an artificial phylogenetic data base that proves base-pairing of nucleotides at five positions within helix 1 and three positions within helix 2 and reveals a significant sequence bias at 3 bp (c+4.G6, c-2.G11, and a-3.U12). This sequence bias was confirmed at two sites by measuring relative cleavage rates of all 16 possible dinucleotide combinations at base pairs c+4.G6 and c-2.G11.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

Motif-biased protein sequence alignment.

A method was developed for pairwise protein sequence alignment to emulate the effect of structural knowledge or multiple sequences. Runs of matches of the preferred length were emphasized through the use of a product-bias allowing short motifs to influence the alignment to a degree that was a realistic reflection of their infrequency of occurrence. This gave motifs a locally high scoring match, making their alignment relatively less sensitive to the value of the gap penalty. This property should be a great advantage when a large number of sequence comparisons are made with a fixed set of parameter values, as typically occurs in the scan of a sequence databank with a probe or in the development of a multiple alignment.

Algorithms↗

Base-specific sequences that bias somatic hypermutation deduced by analysis of out-of-frame human IgVH genes.

Somatic hypermutation introduces mutations into IgV genes during affinity maturation of the B cell response. Mutations are introduced nonrandomly, and are generally targeted to the complementarity determining regions (CDRs). Subsequent selection against mutations that result in lower affinity or nonfunctional Ig increases the relative number of mutations in the CDRs. Investigation of somatic hypermutation is hampered by the effects of selection. We have avoided this by studying out-of-frame human IgVH4.21 and 251 genes, which, being unused alleles, are unselected. By comparison of the frequency of A, C, G, and T nucleotides at positions -3 to +3 around mutated or unmutated A, C, and G nucleotides, we have identified flanking sequences that most commonly surround mutated bases. Distinct trends in flanking sequences that were unique for each base were observed. Statistically significant trends that were common to both IgVH4.21 and 251 were used to deduce motifs that bias somatic hypermutation. The motifs deduced from this data, with targeted bases in regular type, are AANB, WDCH, and DGHD (where W = A/T, B = C/G/T, D = A/G/T, H = A/C/T, and N = any base). Mutations from C and G in two further groups of out-of-frame human IgVH genes, not used in the deduction of the motifs, occurred significantly within the motifs for C and G. The proposed target sequence for G is within the reverse complement of the target sequence for C, suggesting that the hypermutation mechanism may target only G or C. The mutation in the complementary base would appear on the other strand following replication.

Adenine↗

A screen for conserved sequences with biased base composition identifies noncoding RNAs in the A-T rich genome of Plasmodium falciparum.

Noncoding RNAs (ncRNAs) such as snRNAs, snoRNAs and microRNAs play important roles in transcription and translation control. These ncRNAs have yet to be discovered in the malarial parasite Plasmodium falciparum, an organism in which these basic biological processes are poorly understood. Inspired by a report by Klein et al., we initiated a bioinformatics screen to uncover several candidate ncRNAs from the parasite genome using two simple criteria: first, elevated GC content in the highly A-T rich intergenic regions of the P. falciparum genome and second, conservation of sequence homology between malaria parasite species. We show that all the annotated tRNAs can be successfully identified in our screen as well as several new candidates that show homology to snRNAs and snoRNAs, and ten candidate ncRNAs of unknown function. Three of the candidate snRNAs, a predicted selenocysteine tRNA and two candidates of unknown function are expressed in asexual stage parasites, further validating the screen. With these results, the biological processes underlying RNA-mediated regulation of transcription, translation and splicing can be studied in an important human pathogen.

Animals↗

Influence of intercodon and base frequencies on codon usage in filarial parasites.

Base frequency, codon usage, and intercodon identity were analyzed in five filarial parasite species representing five Onchocercidae genera. Wucheria bancrofti, Brugia malayi, Onchocerca volvulus, Acanthocheilonema viteae, and Dirofilaria immitis gene sequences were downloaded from NCBI, and analysis was performed using locally designed computer programs and other freely available applications. A clear sequence bias was observed among the nematode species examined. At the nucleotide level, AT basepairs were present in gene sequences at higher frequencies than GC. In addition, codons ending in A or T were used proportionately more than those with G or C in the third-codon position. In addition, the amino acids used most often corresponded to codons ending in AT basepairs. Intercodon base proportion was biased in that A was found most often at N4, second only to T in certain specific cases. Since all of these sequence biases were observed in a relatively consistent fashion among all of the organisms studied, we conclude that sequence bias is a genetic characteristic, which is associated with multiple filarial genera.

Animals↗

Successful recognition of protein folds using threading methods biased by sequence similarity and predicted secondary structure.

Analysis of our fold recognition results in the 3rd Critical Assessment in Structure Prediction (CASP3) experiment, using the programs THREADER 2 and GenTHREADER, shows an encouraging level of overall success. Of the 23 submitted predictions, 20 targets showed no clear sequence similarity to proteins of known 3D structure. These 20 targets can be divided into 22 domains, of which, 20 domains either entirely match a previously known fold, or partially match a substantial region of a known fold. Of these 20 domains, we correctly assigned the folds in 10 cases.

Algorithms↗

Signals determining translational start-site recognition in eukaryotes and their role in prediction of genetic reading frames.

A special methionyl-tRNA (RNAi) is universally required to initiate translation. The conversation of this reactant throughout evolution, as well as its unusual decoding properties, suggested an alternate mechanism for tRNA-mRNA interactions at initiation. We have reported that the sequence of bases neighboring the start codons of many eubacterial genes are complementary not only to the 16S rRNA 3' end and to the anticodon of tRNAi, but, also, have the potential to base-pair the D, T or extended anticodon loops of this tRNAi. The coding properties of tRNAi and mutations that affect translation suggest that these signals may function. This hypothesis explains the observation that unusual triplets can start prokaryotic and mitochondrial genes and predicts the occurrence of other reading frames. Furthermore, it suggests a unifying model of chain initiation based on RNA-RNA contacts and displacements. Here we examine the start domain of 290 eukaryotic genes for their ability to base-pair the tRNAi loops and the 18S rRNA. We observe that both methionine start, and methionine coding regions have the potential to pair with the 18S rRNA, but that the nucleotide distribution about start codons strongly favoured such pairings over that near internal AUGs. The 5' extended anticodon of tRNAi is methylated, and was not represented in the mRNA with high frequency. However, the tetramer AUGg did occur with high frequency in the start domain. A modification of the tRNAi T loop also decreases its base-pairing potential. Interestingly, complementarity to the T loop did not occur with high frequency in the start sites. The early coding region, 10 to 34 nucleotides 3' to the initiator AUG, is complementary to the tRNAi D loop in many cases, while no such affinity is found near internal AUGs. The nucleotides around initiator AUGs were heavily biassed toward the sequence gccaccAUGgcg. No such tendency was noted around internal AUGs. Although the role of this sequence bias is unclear, the sequence gccaccAUGg has been shown by Kozak to promote initiation. Another distinguishing feature was a C-rich tract 7 to 34 nucleotides 5' to the initiator AUGs. Ability to pair with more than eight bases of the start consensus sequence, matching of 6 or 7 nucleotides to the D loop on the 3' side, an C-richness on the 5' side were used as criteria for distinguishing start AUGs.(ABSTRACT TRUNCATED AT 400 WORDS)

Animals↗

Codon usage in mammalian genes is biased by sequence slippage mechanisms.

The codons for some conserved amino acids are found to be the same between homologous genes from different species when the statistics of codon usage would suggest that they should be different. I examine whether this 'coincidence' of codon usage could be due to genetic mechanisms homogenising the DNA around specific sites. This paper describes the further analysis of the coincident codons in 19 genes (a total of 96 homologues) for slippage. Coincident codons arise in contexts of increased sequence simplicity, and have a high chance of occurring within sequences similar to the recombination-prone minisatellite 'core' sequence. This suggests a role of genetic homogenisation in their generation.

Amino Acid Sequence↗

A novel sensitive method for the detection of user-defined compositional bias in biological sequences.

MOTIVATION: Most biological sequences contain compositionally biased segments in which one or more residue types are significantly overrepresented. The function and evolution of these segments are poorly understood. Usually, all types of compositionally biased segments are masked and ignored during sequence analysis. However, it has been shown for a number of proteins that biased segments that contain amino acids with similar chemical properties are involved in a variety of molecular functions and human diseases. A detailed large-scale analysis of the functional implications and evolutionary conservation of different compositionally biased segments requires a sensitive method capable of detecting user-specified types of compositional bias. RESULTS: We present BIAS, a novel sensitive method for the detection of compositionally biased segments composed of a user-specified set of residue types. BIAS uses the discrete scan statistics that provides a highly accurate correction for multiple tests to compute analytical estimates of the significance of each compositionally biased segment. The method can take into account global compositional bias when computing analytical estimates of the significance of local clusters. BIAS is benchmarked against SEG, SAPS and CAST programs. We also use BIAS to show that groups of proteins with the same biological function are significantly associated with particular types of compositionally biased segments.

Algorithms↗

Properties Governing Native State Entanglements and Relationships to Protein Function.

Non-covalent lasso entanglements are structural motifs found in a majority of globular proteins, and their misfolding has been linked to a range of biological consequences. Here, we characterize these motifs' structural and physicochemical properties, sequence biases, functional site correlations, and universal features across E. coli, S. cerevisiae, and H. sapiens. We find that the crossing residues, which pierce the plane of the entanglement loop, are 11-times more likely to be a &#x3b2;-strand than an &#x3b1;-helix or random coil, and that around this position the protein sequence is 2.5-times more likely to be composed of a stretch of all hydrophobic residues (most often Val, Ile, or Phe) compared to other sequence motifs. Functionally, crossing residues are enriched at enzyme active sites in S. cerevisiae and small molecule binding residues across all species to degrees greater than expected by random chance. Metal binding residues are enriched in these entanglements in H. sapiens. Increasing statistical power by pooling together these species data, we find RNA-binding residues are enriched in these entanglement components. On the other hand, there is a spatial depletion of crossing residues at sites involved in protein binding. Using machine learning, we identified eight robust features predictive of these entanglements, achieving AUROC scores of 0.8 across species. These results are significant because they suggest a direct role for components of native entanglements in particular protein functions, as well as identifying strong secondary structure and sequence preferences in native entanglements.

Humans↗

An analysis of structural instances of low complexity sequence segments.

Amino acid sequence databases contain many low complexity, compositionally biased sequence segments. However, only a limited number of relatively short instances of these segments occur in proteins of known structure. An analysis is presented of structural instances of these low complexity sequence segments in the Brookhaven Protein Data Bank with regard to preferences for sequence composition, secondary structural conformation and the local atomic environment. The complexity varies almost linearly with segment length, reflecting the absence of very long, low complexity segments in the structural database. The low complexity segments identified are not disordered and have temperature factors which are generally the same as the rest of the protein. It is observed that these segments are predominantly exposed and either helical or coiled, in excess of what would be expected by chance. Secondary structure prediction methods perform well in correctly predicting those low complexity segments which are helical but poorly in correctly predicting segments that are strands.

Amino Acid Sequence↗

Comparative genomics reveals long, evolutionarily conserved, low-complexity islands in yeast proteins.

Eukaryotic proteomes abound in low-complexity sequences, including tandem repeats and regions with significantly biased amino acid compositions. We assessed the functional importance of compositionally biased sequences in the yeast proteome using an evolutionary analysis of 2838 orthologous open reading frame (ORF) families from three Saccharomyces species (S. cerevisiae, S. bayanus, and S. paradoxus). Sequence conservation was measured by the amino acid sequence variability and by the ratio of nonsynonymous-to-synonymous nucleotide substitutions (K(a)/K(s)) between pairs of orthologous ORFs. A total of 1033 ORF families contained one or more long (at least 45 residues), low-complexity islands as defined by a measure based on the Shannon information index. Low-complexity islands were generally less conserved than ORFs as a whole; on average they were 50% more variable in amino acid sequences and 50% higher in K(a)/K(s) ratios. Fast-evolving low-complexity sequences outnumbered conserved low-complexity sequences by a ratio of 10 to 1. Sequence differences between orthologous ORFs fit well to a selectively neutral Poisson model of sequence divergence. We therefore used the Poisson model to identify conserved low-complexity sequences. ORFs containing the 33 most conserved low-complexity sequences were overrepresented by those encoding nucleic acid binding proteins, cytoskeleton components, and intracellular transporters. While a few conserved low-complexity islands were known functional domains (e.g., DNA/RNA-binding domains), most were uncharacterized. We discuss how comparative genomics of closely related species can be employed further to distinguish functionally important, shorter, low-complexity sequences from the vast majority of such sequences likely maintained by neutral processes.

Amino Acid Sequence↗

Recognition and classification of histones using support vector machine.

Histones are DNA-binding proteins found in the chromatin of all eukaryotic cells. They are highly conserved and can be grouped into five major classes: H1/H5, H2A, H2B, H3, and H4. Two copies of H2A, H2B, H3, and H4 bind to about 160 base pairs of DNA forming the core of the nucleosome (the repeating structure of chromatin) and H1/H5 bind to its DNA linker sequence. Overall, histones have a high arginine/lysine content that is optimal for interaction with DNA. This sequence bias can make the classification of histones difficult using standard sequence similarity approaches. Therefore, in this paper, we applied support vector machine (SVM) to recognize and classify histones on the basis of their amino acid and dipeptide composition. On evaluation through a five-fold cross-validation, the SVM-based method was able to distinguish histones from nonhistones (nuclear proteins) with an accuracy around 98%. Similarly, we obtained an overall >95% accuracy in discriminating the five classes of histones through the application of 1-versus-rest (1-v-r) SVM. Finally, we have applied this SVM-based method to the detection of histones from whole proteomes and found a comparable sensitivity to that accomplished by hidden Markov motifs (HMM) profiles.

Amino Acid Sequence↗

Unisexuality and molecular drive: bag320 sequence diversity in bacillus taxa (insecta phasmatodea).

Satellite DNA variability follows a pattern of concerted evolution through homogenization of new variants by genomic turnover mechanisms and variant fixation by chromosome redistribution into new combinations with the sexual process. Bacillus taxa share the same Bag320 satellite family and their reproduction ranges from strict bisexuality (B. grandii) to automictic (B. atticus) and apomictic (B. whitei = rossius/ grandii; B. lynceorum = rossius/grandii/atticus) unisexuality. Thelytokous reproduction clearly allows uncoupling of homogenization from fixation. Both trends and absolute values of satellite variability were analyzed in all Bacillus taxa but B. rossius, on 906 sequenced monomers at all level of comparisons: intraspecimen, intrapopulation, interpopulation, intersubspecies, and interspecies. For unisexuals, allozymic and mitochondrial clones were also taken into account. Different reproductive modes (sexual/parthenogenetic) appear to explain observed variability trends, supporting Dover's hypothesis of sexuality acting as a driving force in the fixation of sequence variants, but the present analyses also highlight current spreading of new variants in B. grandii maretimi specimens and point to a biased sequence inheritance at the time of hybrid onset in the apomictic hybrids B. whitei and B. lynceorum. Evidence of biased gene conversion events suggests that, given enough time, sequence homogenization can take place in a unisexual such as B. lynceorum. On the contrary, the absolute values of sequence diversity in each taxon are linked to the species' range, time of divergence, and repeat copy number and, possibly, to transposon features. Satellite dynamics appears therefore to be the outcome of both general molecular processes and specific organismal traits.

Animals↗

General method for sequence-independent site-directed chimeragenesis.

We have developed a simple and general method that allows for the facile recombination of distantly related (or unrelated) proteins at multiple discrete sites. To evaluate the sequence-independent site-directed chimeragenesis (SISDC) method, we have recombined beta-lactamases TEM-1 and PSE-4 at seven sites, examined the quality of the chimeric genes created, and screened the library of 2(8) (256) chimeras for functional enzymes. Probe hybridization and sequencing analyses revealed that SISDC generated a random library with little sequence bias and in which all targeted fragments were recombined in the desired order. Sequencing the genes from clones having functional lactamases identified 14 unique chimeras. These chimeras are characterized by a lower level of disruption, as calculated by the SCHEMA algorithm, than the library as a whole. These results illustrate the use of SISDC in creating designed chimeric protein libraries and further illustrate the ability of SCHEMA to identify chimeras whose folded structures are likely not to be disrupted by recombination.

Algorithms↗

A BAC library and paired-PCR approach to mapping and completing the genome sequence of Sulfolobus solfataricus P2.

The original strategy used in the Sulfolobus solfataricus genome project was to sequence non overlapping, or minimally overlapping, cosmid or lambda inserts without constructing a physical map. However, after only about two thirds of the genome sequence was completed, this approach became counter-productive because there was a high sequence bias in the cosmid and lambda libraries. Therefore, a new approach was devised for linking the sequenced regions which may be generally applicable. BAC libraries were constructed and terminal sequences of the clones were determined and used for both end mapping and PCR screening. The PCR approaches included a novel chromosome walking method termed "paired-PCR". 21 gaps were filled by BAC end sequence analyses and 6 gaps were filled by PCR including three large ones by paired-PCR. The complete map revealed that 0.9 Mb remained to be sequenced and 34 BAC clones were selected for walking over small gaps and preparing template libraries for larger ones. It is concluded that an optimal strategy for sequencing microorganism genomes involves construction of a high-resolution physical map by BAC end analyses, PCR screening and paired-PCR chromosome walking after about half the genome sequence has been accumulated.

Chromosomes, Artificial, Bacterial↗

Structure and sequence determinants required for the RNA editing of ADAR2 substrates.

ADAR2 is a double-stranded RNA-specific adenosine deaminase involved in the editing of mammalian RNAs by the site-specific conversion of adenosine to inosine. We have demonstrated previously that ADAR2 can modify its own pre-mRNA, leading to the creation of a proximal 3'-splice junction containing a non-canonical adenosine-inosine (A-I) dinucleotide. Alternative splicing to this proximal acceptor shifts the reading frame of the mature mRNA transcript, resulting in the loss of functional ADAR2 expression. Both evolutionary sequence conservation and mutational analysis support the existence of an extended RNA duplex within the ADAR2 pre-mRNA formed by base-pairing interactions between regions approximately 1.3-kilobases apart in intron 4 and exon 5. Characterization of ADAR2 pre-mRNA transcripts isolated from adult rat brain identified 16 editing sites within this duplex region, and sites preferentially modified by ADAR1 and ADAR2 have been defined using both tissue culture and in vitro editing systems. Statistical analysis of nucleotide sequences surrounding edited and non-edited adenosine residues have identified a nucleotide sequence bias correlating with ADAR2 site preference and editing efficiency. Among a mixed population of ADAR substrates, ADAR2 preferentially favors its own transcript, yet mutation of a poor substrate to conform to the defined nucleotide bias increases the ability of that substrate to be modified by ADAR2. These data suggest that both sequence and structural elements are required to define adenosine moieties targeted for specific ADAR2-mediated deamination.

Adenosine Deaminase↗