PubMed HealthSearch

SEARCH · PubMed Health

Results for “recognition code”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Assessment of protein coding measures.

A number of methods for recognizing protein coding genes in DNA sequence have been published over the last 13 years, and new, more comprehensive algorithms, drawing on the repertoire of existing techniques, continue to be developed. To optimize continued development, it is valuable to systematically review and evaluate published techniques. At the core of most gene recognition algorithms is one or more coding measures--functions which produce, given any sample window of sequence, a number or vector intended to measure the degree to which a sample sequence resembles a window of 'typical' exonic DNA. In this paper we review and synthesize the underlying coding measures from published algorithms. A standardized benchmark is described, and each of the measures is evaluated according to this benchmark. Our main conclusion is that a very simple and obvious measure--counting oligomers--is more effective than any of the more sophisticated measures. Different measures contain different information. However there is a great deal of redundancy in the current suite of measures. We show that in future development of gene recognition algorithms, attention can probably be limited to six of the twenty or so measures proposed to date.

Algorithms

Nuclear volume control by nucleoskeletal DNA, selection for cell volume and cell growth rate, and the solution of the DNA C-value paradox.

The 40,000-fold variation in eukaryote haploid DNA content is unrelated to organismic complexity or to the numbers of protein-coding genes. In eukaryote microorganisms, as well as in animals and plants, DNA content is strongly correlated with cell volume and nuclear volume, and with cell cycle length and minimum generation time. These correlations are simply explained by postulating that DNA has 2 major functions unrelated to its protein-coding capacity: (1) the control of cell volume by the number of replicon origins, and (2) the determination of nuclear volume by the overall bulk of the DNA: cell growth rates are determined by the cell volume and by the area of the nuclear envelope available for nucleocytoplasmic transport of RNA, which in turn depends on the nuclear volume and therefore on the DNA content. During evolution nuclear volume, and therefore DNA content, has to be adjusted to the cell volume to allow reasonable growth rates. The great diversity of cell volumes and growth rates, and therefore of DNA contents, among eukaryotes results from a varying balance in different species between r-selection, which favours small cells and rapid growth rates and therefore low DNA C-values, and K-selection which favours large cells and slow growth rates and therefore high DNA C-values. In multicellular organisms cell size needs to vary in different tissues: size differences between somatic cells result from polyteny, endopolyploidy, or the synthesis of nucleoskeletal RNA. Conflict between the need for large ova and small somatic cells explains why lampbrush chromosomes, nurse cells, chromatin diminution and chromosome elimination evolved. Similar evolutionary considerations clarify the nature of polygenes, the significance of the distribution of haploidy, diploidy and dikaryosis in life cycles and of double fertilization in angiosperms, and of heteroploidy despite DNA constancy in cultured cells, and other puzzles in eukaryote chromosome biology. Eukaryote DNA can be divided into genic DNA (G-DNA), which codes for proteins (or serves as recognition sites for proteins involved in transcription, replication and recombination), and nucleoskeletal DNA (S-DNA) which exists only because of its nucleoskeletal role in determining the nuclear volume (which it shares with G-DNA, and performs not only directly, but also indirectly by coding for nucleoskeletal RNA). Mechanistic and evolutionary implications of this are discussed.

Animals

Structure of rat calmodulin processed genes with implications for a mRNA-mediated process of insertion.

Two distinct processed calmodulin genes of rat (lambda SC8 and lambda SC9) were identified, cloned and their DNA sequences determined. The existence of direct repeats of 19 base-pairs for lambda SC8 or 9 base-pairs for lambda SC9 at both ends of the coding plus non-coding regions suggested a possible involvement of a mRNA-mediated process of insertion. Total genomic Southern hybridization suggested the existence of at least three different calmodulin-related genes in the rat genome. The other gene was the bona fide calmodulin gene (lambda SC4) which was split into at least five exons. lambda SC9 contained insertions of one nucleotide and two 17 base-pair direct repeats in the coding region. These insertions cause frameshift mutations probably preventing it from encoding a functional calmodulin. It also carried an insertion of a rat middle repetitive sequence, identifier sequence (IDS: Sutcliffe et al., 1982) in the 3'-non-coding region. Otherwise, it consisted of an almost identical DNA sequence to that of the bona fide calmodulin gene (lambda SC4), including the 3'-non-coding region down to the poly(A) recognition signal, A-A-T-A-A-A. On the other hand, lambda SC8 did not possess frameshift mutations in the coding region, and hence was capable of encoding a functional protein. In fact, a probe specific to the lambda SC8 sequence identified a band in Northern blotting whose size was 300 nucleotides smaller than that of authentic calmodulin mRNA. Comparison of the nucleotide sequences showed that only the coding regions of these two processed genes were homologous, indicating that the divergence of these two processed genes from the common ancestor calmodulin was an ancient event.

Amino Acid Sequence

Cloning and sequencing of the yeast Saccharomyces cerevisiae SEC1 gene localized on chromosome IV.

The SEC1 gene of yeast Saccharomyces cerevisiae was cloned by complementing the temperature-sensitive mutation of sec1-1 at 37 degrees C, and its nucleotide sequence was determined. SEC1 is a single copy gene and encodes a protein of 724 amino acids and 83,490 daltons with a predicted pI value of 6.11. Hydrophobicity plotting showed no clearly hydrophobic regions suggesting a soluble nature for the protein. Amino acid sequence comparisons revealed no obvious homologies with the proteins in the SWISSPROT databank. Two consensus sequence for the cdc2 encoded protein kinase recognition site were revealed within Sec1p. The codon usage suggests a low expression level for SEC1. The 5' non-translated region contains two TATA-like sequences at -52 and -215 nucleotides from the translation start site. Two potential regulatory sequences for DNA binding proteins were found in the non-coding 5' region: a HAP2/HAP3 consensus recognition sequence at nucleotide-154 and a BAF1 consensus recognition sequence at nucleotide-136. The SEC1 specific probe detected a 2400 nucleotides long transcript, which was in reasonable agreement with the 2172 nucleotides long open reading frame.

Amino Acid Sequence

Avoidance of DNA methylation. A virus-encoded methylase inhibitor and evidence for counterselection of methylase recognition sites in viral genomes.

The ocr+ gene of bacterial virus T7 codes for the first protein recognized to inhibit a specific group of DNA methylases. The recognition sequences of several other DNA methylases, not susceptible to Ocr inhibition, are significantly suppressed in the virus genome. The bacterial virus T3 encodes an Ado-Met hydrolase, destroying the methyl donor and causing T3 DNA to be totally unmethylated. These observations could stimulate analogous investigations into the regulation of DNA methylation patterns of eukaryotic viruses and cells. For instance, an underrepresentation of methylation sites (5'-CG) is also true for animal DNA viruses. Moreover, we were able to disclose some novel properties of DNA restriction-modification enzymes concerning the protection of DNA recognition sequences in which only one strand can be methylated (e.g., type III enzyme EcoP15) and the primary resistance of (unmethylated) DNA recognition sites towards type II restriction endonuclease EcoRII.

Base Sequence

Evolution of the major histocompatibility complex.

The major histocompatibility complex is a group of closely linked loci that code for molecules used by T-lymphocytes as context for the recognition of antigens. The loci fall into two classes: I, coding for molecules used as context by cytotoxic T-lymphocytes and II, used as context by helper and other regulatory T-cells. The Mhc is present in all mammals and perhaps all vertebrates. Some of the Mhc loci are highly polymorphic, while others are not. This article will summarize what is known about the genetic organization of the Mhc in different species and will discuss the selection pressures acting on the individual loci and the tempo and the mechanisms of their evolution.

Animals

Sensory and meaning features in stimulus recognition and associative retrieval.

Three experiments addressed the problem of isolating the effects of sensory similarity on subprocesses involved in coding paired associates. In the first, the standard recognition-recall procedure was used and stimulus similarity, concreteness, and frequency were varied. However, because of concern with the validity of this recognition procedure as a measure of functional stimulus contact, an alternative was developed. This alternative led to the second study in which only stimulus similarity was manipulated. In the third experiment, similarity was varied, and the pairs were either associatively compatible, unrelated, or incompatible. The results using the new procedure indicated that similarity consistently disrupted functional stimulus contact but not associative retrieval. By contrast, associative relatedness facilitated both subprocesses.

Adolescent

A cluster of mutations in HLA-A2 alpha 2 helix abolishes peptide recognition by T cells.

In order to investigate the regions of HLA-A2 that control peptide-specific cytotoxic T lymphocyte (CTL) recognition, 37 HLA-A2 genes coding for 50 point mutations that span the alpha 2 helix were synthesized by the technique of saturation mutagenesis. Twenty-nine of these genes, which code for 41 point mutations, were transfected into C1R cells and used as targets in cytotoxicity assays, in the presence of influenza-A matrix peptide 58-68 with specific CTL as effectors. All the transfectants were recognized fully by matrix peptide-specific CTL apart from those with amino acid substitutions at positions 152, 154, 155, 156, or 161, which led to a total loss of recognition and those with mutations at residue 27 or a double mutation at 138 and 150, which were recognized in an intermediate manner. The clustering of the crucial residues that emerges may reflect direct interaction of their side-chains with peptide or the CTL receptor.

Cell Line

The major histocompatibility complex determines susceptibility to cytotoxic T cells directed against minor histocompatibility antigens.

Cytotoxic cells were generated by immunizing one strain of mouse with cells from an allogeneic strain which carries the same H-2 region. The effector cells assayed in a 4 h 51Cr release assay were shown to be T cells and indistinguishable, except in specificity, from cytotoxic T cells directed at H-2 alloantigens. Although the genetic differences between responder and stimulator cells responsible for the immunization did not code in H-2, the H-2 complex did restrict susceptibility of target cells. For example, BALB.B cytotoxic cells (H-2b) immunized against and capable of lysing C57BL/6 cells (H-2b) would not lyse B6.C/H-2d target cells. C57BL/6 and B6.C/H-2d are congenic and differ in the H-2 region. Two hypotheses are considered to explain the H-2 restriction of susceptibility to cytotoxic T cells generated by an H-2 identical alloimmunization. (a) The dual (self) recognition hypothesis states that the cytotoxic cell has two recognition units, one for H-2-coded structures and another clonally restricted receptor for the minor alloantigen. (b) The interaction antigen hypothesis states that all the surface alloantigenic determinants recognized by cytotoxic T cells are the result of interaction between H-2- and non-H-2-coded gene products. Two lines of evidence, one with F1 effector cells and the other a cold target competition experiment, are presented which argue strongly in favor of the interaction antigen hypothesis. The regions of H-2 required to be histocompatible were mapped to the D region and to the left of IC, probably the K region. These results, and recent work on the response to virus-infected and TNP-modified syngeneic cells, suggest that cytotoxic cells are restricted in specificity to preferentially recognizing alterations in structures that are coded in the major histocompatibility complex.

Animals

Application of learning techniques to splicing site recognition.

Most genes of eukaryotic genomes are disrupted by introns. The application of a learning technique which uses both statistic and syntactic analysis lead to the establishment of logical rules enabling the recognition of intron/exon junctions between uncoding and coding sequences. The rules were tested on rat actin gene sequences containing some or all of the introns and 50 exon nucleotides on either side of the intron. The results show good recognition of the excision site. This recognition is more ambiguous when the sequence is short; for the acceptor sequence it presents a good selection. The learning achieved with both the donor and acceptor sequence does not lead to recognition. This result indicates that it is not the relationship between donor and acceptor sites in the same intron which determines sequence selection or the splicing mechanism.

Base Sequence

Control of replication of plasmid R1: translation of the 7k reading frame in the RepA mRNA leader region counteracts the interaction between CopA RNA and CopT RNA.

Replication of IncFII plasmids is regulated through the expression of a gene, repA. The RepA protein is rate-limiting for initiation of replication. The main negative control is exerted by a countertranscript, CopA RNA, that binds to the complementary region of the RepA mRNA, thereby inhibiting the formation of the RepA protein. The target region for CopA RNA, CopT, is located upstream of the RepA coding region. An open reading frame for a putative 7k protein overlaps the CopT sequence. Here we show by using lacZ fusions that the 7k gene is expressed. We constructed a translation start mutation in order to abolish formation of the 7k protein. This resulted in a 10-fold decrease in repA expression. The 7k protein produced in trans did not reverse this effect, so the 7k protein per se does not control expression of repA. However, translation of the 7k coding sequence must influence CopA/CopT RNA recognition, since ribosomes will transiently disrupt the target hairpin. We propose here a novel mechanism that affects the level of gene expression: the 7k region of the RepA mRNA is a leader sequence that is involved in expression of the downstream gene; translation of the 7k region competes with a negative control system involving RNA-RNA interaction.

Amino Acid Sequence

Expression in Escherichia coli of the genes coding for reaction center subunits from Rhodobacter sphaeroides: wild-type proteins and fusion proteins containing one or four truncated domains from Staphylococcus aureus protein A at the carboxy-terminus.

Gene cassettes were constructed containing Rhodobacter sphaeroides puhA, pufM and pufL sequences with synthetic 5' ends for production in Escherichia coli of the H, M and L subunits of the photosynthetic reaction center. In addition, gene cassettes coding for fusion proteins with proteinase recognition site(s) between the amino-terminal part of H, M or L subunits, and the carboxy-terminal part consisting of one (B') or four (D'ABC') domains of Staphylococcus aureus protein A were constructed. A modified expression vector pDS12/RBSII containing the T5 promoter PN25, the lac operator, and a newly inserted E. coli lipoprotein ribosome-binding site was used. Inducible synthesis of plasmid-encoded polypeptides was accompanied by reduced growth. The products comigrated with R. sphaeroides reaction center subunits H, M and L. They were identified by Western blot experiments using antibodies raised against reaction center proteins. The hybrid protein containing the reaction center H subunit fused to the single domain B' was not detected by nonspecific antisera. In contrast, the three fusion proteins containing domains D'ABC' were identified using nonspecific antisera. This indicated that domains D'ABC' were sufficient to bind to the Fc part of IgG molecules, whereas domain B' was not sufficient. This property was used to purify all three fusion proteins with domains D'ABC' by affinity chromatography from the membrane fraction of E. coli cells.

Amino Acid Sequence

[Congruence of monotonous DNA/protein polymers].

The comparative study of the monotonous polynucleotides and polypeptides reveals an univocal congruence between: poly-C and poly-PRO;poly-G and poly-GLY;poly-T and poly-PHE;poly-A and poly-LYS. The hypothesis is presented of a code of congruence ruling the specific recognition of the wide groove of the DNA by proteins.

Chemical Phenomena

Word attributes and lateralization revisited: implications for dual coding and discrete versus continuous processing.

Three attributes of words are their imageability, concreteness, and familiarity. From a literature review and several experiments, I previously concluded (Boles, 1983a) that only familiarity affects the overall near-threshold recognition of words, and that none of the attributes affects right-visual-field superiority for word recognition. Here these conclusions are modified by two experiments demonstrating a critical mediating influence of intentional versus incidental memory instructions. In Experiment 1, subjects were instructed to remember the words they were shown, for subsequent recall. The results showed effects of both imageability and familiarity on overall recognition, as well as an effect of imageability on lateralization. In Experiment 2, word-memory instructions were deleted and the results essentially reinstated the findings of Boles (1983a). It is concluded that right-hemisphere imagery processes can participate in word recognition under intentional memory instructions. Within the dual coding theory (Paivio, 1971), the results argue that both discrete and continuous processing modes are available, that the modes can be used strategically, and that continuous processing can occur prior to response stages.

Adult

Cloning and expression in E. coli of a synthetic gene for the bacteriocidal protein caltrin/seminalplasmin.

A synthetic gene coding for the bacteriocidal protein caltrin/seminalplasmin was constructed and expressed in Escherichia coli as a fusion with beta-galactosidase. The gene was designed with a recognition site for the plasma protease, Factor Xa, coded for immediately prior to the N-terminus of caltrin. The beta-galactosidase-caltrin fusion protein was cleaved with Factor Xa to give caltrin, which was identified by its size on SDS-PAGE, its ability to react with an antiserum raised to the N-terminal nonapeptide of caltrin and its N-terminal amino acid sequence. After partial purification, synthetic caltrin was found to be active in an assay involving inhibition of growth of E.coli.

Antineoplastic Agents

Structure and expression of a gene encoding heat-shock protein Hsp70 from the Oomycete fungus Bremia lactucae.

A gene encoding a protein homologous to a 70-kDa heat-shock protein (Hsp70) was isolated from Bremia lactucae and its structure and pattern of expression were determined. This is the first report on the structure of a protein-coding gene from an Oomycete fungus. The cloned gene is a member of a small multigene family. The level of hsp70 mRNA in germlings increases from a low constitutive level in response to heat or cold treatment. A high level of the mRNA is also detected in spores. The hsp70 gene is expressed as a primary transcript of 2241 nucleotides (nt) and contains a continuous open reading frame of 2025 nt. Near the C terminus of the coding sequence is an unusual region that contains repeated enhancer-like sequences. This insert has not been described in other hsp70 genes and is not an intron. Upstream from the 5' terminus of the mRNA are multiple CCAAT motifs, a sequence similar to a consensus heat-shock regulatory element, and an A + T-rich putative 'TATA' box. A canonical polyadenylation recognition sequence is present downstream from the coding sequence. The deduced amino acid sequence is equally similar to yeast and maize Hsp70, providing further evidence of the dissimilarity between Oomycetes and true fungi. The cloning of this gene is part of our strategy to develop a transformation system for B. lactucae.

Amino Acid Sequence

Protein tertiary structure recognition using optimized Hamiltonians with local interactions.

Protein folding codes embodying local interactions including surface and secondary structure propensities and residue-residue contacts are optimized for a set of training proteins by using spin-glass theory. A screening method based on these codes correctly matches the structure of a set of test proteins with proteins of similar topology with 100% accuracy, even with limited sequence similarity between the test proteins and the structural homologs and the absence of any structurally similar proteins in the training set.

Amino Acids