PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “start codons”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

The 5' leader of plant PgiC has an intron: the leader shows both the loss and maintenance of constraints compared with introns and exons in the coding region.

PgiC, a complex gene with 23 coding exons and 22 intervening introns, encodes the cytosolic isozyme of phosphoglucose isomerase (EC 5.3.1.9) in higher plants. Here, we report RNA ligase-mediated rapid amplification of cDNA ends experiments that showed that PgiC in Clarkia (Onagraceae) and Arabidopsis thaliana has an intron in the 5' leader. Comparison of the EMBL accessions of the cDNA and genomic sequences showed that this is also the case in rice (Oryza sativa), suggesting that a leader intron is generally present in higher plant PgiC. The intron is bounded by consensus 5'-GT and AG-3' splice sites but showed alternative splicing in Clarkia, resulting in mature transcripts that differ by 8-19 nt in length. The intron is located 18 or 10 nt upstream of the start codon in Clarkia, 2 nt upstream in Arabidopsis, and 9 nt in rice. PgiC in Clarkia was duplicated before the divergence of the extant species, many of which have two expressed genes PgiC1 and PgiC2. Full-length transcripts of both genes identified the transcription start and made it possible to identify the leader intron and leader exon (between the transcription start and leader intron) from previously obtained genomic sequences of both genes in other Clarkia species. These data permit the comparison of evolution in the leader exon and intron with the exons and introns of the coding region, a topic that has not been studied previously. Both the leader exon and the leader intron resemble introns of the coding region in base substitution rate and accumulation of gaps. But the leader intron splice junctions are not strictly conserved in position as are those of the coding region introns. Also, in base composition, the leader intron resembles the other introns, whereas the leader exon more nearly resembles the coding exons. A difference in base composition between coding exons and flanking introns is known to be important for the recognition of splice sites. Thus, the marked difference in base composition between the leader exon and leader intron is probably maintained by selection despite a high rate of sequence divergence.

5' Untranslated Regions↗

Sequence analysis of a mosquito ribosomal protein rpL8 gene and its upstream regulatory region.

The gene encoding Aedes albopictus ribosomal protein L8 was isolated using a cDNA probe. Based on the deduced amino acid sequence, rpL8 has a mass of 28,605 Da, a pI of 11.97, and contains 9.6% Arg and 11.9% Lys. The rpL8 gene spans 1229 nucleotides, and contains three exons measuring 73, 150, and 648 nucleotides. The first intron is 293 nucleotides long and interrupts an 85-nucleotide untranslated leader sequence. The AUG codon is located 12 nucleotides downstream of the 5'-end of the second exon. Separating the second and third exon is a 65-nucleotide intron. The major transcription initiation site, identified by primer extension and polymerase stop reactions, mapped 378 nucleotides upstream from the AUG start codon; minor initiation sites were also detected. The DNA sequence upstream of the rpL8 gene was T-rich, but conventional TATA and CAAT boxes were absent. This is the first molecular analysis of a mosquito ribosomal protein gene.

Aedes↗

Cloning and partial characterization of the human tie-2 receptor tyrosine kinase gene promoter.

The Tie-2 receptor plays a key role in vascular development, although little is known about the factors controlling its expression. Here we report the first cloning and characterisation of the 5' regulatory region of human tie-2. Multiple transcription start sites were identified between -414 and -265 bp upstream of the start codon using 5' RACE, fluorescent primer extension, and RNase protection assays. The human tie-2 promoter contains several transcription factor-binding sequences including ets, SP-1, AP-1, and GATA-1, but there are no canonical TATA or CCAAT initiation sequences proximal to the transcription start sites. Human tie-2 reporter constructs demonstrated approximately 10-fold greater activity in endothelial cells compared with fibroblasts. In endothelial cells the tie-2 promoter exhibited 5 and 16% of the activity of human tie-1 (830 bp) and KDR (1.1 kb) promoters, respectively. This promoter will be a useful tool for studying factors that regulate tie-2 expression and targeting the vasculature.

Animals↗

Analysis of the Desulfovibrio gigas transcriptional unit containing rubredoxin (rd) and rubredoxin-oxygen oxidoreductase (roo) genes and upstream ORFs.

Rubredoxin-oxygen oxidoreductase, an 86-kDa homodimeric flavoprotein, is the final component of a soluble electron transfer chain that couples NADH oxidation with oxygen reduction to water from the sulfate-reducing bacterium Desulfovibrio gigas. A 4.2-kb fragment of D. gigas chromosomal DNA containing the roo gene and the rubredoxin gene was sequenced. Additional open reading frames designated as ORF-1, ORF-2, and ORF-3 were also identified in this DNA fragment. ORF-1 encodes a protein exhibiting homology to several proteins of the short-chain dehydrogenase/reductase family of enzymes. The N-terminal coenzyme-binding pattern and the active-site pattern characteristic of short chain dehydrogenase/reductase proteins are conserved in ORF-1 product. ORF-2 does not show any significant homology with any known protein, whereas ORF-3 encodes a protein having significant homologies with the branched-chain amino acid transporter AzlC protein family. Northern blot hybridization analysis with rd and roo-specific probes identified a common 1.5-kb transcript, indicating that these two genes are cotranscribed. The transcription start site was identified by primer extension analysis to be a guanidine 87 bp upstream the ATG start codon of rubredoxin. The transcript size indicates that the rd-roo mRNA terminates downstream the roo-coding unit. Putative -10 and -35 regulator regions of a sigma(70)-type promoter, having similarity with E. coli sigma(70) promoter elements, are found upstream the transcription start site. Rubredoxin-oxygen oxidoreductase and rubredoxin genes are shown to be constitutively and abundantly expressed. Using the data available from different prokaryotic genomes, the rubredoxin genomic organization and the first tentative to understand the phylogenetic relationships among the flavoprotein family are reported in this study.

Amino Acid Sequence↗

Transcription mapping and functional analysis of the protein tyrosine/serine phosphatase (PTPase) gene of the Autographa californica nuclear polyhedrosis virus.

The protein tyrosine/serine phosphatase (PTPase) gene of the Autographa californica nuclear polyhedrosis virus yields two major transcripts of approximate sizes 3.1 and 3.9 kb. Both of these are very late transcripts, which accumulate to maximal levels more than 30 hr postinfection. The smaller transcript initiates at the T of an ATAAG sequence that lies 22 bp upstream of the putative start codon of the gene. The larger transcript initiates at the first A of a TTAAG sequence that lies approximately 798 bp further upstream. Thus, the larger transcript is made by transcribing through the entire hr1 region, which lies just 60 bp upstream of the putative translation start site. We have expressed the product of this gene as a fusion protein with glutathione-S-transferase and have shown that it has PTPase activity similar to that of the vaccinia virus H1 gene product: It dephosphorylates both protein phosphotyrosines and phosphoserines/phosphothreonines, and it is inhibited by vanadate, but not by okadaic acid.

Amino Acid Sequence↗

Transcriptional mapping of the genes encoding the early enzymes of the cephamycin biosynthetic pathway of Streptomyces clavuligerus.

Isopenicillin-N synthase (IPNS) of Streptomyces clavuligerus is encoded by the pcbC gene which is found within the cephamycin biosynthetic gene cluster. pcbC is located directly downstream from lat and pcbAB, which encode the enzymes, lysine epsilon-amino transferase and delta-(L-alpha-aminoadipyl)-L-cysteinyl-D-valine synthetase, respectively. These enzymes act prior to IPNS in the biosynthetic pathway, and the three genes are transcribed in the same direction. Previous pcbC transcriptional studies involving recombinant promoter probe plasmids, Northern analysis and 5' primer extension indicated the presence of a monocistronic 1.2-kb transcript that initiated within pcbAB, 92-bp upstream from the pcbC start codon. S1 nuclease mapping studies have now shown, not only the transcript initiating 92 bp upstream from pcbC, but also a transcript initiating further upstream, possibly including the entire pcbAB gene. Promoter probe analysis and S1 nuclease mapping failed to detect promoter activity or a transcription start point (tsp) directly upstream from pcbAB, suggesting that pcbAB transcripts initiated within or upstream from lat. Northern analysis, to search for a pcbAB transcript, showed no distinct transcript and indicated severely degraded mRNA. Similar results were obtained when Northern analysis was used to search for lat transcripts. Promoter probe analysis to locate the lat promoter indicated that a sequence promoting transcription was present in a 330-bp DNA fragment that extended from 227-bp upstream from the lat structural gene to 103 bp inside the gene.(ABSTRACT TRUNCATED AT 250 WORDS)

Base Sequence↗

Mutations in the structural genes for eukaryotic initiation factors 2 alpha and 2 beta of Saccharomyces cerevisiae disrupt translational control of GCN4 mRNA.

The SUI2 and SUI3 genes of Saccharomyces cerevisiae encode the alpha and beta subunits, respectively, of translation initiation factor eIF-2 (eukaryotic initiation factor 2). Previously isolated mutations in these genes restore expression from his4 mutant alleles lacking an ATG initiation codon. The SUI mutations also lead to increased levels of HIS4 mRNA. We show that the latter phenotype exists because the SUI mutations elevate expression of GCN4, an activator of HIS4 transcription. Increased GCN4 expression in the SUI mutants occurs independently of the GCN2 and GCN3 gene products that are normally required to stimulate translation of GCN4 mRNA under conditions of amino acid starvation. Derepression of GCN4 expression in the SUI mutants requires the multiple AUG codons in the leader of the GCN4 transcript that normally mediate its translational control by amino acid availability. In these respects, the SUI mutations resemble mutations in GCD genes whose products function as translational repressors of GCN4. Thus, in addition to its general role in AUG start codon selection, eIF-2 appears to be an important factor in GCN4 translational control. We also show that deletion of GCN3 in sui2-1 strains is lethal, suggesting that GCN3 contributes to eIF-2 alpha function in addition to its role as a translational activator of GCN4.

Chromosome Deletion↗

A periodic pattern of mRNA secondary structure created by the genetic code.

Single-stranded mRNA molecules form secondary structures through complementary self-interactions. Several hypotheses have been proposed on the relationship between the nucleotide sequence, encoded amino acid sequence and mRNA secondary structure. We performed the first transcriptome-wide in silico analysis of the human and mouse mRNA foldings and found a pronounced periodic pattern of nucleotide involvement in mRNA secondary structure. We show that this pattern is created by the structure of the genetic code, and the dinucleotide relative abundances are important for the maintenance of mRNA secondary structure. Although synonymous codon usage contributes to this pattern, it is intrinsic to the structure of the genetic code and manifests itself even in the absence of synonymous codon usage bias at the 4-fold degenerate sites. While all codon sites are important for the maintenance of mRNA secondary structure, degeneracy of the code allows regulation of stability and periodicity of mRNA secondary structure. We demonstrate that the third degenerate codon sites contribute most strongly to mRNA stability. These results convincingly support the hypothesis that redundancies in the genetic code allow transcripts to satisfy requirements for both protein structure and RNA structure. Our data show that selection may be operating on synonymous codons to maintain a more stable and ordered mRNA secondary structure, which is likely to be important for transcript stability and translation. We also demonstrate that functional domains of the mRNA [5'-untranslated region (5'-UTR), CDS and 3'-UTR] preferentially fold onto themselves, while the start codon and stop codon regions are characterized by relaxed secondary structures, which may facilitate initiation and termination of translation.

3' Untranslated Regions↗

Expression of bovine viral diarrhoea virus glycoprotein E2 by bovine herpesvirus-1 from a synthetic ORF and incorporation of E2 into recombinant virions.

Expression cassettes containing the codons for the pestivirus E (rns) signal peptide (Sig) followed by a chemically synthesized ORF that encoded the bovine viral diarrhoea virus (BVDV) strain C86 glycoprotein E2, a class I membrane glycoprotein, were constructed with and without a chimeric intron sequence immediately upstream of the translation start codon, and incorporated into the genome of bovine herpesvirus-1 (BHV-1). The resulting recombinants, BHV- 1/SigE2(syn) and BHV-1/SigE2(syn)-intron, expressed comparable quantities of glycoprotein E2, and Northern blot hybridizations indicated that the presence of the intron did not increase significantly the steady-state levels of transcripts encompassing the SigE2(syn) ORF. In BHV-1/SigE2(syn)- infected cells, the 54 kDa E2 glycoprotein formed a dimer with an apparent molecular mass of 94 kDa, which was further modified to a 101 kDa form found in the envelope of recombinant virus particles. Penetration kinetics and single-step growth curves indicated that the incorporation of the BVDV E2 glycoprotein in the BHV-1 envelope, which apparently did not require BHV-1-specific signals, interfered with entry into target cells and egress of progeny virions. These results demonstrate that a pestivirus glycoprotein can be expressed efficiently by BHV-1 and incorporated into the viral envelope. BHV-1 thus represents a promising tool for the development of efficacious live and inactivated BHV-1-based vector vaccines.

Amino Acid Sequence↗

The human myosin light chain kinase (MLCK) from hippocampus: cloning, sequencing, expression, and localization to 3qcen-q21.

Myosin light chain kinase (MLCK), a key enzyme in muscle contraction, has been shown by immunohistology to be present in neurons and glia. We describe here the cloning of the cDNA for human MLCK from hippocampus, encoding a protein sequence 95% similar to smooth muscle MLCKs but less than 60% similar to skeletal muscle MLCKs. The cDNA clone detected two RNA transcripts in human frontal and entorhinal cortex, in hippocampus, and in jejunum, one corresponding to MLCK and the other probably to telokin, the carboxy-terminal 154 codons of MLCK expressed as an independent protein in smooth muscle. Levels of expression were lower in brain compared to smooth muscle. We show that within the protein sequence, a motif of 28 or 24 residues is repeated five times, the second repeat ending with the putative methionine start codon. These repeats overlap with a second previously reported module of 12 residues repeated five times in the human sequence. In addition, the acidic C-terminus of all MLCKs from both brain and smooth muscle resembles the C-terminus of tubulins. The chromosomal localization of the gene for human MLCK is shown to be at 3qcen-q21, as determined by PCR and Southern blotting using two somatic cell hybrid panels.

Aged↗

Characterisation of the hydroxystreptomycin phosphotransferase gene (sph) of Streptomyces glaucescens: nucleotide sequence and promoter analysis.

The nucleotide sequence of a 1384 bp fragment containing the coding and promoter sequences of the streptomycin phosphotransferase gene (sph) of the hydroxystreptomycin-producing Streptomyces glaucescens was determined. Evidence for an ATG as translation start codon for sph was derived from a comparison with the amino-terminal amino acid sequence of an aminoglycoside phosphotransferase (aphD gene product) of S. griseus, exhibiting a high degree of amino acid homology to the deduced amino acid sequence of the S. glaucescens sph gene product. Transcriptional start and termination sites for the sph gene were identified by primer extension and/or nuclease S1 mapping experiments. The promoter region of the sph gene appears to be complex since tandemly arranged promoters (orfIp1, orfIp2) initiating transcription of a likely coding region (ORFI) in the opposite direction overlap sph promoter sequences. The presumptive sphp and orfIp1 promoters show considerable sequence similarities in the -10 region to Escherichia coli consensus promoter sequences but no homology to E. coli or Streptomyces -35 regions.

Amino Acid Sequence↗

Nucleotide sequence and analysis of a gene (chiA) for a chitinase from Streptomyces lividans 66.

A chitinase gene (chiA) from Streptomyces lividans was characterized and its nucleotides sequenced. Although the deduced amino acid sequence of chitinase A1 did not show any similarity to those of other Streptomyces chitinases that has been sequenced, the C-terminal part, containing both a putative catalytic domain and type-III-like repeating units, showed a similarity (36%) to that of chitinase D from Bacillus circulans. A site of initiation of transcription was found approximately 51 bp upstream from the GTG initiation codon. The promoter region of the chiA gene was subcloned on a 178-bp fragment into the promoter-probe vector pIJ486, resulting in the chitin stimulated expression of the neomycin resistance gene. One of the deleted subclones, which contained a 114-bp sequence upstream from the translation start codon, retained both chitin stimulated production and glucose repression. Chitin stimulated production was lost in an other deleted mutant containing the 104-bp upstream sequence.

Amino Acid Sequence↗

Genomic characterization of human DSPG3.

DSPG3, the human homolog to chick PG-Lb, is a mejrkp6of the small leucine-rich repeat proteoglycan (SLRP) family, including decorin, biglycan, fibromodulin, and lumican. In contrast to the tissue distribution of the other SLRPs, DSPG3 is predominantly expressed in cartilage. In this study, we have determined that the human DSPG3 gene is composed of seven exons: Exon 2 of DSPG3 includes the start codon, exons 4-7 code for the leucine-rich repeats, exons 3 and 7 contain the potential glycosaminoglycan attachment sites, and exon 7 contains the potential N-glycosylation sites and the stop codon. We have identified two polymorphic variations, an insertion/deletion composed of 19 nucleotides in intron 1 and a tetranucleotide (TATT)n repeat in intron 5. Analysis of 1.6 kb of upstream promoter sequence of DSPG3 reveals three TATA boxes, one of which is 20 nucleotides before the transcription start site. The transcription start site precedes the translation start site by 98 nucleotides. There are 14 potential binding sites for SOX9, a transcription factor present in cartilage, in the promoter, and in the first intron of DSPG3. We have examined the evolution of the SLRP gene family and found that gene products clustered together in the evolutionary tree are encoded by genes with similarities in genomic structure. Hence, it appears that the majority of the introns in the SLRP genes were inserted after the differentiation of the SLRP genes from an ancestral gene that was most likely composed of 2-3 exons.

Amino Acid Sequence↗

An extended RNA code and its relationship to the standard genetic code: an algebraic and geometrical approach.

An algebraic and geometrical approach is used to describe the primaeval RNA code and a proposed Extended RNA code. The former consists of all codons of the type RNY, where R means purines, Y pyrimidines, and N any of them. The latter comprises the 16 codons of the type RNY plus codons obtained by considering the RNA code but in the second (NYR type), and the third, (YRN type) reading frames. In each of these reading frames, there are 16 triplets that altogether complete a set of 48 triplets, which specify 17 out of the 20 amino acids, including AUG, the start codon, and the three known stop codons. The other 16 codons, do not pertain to the Extended RNA code and, constitute the union of the triplets YYY and RRR that we define as the RNA-less code. The codons in each of the three subsets of the Extended RNA code are represented by a four-dimensional hypercube and the set of codons of the RNA-less code is portrayed as a four-dimensional hyperprism. Remarkably, the union of these four symmetrical pairwise disjoint sets comprises precisely the already known six-dimensional hypercube of the Standard Genetic Code (SGC) of 64 triplets. These results suggest a plausible evolutionary path from which the primaeval RNA code could have originated the SGC, via the Extended RNA code plus the RNA-less code. We argue that the life forms that probably obeyed the Extended RNA code were intermediate between the ribo-organisms of the RNA World and the last common ancestor (LCA) of the Prokaryotes, Archaea, and Eucarya, that is, the cenancestor. A general encoding function, E, which maps each codon to its corresponding amino acid or the stop signal is also derived. In 45 out of the 64 cases, this function takes the form of a linear transformation F, which projects the whole six-dimensional hypercube onto a four-dimensional hyperface conformed by all triplets that end in cytosine. In the remaining 19 cases the function E adopts the form of an affine transformation, i.e., the composition of F with a particular translation. Graphical representations of the four local encoding functions and E, are illustrated and discussed. For every amino acid and for the stop signal, a single triplet, among those that specify it, is selected as a canonical representative. From this mapping a graphical representation of the 20 amino acids and the stop signal is also derived. We conclude that the general encoding function E represents the SGC itself.

Amino Acids↗

Isolation and functional characterization of the human gene encoding the myeloid zinc finger protein MZF-1.

The expression of the human myeloid zinc finger gene (MZF-1) by human bone marrow cells is necessary for granulopoiesis. We have analyzed the structure and function of the MZF-1 gene by diagnostic polymerase chain reaction, genomic cloning, and promoter analysis. Comparison of human promyelocytic HL-60 cell cDNA with isolated MZF-1 genomic clones indicated that the human MZF-1 gene is without introns and spans approximately 3 kb. Restriction enzyme mapping and Southern analysis indicated further that the human MZF-1 gene is a single-copy gene. Primer extension studies identified the major transcription start site as a thymidine residue located 1102 bp upstream of the ATG translation start codon. A putative TATA box sequence (TAAAAA) was found at -66 bp and a CCAAT box at -130 bp relative to the transcription initiation site. In HL-60 cells, MZF-1 mRNA levels are increased by granulopoietic inducers including retinoic acid and GM-CSF. DNA upstream of the transcription start site contains tandem-repeated consensus retinoic acid response elements at -666 through -696 bp and paired putative GM-CSF-responsive sequences centered at -50 and -100 bp. CAT reporter gene constructs containing these DNA regions promoted transcription and conferred transcriptional responsiveness to retinoic acid and GM-CSF when transfected into HL-60 cells. Additional putative regulatory binding sites included conserved MZF-1 zinc finger binding sequences, the importance of which was suggested by the enhanced expression of the endogenous MZF-1 gene following vector-driven expression of MZF-1 constructs in K562 myeloblastic leukemia cells. These findings provide a clearer basis for understanding the role of MZF-1 gene expression in myeloid cell growth and differentiation.

Base Sequence↗

IS900 targets translation initiation signals in Mycobacterium avium subsp. paratuberculosis to facilitate expression of its hed gene.

The Mycobacterium avium subsp. paratuberculosis (formerly Mycobacterium paratuberculosis) atypical insertion sequence, IS900, encodes a novel gene on the complementary strand to the putative transposase, p43. This gene requires a promoter, ribosome binding site (RBS) and termination codon to be acquired upon insertion into the M. avium subsp. paratuberculosis genome and hence is designated the hed (host expression-dependent) gene of IS900. Analysis of IS900 insertion sites suggests that this element targets translation initiation signals in M. avium subsp. paratuberculosis, specifically inserting between the RBS and start codon of a putative gene sequence. This aligns the hed initiation codon adjacent to a functional RBS and possibly downstream of an active promoter, driving expression of Hed protein. We have confirmed this unique targeting process by detecting expression of hed in M. avium subsp. paratuberculosis at the level of transcription by reverse transcription-PCR. Further, two Hed-specific antibodies detected Hed translation products in Western blots of protein extracts from M. avium subsp. paratuberculosis. A recombinant form of Hed expressed and purified from Escherichia coli will facilitate studies of IS900 transposition and will also be assessed as a diagnostic antigen for M. avium subsp. paratuberculosis disease. Implications of IS900 insertion in M. avium subsp. paratuberculosis pathogenicity are discussed.

Antibodies, Bacterial↗

Massive overproduction of dihydrofolate reductase in bacteria as a response to the use of trimethoprim.

Among several observations of greatly increased levels of chromosomal dihydrofolate reductase as a cause of resistance to high concentrations of the antifolate drug trimethoprim, in clinically isolated bacteria, one is described here of a strain of Escherichia coli overproducing dihydrofolate reductase several hundredfold. The chromosomally located resistance gene of this strain was isolated, inserted into a plasmid vector, and analyzed for its nucleotide sequence. The structural gene for the overproduced dihydrofolate reductase was found to be identical to that of E. coli K12, with nine exceptions, of which seven resulted in synonymous codon usage. Two transversions resulted in a substitution of Gly or Trp at amino acid position 30, and of Gln for Glu at position 154. Six of the nine base changes resulted in codons more frequently used. The Gly substitution which leads to a less commonly used codon, was thought to relate to the observed threefold increase in Ki for trimethoprim. Furthermore, a C----T transition was found in the -35 region of the promoter, increasing its homology with the E. coli consensus promoter sequence. In the ribosome-binding area of the resistant strain, finally, seven base changes were observed, two of which resulted in a five-base sequence of complementarity with the 3'-end of ribosomal 16S RNA. The distance between the -10 site of the promoter and the start codon for translation was finally increased one base pair by the insertion of an A at position +9 in the resistant strain. These genetic changes towards more efficient transcriptional and translational start sequences and towards increased mRNA expressivity are interpreted to reflect an evolutionary adaptation to the presence of antifolates.

Base Sequence↗

Characterization by cDNA cloning of the mRNA of a new growth factor from bovine seminal plasma: acidic seminal fluid protein.

A cDNA expression library in lambda gt11 prepared from cDNA derived of seminal vesicle tissue was screened by means of monospecific rabbit anti-aSFP IgG. The sequence of clone pTF21, containing an insert of 668 bp comprised an open reading frame from position 7 to 411 terminated by two stop codons. From this sequence a protein of 134 amino acid residues can be deduced. The mature aSFP was preceded by a signal peptide of 20 amino acids length. The protein sequence contains no signal for N-glycosylation. The molecular weight calculated from the amino acid sequence is 12922 Da. The start codon ATG is part of the sequence AAGATGA which fulfills the criteria of an initiation consensus sequence. The coding region was followed by 257bp of the complete 3'-untranslated region (3'UTR). A putative polyadenylation signal AATAAT, although not of the standard type, is observed at position 650. According to Northern analysis, aSFP mRNA is expressed in seminal vesicle tissue, ampulla and weakly in tissue of epididymis, but not in testis or other bovine tissue. aSFP is specified by a single copy gene. Attempts to detect homologies to known protein sequences were not successful.

Amino Acid Sequence↗