PubMed HealthSearch

SEARCH · PubMed Health

Results for “Conserved Sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Intermolecular base-paired interaction between complementary sequences present near the 3' ends of 5S rRNA and 18S (16S) rRNA might be involved in the reversible association of ribosomal subunits.

Highly conserved sequences present at an identical position near the 3' ends of eukaryotic and prokaryotic 5S rRNAs are complementary to the 5' strand of the m2(6)A hairpin structure near the 3' ends of 18S rRNA and 16S rRNA, respectively. The extent of base-pairing and the calculated stabilities of the hybrids that can be constructed between 5S rRNAs and the small ribosomal subunit RNAs are greater than most, if not all, RNA-RNA interactions that have been implicated in protein synthesis. The existence of complementary sequences in 5S rRNA and small ribosomal subunit RNA, along with the previous observation that there is very efficient and selective hybridization in vitro between 5S and 18S rRNA, suggests that base-pairing between 5S rRNA in the large ribosomal subunit and 18S (16S) rRNA in the small ribosomal subunit might be involved in the reversible association of ribosomal subunits. Structural and functional evidence supporting this hypothesis is discussed.

Base Sequence

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Sequence homology adjacent to the 3' terminal poly(A) of cowpea mosaic virus RNAs.

We have determined the sequence of 80 nucleotides adjacent to 3' poly(A) of both (middle and bottom component) cowpea mosaic virus RNAs, using dideoxynucleotide termination of reverse transcription. Sequence conservation is indicated, there being about 80% homology between the first 65 bases of each RNA. Although both RNAs are polyadenylated, there is no AAUAAA sequence as associated with most known polyadenylated mRNAs of eukaryotes and viruses. However both RNAs are U and A-rich in this region.

Avian Myeloblastosis Virus

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article

Amplification in the leader sequence of late polyoma virus mRNAs.

Ribonuclease T1 fingerprints of the three "late" polyoma virus mRNAs show that oligonucleotides of the leader sequence are present in multiple copies in each mRNA. These oligonucleotides, however, appear unimolar in fingerprints of complete, continuous transcripts of the late strand of the viral DNA. Oligonucleotides which are represented only once in the DNA are thus reiterated in the mature mRNAs. Consequently, when mRNA was hybridized to the leader region of immobilized viral DNA, those copies present in excess of their genomic representation failed to hybridize and were released by RNAase treatment. Analysis of the RNAase-resistant hybrids revealed a series of leader species with complex sequence arrangements. We suggest that these complicated reiterated sequences are generated during the processing of a precursor RNA which extends several times around the genome. This RNA would be shortened by a series of splicing reactions which conserve sequences from the leader region and attach them to a suitable coding sequence.

Animals

Comparative and Subtractive Genomics Analysis of Multidrug-Resistant Klebsiella pneumoniae Strains for Novel Target Identification and Drug Repurposing Strategies.

The rapid rise of multidrug-resistant (MDR) Klebsiella pneumoniae has created a major global health challenge due to the limited availability of conserved therapeutic targets effective across diverse resistant strains. In this study, an integrative computational target-discovery and drug-repurposing framework was applied to six clinically relevant K. pneumoniae strains. Comparative genomic analysis identified 3012 conserved genes, which were subsequently filtered to nine essential, non-host homologous proteins. Among these, three conserved cytoplasmic proteins (accD, cpxR, and mraZ) were prioritized for functional analysis, with acetyl-CoA carboxylase subunit beta (accD) emerging as the most promising therapeutic target based on sequence conservation, predicted essentiality, subcellular localization, and pathway association. Structural assessment supported the reliability of the predicted accD model, whereas consensus binding-site analysis identified key residues suitable for ligand interaction. Virtual screening of FDA-approved drugs followed by molecular docking identified several compounds with favorable binding profiles toward accD. Subsequent molecular dynamics simulations, including root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), hydrogen-bond occupancy, principal component analysis (PCA), and PCA-based free energy landscape (FEL) analyses, consistently identified tenapanor, micafungin, deferoxamine, and cobicistat as the most stable protein-ligand complexes, with tenapanor exhibiting the most favorable overall structural and thermodynamic stability profile. These findings identify accD as a promising therapeutic target in MDR K. pneumoniae and suggest several FDA-approved compounds as potential candidates for drug repurposing. Although experimental validation is needed to confirm their biological activity and therapeutic potential, this study demonstrates the potential of integrating comparative genomics with molecular dynamics analyses to support antimicrobial target identification and drug repurposing against MDR bacterial pathogens.

Klebsiella pneumoniae

Transcription of Ti plasmid-derived sequences in three octopine-type crown gall tumor lines.

Total RNA isolated from three octopine-type crown gall lines contains sequences homologous to specific regions of the tumor-inducing (Ti) plasmid of Agrobacterium tumefaciens strain 15955. A comparison of transcripts in these three tumor lines suggests that tumor cells transcribe various sequences within a sector of plasmid DNA of 13 x 10(6) daltons and that transcription may not be uniform across the plasmid derived sequences (T-DNA). Transcription of T-DNA by octopine-type tumors occurs at four major sites. The levels of transcription occurring at three of these sites appear to vary considerably among the three tumor lines investigated. Part of this variability may reflect differences in the organization and copy number of T-DNA. One of the transcription sites maps within a region of DNA with common sequence homology with all Ti plasmids. Varying amounts of transcript homologous to this region of T-DNA are present in all three tumor lines. It is suggested that transcription of these conserved sequences in the plant may have significance regarding the mechanism of tumorigenesis.

Arginine

Nucleotide sequence of the BK virus DNA segment encoding small t antigen.

The nucleotide sequence from 0.64 to 0.53 map units in the BK virus genome coding for the small t protein has been determined. There is only one open reading frame that can code for a polypeptide of 172 amino acids, the putative small t protein. Beyond this segment, multiple termination codons are present in all three reading frames. There is considerable nucleotide and amino acid sequence homology between this region of BK virus and the analogous region of simian virus 40, especially in the proximal portion from 0.64 to 0.60 map units which is most likely common to the small t and large T BK virus proteins. A comparison of the conserved sequences within the early papovavirus genes both confirms the evolutionary relationship between these viruses and suggests the amino acid composition of the regions required for T antigen functions.

Antigens, Neoplasm

The emergence and diversification of the DUX gene family across placental mammals.

The DUX gene family encodes transcription factors with paired homeodomains. It has critical roles in embryogenesis and disease, including facioscapulohumeral muscular dystrophy (FSHD) and cancer. This study conducts a comparative analysis of the DUX gene family-DUXA, DUXB (including DUXBL), and DUXC (including DUX4 and Dux)-across placental mammals, highlighting their structural diversity within macrosatellite repeat contexts. Using long-read genomes, we explore gene distribution, array patterns, and phylogenetic relationships in various vertebrate species. Our analysis reveals that DUXA and DUXB are highly conserved, with intriguing variations such as intronless forms likely arising from ancestral retrotransposition events. While DUXBL is inconsistently retained across clades, its locus-which in non-placental mammals harbors the ancestral single-homeodomain sDUX gene-served as an evolutionary hub for diversification, giving rise to DUXA, DUXB and DUXC, as well as macrosatellite tandem array structures. Sequence conservation and syntenic analyses demonstrate array adaptability, exemplified by higher-order repeats in orangutans and disrupted patterns of concerted evolution in elephants. Furthermore, analysis of human pseudo-DUX4 arrays indicates their potential role in disease mechanisms, including as possible contributors to rare cases of FSHD, warranting further investigation. This study thus provides insights into DUX-family gene evolution, offering a foundation for future research into developmental roles and disease implications.

Animals

Pseudouridylation of yeast ribosomal precursor RNA.

The pseudouridylation of ribosomal RNA of Saccharomyces carlsbergensis was investigated with respect to its timing during the maturation of rRNA and its sequence specificity. Analysis of 37-S RNA, the common precursor to 17-S, 5.8-S and 26-S rRNA and most probably the primary ribosomal transcript, shows that this RNA molecule contains already most if not all of the 36-37 pseudouridine residues found in the mature rRNAs. Thus pseudouridylation is, like 2'-0-ribosemethylation, an early event in the maturation of rRNA, taking place immediately after, or even during, transcription. The data presented show that the non-conserved sequences of 37-S precursor rRNA contain very few pseudouridine residues if any. The pseudouridine residues within the rRNA sequences are apparently clustered to a certain degree as can inferred from the occurrence of a single oligonucleotide containing 3 pseudouridines, which was obtained by digestion of 26-S rRNA with ribonuclease T1.

Base Sequence

[Genetic analysis of a male with Multiple morphological abnormalities of sperm flagella combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene and a literature review].

OBJECTIVE: To explore the clinical phenotype and genetic etiology of a male with Multiple morphological abnormalities of sperm flagella (MMAF) combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene, with an aim to provide guidance for assisted reproductive technology in his family. METHODS: A man with MMAF combined with sperm head abnormalities who visited Women and Children's Hospital of Ningbo University in October 2024 was selected as study subject. Clinical data of the patient's family were retrospectively collected. Peripheral blood samples were collected from the patient and his spouse, and G-banding karyotyping and whole exome sequencing (WES) were carried out. Candidate variants were validated by Sanger sequencing. Conservation of the DNAH1 protein was queried on the UCSC website. The difference between wild type and variant DNAH1 proteins were analyzed using AlphaFold v3.0.1 and PyMOL v2.5.6. The pathogenicity of variant was rated based on the guidelines from American College of Medical Genetics and Genomics (ACMG). Previous literature was searched using keywords "DNAH1 gene" and "multiple morphological abnormalities of the sperm flagella" on CNKI, Wanfang Data Knowledge Service Platform, and PubMed database to identify cases of MMAF attributed to biallelic DNAH1 gene variants. The retrieval period was set from the establishment of the databases to December 31, 2025. The genotypes and clinical phenotypes of patients with biallelic DNAH1 mutations were analyzed. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: EC2023-094). RESULTS: The 30-year-old patient and his 30-year-old wife had infertility for 2 years. Semen analysis revealed no motile sperm and a 99.0% abnormal morphology rate. Typical MMAF was observed with phase-contrast microscopy. Sperm morphology analysis revealed abnormalities of the head, neck, and tail with an approximate ratio of 9:5:1. The patient's karyotype was 46,XY, and his wife's karyotype was 45,X[4]/47,XXX[1]/46,XX[84]. WES and Sanger sequencing revealed that the patient harbored compound heterozygous variants of the DNAH1 gene, namely c.1435_1444+3del and c.12204_12206del (p.Asn4069del), but their origin remained unidentified. UCSC genome browser query results showed that the amino acid residue at position 4 069 of the DNAH1 protein is highly conserved across various species. Protein structure prediction reveals that, in the wild-type DNAH1 protein, the Asparagine at position 4 069 (Asn4069) can form hydrogen bonds with the Leucine on the main chain at position 4 086 (Leu4086) and the Serine on the side chain at position 4 087 (Ser4087). The c.12204_12206del variant, resulting in deletion of Asn4069, disrupts these hydrogen bonds and does not generate any compensatory interactions. Based on the ACMG guidelines, the c.1435_1444+3del variant was predicted to be likely pathogenic (PM2_Supporting+PVS1), and the c.12204_12206del(p.Asn4069del) variant was rated as likely pathogenic (PM2_Supporting+PM4+PM3+PP4). The couple had elected for in vitro fertilization using donor sperm. During this cycle, 12 oocytes were retrieved, 10 oocytes were successfully fertilized, 1 embryo and 6 blastocysts were obtained. Following the first transfer of a frozen-thawed blastocyst, implantation of an empty gestational sac occurred, which led to a miscarriage. After the second transfer of a high-quality blastocyst, the embryo split into twins following implantation, and the spouse had selected fetal reduction. The gestational age was 33+3 weeks on June 1, 2026. Literature review identified three studies reporting biallelic mutations of the DNAH1 gene in association with MMAF combined with sperm head abnormalities. Together with the patient from this study, a total of 20 patients were included in the analysis. The rate of sperm flagellar abnormalities in these patients was above 80.0%, while the rate of sperm head abnormalities has ranged from 12.0% to 100.0%. In four patients, the genetic basis was unknown. In the remaining 16 patients, 35 mutations were detected, with c.8626-1G>A being the most common (22.9%, 8/35). CONCLUSION: This patient showed MMAF with frequent sperm head defects. Compound heterozygous variants of the DNAH1 gene probably underlay these abnormalities, which in turn has led to his primary infertility. This study revealed the phenotypic variability of MMAF and broadened the mutational spectrum of the DNAH1 gene.

Humans

Semisynthetic horse heart [65-homoserine]cytochrome c from three fragments.

Horse heart cytochrome c was treated with methylsulfonylethyloxycarbonyl succinimide (Msc-ONSu) to give fully N(epsilon)-protected cytochrome c. Treatment of this derivative with a hard base for 15 sec regenerated the native tetrahectapeptide chain. CNBr degradation of the protected compound produced three fragments bearing only protective Msc functions on epsilon-amino groups. The fragment comprising the sequence 81-104 was isolated from the mixture and acylated with N-hydroxysuccinimidyl-t-butyloxycarbonyl-L-methioninate. The resulting pentacosapeptide derivative was partially deprotected by treatment with acid and condensed in good yield (65%) with fully synthetic N(alpha66), N(epsilon72,73,79)- tetra-Msc-cytochrome-c-(66-79)-tetradecapeptide azide. This pathway is preferred because the pentadecapeptide azide derivative 66-80 acylated the N(epsilon)-protected tetracosapeptide sequence 81-104 in an unpredictable manner. Subsequent treatment of the product with a base produced unprotected semisynthetic cytochrome-c-(66-104)-nonatriacontapeptide, which is known to undergo acylation by unprotected [Hse(65)]cytochrome-c-(1-65)-pentahexacontapeptide lactone. The high specificity of this condensation is ascribed to "conformation direction." Semisynthetic [Hse(65)]cytochrome c thus prepared reacts like native cytochrome c with a succinate cytochrome c reductase preparation and with cytochrome c oxidase (ferrocytochrome c:oxygen oxidoreductase, EC 1.9.3.1). This semisynthetic strategy may provide a rapid route for the production of cytochrome c analogs modified in the highly conservative sequence 66-80.

Amino Acid Sequence

Fine structure of ribosomal RNA. II. Distribution of methylated sequences within Xenopus laevis rRNA.

The distribution of methyl groups in rRNA from Xenopus laevis was analyzed by hybridization of rRNA to subfragments of either of two cloned rDNA fragments, X1r11 and X1r12, which together constitute a complete rDNA repeat unit. Using a mixture of 3H-methyl plus 32P-labelled rRNA as probe, the molar yield of methyl groups per rRNA region in hybrid could be calculated. For this calculation the length of the rRNA coding region in each DNA subfragment is needed, which was determined for X1r11 subfragments by the nuclease S1 mapping method of Berk and Sharp. The results show that both in 18S and 28S rRNA the methyl groups are nonrandomly distributed. For 18S rRNA, clustering was found within a 3' terminal fragment of 310 nucleotides. For 28S rRNA, clustering of methyl groups was found within a region of 750 nucleotides in length, which ends 500 nucleotides from the 3' end. In contrast, the 28S rRNA 5' terminal region of 900 nucleotides is clearly undermethylated. The general position of methyl groups in 28S rRNA correlates with the location of evolutionarily conserved sequences in this molecule, as recently determined in our laboratory.

Animals

The auxin gatekeepers: Evolution and diversification of the YUCCA family.

The critically important YUCCA (YUC) gene family is highly conserved and specific to the plant kingdom, primarily responsible for the final and rate-limiting step for indole-3-acetic acid (IAA) biosynthesis. IAA is an essential phytohormone, involved in virtually all aspects of plant growth and development. In addition, IAA is involved in fine-tuning plant responses to biotic and abiotic interactions and stresses. While the YUC gene family has significantly expanded throughout the plant kingdom, a detailed analysis of the evolutionary patterns driving this diversification has not been performed. Here, we present a comprehensive phylogenetic analysis of the YUC family, combining YUCs from species representing key evolutionary plant lineages. The evolutionary history of YUCs is complex and suggests multiple recruitment events via horizontal gene transfer from bacteria. We identify and hierarchically classify the YUC family into an early diverging grade, five distinct classes and 41 subclasses. Angiosperm YUC diversity and expansion are explained in the context of protein sequence conservation, as well as spatial and gene expression patterns. The presented YUC gene landscape offers new perspectives on the distribution and evolutionary trends of this crucial family, which facilitates further YUC characterization within plant development and response to environmental change.

Indoleacetic Acids

Comparative analysis of conserved non-coding elements identifies gene regulatory networks rewired during the water-to-land transition in vertebrates.

The conquest of land by vertebrates has been a pivotal moment in evolutionary history. Adapting to the new habitats necessitated numerous changes in vertebrate anatomy and physiology, creating an enduring imprint on the developmental gene regulatory networks (GRNs) of tetrapods. The increase of high-quality genomic resources over the past decade has made it possible to study the genomic legacy of the water-to-land transition. While much attention has been given to the highly conserved non-coding elements (CNEs) of the genome that share high levels of similarity across evolutionarily diverged clades, recent evidence suggests that perhaps comparable attention should be given to "missing" CNE-s, conserved sequence patches present in extant stem gnathostomes and actinopterygian fishes that have become undetectable in tetrapods during the adaptation to terrestrial life, whether through true sequence loss or divergence beyond alignability. These sequences could help us reveal the relaxation of certain developmental constraints, related to the aquatic lifestyle, that made reaching new adaptive peaks in the developmental landscape possible. In this paper, we search for such CNEs and characterize them in comparison with pan-Gnathostome CNEs, using the zebrafish (Danio rerio) genome as a reference. Our results suggest that the rewiring of developmental networks related to pigmentation and muscle structure formation has left the largest genomic imprint. We also find that components of canonical Wnt and Hedgehog signalling, are enriched among CNEs retained in fish.

cis-regulatory evolution

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8 billion predictions for 1.4 million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Genomic evolution of EGF-CFC genes in deuterostomes.

BACKGROUND: EGF-CFC proteins are a bilaterian innovation, but they are best known for their roles in Nodal signaling during gastrulation and left-right patterning in vertebrates. Species with multiple family members show evidence of functional specialization. For example, in mouse, Cripto is required for gastrulation, whereas CFC1 is involved in left-right patterning. However, members of the EGF-CFC family across model organisms exhibit limited sequence conservation beyond the EGF-CFC domain, posing challenges for determining their evolutionary history and functional conservation. RESULTS: In this study, we describe the evolutionary history of the EGF-CFC family of proteins across several branches of deuterostomes, with a particular focus on vertebrates. We trace the EGF-CFC gene family from a single gene in the deuterostome ancestor through its expansion and functional specialization in tetrapods, and subsequent gene loss and translocation in eutherian mammals. Mouse Cripto and CFC1, zebrafish Tdgf1, and each Xenopus EGF-CFC gene (Tdgf1, Tdgf1.2 and Cripto.3) are all descendants of the ancestral deuterostome Tdgf1 gene. CONCLUSIONS: We propose that subsequent to EGF-CFC family expansion in tetrapods, Tdgf1B (Xenopus Tdgf1.2) acquired specialization in the left-right patterning cascade, and then after its translocation in eutherians to a different chromosomal location, CFC1 has maintained that specialization.

Animals

Molecular structure and flanking nucleotide sequences of the natural chicken ovomucoid gene.

Five independent clones containing the natural chicken ovomucoid gene have been isolated from a chicken gene library. One of these clones, CL21, contains the complete ovomucoid gene and includes more than 3 kb of DNA sequences flanking both termini of the gene. Restriction endonuclease mapping, electron microscopy and direct DNA sequencing analyses of this clone have revealed that the ovomucoid gene is 5.6 kb long and codes for a messenger RNA of 821 nucleotides. The structural gene sequence coding Ifor the mature messenger RNA is split into at least eight segments by a minimum of seven intervening sequences of various sizes. The shortest structural gene segment is only 20 nucleotides long. All seven intervening sequences are located within the peptide coding region of the gene, and the sequences at the 5' and 3' untranslated regions of the mRNA are not interrupted by intervening sequences. The DNA sequences of the regions flanking the 5' and 3' termini of the gene have been determined. Thirty nucleotides before the start of the messenger RNA coding sequence is the heptanucleotide TATATAT, which is also present in a similar location relative to the chicken ovalbumin gene and other unique sequence eucaryotic genes. This sequence resembles that of the Pribnow box in procaryotic genes where a promoter function has been implicated. Seven nucleotides past the 3' end of the gene is the tetranucleotide TTGT, a sequence found to be present at identical locations as either TTTT or TTGT in other eucaryotic genes that have been sequenced. These conserved DNA sequences flanking eucaryotic genes may serve some regulator function in the expression of these genes.

Animals