PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Complete genome characterization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Sequence determination of the Crimean-Congo hemorrhagic fever virus L segment.

Crimean-Congo hemorrhagic fever (CCHF) virus is highly pathogenic for humans and remains the only Category A virus for which full sequence information is currently unavailable. In this study we completed CCHF genome characterization by determining the L segment sequence using Dugbe and CCHF virus-specific oligonucleotides. Sequence alignments revealed the presence of four previously described conserved regions in all Bunyaviridae polymerases. Interestingly, additional regions containing putative Ovarian Tumor (OTU)-like cysteine protease and helicase domains were identified in the L segments of CCHF and Dugbe viruses, suggesting an autoproteolytic cleavage process for nairovirus L proteins.

Amino Acid Sequence↗

Molecular characterization of a new gene, CEAL1, encoding for a carcinoembryonic antigen-like protein with a highly conserved domain of eukaryotic translation initiation factors.

Carcinoembryonic antigen (CEA) is a complex immunoreactive glycoprotein belonging to a large and heterogeneous group of cross-reacting proteins known as the CEA gene family, which contains 29 genes/pseudogenes. CEA is used as a valuable serum tumor marker for monitoring response to therapy in patients with various solid tumors. Through the positional cloning approach we have identified and characterized a CEA-like gene (CEAL1), a novel member of the CEA multigene family. We have characterized the complete genomic structure of CEAL1, as well as one alternative splice variant and determined its chromosomal localization. The new gene is comprised of eight exons, with seven intervening introns and it is localized to chromosome 19q13.2 between the markers D19S574 and D19S219, approximately 60 kb upstream of the BCL3 gene. The protein-coding region of the gene is formed of 903 bp, encoding for a 300-amino-acid polypeptide with a predicted molecular weight of 32.6 kDa and isoelectric point of 5.74. The CEAL1 protein contains two Immunoglobulin-like (Ig-like) transmembrane domains, which are present in most of the CEA proteins, as well as one highly conserved domain of eukaryotic translation initiation factors. The identified alternative spliced variant has one more exon of 134 bp. This splice variant is expected to encode for a truncated protein of 142 amino acids with the eIF5A domain and without Ig homology domain. CEAL1 mRNA is expressed in a variety of tissues, but highest levels are found in the prostate, uterus, fetal brain, mammary, adrenal gland, skeletal muscle, small intestine and kidney. CEAL1 is highly expressed in BT-474, BT20, T47D and, at much lower levels, in MCF7 breast cancer cell lines. The new gene is also highly expressed in the LNCaP prostate cancer cell line. The CEAL1 gene was found to be down-regulated by dexamethasone in BT-474 breast cancer cell lines. Our data suggest that this gene is overexpressed in a subset of ovarian cancers which are clinically more aggressive.

Alternative Splicing↗

New mammalian selenocysteine-containing proteins identified with an algorithm that searches for selenocysteine insertion sequence elements.

Mammalian selenium-containing proteins identified thus far contain selenium in the form of a selenocysteine residue encoded by UGA. These proteins lack common amino acid sequence motifs, but 3'-untranslated regions of selenoprotein genes contain a common stem-loop structure, selenocysteine insertion sequence (SECIS) element, that is necessary for decoding UGA as selenocysteine rather than a stop signal. We describe here a computer program, SECISearch, that identifies mammalian selenoprotein genes by recognizing SECIS elements on the basis of their primary and secondary structures and free energy requirements. When SECISearch was applied to search human dbEST, two new mammalian selenoproteins, designated SelT and SelR, were identified. We determined their cDNA sequences and expressed them in a monkey cell line as fusion proteins with a green fluorescent protein. Incorporation of selenium into new proteins was confirmed by metabolic labeling with (75)Se, and expression of SelT was additionally documented in immunoblot assays. SelT and SelR did not have homology to previously characterized proteins, but their putative homologs were detected in various organisms. SelR homologs were present in every organism characterized by complete genome sequencing. The data suggest applicability of SECISearch for identification of new selenoprotein genes in nucleotide data bases.

3' Untranslated Regions↗

Codetection of a mixed population of candHPV62 containing wild-type and disrupted E1 open-reading frame in a 45-year-old woman with normal cytology.

We have cloned, sequenced, and characterized the complete genome of a novel human papillomavirus (HPV), candHPV62. During cloning, 2 candHPV62 viral isolates were recovered from a single cervical sample; 1 had all anticipated HPV open-reading frames (ORFs) intact, whereas the other exhibited an E1 frame-shift mutation. Further experiments indicated that the 2 strains were equivalent in abundance. It appears that an early mutation occurred within the E1 ORF, which was transcomplemented by an intact E1 protein. A search of the HPV database identified disruption of the E1 ORF in the cloned reference isolates of HPV16, HPV53, HPV56, and HPV72. These data suggest that disruption of the E1 ORF in genital HPVs is not uncommon.

Base Sequence↗

Isolation and characterization of the complete complementary and genomic DNA sequences of human serum amyloid P component.

Complementary and genomic DNA clones corresponding to the human serum amyloid P component (SAP) mRNA have been isolated and analyzed. The nucleotide sequences of the cDNA and the corresponding regions of the genomic SAP DNA reported here were identical, and revealed that after coding for a signal peptide of 19 amino acids and the first two amino acids of the mature SAP protein, there is one small intron of 115-base pairs (bp), followed by a nucleotide sequence coding for the remaining 202 amino acid residues. The SAP gene has an ATATAAA sequence 29-bp upstream from the cap site, but there is no CAAT box-like sequence. A possible polyadenylation signal sequence, ATTAAA, was found to be located 28-bp upstream from the polyadenylation site. A comparison of the genomic SAP DNA sequence with that of human C-reactive protein (CRP) revealed a striking overall homology which was not uniform: several highly conserved regions were bounded by non-homologous regions. This comparison provides further support for the hypothesis that SAP and CRP are products of a gene duplication event.

Amino Acid Sequence↗

Felis domesticus papillomavirus, isolated from a skin lesion, is related to canine oral papillomavirus and contains a 1.3 kb non-coding region between the E2 and L2 open reading frames.

We have characterized the complete genome (8300 bp) of an isolate of Felis domesticus papillomavirus (FdPV) from a domestic cat with cutaneous papillomatosis. A BLAST homology search using the nucleotide sequence of the L1 open reading frame demonstrated that the FdPV genome was most closely related to canine oral papillomavirus (COPV). A 384 bp non-coding region (NCR) was found between the end of L1 and the beginning of E6, and a 1.3 kbp NCR was located between the end of E2 and the beginning of L2. Phylogenetic analysis placed FdPV in the E3 clade with COPV. Both viruses contain the atypical second NCR, which has no homology with sequences in existing databases.

Animals↗

The Conjugative Megaplasmid pMD9A Mediates Transferring Antibiotic Resistance Genes.

Pseudomonas asiaticais an emerging opportunistic pathogen with a broad host range. Current evidence suggests that some isolates exhibit multidrug resistance, which may complicate treatment. In this study, a multidrug-resistant P. asiatica strain MD9 was isolated from aquaculture water. We aimed to characterize its complete genome sequence and investigate the role of its conjugative megaplasmid pMD9A in the horizontal transfer of antibiotic resistance genes. The genome of MD9 consists of one circular chromosome (5,956,782 bp, with a G + C content of 62.5%) and one circular megaplasmid, pMD9A (455,169 bp, with a G + C content of 56.5%). Genome annotation identified 65 antibiotic resistance genes and 148 putative virulence factor-encoding genes in the MD9 genome. The megaplasmid pMD9A carries 29 antibiotic resistance genes conferring resistance to β-lactams, chloramphenicol/florfenicol, aminoglycosides, and macrolides. A class 1 integron (intI1) and multiple autonomous conjugative transfer elements were identified in pMD9A. Conjugation experiments demonstrated that the β-lactam resistance gene blaOXA-246 could be horizontally transferred from the donor MD9 strain to the recipient Escherichia coli 25DN strain. The megaplasmid pMD9A not only carries a broad array of antibiotic resistance genes, but also facilitates their horizontal spread among environmental bacteria, thereby potentially contributing to the dissemination of multidrug-resistant bacteria.

Pseudomonas asiatica↗

Assembly and characterization of the first complete mitochondrial genome of Epimedium sagittatum (Sieb. et Zucc.) Maxim (Berberidaceae):an invaluable traditional Chinese medicine.

BACKGROUND: Epimedium sagittatum (Sieb. et Zucc.) Maxim is an invaluable traditional Chinese medicine plant known for its properties of tonifying kidney yang, strengthening bones and muscles, and dispelling rheumatism. The chloroplast (cp) genome of E. sagittatum have been sequenced, offering critical insights for breeding and phylogenetic research. However, the mitochondrial (mt) genome of E. sagittatum remains uncharacterized, limiting comprehensive insights into its genomic evolution. RESULTS: In this study, we assembled the first complete mt genome of E. sagittatum employing Illumina and Nanopore sequencing technology and subsequently investigated comparative analysis with its closely related species. The mt genome of E. sagittatum was assembled as a multi-branched structure with a length of 339,191 bp, within a GC content of 46.91%. Our annotation results have shown 39 protein-coding genes (PCGs), 22 tRNA genes, three rRNA genes and four pseudogenes in the E. sagittatum mt genome. The analysis of sequence repeats has detected 79 simple sequence repeats (SSRs), 10 tandem repeats and 255 dispersed repeats in the E. sagittatum mt genome. A total of 720 C to U RNA editing sites of the 34 PCGs was predicted in E. sagittatum. The codons exhibited a strong preference for A or U bases in the E. sagittatum mt genome. The analysis of nucleotide diversity (Pi) highlighted differences in genetic variability across the tested genes, with atp9 gene exhibiting the highest genetic variation. Selection pressure analysis showed that most genes were affected by negative selection during evolution, whereas ccmB, rps10, and rps12 underwent positive selection in different plants. Additionally, a Bayesian phylogenetic tree showed that E. sagittatum was closely related to E. wushanense and E. pubescens. In total of 14 homologous fragments totaling 8,954 bp were identified between the cp and mt genomes of E. sagittatum. CONCLUSIONS: This study presents the first assembled and annotated mt genome of E. sagittatum, which provides a valuable genetic resource for the Epimedium genus and lays the foundation for investigating the phylogenetic relationship and genetic variation of this invaluable medicinal plant.

Epimedium↗

Completion of molecular characterization of Toscana phlebovirus genome: nucleotide sequence, coding strategy of M genomic segment and its amino acid sequence comparison to other phleboviruses.

The M RNA segment of Toscana (TOS) phlebovirus was cloned and the complete nucleotide sequence determined. The M RNA segment is 4215 nucleotides in length, and it contains a single major open reading frame (ORF) in the viral-complementary sequence, between nucleotides 18 and 4034, which can encode for a polyprotein of 1339 amino acids (Mr 149 kDa). The viral segment is expressed via a unique mRNA containing 10-14 non-templated nucleotides at the 5' end and it is truncated at the 3' end by about 140 nucleotides in a purine-rich region. In M predicted amino acid sequences, several hydrophobic regions have been identified. They could function as a signal sequence or a transmembrane region for the different proteins. Comparison of the deduced amino acid sequence of M precursor product revealed 38, 36, and 25% identity and 58, 56, and 47% similarity with those of Rift Valley fever (RVF), Punta Toro (PT) and Unkuniemi (UUK) viruses, respectively. Residues conserved among the proteins are mainly located at the COOH-portion of the precursor, while the major divergence is in the NSm coding regions. Based on sequence comparison and similarity of hydropathic pattern of TOS M segment with other phleboviruses the N-termini of TOS GN and GC glycoproteins were placed at residues 297 and 936 of the precursor.

Amino Acid Sequence↗

Complete DNA sequence of the mitochondrial genome of the sea-slug, Aplysia californica: conservation of the gene order in Euthyneura.

We have sequenced and characterized the complete mitochondrial genome of the sea slug, Aplysia californica, an important model organism in experimental biology and a representative of Anaspidea (Opisthobranchia, Gastropoda). The mitochondrial genome of Aplysia is in the small end of the observed sizes of animal mitochondrial genomes (14,117 bp, NCBI Accession No. NC_005827). The Aplysia genome, like most other mitochondrial genomes, encodes genes for 2 ribosomal subunit RNAs (small and large rRNAs), 22 tRNAs, and 13 protein subunits (cytochrome c oxidase subunits 1-3, cytochrome b apoenzyme, ATP synthase subunits 6 and 8, and NADH dehydrogenase subunits 1-6 and 4L). The gene order is virtually identical between opisthobranchs and pulmonates, with the majority of differences arising from tRNA translocations. In contrast, the gene order from representatives of basal gastropods and other molluscan classes is significantly different from opisthobranchs and pulmonates. The Aplysia genome was compared to all other published molluscan mitochondrial genomes and phylogenetic analyses were carried out using a concatenated protein alignment. Phylogenetic analyses using maximum likelihood based analyses of the well aligned regions of the protein sequences support both monophyly of Euthyneura (a group including both the pulmonates and opisthobranchs) and Opisthobranchia (as a more derived group). The Aplysia mitochondrial genome sequenced here will serve as an important platform in both comparative and neurobiological studies using this model organism.

Animals↗

The complete Chloroplast Genome of Dianthus Helenae, an Endemic Species with Medicinal Potential from the Nuratau Mountains, Uzbekistan.

Dianthus helenae Vved. is an endemic medicinal species of the Nuratau Mountains, Uzbekistan, and its genomic resources have remained largely unavailable. In this study, we sequenced, assembled, and characterized the complete chloroplast genome of D. helenae and evaluated its phylogenetic position within Dianthus. The plastome exhibited a typical circular quadripartite structure with a total length of 149,567 bp, comprising a large single-copy (LSC) region of 82,856 bp, a small single-copy (SSC) region of 17,105 bp, and a pair of inverted repeats (IRs) of 24,803 bp each. The genome contained the typical set of chloroplast genes, including protein-coding genes, transfer RNAs, and ribosomal RNAs, with duplicated genes located in the IR regions. Phylogenetic analysis based on complete chloroplast genome sequences strongly supported the placement of D. helenae within Dianthus and recovered it as a distinct lineage relative to other sampled species. Sliding window analysis of nucleotide diversity revealed uneven sequence variation across the plastome, with higher variability in the SSC and LSC regions than in the IRs. Several highly variable loci, including trnK-UUU , rps16-trnQ-UUG , rpl32, ycf1, and ndh-associated regions, were identified as potential molecular markers. These results provide an important genomic resource for Dianthus and establish a foundation for future phylogenetic, taxonomic, conservation, and molecular identification studies of this endemic Central Asian species.

Genome, Chloroplast↗

Rat copper/zinc superoxide dismutase gene: isolation, characterization, and species comparison.

A 13 kb rat Cu/ZnSOD genomic clone has been purified from a rat liver genomic library and completely characterized by restriction mapping, detailed sequencing and Southern blot analysis. This gene spans approximately 6 kb and contains five exons and four introns. Comparison of rat, mouse, and human Cu/ZnSOD genes reveals a high conservation in genomic organization and exon-intron junctions, including an unusual 5'GC donor sequence at the first intron. The gene contains a TATA box as well as an inverted CCAAT box, a feature common to both the mouse and human genes. Furthermore, several repeats were identified in the 5' promoter region of this gene, and these regulatory elements are also strikingly conserved in these three species.

Amino Acid Sequence↗

Viral genome organizer: a system for analyzing complete viral genomes.

The viral genome organizer (VGO) is designed to simplify the characterization and annotation of complete viral genomes (particularly those of large poxviruses) and to help researchers discover new genes and detect gene fragmentation. VGO is based on Genotator [Harris, N.L., 1997. Genome Res. 7, 754-762], an annotation workbench designed for the analysis of eukaryotic genomic sequences. VGO automates a number of database search routines (FASTA, BLASTP, PSI-BLAST and TBLASTN), processes the results through a multiple-alignment viewer (MView; [Brown, N.P., Leroy, C., Sander, C. , 1998. Bioinformatics 14, 380-381]) and serves to manage the hundreds of DNA, protein and database search results files that must be organized when dealing with large complete poxviral genomes. It also directs the generation a self-dotplot of the genome by Dotter [Sonnhammer, E.L.L., Durbin, R., 1995. A dot-matrix program with dynamic threshold control suited for genomic DNA and protein sequence analysis. Gene 167: GC1-10. http://www.sanger.ac. uk/Software/Dotter/] to uncover repeated genes and sequences and provides Internet links to programs for generation of restriction maps and analysis of potential PCR primers. The user-friendly graphical interface displays DNA and protein sequences, links to search results, ORFs, stop-start codons, restriction sites and flags of database searches. Currently, VGO and associated programs run in an X-windows environment on commonly available UNIX machines.

Databases, Factual↗

Complete mitochondrial genomes of eight cyclophyllidean tapeworms: genome pattern and phylogenetic analysis.

Cyclophyllidean tapeworms are widespread parasites of significant medical and veterinary importance. However, mitochondrial (mt) genomic resources for cyclophyllideans from China, particularly those recovered from wildlife hosts, remain comparatively limited. In this study, we sequenced and characterized the complete mt genomes of eight cyclophyllidean isolates collected from diverse wild and domestic hosts in China, including two Hymenolepis sp. isolates and two Raillietina sp. isolates from China, and four additional isolates of previously sequenced Taenia species. The circular mt genomes ranged from 13,387 to 14,021 bp in length, encoding 36 typical genes with variable non-coding regions. Comparative analysis revealed highly conserved gene composition and mostly conserved mt architecture, with localized rearrangement patterns detected among the cyclophyllidean lineages examined. In particular, all sampled Taeniidae exhibited a consistent trnL1-trnS2 arrangement, whereas the examined non-Taeniidae families showed the trnS2-trnL1 arrangement, confirming and extending, across additional wildlife-associated isolates, a previously proposed family-associated gene-order marker within Cyclophyllidea. Phylogenetic analyses based on concatenated amino acid sequences of the 12 protein-coding genes placed the eight isolates within their expected families, in topologies broadly consistent with previous mitogenomic studies. These data provide additional Chinese mitogenomic references, especially for underrepresented wildlife-associated isolates, and support family-associated gene-order patterns in Cyclophyllidea.

Animals↗

Complete genome of a JC virus genotype type 6 from the brain of an African American with progressive multifocal leukoencephalopathy.

OBJECTIVES: The major genotypes of the human polyomavirus JC (JCV) include type 1 (European), type 2 (Asian), type 3 (African), and type 4 (United States). Here we report characterization of the complete genome of a genotype obtained from the brain of an African American with systemic lupus erythematosus (SLE) and progressive multifocal leukoencephalopathy (PML). STUDY DESIGN/METHODS: DNA extracted from JCV-infected brain tissue was subjected to whole-genome polymerase chain reaction (PCR) amplification and direct cycle sequencing. Relations to other JCV genotypes and the predicted amino acid sequence were analyzed. RESULTS: This African-American type 6 strain (#601) differs from strains of all other genotypes in about 2% of its DNA sequence. The length of the total coding region of strain #601 is increased to 4855 bp by the insertion of a single nucleotide in the large T-antigen intron. This strain, originally placed with the type 2 group on the basis of its sequence in the VT-intergenic region, is very closely related to strains recently identified in the urine of individuals from Ghana, West Africa. CONCLUSIONS: This is the first example of an African JCV genotype identified in the brain of an African-American PML patient. The extent of sequence divergence of JCV type 6 suggests a split of type 6 strains before the separation of types 2 and 3. These findings confirm that distinctive African genotypes of JCV have been maintained in the African-American population and that they are capable of causing PML.

Adult↗

Postgenomics: Proteomics and Bioinformatics in Cancer Research.

Now that the human genome is completed, the characterization of the proteins encoded by the sequence remains a challenging task. The study of the complete protein complement of the genome, the "proteome," referred to as proteomics, will be essential if new therapeutic drugs and new disease biomarkers for early diagnosis are to be developed. Research efforts are already underway to develop the technology necessary to compare the specific protein profiles of diseased versus nondiseased states. These technologies provide a wealth of information and rapidly generate large quantities of data. Processing the large amounts of data will lead to useful predictive mathematical descriptions of biological systems which will permit rapid identification of novel therapeutic targets and identification of metabolic disorders. Here, we present an overview of the current status and future research approaches in defining the cancer cell's proteome in combination with different bioinformatics and computational biology tools toward a better understanding of health and disease.

Journal Article↗

Genomic characterization and expression of mouse prestin, the motor protein of outer hair cells.

We previously identified the gerbil gene (gPres) that encodes prestin, the putative motor protein responsible for outer hair cell (OHC) electromotility. Here we report the cloning and characterization of the complete genomic structure of the mouse Prestin (mPres) gene. We performed 5'- and 3'-RACE to determine the size and identity of the full-length mRNA transcript. The mPres gene encodes a protein 96% identical to the gPres gene product. The prestin open reading frames are 91% identical at the nucleotide level. Using an antibody raised against the N-terminus of gerbil Prestin, we observed mPrestin expression by immunofluorescence in the lateral membrane of mouse OHCs and found no detectable expression elsewhere in the organ of Corti. On the basis of the available genomic sequence from mouse Chromosome (Chr) 5, we concluded that the mPres gene is centromerically related to and resides within 19 kb of, the Reln gene. We were also able to characterize the exon/intron junctions of mPres by using cDNA/genomic sequence comparisons, as well as exon-exon PCR and sequencing. The mPres gene has 18 exons that encode protein and two exons in the 5' UTR. A CpG island, located at the start of exon 1, is a potential transcription start site. Sequence analysis of the ~500 bp upstream from exon 1 revealed multiple potential transcription factor binding sites, including both TATA and GC boxes, as well as other regulatory-element binding sites.

Amino Acid Sequence↗

Complete sequence and characterization of the channel catfish mitochondrial genome.

In order to support analysis of channel catfish populations and genetic improvement programs, the channel catfish, Ictalurus punctatus, mitochondrial genome was completely sequenced and revealed gene structure and gene order common to vertebrates. Nucleotide sequence comparisons of cytochrome b (Cytb) and cytochrome c oxidase subunit 1 (COI) demonstrated genetic separation of the genera Ictalurus, Pylodictis and Ameiurus consistent with the taxonomic classification within Ictaluridae. The ictalurid Cytb nucleotide sequences were significantly different from a putative channel catfish Cytb sequence in GenBank. Genetic relationships based on mitochondrial DNA sequences indicated the value of channel catfish in genomic comparisons between teleosts. Pairwise alignment of DNA sequences revealed conservation of regulatory sequences in the D-loop region with other vertebrates. Analysis of D-loop sequences in commercial populations and a research strain revealed 28 polymorphic sites and 33 D-loop haplotypes. Sequence analysis revealed clustering of haplotypes within commercial farms and the USDA103 research line, but D-loop haplotypes were not sufficient to discriminate the USDA103 fish from commercial catfish.

Analysis of Variance↗