PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

An Integrated Sequence-Structure Database incorporating matching mRNA sequence, amino acid sequence and protein three-dimensional structure data.

We have constructed a non-homologous database, termed the Integrated Sequence-Structure Database (ISSD) which comprises the coding sequences of genes, amino acid sequences of the corresponding proteins, their secondary structure and straight phi,psi angles assignments, and polypeptide backbone coordinates. Each protein entry in the database holds the alignment of nucleotide sequence, amino acid sequence and the PDB three-dimensional structure data. The nucleotide and amino acid sequences for each entry are selected on the basis of exact matches of the source organism and cell environment. The current version 1.0 of ISSD is available on the WWW at http://www.protein.bio.msu.su/issd/ and includes 107 non-homologous mammalian proteins, of which 80 are human proteins. The database has been used by us for the analysis of synonymous codon usage patterns in mRNA sequences showing their correlation with the three-dimensional structure features in the encoded proteins. Possible ISSD applications include optimisation of protein expression, improvement of the protein structure prediction accuracy, and analysis of evolutionary aspects of the nucleotide sequence-protein structure relationship.

Algorithms↗

Characterization of DNA sequence-common and sequence-specific proteins binding to cis-acting sites for cleavage of the terminal a sequence of the herpes simplex virus 1 genome.

The terminal 500-base-pair alpha sequence of the herpes simplex virus 1 genome contains signals for cleavage (Pac1 and Pac2) of unit-length DNA molecules from concatemers in unique stretches of sequences designated Ub and Uc, respectively, and a cis site for cleavage designated DR1. We report that nuclear extracts from infected cells contain factors which form two DNA-virus-specific protein complexes with components of the a sequence. Purification of the factors forming the V2 complex yielded a protein with an apparent molecular weight of 82,000 binding to DNA in a non-sequence-specific manner. Addition of Mg2+ to the purified protein-DNA probe mixture resulted in exonucleolytic degradation of the DNA. The protein was identified as the virus-specific DNase with monoclonal antibody specific for the viral enzyme. The purification of the proteins forming the V4 complex yielded two proteins with molecular weights of greater than 250,000 and 140,000 corresponding to infected cell protein 1 and to an as yet unidentified protein, respectively. These proteins formed two DNA sequence-common bands with a number of DNA probes and one sequence-specific band with probes containing both Pac2 and DR1 but not with probes containing either site alone or Pac1 and DR1. Since the DNA probe containing Pac2 and DR1 inserted into viral genome or into amplicons induced specific cleavage of the DR1 sequence whereas the nonreactive probes failed to induce the cleavage, the formation of this sequence-specific DNA-protein complex is significant and may reflect a DNA-protein interaction essential for cleavage. The possible role of the proteins identified in this study for the cleavage-packaging of viral DNA into capsids is presented.

DNA Probes↗

Innovations in non-isotopic DNA sequencing: using an electrotransfer unit to blot sequencing gels and an automated membrane processor for detecting DNA sequences.

As alternatives to radiolabeled DNA sequencing, chemiluminescent and chromogenic sequencing methods can be comparable in both sensitivity and resolution. Chemiluminescent/chromogenic detection procedures are safer because they completely eliminate the handling and use of radioisotopes. One method involves standard dideoxy DNA sequencing reactions that are initiated with biotinylated primers, separated by gel electrophoresis, transferred onto nylon membrane and detected utilizing chemiluminescent 1,2-dioxetane substrates for alkaline-phosphatase. Alkaline phosphatase is linked to the biotinylated sequencing products by a streptavidin/alkaline phosphatase conjugate (SAAP). In this paper we describe an optimized procedure for transferring sequencing gels. The procedure is based on a semidry method developed at Hoefer Laboratories using the GeneSweep Sequencing Gel Transfer Unit. Transfer is rapid, uniform and reliable from gel to gel. We also describe automation of the development process using a fully programmable Gel/Membrane Processor that automates delivery, incubation and disposal of reagents. All crucial points for electrotransfer of sequencing gels and the detection of biotinylated DNA sequencing reaction products are discussed.

Alkaline Phosphatase↗

Cloning, sequencing and structural analysis of 976 base pairs of the promoter sequence for the rat lipoprotein lipase gene. Comparison with the mouse and human sequences.

We cloned and sequenced the -976bp promoter of the rat lipoprotein lipase LPL gene. The sequence was compared with the mouse and human sequences. The homology between the rat and mouse LPL nucleotide sequences was not quite as strong in the promoter sequence as in the coding sequence. Among the 976nt promoter there were 118 divergences, i.e. 11.8%, compared to only 5.6% for the LPL coding region. However, within the 200nt immediately 5' to the transcriptional start site (proximal promoter), the divergence was only 4%. New potential cis-elements (such as CACCC, GATA, GC and GA boxes, IRS, Krox, MEF 2, E-box, CCArGG and 1/2 VDRE) were identified in the rat, mouse or human LPL gene.

Animals↗

Comparison of the nucleotide sequence of cloned DNA coding for an apolipoprotein (apo VLDL-II) from avian blood and the amino acid sequence of an egg-yolk protein (apovitellenin I): equivalence of the two sequences.

We have compared the amino acid sequences of two low-molecular-weight avian apoproteins: apoVLDL-II from very low-density lipoproteins of hen plasma and apovitellenin I from hen egg yolk. The sequence of White Leghorn apoVLDL-II was derived from the nucleotide sequence of cloned apoVLDL-II DNA (Chan et al., 1980). The sequenator was used to determine the amino acid sequence of apovitellenin I from two breeds of hen (White Leghorn and Australorp). The sequences from the two breeds were not only identical, but they also completely matched the predicted sequence derived from the apoVLDL-II DNA sequence. The identity reported here establishes that this protein is transported intact from the blood to the egg yolk.

Amino Acid Sequence↗

Diversity in HIV-1 envelope V1-V3 sequences early in infection reflects sequence diversity throughout the HIV-1 genome but does not predict the extent of sequence diversity during chronic infection.

Differences in the extent of genetic diversity have been observed in human immunodeficiency virus type-1 (HIV-1) envelope sequences early in infection, and this has been linked to gender and to modifiable exogenous factors such as hormonal contraceptive use and genital tract infections. But it is unclear whether envelope diversity is indicative of diversity in other regions of the viral genome, and thus whether it adequately reflects whether multiple or a single virus initiated the infection. Here we show that six women with homogeneous envelope V1-V3 sequences during primary infection also had homogeneous gag and polymerase (pol) sequences at the same time. On the other hand, six women with multiple envelope sequences had diverse gag and pol genotypes during a similar interval after infection. This suggests that envelope sequences reflect sequence diversity throughout the viral genomes present early in infection and thus provide an indication of whether a single virus or multiple viruses initiated the infection. Analysis of HIV-1 sequences from about 3 years after infection revealed that the level of diversity and diversification was similar between the women in the two groups.

Chronic Disease↗

Sequence 'minimization': exploring the sequence landscape with simplified sequences.

The challenges of protein engineering arise, in part, from the enormous number of possible sequences and the almost unimaginably small fraction of such sequences that can be studied experimentally or computationally. Fortunately, not all possibilities need to be considered because many different sequences can adopt the same structure. Of the vast number of sequences that fold into a given conformation, some are 'simpler' than the sequences of typical proteins. Studying protein sequences that are simpler helps focus attention on the principal determinants of structure. Recent examples of this strategy are the simplification of protein surfaces and cores, the use of a binary 'code' for protein design and the structural analysis of random simple sequences.

Alanine↗

Prediction of the coding sequences of mouse homologues of KIAA gene: IV. The complete nucleotide sequences of 500 mouse KIAA-homologous cDNAs identified by screening of terminal sequences of cDNA clones randomly sampled from size-fractionated libraries.

We have been conducting a mouse cDNA project to predict protein-coding sequences of mouse homologues of human KIAA and FLJ genes since 2001. As an extension of these projects, we herein present the entire sequences of 500 mKIAA cDNA clones and 4 novel cDNA clones that were incidentally identified during this project. We have isolated cDNA clones from the size-fractionated mouse cDNA libraries derived from 7 tissues and 3 types of cultured cells. The average size of the 504 cDNA sequences reached 4.3 kb and that of the deduced amino acid sequences from these cDNAs was 807 amino acid residues. We assigned the integrity of CDSs from the comparison with the corresponding human KIAA cDNA sequences. The comparison of mouse and human sequences revealed that two different human KIAA cDNAs are derived from single genes. Furthermore, 3 out of 4 proteins encoded in the novel cDNA clones showed moderate sequence similarity with human KIAA proteins, thus we could obtain new members of KIAA protein families through our mouse cDNA projects.

Animals↗

The sequence of the N and L genes of rinderpest virus, and the 5' and 3' extra-genic sequences: the completion of the genome sequence of the virus.

We have sequenced the nucleocapsid (N) and polymerase (L) genes of the vaccine strain of rinderpest, and the 5' and 3' terminal domains of the genome. Together with previously published data, this completes the sequence of the entire genome of rinderpest virus. The L gene is identical in length to that of measles virus, encoding a 2183 amino acid protein with a calculated molecular weight of 248,100. The L protein sequence of three morbilliviruses is highly conserved, greater than 76% of residues being identical or conserved in all sequences. The N protein was, as for other sequenced genes, essentially identical to that of the virulent parent. The viral genome is 15,881 bases in length, similar to that of measles virus and slightly longer than that of canine distemper virus. The terminal sequences of the genome and those at the gene boundaries were compared to the analogous regions of other morbilliviruses and representatives of related groups of paramyxoviruses.

Amino Acid Sequence↗

DNA sequencing and comparative sequence analysis reveal that the Escherichia coli genomic DNA may replace the target DNA during molecular cloning: evidence for the erroneous assembly of E. coli DNA into database sequences.

DNA sequencing and similarity search of databases provide experimental evidence that portions of the host Escherichia coli genome may get ligated into the cloning vector, resulting in clones containing nontargeted inserts. Several lines of evidence suggest that this non-targeted ligation, as observed by us while subcloning troponin I cDNA, is presumably due to a recombination-mediated mechanism by which host DNA replaces the target DNA in the cloning vector. The E. coli genome mapping to 64-65 min and 92.8-00.1 min, the latter containing insertion sequences, appears to be the hotspot regions involved in this process. We examined the possibility that some sequences reported in the databases may also contain genomic sequences of E. coli. A search of current databases revealed that a rat hepatic glutathione transporter cDNA contains a 2.2-kb-long portion of the E. coli genome that has been wrongly assembled into its 5' untranslated and coding regions. In addition, about 30 sequences in databases, including a Yersinia pestis toxin gene, showed relatively high sequence identity with those portions of the E. coli genome that were present in the nonauthentic clones.

Animals↗

Prediction of the coding sequences of mouse homologues of KIAA gene: I. The complete nucleotide sequences of 100 mouse KIAA-homologous cDNAs identified by screening of terminal sequences of cDNA clones randomly sampled from size-fractionated libraries.

We have been conducting a human cDNA project to predict protein-coding sequences in long cDNAs (> 4 kb) since 1994. The number of these newly identified human genes exceeds 2000 and these genes are known as KIAA genes. As an extension of this project, we herein report characterization of cDNAs derived from mouse KIAA-homologous genes. A primary aim of this study was to prepare a set of mouse. KIAA-homologous cDNAs that could be used to analyze the physiological roles of KIAA genes in mice. In addition, comparison of the structures of mouse and human KIAA cDNAs might enable us to evaluate the integrity of KIAA cDNAs more convincingly. In this study, we selected mouse KIAA-homologous cDNA clones to be sequenced by screening a library of terminal sequences of mouse cDNAs in size-fractionated libraries. We present the entire sequences of 100 cDNA clones thus selected and predict their protein-coding sequences. The average size of the 100 cDNA sequences reached 5.1 kb and that of mouse KIAA-homologous proteins predicted from these cDNAs was 989 amino acid residues.

Animals↗

Two DNA-binding factors recognize specific sequences at silencers, upstream activating sequences, autonomously replicating sequences, and telomeres in Saccharomyces cerevisiae.

Two DNA-binding factors from Saccharomyces cerevisiae have been characterized, GRFI (general regulatory factor I) and ABFI (ARS-binding factor I), that recognize specific sequences within diverse genetic elements. GRFI bound to sequences at the negative regulatory elements (silencers) of the silent mating type loci HML E and HMR E and to the upstream activating sequence (UAS) required for transcription of the MAT alpha genes. A putative conserved UAS located at genes involved in translation (RPG box) was also recognized by GRFI. In addition, GRFI bound with high affinity to sequences with the (C1-3A)-repeat region at yeast telomeres. Binding sites for GRFI with the highest affinity appeared to be of the form 5'-(A/G)(A/C)ACCCANNCA(T/C)(T/C)-3', where N is any nucleotide. ABFI-binding sites were located next to autonomously replicating sequences (ARSs) at controlling elements of the silent mating type loci HMR E, HMR I, and HML I and were associated with ARS1, ARS2, and the 2 micron plasmid ARS. Two tandem ABFI binding sites were found between the HIS3 and DED1 genes, several kilobase pairs from any ARS, indicating that ABFI-binding sites are not restricted to ARSs. The sequences recognized by ABFI showed partial dyad-symmetry and appeared to be variations of the consensus 5'-TATCATTNNNNACGA-3'. GRFI and ABFI were both abundant DNA-binding factors and did not appear to be encoded by the SIR genes, whose products are required for repression of the silent mating type loci. Together, these results indicate that both GRFI and ABFI play multiple roles within the cell.

Chromosomes↗

Amino acid sequence analysis of the H-2Kk alloantigen: complete sequence of residues 1-98 and partial sequence from 99 to 263.

The H-2Kk molecule was purified by immunoprecipitation from the glycoprotein fraction of Nonidet P-40 extracts of RDM 4 mouse tumor cells. Cyanogen bromide cleavage of the major papain fragment yielded three peptides, the largest of which consisted of three disulfide-linked peptides which could be separated after reduction and alkylation. These peptides were readily aligned by their homology to similar fragments derived from other H-2 class I molecules. Amino acid sequence analyses of the two nondisulfide-linked peptides, peptide E (residues 1-52) and peptide D (53-98), yielded the following NH2-terminal sequence for the H-2Kk molecule: [sequence in text]. Comparison of this sequence with those of other H-2 class I molecules revealed that: (1) Lys-19, Val-55, Glu-56, Asn-63 and Ile-73 are unique to the H-2Kk molecule; and (2) H-2Kk shares 79-83% homology in this region with other mouse class I molecules. Partial NH2-terminal amino acid sequences are also reported for the three disulfide-linked peptides. Several discrepancies from previously reported partial sequences of the H-2Kk molecule were detected.

Amino Acid Sequence↗

Representation of cloned genomic sequences in two sequencing vectors: correlation of DNA sequence and subclone distribution.

Representation of subcloned Caenorhabditis elegans and human DNA sequences in both M13 and pUC sequencing vectors was determined in the context of large scale genomic sequencing. In many cases, regions of subclone under-representation correlated with the occurrence of repeat sequences, and in some cases the under-representation was orientation specific. Factors which affected subclone representation included the nature and complexity of the repeat sequence, as well as the length of the repeat region. In some but not all cases, notable differences between the M13 and pUC subclone distributions existed. However, in all regions lacking one type of subclone (either M13 or pUC), an alternate subclone was identified in at least one orientation. This suggests that complementary use of M13 and pUC subclones would provide the most comprehensive subclone coverage of a given genomic sequence.

Animals↗

Identification of Campylobacter spp. and discrimination from Helicobacter and Arcobacter spp. by direct sequencing of PCR-amplified cpn60 sequences and comparison to cpnDB, a chaperonin reference sequence database.

A robust method for the identification of Campylobacter spp. based on direct sequencing of PCR-amplified partial cpn60 sequences and comparison of these to a reference database of cpn60 sequences is reported. A total of 53 Campylobacter isolates, representing 15 species, were identified and distinguished from phenotypically similar Helicobacter and Arcobacter strains. Pairwise cpn60 sequence identities between Campylobacter spp. ranged from 71 to 92 %, with most between 71 and 79 %, making discrimination of these species obvious. The method described overcomes limitations of existing PCR-based methods, which require time-consuming and complex post-amplification steps such as the cloning of amplification products. The results of this study demonstrate the potential for use of the reference chaperonin sequence database, cpnDB, as a tool for identification of bacterial isolates based on cpn60 sequences amplified with universal primers.

Arcobacter↗

Amino-acid sequence of lac repressor from Escherichia coli. Isolation, sequence analysis and sequence assembly of tryptic peptides and cyanogen-bromide fragments.

The lac repressor from Escherichia coli, composed of four identical subunits with a molecular weight of 37160, was carboxymethylated and fragmented by tryptic digestion and cyanogen bromide treatment. Using ion-exchange chromatography, gel filtration and preparative thin-layer electrophoresis and chromatography 29 of the 30 tryptic peptides were isolated in pure form. Direct Edman degradation and the dansyl-Edman technique were used to determine the sequence of the small tryptic peptides. Special emphasis was put on the sequence determination of the six large tryptic fragments which together account for 177 residues, corresponding to 51% of the repressor subunit with its 347 residues. The large tryptic fragments were analyzed after fragmentation with chymotrypsin, thermolysin and dipeptidyl aminopeptidase I. Thus the sequence of all 30 tryptic peptides could be deduced. The complete sequences of all cyanogen bromide fragments were deduced from peptides obtained by tryptic, chymotryptic and thermolytic digestion of the individual fragments and by automated stepwise Edman degradation of lac repressor and of the large cyanogen bromide fragments. The order of the cyanogen bromide fragments was given by overlapping tryptic peptides. The resulting amino acid composition of the monomer is Asp15, Asn11, Thr18, Ser30, Glu14, Gln27, Pro13, Gly22, Ala44, Cys3, Val33, Met9, Ile17, Leu40, Tyr8, Phe4, Trp2, Lys11, His7, Arg19. The sequence of lac repressor shows no similarities with that of other proteins known to bind to DNA or RNA. The N-terminal 55 residues contain two homologous regions. This part of the sequence which is involved in lac operator binding might have been formed by gene duplication.

Amino Acid Sequence↗

Prediction of the coding sequences of mouse homologues of KIAA gene: III. the complete nucleotide sequences of 500 mouse KIAA-homologous cDNAs identified by screening of terminal sequences of cDNA clones randomly sampled from size-fractionated libraries.

We have conducted a human cDNA project to predict protein-coding sequences (CDSs) in large cDNAs (> 4 kb) since 1994, and the number of newly identified genes, known as KIAA genes, already exceeds 2000. The ultimate goal of this project is to clarify the physiological functions of the proteins encoded by KIAA genes. To this end, the project has recently been expanded to include isolation and characterization of mouse KIAA-counterpart genes. We herein present the entire sequences and the chromosome loci of 500 mKIAA cDNA clones and 13 novel cDNA clones that were incidentally identified during this project. The average size of the 513 cDNA sequences reached 4.3 kb and that of the deduced amino acid sequences from these cDNAs was 816 amino acid residues. By comparison of the predicted CDSs between mouse and human KIAAs, 12 mKIAA cDNA clones were assumed to be differently spliced isoforms of the human cDNA clones. The comparison of mouse and human sequences also revealed that four pairs of human KIAA cDNAs are derived from single genes. Notably, a homology search against the public database indicated that 4 out of 13 novel cDNA clones were homologous to the disease-related genes.

Animals↗

[MRI of the regions of the inner ear and cerebellopontine angle using a 3D T2-weighted turbo spin-echo sequence. Comparison with conventional 2D T2-weighted turbo spin-echo sequences and T1-weighted spin-echo sequences].

PURPOSE: To assess the value of a three-dimensional (3D) T2-weighted turbo spin-echo sequence (3D T2-TSE) in comparison to conventional two-dimensional (2D) T2-weighted TSE and unenhanced and enhanced T1-weighted spin-echo sequences (SE) in imaging anatomic structures and pathologic changes of the inner ear and cerebellopontine angle. PATIENTS AND METHODS: The inner ear and cerebellopontine angle were investigated by MRI in three healthy volunteers and 18 patients performing a 2D T2-weighted turbo spin-echo sequence and a 3D T2-TSE in the axial plane. In the patient study, 2D T1-weighted SE sequences both before and after the i.v. injection of gadopentetate dimeglumine in both the axial and coronal plane were performed in addition. RESULTS: Only the 3D T2-TSE enabled an accurate imaging of the anatomic structures. In cases of pathology, the 3D T2-TSE provided additional information to the performed 2D sequences. The combination of the 3D T2-TSE with unenhanced and enhanced 2D T1-weighted SE enabled the most accurate diagnosis in cases of pathology. CONCLUSIONS: Accurate depiction of anatomic structures of the inner ear and cerebellopontine angle could be obtained by 3D T2-TSE only. The most accurate diagnosis in cases of pathology was provided by the combination of the 3D T2-TSE with unenhanced and enhanced 2D T1-weighted spin-echo sequences.

Adult↗