PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Base Sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

A novel HLA-B allele, B*5612, identified by sequence-based typing method.

HLA-B*5612, found in Taiwan using sequence-based typing method, was identical to HLA-B*5502 in exon 2 but differed in exon 3 by 10 nucleotide substitutions at positions 353-420 leading to five amino acid change at codon 94, 95, 97, 103 and 116. As this sequence motif was not found in the Asian population, it is likely that HLA-B*5612 is the product of a complex mechanism of implying a dual gene conversion event.

Base Sequence↗

Characterization of DNA primary sequences based on the average distances between bases.

We outline numerical characterization of DNA primary sequence based on calculation of the average distance between pairs of nucleic acid bases. This leads to a representation of DNA by a condensed 4 x 4 symmetrical matrix, the elements of which give the average separation between pair of bases X, Y in DNA (X, Y = A, C, G, T). As an invariant of choice we consider the leading eigenvalue of the derived 4 x 4 matrix. Additional structurally related invariants were obtained by constructing additional "higher order" 4 x 4 matrices derived from the initial 4 x 4 matrix by raising its elements to higher powers. Suitably normalized leading eigenvalue of these matrices offer a novel characterization of DNA primary sequences, referred to as "DNA profiles". The approach is illustrated on exon 1 of human beta-globin gene.

DNA↗

The map problem: a comparison of genetic and sequence-based physical maps.

The genetic order of autosomal genome-scan markers from Marshfield panels 9 and 10 were compared with their physical order, on the basis of the assembled nonredundant human genome sequence from the Human Genome Project-Santa Cruz (HGP-sc; October 2000 and April 2001 releases) and Celera (CEL; February 2001 release) databases. The genetic order of 96% of the markers on the Marshfield map for panel 10 is supported by a likelihood ratio of > or = 3 (odds ratio of 1,000:1). Inconsistencies with the genetic panel 10 map were found for 5% and 2% of the markers in the CEL and HGP-sc sequences, respectively. These inconsistencies consisted of both positional and chromosomal-assignment disagreements. For the majority of these inconsistent markers, the genetic order was supported by a likelihood ratio of > or = 3, and the physical order in the other assembly matched the genetic order. The majority of the inconsistencies between the physical- and genetic-map order point to errors in the physical-map order. A Web site is made available that displays inconsistencies for genetic markers from Marshfield panels 9 and 10 between their genetic-map positions and sequence-based physical-map positions, as well as inconsistencies between their sequence-based physical position. This Web site also contains genetic-map distances, physical-map positions from the Celera and Human Genome Project sequence, and likelihood-ratio support for the genetic maps.

Chromosome Mapping↗

High-resolution sequence-based typing strategy for HLA-DQA1 using SSP-PCR and subsequent genotyping analysis with novel spreadsheet program.

We present a new sequence-based typing (SBT) strategy for the polymorphic HLA-DQA1 locus that is based on sequence-specific primer - polymerase chain reaction (SSP-PCR) amplification from genomic DNA. This method allows high-resolution genotyping in the second exon of the DQA1 gene. This gene presents a unique situation in which half of the known alleles contain an inframe three base pair deletion of codon 56. This deletion confounds direct SBT methodologies of heterozygous individuals containing both a deletion and nondeletion allele. The primary HLA haplotype associated with type 1 diabetes susceptibility is DR3/DR4. The DQA1 genotype for these two haplotypes are DQA1 *0501, a non-deletion allele and *0301, a deletion allele, thus creating a situation that cannot be resolved using a direct sequencing approach. Our group-specific SBT strategy isolates the deletion alleles from the nondeletion alleles, allowing them to be resolved by direct sequencing. Additionally, we present a novel spreadsheet program that accurately assigns the genotype of both homozygous and heterozygous persons.

Alleles↗

Mapping Ds insertions in barley using a sequence-based approach.

A transposon tagging system, based upon maize Ac/Ds elements, was developed in barley (Hordeum vulgaresubsp. vulgare). The long-term objective of this project is to identify a set of lines with Ds insertions dispersed throughout the genome as a comprehensive tool for gene discovery and reverse genetics. AcTPase and Ds-bar elements were introduced into immature embryos of Golden Promise by biolistic transformation. Subsequent transposition and segregation of Ds away from AcTPase and the original site of integration resulted in new lines, each containing a stabilized Ds element in a new location. The sequence of the genomic DNA flanking the Ds elements was obtained by inverse PCR and TAIL-PCR. Using a sequence-based mapping strategy, we determined the genome locations of the Ds insertions in 19 independent lines using primarily restriction digest-based assays of PCR-amplified single nucleotide polymorphisms and PCR-based assays of insertions or deletions. The principal strategy was to identify and map sequence polymorphisms in the regions corresponding to the flanking DNA using the Oregon Wolfe Barley mapping population. The mapping results obtained by the sequence-based approach were confirmed by RFLP analyses in four of the lines. In addition, cloned DNA sequences corresponding to the flanking DNA were used to assign map locations to Morex-derived genomic BAC library inserts, thus integrating genetic and physical maps of barley. BLAST search results indicate that the majority of the transposed Ds elements are found within predicted or known coding sequences. Transposon tagging in barley using Ac/Ds thus promises to provide a useful tool for studies on the functional genomics of the Triticeae.

Base Sequence↗

The amino acid sequence of a crystal protein from Bacillus thuringiensis deduced from the DNA base sequence.

We have determined the nucleotide sequence of a 4222-base segment of DNA which contains the promoter, the coding region, and the terminator of a crystal protein gene cloned from a Bacillus thuringiensis plasmid. A sequence of 1176 amino acids encoding a Mr 133,500 peptide was deduced from the single open reading frame. This protein-coding region was analyzed for codon usage, predicted hydropathy, and predicted secondary structure. Examination of the base sequence revealed the presence of several inverted and direct repeats located in both the coding and noncoding regions. S1 nuclease mapping was used to locate the transcription termination point at a site following a potentially very stable stem-and-loop structure.

Amino Acid Sequence↗

Nucleic acid sequence-based amplification, a new method for analysis of spliced and unspliced Epstein-Barr virus latent transcripts, and its comparison with reverse transcriptase PCR.

Nucleic acid sequence-based amplification (NASBA) assays were developed for direct detection of Epstein-Barr virus (EBV) transcripts encoding EBV nuclear antigen 1 (EBNA1), latent membrane proteins (LMP) 1 and 2, and BamHIA rightward frame 1 (BARF1) and for the noncoding EBV early RNA 1 (EBER1). The sensitivities of all NASBAs were at least 100 copies of specific in vitro-generated RNA. Furthermore, 1 EBV-positive JY cell in a background of 50,000 EBV-negative Ramos cells (the relative sensitivity) was detected by using the EBNA1, LMP1, and LMP2 NASBA assays. The relative sensitivity of the EBER1 NASBA was 100 EBV-positive cells, which was probably related to the loss of small RNA molecules during the isolation. The BARF1 and LMP2 NASBAs were evaluated on clinical material. BARF1 expression was found in 6 of 7 nasopharyngeal carcinomas (NPC) but in 0 of 22 Hodgkin's disease (HD) cases, whereas LMP2 expression was found in 7 of 7 NPCs and in 17 of 22 HD cases. For detection of EBNA1 transcripts in HLs (n = 12) and T- and B-cell non-Hodgkin's lymphomas (n = 3 and n = 2, respectively), NASBA was compared with reverse transcriptase (RT) PCR. Two samples were positive only with NASBA, and two other samples were positive only with RT-PCR; for all other samples, the RT-PCR and NASBA results were in agreement. We conclude that NASBA is suitable for sensitive and specific detection of the above-mentioned EBV transcripts, regardless of their splicing patterns and the presence of EBV DNA. The EBNA1, LMP2, and BARF1 NASBAs developed in this study proved to be reliable assays for detection of the corresponding transcripts in EBV-positive clinical material.

Base Sequence↗

Comparison of serological and sequence-based methods for typing feline calcivirus isolates from vaccine failures.

Feline calicivirus (FCV) can be typed by exploiting antigenic differences between isolates or, more recently, by the sequence analysis of a hypervariable region of the virus's capsid gene. These two methods were used to characterise FCV isolates from 20 vaccine failures which occurred after the use of a commercial, live-attenuated vaccine. Using virus neutralisation, the isolates showed a spectrum of relatedness to the vaccine; depending on the criterion adopted for identity, 10 to 40 per cent of them appeared to be similar to the vaccine virus. Using sequence analysis, the isolates fell into one of two categories; 20 per cent had a similar sequence to the vaccine (0-67 to 2-67 per cent distant), and the remainder had a dissimilar sequence (21-3 to 36-0 per cent distant). Sequence analysis identified one cat that appeared to be infected with two distinct FCVs. The serological and sequence-based typing methods gave the same result in 80 to 95 per cent of individual cases, depending on the criterion adopted for serological identity. It is suggested that molecular typing is a more definitive method for characterising the relatedness of FCV isolates.

Animals↗

A sequence-based filtering method for ncRNA identification and its application to searching for riboswitch elements.

MOTIVATION: Recent studies have uncovered an "RNA world", in which non coding RNA (ncRNA) sequences play a central role in the regulation of gene expression. Computational studies on ncRNA have been directed toward developing detection methods for ncRNAs. State-of-the-art methods for the problem, like covariance models, suffer from high computational cost, underscoring the need for efficient filtering approaches that can identify promising sequence segments and speedup the detection process. RESULTS: In this paper we make several contributions toward this goal. First, we formalize the concept of a filter and provide figures of merit that allow comparison between filters. Second, we design efficient sequence based filters that dominate the current state-of-the-art HMM filters. Third, we provide a new formulation of the covariance model that allows speeding up RNA alignment. We demonstrate the power of our approach on both synthetic data and real bacterial genomes. We then apply our algorithm to the detection of novel riboswitch elements from the whole bacterial and archaeal genomes. Our results point to a number of novel riboswitch candidates, and include genomes that were not previously known to contain riboswitches. AVAILABILITY: The program is available upon request from the authors.

Algorithms↗

Investigation of killer cell immunoglobulin-like receptor KIR2DL4 diversity by sequence-based typing in Chinese population.

Human killer cell immunoglobulin-like receptors (KIRs) play an important role in controlling natural killer (NK) cell function. Here, polymerase chain reaction sequence-based typing (PCR-SBT) procedures identifying alleles of the KIR2DL4 gene have been established. The method was designed around the specific amplification of exon 3 to exon 5 and exon 7 to exon 9 of the KIR2DL4 gene and produce discrimination of KIR2DL4 alleles. Genomic DNAs from 83 healthy unrelated Chinese Han individuals were typed for KIR2DL4 alleles by this method. Each sample was assigned to the putative KIR2DL4 allele combination according to the nucleotide polymorphism profiles of all KIR2DL4 alleles. Twenty-one different genotypes and seven KIR2DL4 alleles were observed in the population, with KIR2DL4*00102 having the highest frequency, 0.5. Five individuals bear a recombinant allele KIR3DP*004 that associated with three putative KIR2DL4 alleles. Our data demonstrated that the established PCR-SBT method for KIR2DL4 allele typing was reliable, and Chinese Han population is distinct in KIR2DL4 allele frequencies in comparison to some other populations.

Alleles↗

Characterization of a new HLA-DRB1*01 allele (HLA-DRB1*010203) in a Caucasian Italian family by using sequence-based typing.

We report the identification of an HLA-DRB1*01 nucleotide sequence variant in three members of a Caucasian Italian family by using sequence-based typing. The nucleotide sequence of exon 2 observed in the new allele is identical to that of HLA-DRB1*010201 except in position 189 (codon 34) where the adenine of the consensus was replaced by a guanine and it was designated officially as HLA-DRB1*010203* by the WHO Nomenclature Committee.

Alleles↗

Evolutionary relationships of the primate papovaviruses: base sequence homology among the genomes of simian virus 40, stump-tailed macaque virus, and SA12 virus.

Physical maps of the genomes of the two newly discovered primate papovaviruses, SA12 and stump-tailed macaque virus (STMV), were generated by restriction endonuclease analysis. The base sequence homologies among the genomes of SA12, stump-tailed macaque virus, and simian virus 40 (SV40) were studied by heteroduplex analysis. Heteroduplexes between SA12 and SV40 DNAs and stump-tailed macaque virus and SV40 DNAs were constructed and mounted for electron microscopy in various amounts of formamide to achieve a range of effective temperatures. At each effective temperature, the regions of duplex DNA in the heteroduplexes were measured and localized on the SV40 physical and functional maps. By analyzing the data from this study and rom our previous study (N. Newell, C. J. Lai, G. Khoury, and T. J. Kelly Jr., J. Virol. 25:193-201, 1978) on the base sequence homology between the genomes of BK virus and SV40, some general conclusions have been drawn concerning the evolutionary relationships among the genomes of the primate papovaviruses. The extent of homology among the viral genomes does not reflect the phylogenetic relationships of their hosts. At comparable effective temperatures Tm - 33 degrees C), the heteroduplexes between the DNAs of BK virus and SV40 contained the largest amount of duplex (about 90%). The heteroduplexes made between SA12 and SV40 DNAs were slightly less homologous, containing about 80% duplex. The heteroduplexes made between SV40 and stump-tailed macaque virus DNAs were only 20% duplex under the same conditions. When the various heteroduplexes were mounted for microscopy at effective temperatures greater than Tm - 33 degrees C, the fraction of the duplex DNA decreased in each case, indicating the existence of considerable base mismatching in the homologous regions. When specific coding or noncoding regions of the viral genomes were compared, the data indicated that the extent of sequence divergence differed markedly from one region to another. In all the heteroduplexes studied, there were two regions, located near the junctions between early and late regions on the SV40 map, which were essentially nonhomologous. All of the heteroduplexes studied showed significantly greater homology in the late region than in early region. Within the late region, the sequences coding for the major capsid polypeptide, VP1, were the most highly conserved.

Base Sequence↗

Identification of a novel HLA-DPB1 allele, DPB1*02014, by sequence-based typing.

In this report we describe the identification of a novel DPB1 allele, DPB1*02014, found in an Italian Caucasian individual. The new allele was detected during routine HLA sequence-based typing (SBT) for an individual undergoing bone marrow transplantation. DPB1*02014 was identical to DPB1*02012 except for a single nucleotide substitution in codon 72 (GTG-->GTT). This nucleotide change represents a synonimous mutation, as both triplets code for a valine. This new allele has been submitted to GenBank and assigned the accession number AF326565. The WHO Nomenclature Committee has officially assigned the name DPB1*02014.

Base Sequence↗

RNA amplification by nucleic acid sequence-based amplification with an internal standard enables reliable detection of Chlamydia trachomatis in cervical scrapings and urine samples.

In the present study, the suitability of RNA amplification by nucleic acid sequence-based amplification (NASBA) for the detection of Chlamydia trachomatis infection was investigated. When comparing different primer sets for their sensitivities in NASBA, use of both the plasmid and omp1 targets resulted in a detection limit of 1 inclusion-forming unit (IFU), while the 16S rRNA appeared to be the most sensitive RNA target for amplification (10(-3) IFU). In contrast, for DNA amplification by PCR, the plasmid target was optimal (10(-2) IFU), which is 10 times less sensitive than rRNA NASBA. To exclude false negativity in NASBA detection because of inhibition of amplification and/or inefficient sample preparation, an internal standard was developed. The internal control was added prior to sample preparation. This 16S rRNA NASBA with an internal control was compared with a plasmid DNA PCR by using a group of C. trachomatis-negative (n = 41) and -positive (n = 37) cervical scrapings, as determined by enzyme immunoassay (EIA). In addition, urine samples from the EIA-positive women were tested (n = 17). Both NASBA and PCR assays were able to detect C. trachomatis in all EIA-positive cervical scrapings, the corresponding urine samples, and two samples from the EIA-negative group. The internal NASBA standard was found clearly in all EIA-negative samples. In conclusion, these results indicate that detection of C. trachomatis by RNA amplification by NASBA with an internal standard is a suitable and highly sensitive detection method, with potential use in the diagnosis of urogenital C. trachomatis infections with cervical scrapings as well as urine specimens.

Bacterial Outer Membrane Proteins↗

Comparison of the NucliSens Basic kit (Nucleic Acid Sequence-Based Amplification) and the Argene Biosoft Enterovirus Consensus Reverse Transcription-PCR assays for rapid detection of enterovirus RNA in clinical specimens.

Samples were tested for enterovirus by nucleic acid sequence-based amplification (NASBA) (NucliSens Basic kit; BioMerieux), reverse transcription-PCR (RT-PCR) (Enterovirus Consensus RT-PCR kit; Argene Biosoft), and virus isolation. Eighty-two samples were tested, and 44 were positive, 34 by both NASBA and RT-PCR and 5 each by NASBA or RT-PCR only. Two nasopharyngeal samples positive only by RT-PCR were determined to be rhinovirus. Of 42 enterovirus-positive samples, NASBA detected 39 (92.9%) and RT-PCR detected 37 (88.1%). The NucliSens Basic kit and the Argene Biosoft RT-PCR had comparable sensitivities for detection of enterovirus RNA, and both molecular methods were more sensitive than culture, which detected only 60.5% of positive samples. NASBA could be completed in 6.5 h versus 9 h for the Argene Biosoft RT-PCR kit.

Base Sequence↗

A sequencing-based typing method for HLA-DQA1 alleles.

Sequencing-based typing (SBT) is the most comprehensive method for characterizing human leukocyte antigen gene polymorphisms. Development of a SBT method for DQA1 is hampered because of a deletion of codon 56 in nearly half of the known DQA1 alleles. Sequence electropherograms of heterozygous samples comprising a deletion allele and a non-deletion allele display misalignment after codon 56 because of a three base-pair shift in the deletion allele. To overcome this problem, we have designed three group-specific primer sets to selectively amplify the deletion alleles from the nondeletion alleles. DNA samples are initially polymerase chain reaction (PCR)-typed using these primer sets along with an internal positive control primer set specific to growth hormone gene 1 (hGH1). The positive group-specific PCR reactions were selectively repeated without hGH1 control primers, and the amplicons were used as template in sequencing reactions. The sequence data were analyzed to obtain DQA1 types using ABI MatchTools software as well as the newly available Conexio Genomics Assign SBT Genotyping Software. The method was validated using a panel of reference DNA from the University of California, Los Angeles, International DNA Exchange Program. We conclude that the present SBT method is a technically simple and robust procedure to characterize the sequence polymorphisms in exon 2 of DQA1 gene.

Alleles↗

DNA cross-linking by dehydromonocrotaline lacks apparent base sequence preference.

Pyrrolizidine alkaloids (PAs) are ubiquitous plant toxins, many of which, upon oxidation by hepatic mixed-function oxidases, become reactive bifunctional pyrrolic electrophiles that form DNA-DNA and DNA-protein cross-links. The anti-mitotic, toxic, and carcinogenic action of PAs is thought to be caused, at least in part, by these cross-links. We wished to determine whether the activated PA pyrrole dehydromonocrotaline (DHMO) exhibits base sequence preferences when cross-linked to a set of model duplex poly A-T 14-mer oligonucleotides with varying internal and/or end 5'-d(CG), 5'-d(GC), 5'-d(TA), 5'-d(CGCG), or 5'-d(GCGC) sequences. DHMO-DNA cross-links were assessed by electrophoretic mobility shift assay (EMSA) of 32P endlabeled oligonucleotides and by HPLC analysis of cross-linked DNAs enzymatically digested to their constituent deoxynucleosides. The degree of DNA cross-links depended upon the concentration of the pyrrole, but not on the base sequence of the oligonucleotide target. Likewise, HPLC chromatograms of cross-linked and digested DNAs showed no discernible sequence preference for any nucleotide. Added glutathione, tyrosine, cysteine, and aspartic acid, but not phenylalanine, threonine, serine, lysine, or methionine competed with DNA as alternate nucleophiles for cross-linking by DHMO. From these data it appears that DHMO exhibits no strong base preference when forming cross-links with DNA, and that some cellular nucleophiles can inhibit DNA cross-link formation.

Amino Acids↗