PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Second generation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Antisense and/or immunostimulatory oligonucleotide therapeutics.

Antisense technology, which is based on a simple and rational principle of Watson-Crick complementary base pairing of a short oligonucleotide with the targeted mRNA to downregulate the disease-causing gene product, has progressed tremendously in the last two decades. Antisense oligonucleotides targeted to a number of cancer-causing genes are being evaluated in human clinical trials. While the first-generation phosphorothioate antisense oligonucleotides are in clinical trials, a number of factors, including sequence motifs that could lead to unwanted mechanisms of action and side effects, have been identified. The severity of the side effects of first-generation antisense oligonucleotides is mostly dependent on the presence of certain sequence motifs, such as CpG dinucleotides. A number of second-generation chemical modifications have been proposed to overcome the limitations of the first-generation antisense oligonucleotides. The safety and efficacy of several second-generation mixed-backbone antisense oligonucleotides are being evaluated in clinical trials. The immune stimulation affects observed with CpG-containing antisense oligonucleotides are being exploited as a novel therapeutic modality, with several CpG oligonucleotides being evaluated in clinical trials. A number of medicinal chemistry studies performed to date suggest that the immunomodulatory activity of CpG oligonucleotides can be fine-tuned by site-specific incorporation of chemical modifications in order to design disease-specific oligonucleotide therapeutics.

Adjuvants, Immunologic↗

Apparent movement of successively generated subjective figures.

In the present studies a pair of random-dot frames was constructed so that two areas in the first frame (f1) were correlated with two areas in the second frame (f2). The alternation of the pair of frames (an f1--f2 sequence) gave rise to two subjective figures. When two pairs of randomdot frames (an f1--f2 sequence and an f3--f4 sequence), each of which produced two subjective figures in different locations, were thmeselves alternated, the subjective figures from the f1--f2 sequence interacted with the subjective figures from the f3--f4 sequence to produce apparent movement. With any one of the four general kinds of displays which we constructed, subjects usually perceived only one of two types of subjective-figure movement. The type of movement that was perceived with a given display depended primarily upon the degree of change (across the interval between an f1--f2 and an f3--f4 sequence) of the internal structure of the successively generated subjective figures. Relative intensity differences between the subjective figures and their backgrounds influenced the type of apparent movement seen, whereas variations in the density of elements in a display did not. We tentatively propose a two-stage model to explain the apparent movement of the subjective figures: the first stage is assumed to generate the subjective figures by means of a cross-correlation of the intensity distributions of the two frames within an f1--f2 sequence and within an f3--f4 sequence; on the basis of inputs from the first stage, the second stage generates apparent movement signals for the subjective figures.

Discrimination, Psychological↗

High HIV type 1 subtype diversity and few drug resistance mutations among seropositive people detected during the 2005 second generation HIV surveillance in Madagascar.

Subtype determination and detection of drug resistant-associated mutations (DRM) were performed on 31 HIV-1 Western blot-positive sera during the 2005 second-generation HIV surveillance in Madagascar. Amplification and sequencing of at least one of the partial reverse transcriptase, protease, and partial envelope genes were successful for all strains. All three gene sequences were obtained for 28 strains. A high degree of subtype or circulating recombinant forms (CRF) was observed for these 28 strains: A-A1 (eight cases), CRF02_AG (six cases), B (five cases), C (three cases), CRF06_cpx (three cases), CRF10_CD, BC()CRF, and unique RF (one case each). According to the ANRS September 2005 DRM list and algorithm, no DRM was detected in the reverse transcriptase and only one strain bore three major DRM in the protease M46I, I84V, and L90M leading to resistance to indinavir, saquinavir, nelfinavir, atazanavir/ritonavir, and possibly lopinavir.

Adult↗

Transcription in vivo and in vitro of the histone-encoding gene hmfB from the hyperthermophilic archaeon Methanothermus fervidus.

Immediately upstream of the hmfB gene, in a DNA fragment cloned from Methanothermus fervidus, are two identical tandemly repeated copies of a 73-bp sequence that contain the sequence 5'TTTATATA, which conforms precisely to the consensus TATA box element proposed for methanogen promoters. By using this duplicated region as the template DNA and a cell-free transcription system derived from Methanococcus thermolithotrophicus, transcription in vitro was found to initiate at two identical sites 73 bp apart, each 25 bp downstream from a TATA box, thus providing strong evidence for the functional conservation of this transcriptional signal in two phylogenetically very diverse methanogens. Transcription of the hmfB gene in vivo in M. fervidus was found to occur at only one of these sites, and consistent with this observation, recloning and sequencing of this intergenic region after its amplification by the polymerase chain reaction demonstrated that the genome of M. fervidus contains only one copy of the 73-bp sequence upstream of the hmfB gene. Since the second copy of the 73-bp sequence, presumably generated artifactually during the original hmfB cloning, functioned equally well as a promoter in the M. thermolithotrophicus transcription system, all information needed by the heterologous RNA polymerase to initiate transcription accurately in vitro must be present within this sequence. The hmfB gene encodes HMf-2, one of the two subunits of HMf, an abundant DNA binding protein in M. fervidus which binds to DNA molecules in vitro, forming nucleosomelike structures. Cell-free transcription was inhibited by adding HMf or eucaryotic core histones at protein-to-DNA mass ratios of 0.3:1 and 1:1, respectively, whereas the archael histonelike protein HTa from Thermoplasma acidophilum inhibited transcription in vitro only at much higher protein-to-DNA mass ratios and the bacterial histonelike protein HU from Escherichia coli had no detectable effect on transcription.

Archaeal Proteins↗

The ultrastructural study of microgametogenesis of Eimeria scinci Phisalix 1923 (Apicomplexa: Eimeriidae) infecting the sandfish lizard, Scincus mitranus Anderson 1871 in Saudi Arabia.

The ultrastructure of microgametogenesis of Eimeria scinci Phisalix 1923 was described for the first time in the gall bladder epithelium of experimentally infected sandfish lizards, Scincus milranus Anderson 1871 from Al-Baha region in Saudi Arabia. Recorded sequence of events started as sexually differentiated second generation merozoites transformed into microgamonts, where most of the apicomplexan organelles have been disappeared gradually. Microgamonts were recognizable by the presence of peripherally arranged nuclei and the presence of one or two centrioles between each nucleus and the limiting membrane of the gamont. Early microgamonts were surrounded by a very narrow parasitophorous vacuoles, which widened during development and contained a few intravacuolar folds and tubules. Differentiation of microgametes began by elevations of the limiting membrane appeared above the centrioles and the segregation of nuclear content into a dense osmiophilic portion and an electron-pale one. A gradual protrusion of the dense portion of the nucleus and the developing flagella into the parasitophorous vacuole was proceeded. Fully developed microgametes become detached and occupied the parasitophorous vacuole along with the residual body of the mother microgamont. Each microgamete had at least two flagella; an anterior perforatorium; a dense elongate nucleus and an anteriorly located tubular mitochondrion. The (9x2+2) pattern of flagella was detected in transverse sections.

Animals↗

Neutrophil stimulation: receptor, membrane, and metabolic events.

In the neutrophil, binding of ligands to their appropriate receptors initiates a sequence of events culminating in the physiological responses of aggregation, degranulation, and superoxide anion generation. Calcium has been proposed as a second messenger in the activation sequence of the neutrophil. Increments in cytosolic free calcium are one of the first measurable events subsequent to receptor occupancy, followed by enhanced plasmalemmal permeability to calcium, a process that may serve to enhance the physiological responses. In contrast to calcium, cyclic AMP (cAMP) does not act as a signal in the activation sequence of the neutrophil. Increments in cAMP that are triggered by complete secretagogues may act as an inhibitory feedback mechanism. Protein kinases, both cAMP- and calcium/phospholipid-sensitive enzymes, may play a role in the activation sequence. Phosphorylation of proteins occurs during neutrophil activation. A role for phosphatidylinositol/phosphatidic acid turnover in calcium gating has been proposed. In addition, modulation of phospholipids could serve to activate a protein kinase C. Finally, phospholipids can serve as a source for arachidonic acid, which is metabolized by a 5-lipoxygenase pathway in the neutrophil. Products of this pathway, such as leukotriene B4, may serve to mediate or modulate the activation sequence.

Animals↗

Transcriptional and translational analysis of the human theta globin gene.

The human theta-globin gene in man appears to be functional, based on its sequence and evolutionary conservation. However its physiological role is unknown and furthermore its deletion in some individuals appears to have no effect on erythroid development. We have therefore analysed the transcriptional and translational competence of the theta globin gene to assess whether or not it is a silent or active globin gene. First, we demonstrate that theta globin mRNA is correctly spliced, by sequencing its cDNA. Second, using this theta cDNA, we generated synthetic theta globin mRNA and were able to demonstrate that this mRNA is translated into theta globin protein in wheat germ in vitro translation extracts. Similarly, the theta globin gene transfected into an erythroid cell line produces a protein product that comigrates with theta globin. Finally, we analysed the unusual promoter of the theta globin gene. The GC rich sequence directly adjacent to the multiple cap sites of theta globin mRNA functions as a promoter element in both erythroid and non-erythroid cell lines, while the more usual CCAAT and ATA box regions (found in all other globin genes) which are displaced by the GC rich promoter sequence, do not possess detectible promoter activity. Taken together, these results suggest that theta globin may have some as yet undetermined role in human erythropoiesis.

Amino Acid Sequence↗

Immunopharmacological and antitumor effects of second-generation immunomodulatory oligonucleotides containing synthetic CpR motifs.

Bacterial and synthetic DNAs containing CpG dinucleotides in specific sequence contexts (CpG DNA) activate the vertebrate immune system and produce potent Th1 immune responses. Recently, we reported immunomodulatory oligonucleotides (IMOs) containing 3'-3'-attached novel structures (immunomers) and synthetic immunostimulatory CpR (R=2'-deoxy-7-deazguanosine) dinucleotides show potent stimulatory activity with distinct cytokine secretion profiles. In the present study, we evaluated in vivo immunopharmacological and antitumor properties of second-generation immunomodulatory oligonucleotides (IMOs) either alone or in combination with chemotherapeutic agents. Repeated peritumoral administration of IMOs at 1 mg/kg to mice bearing established subcutaneous CT26 colon tumor or B16.F0 melanoma resulted in complete regression or strong inhibition of tumor growth. Direct peritoneal injection of IMOs at 2.5 mg/kg to mice bearing peritoneally implanted ascites CT26 or B16.F0 tumors completely eradicated or inhibited tumor growth. Treatment of mice bearing beta-gal expressing CT26.CL25 tumor with IMOs resulted in a significant tumor-specific CTL responses compared with treatment with a control non-CpG DNA or PBS. These responses correlated with IFN-gamma, but not IL-4 secreted in IMO, treated mice. A 5-fold increase in beta-gal specific IgG2a antibodies was found in mice, significantly increasing the IgG2a/IgG1 ratio. IMOs showed similar antitumor activity in both wt and IL-6 knockout (ko) C57BL/6 mice but failed to elicit activity in IL-12 ko C57BL/6 mice. Tumor-free mice from the IMO treatment group rejected the same tumor cell rechallenge, suggesting an adaptive immune response against these cells. Moreover, naïve mice quickly developed specific antitumor response without IMO treatment following adoptive transfer of splenocytes obtained from tumor free mice from the IMO treated group. Additionally, the co-administration of IMOs with chemotherapeutic agents, docetaxel and doxorubicin, resulted in synergistic antitumor effects in both B16.F0 melanoma and 4T1 breast carcinoma models. These results demonstrate potent antitumor activity of second-generation IMO compounds containing a synthetic CpR stimulatory motif in a broad spectrum of tumor models through induction of strong Th1 immune responses. IMO treatment resulted in the development of tumor-specific memory immune responses. No treatment-related toxicity was observed in mice at the doses and treatment schedules studied.

Adoptive Transfer↗

Nucleotide sequence-based multitarget identification.

MULTIGEN technology (T. Vinayagamoorthy, U.S. patent 6,197,510, March 2001) is a modification of conventional sequencing technology that generates a single electropherogram consisting of short nucleotide sequences from a mixture of known DNA targets. The target sequences may be present on the same or different nucleic acid molecules. For example, when two DNA targets are sequenced, the first and second sequencing primers are annealed to their respective target sequences, and then a polymerase causes chain extension by the addition of new deoxyribose nucleotides. Since the electrophoretic separation depends on the relative molecular weights of the truncated molecules, the molecular weight of the second sequencing primer was specifically designed to be higher than the combined molecular weight of the first sequencing primer plus the molecular weight of the largest truncated molecule generated from the first target sequence. Thus, the series of truncated molecules produced by the second sequencing primer will have higher molecular weights than those produced by the first sequencing primer. Hence, the truncated molecules produced by these two sequencing primers can be effectively separated in a single lane by standard gel electrophoresis in a single electropherogram without any overlapping of the nucleotide sequences. By using sequencing primers with progressively higher molecular weights, multiple short DNA sequences from a variety of targets can be determined simultaneously. We describe here the basic concept of MULTIGEN technology and three applications: detection of sexually transmitted pathogens (Neisseria gonorrhoeae, Chlamydia trachomatis, and Ureaplasma urealyticum), detection of contaminants in meat samples (coliforms, fecal coliforms, and Escherichia coli O157:H7), and detection of single-nucleotide polymorphisms in the human N-acetyltransferase (NAT1) gene (S. Fronhoffs et al., Carcinogenesis 22:1405-1412, 2001).

Animals↗

Sequence-specific and/or stereospecific constraints of the U3 enhancer elements of MCF 247-W are important for pathogenicity.

The oncogenic potential of many nonacute retroviruses is dependent on the duplication of the enhancer sequences present in the unique 3' (U3) region of the long terminal repeat (LTR). In a molecular clone (MCF 247-W) of the murine leukemia virus MCF 247, a leukemogenic mink cell focus-inducing (MCF) virus, the U3 enhancer sequences are tandemly repeated in the LTR. We mutated the enhancer region of MCF 247-W to test the hypothesis that the duplicated enhancer sequences of this virus have a sequence-specific and/or a stereospecific role in enhancer function required for transformation. In one virus, we inserted 14 nucleotide bp into the novel sequence generated at the junction of the two enhancers to generate an MCF virus with an interrupted enhancer region. In the second virus, only one copy of the enhancer sequences was present. This second virus also lacked the junction sequence present between the two enhancers of MCF 247-W. Both viruses were less leukemogenic and had a longer mean latency period than MCF 247-W. These data indicate that the sequence generated at the junction of the two enhancers and/or the stereospecific arrangement of the two enhancer elements are required for the full oncogenic potential of MCF 247-W. We analyzed proviral LTRs within the c-myc locus in tumor DNAs from mice injected with the MCF virus with the interrupted enhancer region. Some of the proviral LTRs integrated upstream of c-myc contain enhancer regions that are larger than those of the injected virus. These results are consistent with the suggestion that the virus with an interrupted enhancer changes in vivo to perform its role in the transformation of T cells.

Animals↗

Cleavable CD40Ig fusion proteins and the binding to sgp39.

Recombinant immunoglobulin (Ig) fusion proteins of cell surface and intracellular proteins have wide applications. For example, fusion proteins have been used in the isolation, identification and study of ligands and the effects of binding or blocking a receptor-ligand pair, either in vivo or in vitro. For some applications, removal of the immunoglobulin Fc region is advantageous. We have developed two vectors for the expression of Ig fusion proteins that contain recognition sequences for protease cleavage using thrombin. In one vector, the sequence encoding the thrombin cleavage site is located at the junction of the DNA fragment encoding the protein or protein fragment to be studied and the hinge and constant regions of the immunoglobulin, allowing the generation of a monomeric form of the protein of interest. In the second vector, the sequence encoding the thrombin cleavage site is located between the sequences encoding the hinge and constant regions of the immunoglobulin, allowing for the generation of covalent dimers of the recombinant protein without the constant Fc domains. We have used these vectors to produce the constructs encoding two forms of the extracellular domain of CD40, CD40ThrIg and CD40HinThrIg, allowing production of a monomeric and dimeric form of recombinant CD40. Cleavage is efficient and complete. Following cleavage, there was no detectable binding of the monomeric form of CD40 to a soluble form of gp39, the ligand of CD40, while the dimeric form was able to bind. These vectors have been constructed to allow facile substitution with other sequences to generate cleavable forms of other proteins of interest.

Base Sequence↗

Design and characterization of alpha-melanotropin peptide analogs cyclized through rhenium and technetium metal coordination.

alpha-Melanocyte stimulating hormone (alpha-MSH) analogs, cyclized through site-specific rhenium (Re) and technetium (Tc) metal coordination, were structurally characterized and analyzed for their abilities to bind alpha-MSH receptors present on melanoma cells and in tumor-bearing mice. Results from receptor-binding assays conducted with B16 F1 murine melanoma cells indicated that receptor-binding affinity was reduced to approximately 1% of its original levels after Re incorporation into the cyclic Cys4,10, D-Phe7-alpha-MSH4-13 analog. Structural analysis of the Re-peptide complex showed that the disulfide bond of the original peptide was replaced by thiolate-metal-thiolate cyclization. A comparison of the metal-bound and metal-free structures indicated that metal complexation dramatically altered the structure of the receptor-binding core sequence. Redesign of the metal binding site resulted in a second-generation Re-peptide complex (ReCCMSH) that displayed a receptor-binding affinity of 2.9 nM, 25-fold higher than the initial Re-alpha-MSH analog. Characterization of the second-generation Re-peptide complex indicated that the peptide was still cyclized through Re coordination, but the structure of the receptor-binding sequence was no longer constrained. The corresponding 99mTc- and 188ReCCMSH complexes were synthesized and shown to be stable in phosphate-buffered saline and to challenges from diethylenetriaminepentaacetic acid (DTPA) and free cysteine. In vivo, the 99mTcCCMSH complex exhibited significant tumor uptake and retention and was effective in imaging melanoma in a murine-tumor model system. Cyclization of alpha-MSH analogs via 99mTc and 188Re yields chemically stable and biologically active molecules with potential melanoma-imaging and therapeutic properties.

Amino Acid Sequence↗

Overexpression of BsoBI restriction endonuclease in E. coli, purification of the recombinant BsoBI, and identification of catalytic residues of BsoBI by random mutagenesis.

BsoBI is a type II restriction enzyme found in Bacillus stearothermophilus JN209 that recognizes the symmetric sequence 5'-CYCGRG-3' (Y=C or T; R=A or G) and cleaves between the first and second base to generate a four-base 5' extension. The cloning and sequencing of BsoBI restriction-modification system has been described by Ruan et al. [Mol. Gen. Genet. 252 (1996) 695-699]. Here we report the overexpression of BsoBI restriction endonuclease gene in E. coli by insertion of the endonuclease gene into an expression vector pRRS. The recombinant BsoBI was purified to homogeneity and its N-terminus sequence was determined. It has the same N-terminal aa sequence as the native enzyme. The constitutive expression of BsoBI from pRRS is lethal to E. coli in the absence of the cognate methylase. The bsoBIR gene was mutagenized with either hydroxylamine or by error-prone polymerase chain reaction in vitro and transferred into E. coli via plasmid vectors in the absence of the cognate methylase. Surviving transformants were selected that carry BsoBI variants which lost endonuclease activity. DNA sequencing of the mutant alleles revealed that G123, D124, D212, D246, E252 and H253 are important residues for enzymatic activity. An electrophoretic mobility shift assay was used to identify binding-proficient and cleavage-deficient variants. Seven variants I95M&D124Y, G123R, D212N, K207R&D212V, D246N, D246G and E252K can still bind DNA despite the loss of cleavage activity. Thus, residues D124, D212, D246 and E252 may be located near or within the catalytic center, and are likely involved in metal ion binding.

Amino Acid Sequence↗

Erythropoietin structure-function relationships: high degree of sequence homology among mammals.

To investigate structure-function relationships of erythropoietin (Epo), we have obtained cDNA sequences that encode the mature Epo protein of a variety of mammals. A first set of primers, corresponding to conserved nucleotide sequences between mouse and human DNAs, allowed us to amplify by polymerase chain reaction (PCR) intron 1/exon 2 fragments from genomic DNA of the hamster, cat, lion, dog, horse, sheep, dolphin, and pig. Sequencing of these fragments permitted the design of a second generation of species-specific primers. RNA was prepared from anemic kidneys and reverse-transcribed. Using our battery of species-specific 5' primers, we were able to successfully PCR-amplify Epo cDNA from Rhesus monkey, rat, sheep, dog, cat, and pig. Deduced amino acid sequences of mature Epo proteins from these animals, in combination with known sequences for human, Cynomolgus monkey, and mouse, showed a high degree of homology, which explains the biologic and immunological cross-reactivity that has been observed in a number of species. Human Epo is 91% identical to monkey Epo, 85% to cat and dog Epo, and 80% to 82% to pig, sheep, mouse, and rat Epos. There was full conservation of (1) the disulfide bridge linking the NH2 and COOH termini; (2) N-glycosylation sites; and (3) predicted amphipathic alpha-helices. In contrast, the short disulfide bridge (C29/C33 in humans) is not invariant. Cys33 was replaced by a Pro in rodents. Most of the amino acid replacements were conservative. The C-terminal part of the loop between the C and D helices showed the most variation, with several amino acid substitutions, deletions, and/or insertions. Calculations of maximum parsimony for intron 1/exon 2 sequences as well as coding sequences enabled the construction of cladograms that are in good agreement with known phylogenetic relationships.

Animals↗

Heterodyned fifth-order two-dimensional IR spectroscopy: third-quantum states and polarization selectivity.

A heterodyned fifth-order two-dimensional (2D) IR spectrum of a model coupled oscillator system, Ir(CO)(2)(C(5)H(7)O(2)), is reported. The spectrum is generated by a pulse sequence that probes the eigenstate energies up to the second overtone and combination bands, providing a more rigorous potential-energy surface of the coupled carbonyl local modes than can be obtained with third-order spectroscopy. Furthermore, the pulse sequence is designed to generate and then rephase a two-quantum coherence so that the spectrum is line narrowed and the resolution improved for inhomogeneously broadened systems. Features arising from coherence transfer processes are identified, which are more pronounced than in third-order 2D IR spectroscopy because the transition dipoles of the second overtone and combination states are not rigorously orthogonal, relaxing the polarization constraints on the signal intensity for these features. The spectrum provides a stringent test of cascading signals caused by third-order emitted fields and no cascading is observed. In the Appendix, formulas for calculating the signal intensities for resonant fifth-order spectroscopies with arbitrarily polarized pulses and transition dipoles are reported. These relationships are useful for interpreting and designing polarization conditions to enhance specific spectral features.

Journal Article↗

Immunogenicity of HIV type 1 gp120 CD4 binding site phage mimotopes.

The conserved domain of the CD4 binding site (CD4bs) on the human immunodeficiency virus type 1 (HIV- 1) envelope represents a potential target for vaccine development. Here we describe selection of peptide mimotopes by panning a phage peptide library on the HIV-1 CD4bs-specific, broadly neutralizing anti-HIV-1 monoclonal antibody, IgG(1) b12. We identified an initial consensus sequence for IgG1 b12 binding (M/VThetaSD, where Theta represents an aromatic amino acid). A molecular evolution approach, using second- and third-generation libraries, led us to identify a refined consensus sequence (GLLVWSDEL). The resulting IgG1 b12 phage mimotopes compete with gp160 for the IgG1 b12 antigen-binding site, but the phage coat protein (pIII) may play an important structural role, since both free peptides and KLH-conjugated peptides have no detectable binding activity. Mice immunized with IgG1 b12 phage mimotopes elicited a weak but persistent humoral response directed against the HIV-1 envelope. An antibody fragment was isolated from the antibody repertoires of these animals. It is noteworthy that while it has a relatively low affinity for HIV-1 gp160, the antibody targets an epitope that overlaps with that of IgG1 b12. Our data therefore suggest that engineered IgG1 b12 mimotopes share immunogenic features with the CD4bs. However, these peptidic structures will require further improvement in order to generate broad specificity neutralizing antibodies like IgG1 b12.

Amino Acid Sequence↗

Prediction of protein subcellular localization.

Because the protein's function is usually related to its subcellular localization, the ability to predict subcellular localization directly from protein sequences will be useful for inferring protein functions. Recent years have seen a surging interest in the development of novel computational tools to predict subcellular localization. At present, these approaches, based on a wide range of algorithms, have achieved varying degrees of success for specific organisms and for certain localization categories. A number of authors have noticed that sequence similarity is useful in predicting subcellular localization. For example, Nair and Rost (Protein Sci 2002;11:2836-2847) have carried out extensive analysis of the relation between sequence similarity and identity in subcellular localization, and have found a close relationship between them above a certain similarity threshold. However, many existing benchmark data sets used for the prediction accuracy assessment contain highly homologous sequences-some data sets comprising sequences up to 80-90% sequence identity. Using these benchmark test data will surely lead to overestimation of the performance of the methods considered. Here, we develop an approach based on a two-level support vector machine (SVM) system: the first level comprises a number of SVM classifiers, each based on a specific type of feature vectors derived from sequences; the second level SVM classifier functions as the jury machine to generate the probability distribution of decisions for possible localizations. We compare our approach with a global sequence alignment approach and other existing approaches for two benchmark data sets-one comprising prokaryotic sequences and the other eukaryotic sequences. Furthermore, we carried out all-against-all sequence alignment for several data sets to investigate the relationship between sequence homology and subcellular localization. Our results, which are consistent with previous studies, indicate that the homology search approach performs well down to 30% sequence identity, although its performance deteriorates considerably for sequences sharing lower sequence identity. A data set of high homology levels will undoubtedly lead to biased assessment of the performances of the predictive approaches-especially those relying on homology search or sequence annotations. Our two-level classification system based on SVM does not rely on homology search; therefore, its performance remains relatively unaffected by sequence homology. When compared with other approaches, our approach performed significantly better. Furthermore, we also develop a practical hybrid method, which combines the two-level SVM classifier and the homology search method, as a general tool for the sequence annotation of subcellular localization.

Algorithms↗