PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “second generation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

A novel second-generation HIV-1 circulating recombinant form (CRF183_0107) identified among men who have sex with men in China.

OBJECTIVE: This study aimed to report a novel HIV-1 circulating recombinant form (CRF) identified among men who have sex with men (MSM) in China. DESIGN: Viral sequences were isolated from MSM patients, and the recombination and evolutionary histories of this CRF were elucidated through phylogenetic and Bayesian analyses. METHODS: Near full-length genomes (NFLGs) and partial genome sequences were amplified from RNA extracted from plasma samples of three HIV-1 seropositive MSM in Heilongjiang Province, China. Phylogenetic analysis was conducted using FastTree v2.1.9, and recombination analysis was performed using Simplot v3.5.1. The emergence time of the novel CRF was estimated by Bayesian evolutionary analysis using BEAST v1.10.4. RESULTS: Two NFLGs and two partial genome segments were successfully obtained from three MSM participants. This novel CRF was characterized by 12 mosaic gene segments, comprising 6 segments from the CRF01_AE cluster 4 and 6 segments from the CRF07_BC cluster N, and thus was designated as CRF183_0107. The estimated time of origin for the CRF01_AE and CRF07_BC components within CRF183_0107 were approximately 2009.2 and 2012.5, respectively. CONCLUSION: A novel second-generation HIV-1 recombinant, named CRF183_0107, was identified within the MSM population in China. This CRF exemplified the recombination events occurring between the CRF01_AE cluster 4 and the CRF07_BC cluster N during 2009-2012.

Humans↗

Subclustering of human immunoglobulin kappa light chain variable region genes.

The human immunoglobulin kappa light chain (IgK) locus includes multiple variable region gene segments (Vk) that can be divided into four subgroups. Oligonucleotide primers were designed to amplify specifically gene segments of the VkI, VkII, and VkIII subgroups using the polymerase chain reaction (PCR). Product sequences were subcloned, sequenced, and compared. Phylogenetic analyses of sequences within each subgroup indicate that some subgroups can be subdivided further into "sub-subgroups." The history of Vk segment duplications apparently includes at least two separate periods, the first giving rise to the subgroups and the second generating further complexity within each subgroup. Duplications of large pieces of DNA (demonstrated by others through pulsed-field gel electrophoresis) also played a role. Rates of synonymous and non-synonymous base changes between pairs of sequences suggest that natural selection has played a major role in the evolution of the Vk variable gene segments, leading to sequence conservation in some regions and to increased diversity in others.

Base Sequence↗

Expression and antigenecity of human immunodeficiency virus type-1 transmembrane protein gp41 in insect cells.

The HIV-1 transmembrane protein, gp41, is processed together with the envelope glycoprotein, gp120, from the same precursor, gp160, during the virus maturation. We used a baculovirus expression system to demonstrate that gp41 could be properly expressed without the preceding gp120 sequence. Two constructs with slight differences in the N-terminal region of gp41 were generated: one with a deletion of the first 7 hydrophobic residues of gp41, which have been suggested to be in a region important for membrane fusion and penetration, whereas the second with a complete sequence of gp41 except that a nonconserved leucine was substituted with a glutamine during DNA manipulation. Results from Western blotting with specific antisera confirm the gp41 identity. The sizes of gp41 were sensitive to tunicamycin treatment, indicating that N-linked glycosylation did occur. Further immunoblotting analyses with 90 different serum samples from HIV-1-infected individuals gave similar reaction patterns, suggesting that gp120 as well as the N-terminal region of gp41 are not critical for the expression and antigenecity of gp41. These eucaryotic constructs should provide valuable gp41 sources for detailed characterization of gp41 functions.

Amino Acid Sequence↗

Genetic analysis of HIV-1 strains in rural eastern Cameroon indicates the evolution of second-generation recombinants to circulating recombinant forms.

The HIV-1 genetic diversity in most parts of Cameroon is well described and shown to be very broad. However, little is known about the composition of the HIV-1 epidemic in the rural parts of eastern Cameroon. Therefore, we investigated 25 specimens from this region for their subtypes in gag, pol, and env gene fragments. Along with genetic material of subtypes A1, C, G, CRF01_AE, CRF02_AG, and CRF11_cpx, we also identified a large number (24%, 6/25) of distinct env sequences within the subtype A radiation. CRF02_AG was the predominant genetic form in all genes studied. Half of the specimens studied were considered "pure" based on concordant subtypes in the genes studied, whereas the other half were unique recombinant forms (URFs). Except for 1 URF, all were second-generation recombinants (SGRs), 90% of which contained genetic material of CRF02_AG in at least 1 gene. Notably, we identified individuals from 3 different villages infected with CRF01_AE(gag)CRF02_AG(pol)A(env) strains, which is indicative of the evolution of this URF to a circulating recombinant form (CRF). In addition, we identified a CRF02_AG(pol)C(env) recombinant infecting a man and a woman living in the same village, suggesting horizontal transmission of this recombinant. The current study emphasizes the power of HIV-1 recombination through the generation of SGRs and the evolution of URFs into CRFs. These findings suggest that, in a region where a predominant HIV-1 strain cocirculates among several subtypes, recombination could eventually decrease the proportion of this strain over time, such as CRF02_AG in Cameroon.

Cameroon↗

Characterization of the structure and melting of DNAs containing backbone nicks and gaps.

A DNA molecule containing a gap (a missing phosphate) has been examined and compared to two other molecules of the same sequence, one containing a nick (a phosphorylated gap) and the other a normal duplex containing no break in the backbone. A second gapped sequence was also compared to a normal duplex of the same sequence. The molecules containing nicks or gaps were generated as dumbbell molecules, short helices closed by a loop at each end. The dumbbells were formed by the association of two hairpins with self-complementary dangling 5'-ends. Nuclear magnetic resonance was used to monitor the melting transition and to probe structural differences between molecules. Under the conditions used here no change in stability was observed upon phosphorylation of the gap. Structural changes upon phosphorylation of a gap or closure of a nick were minimal and were localized to the region immediately around the gap or nick. Two transitions can be observed as a gapped or nicked molecule melts, although the resolution of the two transitions varies with the salt concentration. At moderate to high salt (greater than or equal to 30 mM) the molecule melts essentially all at once. At low salt the two transitions occur at temperatures that differ by as much as 15 degrees C. In addition, comparison with other NMR melting studies indicates that the duplex formed by the overlap of the dangling ends of the hairpins is stabilized relative to a free duplex of the same sequence, probably by stacking onto the hairpin stem.

Base Composition↗

Identification of two splice variant forms of type-IVB cyclic AMP phosphodiesterase, DPD (rPDE-IVB1) and PDE-4 (rPDE-IVB2) in brain: selective localization in membrane and cytosolic compartments and differential expression in various brain regions.

In order to detect the two splice variant forms of type-IVB cyclic AMP phosphodiesterase (PDE) activity, DPD (type-IVB1) and PDE-4 (type-IVB2), anti-peptide antisera were generated. One set ('DPD/PDE-4-common'), generated against a peptide sequence found at the common C-terminus of these two PDEs, detected both PDEs. A second set was PDE-4 specific, being directed against a peptide sequence found within the unique N-terminal region of PDE-4. In brain, DPD was found exclusively in the cytosol and PDE-4 exclusively associated with membranes. Both brain DPD and PDE-4 activities, isolated by immunoprecipitation, were cyclic AMP-specific (KmcyclicAMP: approximately 5 microM for DPD; approximately 4 microM for PDE-4) and were inhibited by low rolipram concentrations (K1rolipram approximately 1 microM for both). Transient expression of DPD in COS-1 cells allowed identification of an approx. 64 kDa species which co-migrated on SDS/PAGE with the immunoreactive species identified in both brain cytosol and membrane fractions using the DPD/PDE-4-common antisera. The subunit size observed for PDE-4 (approx. 64 kDa) in brain membranes was similar to that predicted from the cDNA sequence, but that observed for DPD was approx. 4 kDa greater. Type-IV, rolipram-inhibited PDE activity was found in all brain regions except the pituitary, where it formed between 30 and 70% of the PDE activity in membrane and cytosolic fractions when assayed with 1 microM cyclic AMP, PDE-4 formed 40-50% of the membrane type-IV activity in all brain regions save the midbrain (approx. 20%). DPD distribution was highly restricted to certain regions, providing approx. 35% of the type-IV cytosolic activity in hippocampus and 13-21% in cortex, hypothalamus and striatum with no presence in brain stem, cerebellum, midbrain and pituitary. The combined type-IVB PDE activities of DPD and PDE-4 contributed approx. 10% of the total PDE activity in most brain regions except for the pituitary (zero) and the mid-brain (approx. 3%. The isolated cDNAs for DPD and PDE-4 appear to reflect transcription products which are expressed in vivo in brain. The unique N-terminal domain of PDE-4 is suggested to target this PDE to membranes in brain. Type-IVB PDEs are differentially expressed in various brain regions, indicating that there are tissue-specific controls on both the expression of the gene and the splicing of its products.

3',5'-Cyclic-AMP Phosphodiesterases↗

A coalescent approach to the polymerase chain reaction.

A versatile algorithm is developed to model PCR on a computer. The method is based on a modification of the coalescent process and provides a general framework to analyse data from PCR. It allows for incorporation of the dynamics of the replication process as described in terms of the number of starting template molecules and cycle-dependent PCR efficiency. The simulation method generates, as a first step, the genealogy of a set of sequences sampled from a final PCR product. In a second step a mutation process is superimposed and the resulting data set is analysed. The efficiency of our algorithm enables us to get reliable approximations of various sample distributions. We demonstrate the relevance of our method with two applications: maximum likelihood estimation of the error rate in PCR and a test of homogeneity of the template.

Algorithms↗

Evaluation of arrayed primer extension for TP53 mutation detection in breast and ovarian carcinomas.

Mutations in the tumor suppressor gene TP53 are associated with a wide range of different cancers and may have prognostic and therapeutic implications. Methods for rapid and sensitive detection of mutations in this gene are therefore required. In order to make screening more effective, a commercially available TP53 genotyping microarray from Asper Biotech has been constructed by arrayed primer extension (APEX). The present study is the first report that blindly evaluates the efficiency of the second generation APEX TP53 genotype chip outside the Asper laboratory and compares it to temporal temperature gradient electrophoresis (TTGE) and sequencing of TP53 for mutation detection in ovarian and breast cancer samples. All nucleotides in the TP53 gene from exon 2-9 are included on the chip by synthesis and application of sequence-specific oligonucleotides. The chip was validated by screening 48 breast and 11 ovarian cancer cases, all of which had previously been analyzed by TTGE and sequencing. APEX scored 17 of 20 sequence variants, missing one deletion, one insertion, and a missense mutation. Resequencing efficiency using APEX was 92% for both DNA strands and 99.5% for sense and/or antisense strand. We conclude that the APEX TP53 microarray is a robust, rapid, and comprehensive screening tool for sequence alterations in tumors.

Alleles↗

Apolipoprotein B(Arg3500----Gln) allele specific polymerase chain reaction: large-scale screening of pooled blood samples.

A two-step polymerase chain reaction (PCR) method for the rapid detection of the apolipoprotein B(Arg3500----Gln) mutation in a mixture of pooled blood samples is described. In the first step PCR, a short gene fragment surrounding codon 3500 is amplified. Subsequently the reaction product is subjected to a second amplification in which a mutation-specific primer is used. A PCR product is generated only if the mutant sequence is present in the DNA pool. Individuals carrying the mutation can then be identified by PCR with mutagenic primers and MspI restriction typing, essentially as described by Hansen et al. (J. Lipid Res. 1991. 32: 1229-1233).

Alleles↗

The G protein beta 2 complementary DNA encodes the beta 35 subunit.

Antisera were generated against synthetic peptides that correspond to amino acid sequences deduced from a cDNA (designated beta 2) that encodes a second form of the beta subunit of guanine nucleotide-binding regulatory proteins (G proteins). The specificity of interactions of these antisera with purified G protein beta subunits indicates that the beta 2 cDNA encodes the beta 35 form of this polypeptide. This hypothesis is confirmed by the use of these antisera to detect expression of the beta 2 cDNA in COS-m6 cells.

Base Sequence↗

Formation of the 3' end of U1 snRNA is directed by a conserved sequence located downstream of the coding region.

U1 is a small non-polyadenylated nuclear RNA that is transcribed by RNA polymerase II and is known to play a role in mRNA splicing. The mature 3' end of U1 snRNA is formed in at least two steps. The first step generates precursors of U1 RNA with a few extra nucleotides at the 3' end; in the second step, these precursors are shortened to mature U1 RNA. Here, I have determined the sequences required for the first step. Human U1 genes with various deletions and substitutions near the 3' end of the coding region were constructed and introduced into HeLa cells by DNA transfection. The structure of the RNA synthesized during transient expression of the exogenous U1 gene was analyzed by S1 mapping. The results show that a 13 nucleotide sequence located downstream from the U1 coding region and conserved among U1, U2 and U3 genes of different species is the only sequence required to direct the first step in the formation of the 3' end of U1 snRNA.

Animals↗

Infrared fluorescent detection of PCR amplified gender identifying alleles.

An automated DNA sequencer utilizing high sensitivity infrared (IR) fluorescence technology together with Polymerase Chain Reaction (PCR) methodology was used to detect several sex differentiating loci on the X and Y chromosomes from various samples often encountered in forensic case work. Amplifications of the X-Y homologous amelogenin gene, the alpha-satellite (alphoid) repeat sequences and the X and Y chromosome zinc finger protein genes ZFX and ZFY (ZFX/ZFY) were performed. DNA extracted from various forensic specimens was amplified using either Taq, Tth or ThermoSequenase. Multiplexing using primers for all three loci in one reaction tube was achieved using Tth and ThermoSequenase. Two IR labeling strategies for detection of PCR products were utilized. In the first strategy, one of the PCR primers contained a 19-base extension at its 5' end identical to an IR-labeled universal M13 Forward (-29) primer which was included in the amplification reactions. During PCR the tailed primer generates sequence complementary to the M13 primer which subsequently primes the initial amplification products, thereby generating IR-labeled PCR products. In the second strategy, dATP labeled with an IR dye (IR-dATP) was included in the amplification reaction. During amplification IR-dATP was utilized by the polymerase and incorporated into the synthesized DNA, thus resulting in IR-labeled PCR products. X and Y specific bands were readily detected using both labeling methodologies. Amplified products were electrophoretically resolved using denaturing Long-Ranger gels and detected with an automated detection system using IR laser irradiation. A separation distance of 15 cm allowed run times of less than 2 h from sample loading to detection. Because the gels could be run more than once, at least 120 samples (2 loads x 60 samples/load) can be typed using a single gel.

Alleles↗

Real value prediction of solvent accessibility in proteins using multiple sequence alignment and secondary structure.

The present study is an attempt to develop a neural network-based method for predicting the real value of solvent accessibility from the sequence using evolutionary information in the form of multiple sequence alignment. In this method, two feed-forward networks with a single hidden layer have been trained with standard back-propagation as a learning algorithm. The Pearson's correlation coefficient increases from 0.53 to 0.63, and mean absolute error decreases from 18.2 to 16% when multiple-sequence alignment obtained from PSI-BLAST is used as input instead of a single sequence. The performance of the method further improves from a correlation coefficient of 0.63 to 0.67 when secondary structure information predicted by PSIPRED is incorporated in the prediction. The final network yields a mean absolute error value of 15.2% between the experimental and predicted values, when tested on two different nonhomologous and nonredundant datasets of varying sizes. The method consists of two steps: (1) in the first step, a sequence-to-structure network is trained with the multiple alignment profiles in the form of PSI-BLAST-generated position-specific scoring matrices, and (2) in the second step, the output obtained from the first network and PSIPRED-predicted secondary structure information is used as an input to the second structure-to-structure network. Based on the present study, a server SARpred (http://www.imtech.res.in/raghava/sarpred/) has been developed that predicts the real value of solvent accessibility of residues for a given protein sequence. We have also evaluated the performance of SARpred on 47 proteins used in CASP6 and achieved a correlation coefficient of 0.68 and a MAE of 15.9% between predicted and observed values.

Amino Acids↗

High-throughput generation of sequence indexes from T-DNA mutagenized Arabidopsis thaliana lines.

A pipeline has been created for the characterization of Arabidopsis thaliana mutants by generating flanking sequence tags (FSTs) and optimized for economic, high-throughput production. The GABI-Kat collection of T-DNA mutagenized A. thaliana plants was used as a source of independent transgenic lines. The pipeline included robotized extraction of genomic DNA in a 96-well format, an adapter-ligation PCR method for amplification of plant sequences adjacent to T-DNA borders, automated purification and sequencing of PCR products, and computational trimming of the resulting sequence files. Data quality was significantly improved by (i) restriction digestion of the adaptor-ligation products to reduce trivial sequences caused by co-amplification of fragments derived from the free plasmid, and (ii) the design of the adaptor primers for the second amplification step to enhance selective generation of single PCR fragments, even from lines with multiple T-DNA insertions. Gel-purification was avoided by including these steps, the number of amplification reactions per line was reduced from four to three, and the percentage of lines that yielded at least one FST was increased from 66% to 86%. More than 58,000 FSTs have been submitted to GenBank and are available at http://www.mpiz-koeln.mpg.de/GABI-Kat/.

Arabidopsis↗

Using amino acid patterns to accurately predict translation initiation sites.

The translation initiation site (TIS) prediction problem is about how to correctly identify TIS in mRNA, cDNA, or other types of genomic sequences. High prediction accuracy can be helpful in a better understanding of protein coding from nucleotide sequences. This is an important step in genomic analysis to determine protein coding from nucleotide sequences. In this paper, we present an in silico method to predict translation initiation sites in vertebrate cDNA or mRNA sequences. This method consists of three sequential steps as follows. In the first step, candidate features are generated using k-gram amino acid patterns. In the second step, a small number of top-ranked features are selected by an entropy-based algorithm. In the third step, a classification model is built to recognize true TISs by applying support vector machines or ensembles of decision trees to the selected features. We have tested our method on several independent data sets, including two public ones and our own extracted sequences. The experimental results achieved are better than those reported previously using the same data sets. Our high accuracy not only demonstrates the feasibility of our method, but also indicates that there might be "amino acid" patterns around TIS in cDNA and mRNA sequences.

Algorithms↗

Generating quantitative models describing the sequence specificity of biological processes with the stabilized matrix method.

BACKGROUND: Many processes in molecular biology involve the recognition of short sequences of nucleic-or amino acids, such as the binding of immunogenic peptides to major histocompatibility complex (MHC) molecules. From experimental data, a model of the sequence specificity of these processes can be constructed, such as a sequence motif, a scoring matrix or an artificial neural network. The purpose of these models is two-fold. First, they can provide a summary of experimental results, allowing for a deeper understanding of the mechanisms involved in sequence recognition. Second, such models can be used to predict the experimental outcome for yet untested sequences. In the past we reported the development of a method to generate such models called the Stabilized Matrix Method (SMM). This method has been successfully applied to predicting peptide binding to MHC molecules, peptide transport by the transporter associated with antigen presentation (TAP) and proteasomal cleavage of protein sequences. RESULTS: Herein we report the implementation of the SMM algorithm as a publicly available software package. Specific features determining the type of problems the method is most appropriate for are discussed. Advantageous features of the package are: (1) the output generated is easy to interpret, (2) input and output are both quantitative, (3) specific computational strategies to handle experimental noise are built in, (4) the algorithm is designed to effectively handle bounded experimental data, (5) experimental data from randomized peptide libraries and conventional peptides can easily be combined, and (6) it is possible to incorporate pair interactions between positions of a sequence. CONCLUSION: Making the SMM method publicly available enables bioinformaticians and experimental biologists to easily access it, to compare its performance to other prediction methods, and to extend it to other applications.

Algorithms↗

Toll-like receptor 9: modulation of recognition and cytokine induction by novel synthetic CpG DNAs.

Bacterial and synthetic DNA containing unmethylated 2'-deoxyribo(cytidine-phosphate-guanosine) (CpG) dinucleotides in specific sequence contexts activate the vertebrate innate immune system. A molecular pattern recognition receptor, Toll-like receptor 9 (TLR9), recognizes CpG DNA and initiates the signalling cascade, although a direct interaction between CpG DNA and TLR9 has not been demonstrated yet. TLR9 in different species exhibits sequence specificity. Our extensive structure-immunostimulatory activity relationship studies showed that a number of synthetic pyrimidine (Y) and purine (R) nucleotides are recognized by the receptor as substitutes for the natural nucleotides deoxycytidine and deoxyguanosine in a CpG dinucleotide. These studies permitted development of synthetic YpG, CpR and YpR immunostimulatory motifs, and showed divergent nucleotide motif recognition pattern of the receptor. Surprisingly, we found that synthetic immunostimulatory motifs produce different cytokine induction profiles compared with natural CpG motifs. Importantly, we also found that some of these synthetic immunostimulatory motifs show optimal activity in both mouse and human systems without the need to change sequences, suggesting an overriding of the species-dependent specificity of the receptor by the use of synthetic motifs. In the present paper, we review current understanding of structural recognition and functional modulation of TLR9 receptor by second-generation synthetic CpG DNAs and their potential application as wide-spectrum therapeutic agents.

Animals↗

[Re-amplification of differentially expressed mRNA fragments of head-neck cancers without cloning].

BACKGROUND: mRNA expression of healthy and malignant cells can be compared to each other by employing the "differential display" (DD) technique. Most studies describe sequence analysis of differentially expressed fragments after reamplification by a second round of PCR and subsequent molecular cloning to gain a sufficient amount of DNA for sequencing. The aim of this study was to show whether a sufficient amount of differentially expressed mRNA of squamous cell carcinoma cells of the head and neck region can be generated by PCR alone without cloning steps. MATERIAL AND METHODS: mRNA isolated from cultivated keratinocytes and squamous cell carcinoma cells was reverse transcribed into cDNA which was amplified with PCR. Differentially expressed fragments detected after gel electrophoresis were isolated from the gel and reamplified in a second PCR. The resulting cDNA amounts of the second PCR were suitable for cloning but not for direct sequencing. A third round of PCR with the undiluted final product of the second PCR as template regularly failed. Dilutions of the second PCR products between 1:10 and 1:10(10) were prepared. The third round of PCR was carried out with these various template concentrations. RESULTS: A sufficient amount of differentially expressed fragments for sequencing procedures resulted when dilutions of the second PCR products ranging from 1:10(2) to 1:10(7) were used as templates in the third round of PCR. CONCLUSION: Modifications of PCR parameters provide high DNA copy numbers of differentially expressed mRNA fragments from squamous cell carcinoma cells of the upper aerodigestive tract in amounts that are needed for sequence analysis. This may make it possible to avoid labor-intensive cloning procedures requiring high safety standards.

Base Sequence↗