PubMed HealthSearch

SEARCH · PubMed Health

Results for “second generation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Clinical Utility of Next-Generation Sequencing in Tumors Diagnosed as Lung Squamous Cell Carcinoma: Real-World Data of Diagnostic and Therapeutic Implications.

Lung squamous cell carcinoma (LUSC) is the second most common subtype of non-small cell lung carcinoma (NSCLC), typically associated with a poor prognosis. Unlike lung adenocarcinoma, the application of next-generation sequencing (NGS) in LUSC has lagged because of the long-standing perception of low therapeutic yield, primarily based on highly selected, resected cohorts. We sought to determine the real-world clinical utility of NGS in LUSC. We analyzed an institutional cohort of 576 tumors initially diagnosed as LUSC that underwent NGS profiling. We defined "clinical yield" as either diagnostic reclassification or the identification of a targetable mitogenic alteration. Twenty cases (3.5%) were reclassified, including rediagnosis to cutaneous squamous cell carcinoma, transformed adenocarcinoma (post targeted therapy), and rare entities such as nuclear protein of the testis-rearranged carcinoma and lymphoepithelial carcinoma. Primary mitogenic drivers were identified in 83 cases (14.4% of the total cohort), of which 43 (7.5% of the total cohort) harbored alterations with currently Food and Drug Administration-approved therapies for NSCLC (including KRAS, EGFR, MET, ALK, and ROS1). Overall clinical yield-defined as the sum of diagnostic reclassifications and identification of NSCLC-specific targetable alterations-was 11.0% (63/576). Univariate and multivariate analysis demonstrated that never or light smoking history was the strongest independent predictor of clinical yield, with 57.3% of tumors in this subset being reclassified or harboring a strong driver. Our findings demonstrate that NGS provides significant diagnostic and therapeutic value in a real-world LUSC cohort, challenging the historical premise of low yield. Although clinicodemographic features can help prioritize testing in resource-limited settings, the identification of targetable drivers across all smoking groups supports the universal application of comprehensive NGS for all patients diagnosed with LUSC.

Humans

Influence of cellular sequences on instability of plasmid integration sites in human cells.

To learn more about mechanisms of genome instability in human cells, I investigated DNA sequences that promote high rates of recombination by analyzing rare unstable plasmid integration sites in simian virus 40-transformed human fibroblasts. Previous studies had hypothesized that rearrangement or loss of integrated sequences could be attributed to adjacent cellular DNA. Consistent with this interpretation, a cloned fragment containing both the integrated plasmid and 2.0 kb of adjacent cell DNA from one such unstable integration site in the cell line LM205 demonstrated a much higher incidence of rearrangements when integrated into other chromosome locations than did the original plasmid. To further test this hypothesis, portions of cellular DNA from this region were integrated in duplicate in other locations to determine their ability to promote restriction-fragment-length polymorphism, an indicator of high rates of homologous recombination. Although two types of instability were observed, neither could be attributed solely to the cell sequences being tested in the plasmid. The first type of instability was a transient deletion or amplification of the plasmid DNA soon after integration, which appeared to be a general phenomenon often associated with any type of newly integrated sequence. A second type of instability continued indefinitely for many cell generations, as did that observed in cell line LM205. Because this was rare (one of 78 clones tested), it could not be attributed solely to cell sequences contained within the plasmid. However, the rearrangements in this cell clone occurred exclusively within the cell DNA adjacent to the integration site, again suggesting a role for cis-acting cell sequences in this process. The inability to identify specific cell sequences responsible for instability may therefore indicate that a complex combination of sequences is involved, possibly within both the plasmid and cell DNA.

Base Sequence

Rearrangements in unintegrated retroviral DNA are complex and are the result of multiple genetic determinants.

We used a replication-competent retrovirus shuttle vector based on a DNA clone of the Schmidt-Ruppin A strain of Rous sarcoma virus to characterize rearrangements in circular viral DNA. In this system, circular molecules of viral DNA present after acute infection of cultured cells were cloned as plasmids directly into bacteria. The use of a replication-competent shuttle vector permitted convenient isolation of a large number of viral DNA clones; in this study, over 1,000 clones were analyzed. The circular DNA molecules could be placed into a limited number of categories. Approximately one-third of the rescued molecules had deletions in which one boundary was very near the edge of a long terminal repeat (LTR) unit. Subtle differences in the patterns of deletions in circular DNAs with one versus two copies of the LTR sequence were observed, and differences between deletions emanating from the right and left boundaries of the LTR were seen. A virus with a missense mutation in the region of the pol gene responsible for integration and exhibiting a temperature sensitivity phenotype for replication had a marked decrease in the number of rescued molecules with LTR-associated deletions when infection was performed at the nonpermissive temperature. This result suggests that determinants in the pol gene, possibly in the integration protein, play a role in the generation of LTR-associated deletions. Sequences in a second region of the genome, probably within the viral gag gene, were also found to affect the types of circular viral DNA molecules present after infection. Sequences in this region from different strains of avian sarcoma-leukosis viruses influenced the fraction of circular molecules with LTR-associated deletions, as well as the relative proportion of circular molecules with either one or two copies of the LTR. Thus, the profile of rearrangements in unintegrated viral DNA is complex and dependent upon the nature of sequences in the gag and pol regions.

Animals

Genetic Analysis of Genomic and Methylomic Variation and Identification of Multi-Trait Mutants in Rice Carried on Chang'e-5.

Global food security is facing challenges from population growth to diminishing arable land. Space mutation breeding holds promise for overcoming the variation limitations in conventional breeding; however, the mutagenic effects of the deep-space environment on rice and the transgenerational inheritance patterns of induced variations remain unclear. In this study, rice seeds carried by the Chang'e-5 spacecraft were used as materials. Whole-genome sequencing and whole-genome bisulfite sequencing were performed on the first (SP1) and second generations (SP2) of space-mutagenized plants after their return to Earth. The results showed that the number of genomic variants in the SP2 generation increased significantly compared with SP1, and SNPs, homozygous sites, and variants in coding regions were more heritable. The genome-wide methylation level was elevated in the SP2 generation, and among differentially methylated cytosines, those in the CG context exhibited the highest heritability. Furthermore, large-scale screening for nitrogen efficiency, tolerance to PEG-induced stress, and germination-stage cold resistant mutants was conducted in the SP2 generation, and phenotypic validation was performed in the third generation (SP3). By integrating multi-omics analyses of representative mutants to mine candidate genes, a number of heritable elite mutants were obtained, and seven candidate genes for key traits were identified. This study systematically elucidates the transgenerational inheritance patterns of deep-space-induced variation in rice. The multi-trait mutants obtained provide valuable germplasm resources for gene cloning and breeding applications in rice.

DNA methylation

[Sex differences in memory performance for odors, tone sequences and colors].

Sense of olfaction would seem to be of little importance for human behavior. However, a closer look at this from the psychological point of view reveals many interesting aspects, such as sex differences in olfactory perception, that are of interest to differential psychology. The present study deals with sex differences in the memory for odors; we assume that women will do better here than men while other memory tasks involving acoustical and optical stimuli will show no such differences. Sixty women and 40 men were examined. On the first day, they had to retain 10 odors, 10 random-generated tone-sequences, and 10 colors. On the second day, 20 such stimuli for each memory task were presented, and the subjects had to remember and to tell which were known and which of them were unknown stimuli to them. A significant advantage in the olfaction memory task was found for women, while acoustical and optical memory scores showed no such differences. This expected finding is discussed in two ways. First, the female advantage might result from phylogenetic sources. Second, it might arise because women in general more often than men seem to deal with olfactory cues, so that they might simply have more experience and therefore the greater chance to score higher in an odor memory task.

Adult

Somatic reversion/suppression in Duchenne muscular dystrophy (DMD): evidence supporting a frame-restoring mechanism in rare dystrophin-positive fibers.

Many Duchenne muscular dystrophy (DMD) patients are known to have rare staining dystrophin-positive fibers, termed "revertants." The precise etiology of these rare fibers is unknown. The most likely explanation, however, is somatic mosaicism or somatic reversion/suppression. Immunocytochemistry was performed on serial sections from deleted and nondeleted patients, with a panel of antibodies--9219, 1377, 9218, and Dys-2--that span dystrophin. Both familial and nonfamilial patients possessed revertants. Either the same clusters or individual revertant fibers stained with amino- and carboxyl-terminal antibodies in all 14 DMD patients. In patients with deletions, revertants did not stain with antibodies raised to polypeptide sequences within the deletion. These results indicate that positively staining fibers are not the result of somatic mosaicism in deleted patients. Five of 10 patients without deletions had revertant fibers. In two of these patients, the revertant fibers did not stain with antibody 9218, which was generated against amino acids 2305-2554 and which corresponds to exons 48-52. The remaining antibodies from the panel stained the same fibers on separate serial sections in these two patients. The most likely mechanism giving rise to these positively staining fibers is a second site in-frame deletion. Antibodies generated to polypeptide sequences within deletions can be used to control for the natural occurrence of revertant fibers in myoblast transfer studies and may be useful in the detection of point mutations.

Antibodies

Molecular determinants of antimicrobial resistance in Klebsiella pneumoniae isolates among geriatric patients in Chattogram, Bangladesh: a cross-sectional study.

Klebsiella pneumoniae (KPN) infections pose heightened risks in the geriatric population due to weakened immunity, prevalent comorbidities, potential exposure in long-term care settings, and increased likelihood of antibiotic resistance (ABR). The study focused on the prevalence and antibiotic resistance of KPN infections, the presence of ABR genes in KPN, and the genomic characterization of KPN obtained from geriatric patients in Chattogram. A total of 543 specimens were collected from four hospitals in Chattogram, along with demographic data from hospital records. Genomic DNA was extracted from multi-drug-resistant (MDR) KPN, and the presence of ABR genes, blaTEM-1, sul-1, aadB, blaNDM-1, blaSHV-11, and phoE was identified. To characterize the KPN genomes, two MDR KPN isolates were subjected to whole-genome sequencing (WGS), and the data were analyzed using bioinformatics tools to identify genomic determinants of ABR. KPN exhibited high resistance to ceftazidime (96%), cefuroxime (92%), and cefixime (83%), but sensitivity to colistin (79%) and amikacin (75%). MDR KPN was mostly detected in sputum (36%) and urine (27%) specimens, where the prevalence of ABR genes, blaTEM-1, sul1, aadB, blaNDM-1, and blaSHV-11 were 28.2%, 17%, 6.17%, 56%, and 48% of these strains, respectively. These genomes exhibited distinct profiles for sequence types, ST420 and ST277 in Kpn007 and Kpn016, respectively, and ABR genes (qnrS1, blaCTX-M-15, and blaSHV-27), virulence factors (ybt, iuc1, iro1), and contained both K (K20, K46) and O antigens (O1, O3b). MDR KPN in the geriatric population poses a serious health concern due to their increased vulnerability to infections and limited treatment options, requiring careful management.IMPORTANCEMultidrug resistance (MDR) and the hypervirulence nature of Klebsiella pneumoniae (KPN) in geriatric patients pose a critical health concern in nosocomial infections worldwide and result in high clinical complexity and mortality. The study investigated the factors for KPN infections and analyzed antimicrobial resistance profiles. More than 60% of Klebsiella pneumoniae isolates from geriatric patients were resistant to third- and fourth-generation cephalosporins, and most isolates carried blaNDM-1 and blaSHV-11 genes. Analyzing whole genomes of two KPNs, Kpn007 (ST277) was identified as a hypervirulent strain with aerobactin and yersinia siderophores, contributing to virulence, and Kpn016 (ST420) carried fluoroquinolone (qnrS1), ESBL (blaCTX-M-15 and blaSHV-27) resistance. Both genomes contained K antigens (K20 and K46) and O antigens (O1 and O3b).

Humans

Phage lambda cDNA cloning vectors for subtractive hybridization, fusion-protein synthesis and Cre-loxP automatic plasmid subcloning.

We describe the construction and use of two classes of cDNA cloning vectors. The first class comprises the lambda EXLX(+) and lambda EXLX(-) vectors that can be used for the expression in Escherichia coli of proteins encoded by cDNA inserts. This is achieved by the fusion of cDNA open reading frames to the T7 gene 10 promoter and protein-coding sequences. The second class, the lambda SHLX vectors, allows the generation of large amounts of single-stranded DNA or synthetic cRNA that can be used in subtractive hybridization procedures. Both classes of vectors are designed to allow directional cDNA cloning with non-enzymatic protection of internal restriction sites. In addition, they are designed to facilitate conversion from phage lambda to plasmid clones using a genetic method based on the bacteriophage P1 site-specific recombination system; we refer to this as automatic Cre-loxP plasmid subcloning. The phage lambda arms, lambda LOX, used in the construction of these vectors have unique restriction sites positioned between the two loxP sites. Insertion of a specialized plasmid between these sites will convert it into a phage lambda cDNA cloning vector with automatic plasmid subcloning capability.

Bacteriophage lambda

Rescue of a tk-plasmid from transgenic mice reveals its episomal transmission.

This communication demonstrates the usefulness of the plasmid rescue procedure for recovery of plasmids from transgenic mice. We have microinjected the plasmid pSK1 harbouring the Herpes simplex virus thymidine kinase gene into fertilized mouse oocytes and succeeded in recovering plasmids from newborns by transformation of E. coli either with HindIII cut cellular DNA or with uncut DNA. The majority of the rescued plasmids were indistinguishable from pSK1 by restriction analysis. The rescued plasmids proved to be functionally active in a transient expression assay in mouse Ltk- cells. The pSK1 DNA sequences were inherited by up to 90% of the second generation progeny mice, which is not in agreement with a Mendelian transmission of heterozygous markers integrated into a single site of the chromosome. These data support the assumption that germ line transmission of non-integrated episomal plasmids can occur.

Animals

Reevaluating human gene annotation: a second-generation analysis of chromosome 22.

We report a second-generation gene annotation of human chromosome 22. Using expressed sequence databases, comparative sequence analysis, and experimental verification, we have extended genes, fused previously fragmented structures, and identified new genes. The total length in exons of annotation was increased by 74% over our previously published annotation and includes 546 protein-coding genes and 234 pseudogenes. Thirty-two potential protein-coding annotations are partial copies of other genes, and may represent duplications on an evolutionary path to change or loss of function. We also identified 31 non-protein-coding transcripts, including 16 possible antisense RNAs. By extrapolation, we estimate the human genome contains 29,000-36,000 protein-coding genes, 21,300 pseudogenes, and 1500 antisense RNAs. We suggest that our revised annotation criteria provide a paradigm for future annotation of the human genome.

Animals

Retroviral src gene expression in continuous marrow culture increases the self-renewal capacity of multilineage hematopoietic stem cells.

To define the action of the retroviral src gene on hematopoietic stem cells, C57BL/6 x DBA/2 (B6D2F1) mouse long-term marrow cultures were infected at initiation with Moloney murine leukemia virus (MuLV) pseudotypes of src-recombinant retroviruses with the src gene inserted in the env region of an amphotropic MuLV (src-Ampho), or in the gag region of Moloney MuLV (src-Mo). Other cultures were infected with Friend spleen focus-forming virus polycythemia-inducing strain (SFFVp), Moloney MuLV, or amphotropic MuLV, or were uninfected controls. Harvested nonadherent cells were tested weekly for multilineage, granulocyte-erythroid-megakaryocyte macrophage (CFU-GEMM) colony formation in vitro in recombinant murine IL-3 and erythropoietin, and individual colonies were removed, split 1:2, with half of each replated for in vitro self-renewal and the other half examined morphologically for number of hematopoietic cellular lineages, or tested for release of MuLV and src virus. Cultures infected with src-Ampho, src-Mo, or SFFVp demonstrated a significant increase in cumulative nonadherent cell and CFU-GEMM production. There was prolonged self-renewal over seven serial transfers of individual CFU-GEMM from src virus-infected cultures over seven serial transfers, and five of 61 individual colonies from the second or third generations contained detectable v-src gene sequences, but none released detectable src virus. Self-renewal of CFU-GEMM was similar to that with permanent IL-3-dependent cell line B6SUtA. In contrast, MuLV-infected or control uninfected cultures produced fewer cells, and self-renewal of CFU-GEMM did not exceed three generations. IL-3-dependent clonal hematopoietic progenitor cell lines, derived from each culture group, formed no detectable tumors in vivo; however, each released the original helper and/or transforming virus. Adherent cell lines, derived from src-Ampho-infected cultures released src virus and formed fibro-sarcomas in vivo. The data support the conclusion that src-recombinant virus expression in long-term marrow cultures increases the self-renewal capacity of multilineage hematopoietic stem cells.

Animals

Nucleotide sequence analysis of human beta-globin gene by the quantification method: mutations in 3'-splice junction sequence and beta-thalassemia.

The nucleotide sequence at the intron-exon junction in the human beta-globin gene was analyzed by the quantification method (categorical discriminant analysis) proposed previously. Using the sample score of a 16-nucleotide sequence at a 3'-splice junction, we studied to what extent such a sequence contains the 3'-splice signal. To examine the applicability of our method, we further studied several mutants of beta-thalassemia, where nucleotide changes exist at 3'-splice junction sequences of the first and second introns. Other mutants involve point mutations which generate new 3'-splice signals within the first intron. Experimental results on the abnormal splicing in those mutants could be explained in terms of the sample scores of 16-nucleotide sequences and their locations relative to the branch point.

Base Sequence

Proteolytic processing of the Ada protein that repairs DNA O6-methylguanine residues in E. coli.

In extracts of E. coli treated with an adapting regime of MNNG, the induced 39kd Ada protein having O6-MeG-DNA methyltransferase activity is processed to a 19kd active domain corresponding to the C-terminal half of the intact protein. This proteolytic processing has been followed on Western immunoblots using antisera raised against the 19kd fragment. Initial processing at 25 degrees C or 37 degrees C mainly generates a fragment of mol. wt. 24kd which then undergoes a slower second cleavage to generate the 19kd active domain. Preceding this second cleavage site is a sequence of amino acids Thr- -Gly-Met-Thr- -Lys that also occurs at another site in the N-terminal half of the 39kd methyltransferase. It is proposed that this sequence is a recognition site for proteolytic activity. On the basis of cleavage of the Ada protein at either one or both of these sites, fragments may be generated of mol. wt. 24kd and 19kd containing the active site for O6-methylguanine and O4-methylthymine repair, and 15kd and 20kd, containing the active site for methylphosphotriester repair. These observations explain previous reports by others on the existence in cell extracts of multiple methyltransferase activities of different sizes recognizing O-methyl lesions in DNA. The cellular protease involved is resistant to a wide range of protease inhibitors.

Bacterial Proteins

Cloning and expression of the luxY gene from Vibrio fischeri strain Y-1 in Escherichia coli and complete amino acid sequence of the yellow fluorescent protein.

Vibrio fischeri strain Y-1 (ATCC 33715) emits light with a lambda max of 545 nm rather than the 485-nm emission typical of other strains of V. fischeri. The yellow emission is due to the interaction of the enzyme luciferase with a yellow fluorescent protein (YFP). On the basis of the N-terminal amino acid sequence of YFP, a mixed-sequence oligonucleotide probe was synthesized and used to isolate a 1.6-kbp HindIII fragment containing the first 208 bases of the gene that codes for YFP (luxY). Another synthetic oligonucleotide complementary to bases 167-184 of the YFP coding sequence was used to isolate a second (ca. 1.9 kbp) DNA fragment generated by digestion with both EcoRI and ClaI that contained the remainder of the luxY gene. The intact luxY gene, which encoded a 22,211-dalton polypeptide composed of 194 amino acid residues, was reconstructed from the two primary clones and is contained within a 765-bp SspI-XhoII fragment. Both strands of the entire luxY coding sequence were determined from the reconstructed gene, while the region surrounding the junction used in the reconstruction was also determined from the original partial clones. As with other genes that have been studied from V. fischeri, the luxY gene was unusually AT-rich. The sequence of luxY did not bear any apparent similarity to any of the sequences contained in the current GenBank database. Escherichia coli containing a plasmid with the luxY gene expresses a protein that reacts with antibody raised to authentic YFP.

Amino Acid Sequence

Complete characterization and sequence of an HLA class II DR beta chain cDNA from the DR5 haplotype.

The human major histocompatibility complex includes the DP, DQ, and DR subregions, each of which contains at least one alpha chain gene and two beta chain genes. The products of the alpha chain gene and a beta chain gene from a given subregion combine to form a heterodimer which is found predominantly on the surface of immunocompetent cells, and is essential for effective cell-cell interactions and the generation of an immune response. The beta chain of the DR molecule is highly polymorphic, and it is this polymorphism which is thought to be ultimately responsible for the specific immune responsiveness and disease predisposition conferred by different DR molecules. While the sequences of DR beta chains of the homozygous DR1 cells, homozygous DR2, homozygous DR4, DR3/w6 cells and DR4/w6 genotypes have been partially or completely characterized, no sequence is yet available for the DR beta chain from a homozygous DR5 cell. A cDNA library was therefore constructed from the Swei cell line homozygous for the DR5 haplotype. A beta chain clone was isolated, characterized, and sequenced. Comparison with previously published DR beta chain restriction endonuclease maps and nucleotide sequences demonstrated that this clone was a DR beta chain clone. Comparison of the deduced amino acid sequence with other DR beta chain amino acid sequences shows three regions of variability in the first external domain, corresponding to amino acid residues 9-13, 26-38, and 67-74. The sequence of each of these variable regions in the beta chain from DR5 cells was identical or nearly identical to the sequences of variable regions found in the beta chains of other DR haplotypes, supporting the notion of gene conversion as an evolutionary mechanism generating polymorphism. The second external domain, and transmembrane and intracytoplasmic regions show a high degree of sequence conservation.

Amino Acid Sequence

The amino acid sequence of rat liver glucokinase deduced from cloned cDNA.

Rat liver glucokinase (ATP:D-hexose 6-phosphotransferase, EC 2.7.1.1) was purified to homogeneity, cleaved, and subjected to amino acid sequence analysis. Forty-five percent of the protein sequence was obtained, and this information was used to design oligonucleotide probes to screen a rat liver cDNA library. A 1601-base pair cDNA (GK1) contained an open reading frame that encoded the amino acid sequences found in the peptides used to generate the oligonucleotide probes. A second cDNA was subsequently identified (GK.Z2), which is 2346 base pairs long and corresponds to nearly the entire glucokinase mRNA. Blot transfer analysis of hepatic RNA showed that glucokinase mRNA exists as a single species of about 2400 nucleotides. Four hours of insulin treatment of diabetic rats resulted in a 30-fold induction of this mRNA. GK.Z2 has a long open reading frame which, with the known partial peptide sequence, allowed us to deduce the primary structure of glucokinase. The enzyme is composed of 465 amino acids and has a mass of 51,924 daltons. Glucokinase has 53 and 33% amino acid sequence identities with the carboxyl-terminal domains of rat brain hexokinase I and yeast hexokinase, respectively. If conservative amino acid replacements are also considered, glucokinase is similar to these two enzymes at 75 and 63% of positions, respectively. The putative glucose- and ATP-binding domains of glucokinase were identified, and these regions appear to be highly conserved in the hexokinase family of enzymes.

Amino Acid Sequence

Amplified RNA synthesized from limited quantities of heterogeneous cDNA.

The heterogeneity of neural gene expression and the spatially limited expression of many low-abundance messenger RNAs in the brain has made cloning and analysis of such messages difficult. To generate amounts of nucleic acids sufficient for use in standard cloning strategies, we have devised a method for producing amplified heterogeneous populations of RNA from limited quantities of cDNA. Whole cerebellar RNA was primed with a synthetic oligonucleotide containing the T7 RNA polymerase promoter sequence 5' to a polythymidylate region. After second-strand cDNA synthesis, T7 RNA polymerase was used to generate amplified antisense RNA (aRNA). Up to 80-fold molar amplification has been achieved from nanogram quantities of cDNA. The amplified material is similar in size distribution to the parent cDNA and shows sequence heterogeneity as assessed by Southern and Northern blot analysis. Specific messages for moderate-abundance mRNAs for actin and guanine nucleotide-binding protein (G-protein) alpha subunits have been detected in the amplified material. By using in situ transcription to generate cDNA, sequences for cyclophilin have been detected in aRNA derived from single cerebellar tissue sections. cDNA derived from a single cerebellar Purkinje cell also has been amplified and yields material that hybridizes to cognate whole RNA and mRNA but not to Escherichia coli RNA.

Actins

Mathematical characterization of Chaos Game Representation. New algorithms for nucleotide sequence analysis.

Chaos Game Representation (CGR) can recognize patterns in the nucleotide sequences, obtained from databases, of a class of genes using the techniques of fractal structures and by considering DNA sequences as strings composed of four units, G, A, T and C. Such recognition of patterns relies only on visual identification and no mathematical characterization of CGR is known. The present report describes two algorithms that can predict the presence or absence of a stretch of nucleotides in any gene family. The first algorithm can be used to generate DNA sequences represented by any point in the CGR. The second algorithm can simulate known CGR patterns for different gene families by setting the probabilities of occurrence of different di- or trinucleotides by a trial and error process using some guidelines and approximate rules-of-thumb. The validity of the second algorithm has been tested by simulating sequences that can mimic the CGRs of vertebrate non-oncogenes, proto-oncogenes and oncogenes. These algorithms can provide a mathematical basis of the CGR patterns obtained using nucleotide sequences from databases.

Algorithms