PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Molecular and cellular characterization of SEL-OB/SVEP1 in osteogenic cells in vivo and in vitro.

We describe a novel human gene, named SEL-OB/SVEP1, expressed by skeletal tissues in vivo and by cultured osteogenic cells. The mRNA expression was analyzed on frozen tissues retrieved by laser-capture microscope dissection (LCM) and was detected in osteogenic tissues (periosteum and bone) but not in cartilage or skeletal muscle. The SEL-OB/SVEP1 cDNA of 11,139 bp was in silico translated into a 3574AA protein with expected molecular weight of 370 kDa. The protein is composed of multiple domains including complement control protein (CCP) modules with selectin superfamily signature; sushi and other domains, such as vWA, EGF, PTX, and HYR. Stromal osteogenic cells were analyzed for the protein expression using anti-SEL-OB/SVEP1 for immuno-precipitation and Western blot application confirm the presence of high molecular weight protein. Immuno-histochemistry and fluorescence-activated cell sorting (FACS) were applied to detect SEL-OB/SVEP1 on the surface of stromal cells. ELISA quantified the dependence of protein expression on cell density. Bioinformatic analysis of SEL-OB/SVEP1 revealed domains compositions recognized in cell surface molecules and suggested its role in cell adhesion. Analysis of mesechymal osteogenic cells' adhesion in presence of anti-SEL-OB/SVEP1 antibody demonstrated its interference with initial adhesion stages. In summary, present study describes novel SEL-OB/SVEP1 protein with a unique composition of functional domains, restricted pattern of expression in skeletal cells and demonstrated involvement in attachment of mesenchymal cells. The unusual composition of functional domains puts SEL-OB/SVEP1 in the discrete new group of membrane proteins involved in cell adhesion processes. All together makes SEL-OB/SVEP1 an attractive marker for studying the role of stromal osteogenic cells and their interactions within the bone marrow microenvironment creating a network that regulates the skeletal homeostasis.

Animals↗

ATAC-seq in Emerging Model Organisms: Challenges and Strategies.

The Assay for Transposase-Accessible Chromatin with sequencing (ATAC-seq) is a versatile and widely utilized method for identifying potential regulatory regions, such as promoters and enhancers, within a genome. ATAC-seq has been successfully applied to a wide range of established and emerging model organisms. However, implementing this method in emerging model systems, such as arthropods, can be challenging due to several factors that influence data quality. These factors include the availability of a sufficient amount and quality of tissue or cells, the need for species- and tissue-specific protocol optimization, the completeness and accuracy of the reference genome, and the quality of the genome annotation. In this article, we emphasize the key steps in the ATAC-seq protocol that, based on our experience, have the greatest impact on data quality when adapting this method for emerging model organisms. Specifically, we discuss the importance of nuclei isolation, the incubation conditions of the Tn5 transposase, and PCR amplification of the library. Furthermore, we outline essential quality checkpoints during the bioinformatic analysis of ATAC-seq data to assist in assessing data integrity and consistency. Given that many emerging model organisms may not be readily available in laboratory cultures, we also emphasize the importance of evaluating how different preservation methods affect ATAC-seq data quality. Based on examples in one spider and one ant species, we demonstrate that replication and thorough quality controls at all steps of the protocol and data analysis are essential to assess the usability of ATAC-seq data. Our data highlights the importance of isolating the right number of intact nuclei, as well as ensuring optimal amplification conditions during library preparation to obtain good-quality sequence data for downstream analyses. We recommend using fresh tissue samples if possible because we show that direct cryopreservation of the tissue may affect chromatin integrity. This effect could be avoided or reduced by preserving the homogenate in cell culture medium. Overall, we explain the ATAC-seq protocol and downstream analyses in detail and give step-by-step advice to researchers who are new to the field and want to implement this method. With careful planning and validation, ATAC-seq can reveal the regulatory landscape of a genome and aid in identifying elements that govern gene expression.

Animals↗

Selection and identification of human gonadotropin-releasing hormone promoter binding peptides by phage display-CEMSA.

Specific interactions between transcription factors and cis-acting DNA sequences form the molecular basis of gene expression regulation. Here, we applied phage display technology to DNA-protein interaction studies. A phage-displayed peptide library was used to select Gonadotropin-releasing hormone promoter (GP) binding peptides. After four sequential rounds of biopanning on GP-conjugated magnetic beads, phage clones encoding GQPTPRNAGLPL (B6), SRLNVEPLTTYS (B3), and TTLHWASLTTGR (B11) were enriched. Phages bearing these peptides showed specific binding to GP in solution by capillary electrophoresis mobility shift assay (CEMSA). In addition, some human transcription factors were speculated as the potential transcription factors or co-activators of GnRH gene by bioinformatic analysis. These results suggest that phage display-CEMSA methodology should be a powerful tool to screen and identify site-specific DNA-binding peptides.

Amino Acid Sequence↗

Novel TCOF1 Frameshift Variant and Phenotypic Heterogeneity in a Chinese Family With Treacher Collins Syndrome.

BACKGROUND: Treacher Collins syndrome (TCS) is a congenital craniofacial disorder characterized by malar and mandibular hypoplasia, downward-slanting palpebral fissures, and conductive hearing loss. Pathogenic variants in TCOF1 account for most cases, with POLR1D, POLR1C, and POLR1B also implicated. METHODS: Whole-exome sequencing was performed in a two-generation Chinese family with TCS, followed by Sanger sequencing validation. Clinical features were systematically evaluated, and bioinformatic analyses combined with structural modeling were employed to assess the potential pathogenicity of the identified variant. RESULTS: In this study, a novel heterozygous frameshift variant in TCOF1 (NM_001371623.1:c.1601_1602delCC, p.Pro534Leufs*15) was identified in the proband and his affected father. The proband presented classic TCS features including craniofacial skeletal hypoplasia, downward-slanting palpebral fissures, and conductive hearing loss. He also carried a right-sided preauricular fistula, a nonclassical feature of TCS. The same variant was detected in his affected father with a substantially milder phenotype, indicating marked intrafamilial phenotypic variability. Bioinformatic analysis and structural modeling predicted that this variant produces a severely truncated Treacle protein lacking key functional domains, which is predicted to disrupt nucleolar localization and ribosome biogenesis. CONCLUSION: Our findings expand the variant spectrum of TCOF1, highlight phenotypic heterogeneity in TCS, and reinforce the critical role of molecular diagnosis in distinguishing TCS from phenotypically overlapping craniofacial syndromes.

Humans↗

Solution structures of the putative anti-sigma-factor antagonist TM1442 from Thermotoga maritima in the free and phosphorylated states.

The NMR structures of the unphosphorylated Thermotoga maritima protein TM1442 at pH 4.8 and of the phosphorylated TM1442 (TM1442-P) at pH 7.0 are presented, and a functional interaction of TM1442 with TM0733 is characterized. Although the NMR spectra of TM1442-P at pH 7.0 are of high quality, detailed NMR studies of unphosphorylated TM1442 could be performed only at slightly acidic pH values and high salt concentration. TM1442 is a putative anti-sigma-factor antagonist related to the sigmaF and sigmaB regulation systems in Bacillus subtilis, which is the component in this system that can be phosphorylated. The kinase TM0733, which shows sequence similarity to the GHKL ATPase/kinase superfamily, was identified as the possible anti-sigma-factor of TM1442 using a bioinformatics analysis. Phosphorylation of TM1442 by TM0733 was confirmed by NMR, mass spectroscopy and native gel electrophoresis, and Ser59 was identified as the phosphorylation site using site-directed mutational analysis. The solution structure of TM1442-P at pH 7.0 has the same global fold as free TM1442 at pH 4.8, with an alpha/beta topology consisting of a central four-stranded beta sheet and three alpha helices, but the regular secondary structure elements wrapping the hydrophobic core of the protein undergo subtle conformational changes upon phosphorylation.

Amino Acid Sequence↗

Cell wall proteins in apoplastic fluids of Arabidopsis thaliana rosettes: identification by mass spectrometry and bioinformatics.

Weakly bound cell wall proteins of Arabidopsis thaliana were identified using a proteomic and bioinformatic approach. An efficient protocol of extraction based on vacuum-infiltration of the tissues was developed. Several salts and a chelating agent were compared for their ability to extract cell wall proteins without releasing cytoplasmic contaminants. Of the 93 proteins that were identified, a large proportion (60%) was released by calcium chloride. From bioinformatics analysis, it may be predicted that most of them (87 out of 93) had a signal peptide, whereas only six originated from the cytoplasm. Among the putative apoplastic proteins, a high proportion (67 out of 87) had a basic pI. Numerous glycoside hydrolases and proteins with interacting domains were identified, in agreement with the expected role of the extracellular matrix in polysaccharide metabolism and recognition phenomena. Ten proteinases were also found as well as six proteins with unknown functions. Comparison of the cell wall proteome of rosettes with the previously published cell wall proteome of cell suspension cultures showed a high level of cell specificity, especially for the different members of several large multigenic families.

Arabidopsis↗

A proteomic view of Desulfovibrio vulgaris metabolism as determined by liquid chromatography coupled with tandem mass spectrometry.

Direct LC-MS/MS was used to examine the proteins extracted from exponential or stationary phase Desulfovibrio vulgaris cells that had been grown on a minimal medium containing either lactate or formate as the primary carbon source. Across all four growth conditions, 976 gene products were identified with high confidence, which is equal to approximately 28% of all predicted proteins in the D. vulgaris genome. Bioinformatic analysis showed that the proteins identified were distributed among almost all functional classes, with the energy metabolism category containing the greatest number of identified proteins. At least 154 ORFs originally annotated as hypothetical proteins were found to encode the expressed proteins, which provided verification for the authenticity of these hypothetical proteins. Proteomic analysis showed that proteins potentially involved in ATP biosynthesis using the proton gradient across membrane, such as ATPase, alcohol dehydrogenases, heterodisulfide reductases, and [NiFe] hydrogenase (HynAB-1) of the hydrogen cycling were highly expressed in all four growth conditions, suggesting they may be the primary pathways for ATP synthesis in D. vulgaris. Most of the enzymes involved in substrate-level phosphorylation were also detected in all tested conditions. However, no enzyme involved in CO cycling or formate cycling was detected, suggesting that they are not the primary ATP-biosynthesis pathways under the tested conditions. This study provides the first proteomic overview of the cellular metabolism of D. vulgaris. The complete list of proteins identified in this study and their abundances (peptide hits) is provided in Supplementary Table 1.

Alcohol Dehydrogenase↗

Functional annotation of proteins identified in human brain during the HUPO Brain Proteome Project pilot study.

The HUPO Brain Proteome Project is an initiative coordinating proteomics studies to characterise human and mouse brain proteomes. Proteins identified in human brain samples during the project's pilot phase were put into biological context through integration with various annotation sources followed by a bioinformatics analysis. The data set was related to the genome sequence via the genes encoding identified proteins including an assessment of splice variant identification as well as an analysis of tissue specificity of the respective transcripts. Proteins were furthermore categorised according to subcellular localisation, molecular function and biological process, grouped into protein families and mapped to biological pathways they are known to act in. Involvement in pathological conditions was examined based on association with entries in the online version of Mendelian Inheritance in Man and an interaction network was derived from curated protein-proteininteraction data. Overall a non-redundant set of 1804 proteins was identified in human brain samples. In the majority of cases splice variants could be unambiguously identified by unique peptides, including matches to several hypothetical transcripts of known as well as predicted genes.

Alternative Splicing↗

P704P, P712P, and P775P: A genomic cluster of prostate-specific genes.

BACKGROUND: Discovery of prostate cancer- and tissue-specific genes will lead to an increased understanding of the molecular events associated with the malignant transformation and tumorigenesis of prostate cells. Such understanding will likely result in the development of promising new markers for screening, diagnosis, and prognosis, as well as potential therapeutic approaches for combating this disease. METHODS: A PCR-based subtraction method was combined with a high-throughput microarray screening approach to identify prostate tissue- and/or cancer-specific genes. Northern blot and quantitative real-time PCR were used to confirm prostate specificity. Bioinformatics analysis was performed to determine gene localization and to identify the open reading frame of novel genes. RESULTS: Three novel cDNA clones, P704P, P712P, and P775P, were identified and characterized to be specific for normal and malignant prostate tissues. Furthermore, P712P mRNA expression was found to be androgen responsive in LNCaP cells. Sequences for all three cDNAs were localized to an 80 kb genomic region on chromosome 22. Attempts to identify full-length transcripts did not reveal any apparent open reading frames, indicating that P704P, P712P, and P775P may belong to a novel class of transcripts with specific patterns of gene expression that do not code for translated proteins. CONCLUSIONS: A genomic cluster of prostate-specific genes with no apparent open reading frame has been discovered using a high-throughput approach combining subtraction with microarray. This may represent an important genomic region having possible connections to prostate biology with potential applications in prostate diagnostics and therapy.

Blotting, Northern↗

Sequence-structure-function relationships of a tRNA (m7G46) methyltransferase studied by homology modeling and site-directed mutagenesis.

The Escherichia coli TrmB protein and its Saccharomyces cerevisiae ortholog Trm8p catalyze the S-adenosyl-L-methionine-dependent formation of 7-methylguanosine at position 46 (m7G46) in tRNA. To learn more about the sequence-structure-function relationships of these enzymes we carried out a thorough bioinformatics analysis of the tRNA:m7G methyltransferase (MTase) family to predict sequence regions and individual amino acid residues that may be important for the interactions between the MTase and the tRNA substrate, in particular the target guanosine 46. We used site-directed mutagenesis to construct a series of alanine substitutions and tested the activity of the mutants to elucidate the catalytic and tRNA-recognition mechanism of TrmB. The functional analysis of the mutants, together with the homology model of the TrmB structure and the results of the phylogenetic analysis, revealed the crucial residues for the formation of the substrate-binding site and the catalytic center in tRNA:m7G MTases.

Amino Acid Sequence↗

Evidence for a novel domain of bacterial outer membrane ushers.

Many pathogenic bacteria possess adhesive surface organelles (called pili), anchored to their outer membrane, which mediate the first step of infection by binding to host tissue. Pilus biogenesis occurs via the "chaperone-usher" pathway: the usher, a large outer membrane protein, binds complexes of a periplasmic chaperone with pilus subunits, unloads the subunits from the chaperone, and assembles them into the pilus, which is extruded into the extracellular space. Ushers comprise an N-terminal periplasmic domain, a large transmembrane beta-barrel central domain, and a C-terminal periplasmic domain. Since structural data are available only for the N-terminal domain, we performed an in-depth bioinformatic analysis of bacterial ushers. Our analysis led us to the conclusion that the transmembrane beta-barrel region of ushers contains a so far unrecognized soluble domain, the "middle domain", which possesses a beta-sandwich fold. Two other bacterial beta-sandwich domains, the TT0351 protein from Thermus thermophilus and the carbohydrate binding module CBM36 from Paenibacillus polymyxa, are possible distant relatives of the usher "middle domain". Several mutations reported to abolish in vivo pilus formation cluster in this region, underlining its functional importance.

Bacterial Outer Membrane Proteins↗

Identification of a novel human glutathione S-transferase using bioinformatics.

In searching the expressed sequence tag (EST) data-base of GenBank with coding sequences of 11 known human glutathione S-transferases in conjunction with bioinformatic analysis, we have identified five ESTs that encode a new human glutathione S-transferase (GST) designated GST A4. The cDNA clone (I.M.A.G.E. Consortium cDNA Clone ID 515157) had an insert length of 1279 bp and contains an open reading frame of 666 bp, which encodes a protein of 222 amino acid residues. The GST A4 protein is identical in length to human GST A1 and A2 and is 54% identical to human GST A1 and A2. Sequence comparison with other human GSTs suggests that it is a new GST belonging to the alpha class GSTs. Northern blot analysis and EST database searches have demonstrated that the GST A4 mRNA is expressed at a high level in brain, placenta, and skeletal muscle and much lower in lung and liver. Analysis of the sequence tagged site (STS) database indicated that the GST A4 gene is located on chromosome 6. This STS represents a previously unidentified transcript further confirming the novelty of the new sequence.

Amino Acid Sequence↗

Full-length cDNA cloning of human neuroglobin and tissue expression of rat neuroglobin.

Neuroglobin is a recently discovered respiratory, porphyrin-containing protein that is expressed in the brain of mouse and human. However, the full-length cDNA sequence and genomic organization of human neuroglobin have not been reported. In this paper, the full-length cDNA sequence of human neuroglobin was cloned following bioinformatic analysis and the rapid amplification of cDNA ends (RACE) technique. It was shown that the full-length cDNA sequence (GenBank Accession No. AF422796) of human neuroglobin is 1909 bp in size, and the genomic sequence is 8041 bp in size (GenBank Accession No. AF422797). To further study the characterization of this gene, the coding region of rat neuroglobin (GenBank Accession No. AF333245) was cloned by using degeneracy PCR. The result showed high conservation among human, rat, and mouse neuroglobin. Furthermore, it was demonstrated that NGB was extensively expressed in rat brain by using in situ hybridization and the immunohistochemical technique. Transcription of NGB mRNA was shown to be widely distributed throughout the adult rat brain, including cerebral cortex, hippocampus, thalamus, hypothalamus, olfactory bulb, and cerebellum. NGB protein immunoreactive cells were also widely distributed throughout normal adult rat brain, including cerebral cortex, hippocampus, thalamus, hypothalamus, pons, and cerebellum. It could be seen that the NGB-immunopositive signals were in the cytoplasma and processes of the neuron. These data strongly support the notion that neuroglobin is a highly conserved gene in evolution and is very important in the nervous system, possibly related to the oxygen supply of the neuron.

Aging↗

A new gene family including DSCR1 (Down Syndrome Candidate Region 1) and ZAKI-4: characterization from yeast to human and identification of DSCR1-like 2, a novel human member (DSCR1L2).

A new gene family has been identified on the basis of in-depth bioinformatics analysis of the Down syndrome candidate region 1 (DSCR1) gene, located on 21q22.1. We have determined the complete coding sequences of similar genes in Saccharomyces cerevisiae and Caenorhabditis elegans, as well as that of a novel human gene, named DSCR1L2 (DSCR1-like 2). Peripheral blood leukocyte cDNA sequencing predicts as its product a 241-amino-acid protein highly similar to products of the human genes DSCR1 and ZAKI-4 (HGMW-approved symbol DSCR1L1). The highest level of expression of DSCR1L2 mRNA was found by Northern blot analysis in heart and skeletal muscles, liver, kidney, and peripheral blood leukocytes (three transcripts of 3.2, 5. 2, and 7.5 kb). The gene consists of four exons and spans about 22 kb on chromosome 1 (1p33-p35.3) (Human Chromosome 1, Sanger Centre). Exon/intron organization is highly conserved between DSCR1 and DSCR1L2. Two alternative DSCR1L2 mRNA splicing forms have been recognized, with one lacking 10 amino acids in the middle of the protein. Analysis of expressed sequence tags (ESTs) shows DSCR1L2 expression in fetal tissues (heart, liver, and spleen) and in adenocarcinomas. ESTs related to the murine DSCR1L2 orthologue are found in the 2-cell stage mouse embryo, in developing brain stem and spinal cord, and in thymus and T cells. The most prominent feature identified in the protein family is a central short, unique serine-proline motif (including an ISPPXSPP box), which is strongly conserved from yeast to human but is absent in bacteria. Moreover, homology with the RNA-binding domain was weakly but consistently detected in a stretch of 80 amino acids at the amino-terminus by fine sequence analysis based on tools utilizing both hidden Markov models and BLAST. The identification of this new gene family should allow a better understanding of the functions of the genes belonging to it.

Adaptor Proteins, Signal Transducing↗

Cloning and characterization of Disc1, the mouse ortholog of DISC1 (Disrupted-in-Schizophrenia 1).

We cloned the mouse ortholog of DISC1 (Disrupted-in-Schizophrenia 1), a candidate gene for schizophrenia. Disc1 is 3163 nucleotides long and has 60% identity with the human DISC1. Disc1 encodes 851 amino acids and has 56% identity with the human protein. Disc1 maps to the DISC1 syntenic region in the mouse, and genomic structure is conserved. A Disc1 splice variant deletes a portion of Disc1 beginning at amino acids orthologous to the human truncation. Bioinformatic analysis and cross-species comparisons revealed sequence conservation distributed across the genes and conservation of leucine zipper and coiled-coil domains in both orthologs. In situ hybridization in adult mouse brain revealed a restricted expression pattern, with highest levels in the dentate gyrus of the hippocampus and lower expression in CA1-CA3 of the hippocampus, cerebellum, cerebral cortex, and olfactory bulbs. Identification of Disc1 will facilitate the study of DISC1's function and creation of mouse models of DISC1 disruption.

Alternative Splicing↗

Induction of LPL gene expression by sterols is mediated by a sterol regulatory element and is independent of the presence of multiple E boxes.

Overexpression of the adipocyte differentiation and determination factor-1 (ADD-1) or sterol regulatory element binding protein-1 (SREBP-1) induces the expression of numerous genes involved in lipid metabolism, including lipoprotein lipase (LPL). Therefore, we investigated whether LPL gene expression is controlled by changes in cellular cholesterol concentration and determined the molecular pathways involved. Cholesterol depletion of culture medium resulted in a significant induction of LPL mRNA in the 3T3-L1 preadipocyte cell line, whereas addition of cholesterol reduced LPL mRNA expression to basal levels. Similar to the expression of the endogenous LPL gene, the activity of the human LPL gene promoter was enhanced by cholesterol depletion in transient transfection assays, whereas addition of cholesterol caused a reversal of its induction. The effect of cholesterol depletion upon the human LPL gene promoter was mimicked by cotransfection of expression constructs encoding the nuclear form of SREBP-1a, -1c (also called ADD-1) and SREBP-2. Bioinformatic analysis demonstrated the presence of 3 potential sterol regulatory elements (SRE) and 3 ADD-1 binding sequences (ABS), also known as E-box motifs. Using a combination of in vitro protein-DNA binding assays and transient transfection assays of reporter constructs containing mutations in each individual site, a sequence element, termed LPL-SRE2 (SRE2), was shown to be the principal site conferring sterol responsiveness upon the LPL promoter. These data furthermore underscore the importance of SRE sites relative to E-boxes in the regulation of LPL gene expression by sterols and demonstrate that sterols contribute to the control of triglyceride metabolism via binding of SREBP to the LPL regulatory sequences.

Adipocytes↗

Complete genomic sequence of bacteriophage ul36: demonstration of phage heterogeneity within the P335 quasi-species of lactococcal phages.

The complete genomic sequence of the Lactococcus lactis virulent phage ul36 belonging to P335 lactococcal phage species was determined and analyzed. The genomic sequence of this lactococcal phage contained 36,798 bp with an overall G+C content of 35.8 mol %. Fifty-nine open reading frames (ORFs) of more than 40 codons were found. N-terminal sequencing of phage structural proteins as well as bioinformatic analysis led to the attribution of a function to 24 ORFs (41%). A lysogeny module was found within the genome of this virulent phage. The putative integrase gene seems to be the product of a horizontal transfer because it is more closely related to Streptococcus pyogenes phages than it is to L. lactis phages. Comparative genome analysis with six complete genomes of temperate P335-like phages confirmed the heterogeneity among phages of P335 species. A dUTPase gene is the only conserved gene among all P335 phages analyzed as well as the phage BK5-T. A genetic relationship between P335 phages and the phage-type of the BK5-T species was established. Thus, we proposed that phage BK5-T be included within the P335 species and thereby reducing the number of lactococcal phage species to 11.

Bacteriophages↗

A Practical Approach to High-Throughput and Accurate Mapping-by-Sequencing in Arabidopsis.

Forward-directed genetic screens are extremely powerful in identifying novel genes involved in a specific biological process, including various chromatin regulatory pathways. However, the traditional ways of genetic mapping are time- and cost-demanding. Recently, the whole process was revolutionized by the development of mapping-by-sequencing (MBS) protocols. In MBS, the causal mutations and their positions within genes are identified directly by whole-genome sequencing and bioinformatics analysis of the bulk of mutant plants selected based on the mutant phenotype from a segregating population. MBS increases precision and economizes the mapping. Here, we describe a general protocol and provide practical tips on how to proceed with the mapping-by-sequencing on the example of Arabidopsis forward-directed genetic screen designed to identify mutants sensitive to a specific type of DNA damage. The described protocol is generally applicable to a wide range of genetic screens in various inbreeding species with a reference genome sequence.

Arabidopsis↗