Construction of a normalized Bos taurus and Bos indicus macrophage-specific cDNA library.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Despite intense research on light responses in plants, the consequences of a simple shift from light to darkness remain poorly characterized. We have examined the transcriptome of Arabidopsis thaliana seedling leaves upon a shift from constant light to darkness for between 1 and 8 h, while excluding most effects associated with circadian oscillation. Expression clustering and gene ontology analyses identified about 790 responsive genes implicated in diverse cellular processes. Compared to the better-studied long-term dark adaptation response, the early response to darkness is partially overlapping yet clearly distinct, encompassing early transient, early sustained, and late response clusters. The repressor of photomorphogenesis, COP1 (constitutive photomorphogenic 1), is not a chief regulator of the early response to darkness, in contrast to its well-established role during long-term dark adaptation and etiolation. Only part of the early dark response can be understood as the opposite of the response following a dark-to-light transition and as a response to sugar deprivation. Bioinformatic comparisons with published microarray datasets further suggest that abscisic acid (ABA) signaling plays a prominent role in the early response to darkness, although this effect is not mediated by an increase in the ABA level. The potential basis for the co-regulation by darkness and ABA is discussed in light of sugar and redox signaling.
BACKGROUND: The alignment of two or more protein sequences provides a powerful guide in the prediction of the protein structure and in identifying key functional residues, however, the utility of any prediction is completely dependent on the accuracy of the alignment. In this paper we describe a suite of reference alignments derived from the comparison of protein three-dimensional structures together with evaluation measures and software that allow automatically generated alignments to be benchmarked. We test the OXBench benchmark suite on alignments generated by the AMPS multiple alignment method, then apply the suite to compare eight different multiple alignment algorithms. The benchmark shows the current state-of-the art for alignment accuracy and provides a baseline against which new alignment algorithms may be judged. RESULTS: The simple hierarchical multiple alignment algorithm, AMPS, performed as well as or better than more modern methods such as CLUSTALW once the PAM250 pair-score matrix was replaced by a BLOSUM series matrix. AMPS gave an accuracy in Structurally Conserved Regions (SCRs) of 89.9% over a set of 672 alignments. The T-COFFEE method on a data set of families with <8 sequences gave 91.4% accuracy, significantly better than CLUSTALW (88.9%) and all other methods considered here. The complete suite is available from http://www.compbio.dundee.ac.uk. CONCLUSIONS: The OXBench suite of reference alignments, evaluation software and results database provide a convenient method to assess progress in sequence alignment techniques. Evaluation measures that were dependent on comparison to a reference alignment were found to give good discrimination between methods. The STAMP Sc Score which is independent of a reference alignment also gave good discrimination. Application of OXBench in this paper shows that with the exception of T-COFFEE, the majority of the improvement in alignment accuracy seen since 1985 stems from improved pair-score matrices rather than algorithmic refinements. The maximum theoretical alignment accuracy obtained by pooling results over all methods was 94.5% with 52.5% accuracy for alignments in the 0-10 percentage identity range. This suggests that further improvements in accuracy will be possible in the future.
Clustering is a popular method for analyzing microarray data. Given the large number of clustering algorithms being available, it is difficult to identify the most suitable ones for a particular task. It is also difficult to locate, download, install and run the algorithms. This paper describes a matchmaking system, SemBiosphere, which solves both problems. It recommends clustering algorithms based on some minimal user requirement inputs and the data properties. An ontology was developed in OWL, an expressive ontological language, for describing what the algorithms are and how they perform, in addition to how they can be invoked. This allows machines to "understand" the algorithms and make the recommendations. The algorithm can be implemented by different groups and in different languages, and run on different platforms at geographically distributed sites. Through the use of XML-based web services, they can all be invoked in the same standard way. The current clustering services were transformed from the non-semantic web services of the Biosphere system, which includes a variety of algorithms that have been applied to microarray gene expression data analysis. New algorithms can be incorporated into the system without too much effort. The SemBiosphere system and the complete clustering ontology can be accessed at http://yeasthub2.gersteinlab. org/sembiosphere/.
In modern biology, one of the most important research problems is to understand how protein sequences fold into their native 3D structures. To investigate this problem at a high level, one wishes to analyze the protein landscapes, i.e., the structures of the space of all protein sequences and their native 3D structures. Perhaps the most basic computational problem at this level is to take a target 3D structure as input and design a fittest protein sequence with respect to one or more fitness functions of the target 3D structure. We develop a toolbox of combinatorial techniques for protein landscape analysis in the Grand Canonical model of Sun, Brem, Chan, and Dill. The toolbox is based on linear programming, network flow, and a linear-size representation of all minimum cuts of a network. It not only substantially expands the network flow technique for protein sequence design in Kleinberg's seminal work but also is applicable to a considerably broader collection of computational problems than those considered by Kleinberg. We have used this toolbox to obtain a number of efficient algorithms and hardness results. We have further used the algorithms to analyze 3D structures drawn from the Protein Data Bank and have discovered some novel relationships between such native 3D structures and the Grand Canonical model.
BACKGROUND: Cell adhesion involves interactions of integrins and extracellular proteins, often facilitated by the RGD motif. Only presence of the RGD in a sequence of a protein may be not sufficient for the biological activity (binding to an integrin) and additional biochemical and/or structural studies are essential. MATERIAL/METHODS: Structural criteria that would allow identification biologically active RGD-sites on the base of a spatial structure may assist analysis of function of a protein in the cell. For the first time, computational analysis of RGD-sites in a large non-redundant set of protein structures was done. RESULTS: Out of 3819 protein chains sequences of about 100 contained RGDs. Analysis of the structures of the RGD-'native' proteins has allowed establishing main determinants of the biologically active conformations of the RGD sites: surface accessibility of the whole RGD-sequence and the secondary structure. The criteria, applied to the remaining proteins of the set, identify 23 proteins ( approximately 25%) with potentially active RGD-sites. The results strongly suggest that RGD has a high propensity for being involved in protein-protein interactions and this may explain occurrence of RGDs in intracellular proteins. Results of the analysis suggest (in some cases, confirm) novel integrin-related activities for 7 membrane/extracellular proteins, as well as confirm RGD-facilitated cell attachment for 5 viral proteins. CONCLUSIONS: Only presence of RGD in a sequence is not sufficient to propose biological activity of this site. The results also suggest that the method can be used on large scale: for example, for identifying potential integrin-interacting proteins in an animal genome.
We have analyzed several lots of epidermal growth factor (EGF) purified from murine submaxillary glands including "receptor grade" EGF from Collaborative Research and EGF from Boehringer Mannheim Biochemicals. New England Nuclear uses "receptor grade" EGF to produce 125I-labeled EGF. Though these reagents are reported to be homogeneous, we found them to be a mixture of six species. A method was developed to separate this mixture into its component parts. The individual components were chemically characterized and tested for biological potency. N-terminal sequence analysis of the unfractionated EGF-mixture reveals three different sequences starting with residues 1, 2, or 3 of the mature peptide. Each component exhibited different degrees of mitogenic and EGF receptor binding activity indicating that the N-terminal region contributes to the biological response. The species representing the complete EGF peptide is the most active species in all biological assays. A rapid method for purification of homogeneous complete EGF from commercial EGF preparations is described.
A decade of access to whole-genome sequences has been increasingly revealing about the informational network relating all living organisms. Although at one point there was concern that extensive horizontal gene transfer might hopelessly muddle phylogenies, it has not proved a severe hindrance. The melding of sequence and structural information is being used to great advantage, and the prospect exists that some of the earliest aspects of life on Earth can be reconstructed, including the invention of biosynthetic and metabolic pathways. Still, some fundamental phylogenetic problems remain, including determining the root--if there is one--of the historical relationship between Archaea, Bacteria and Eukarya.
Ordered chondrocyte differentiation and maturation is required for normal skeletal development, but the intracellular pathways regulating this process remain largely unclear. We used Affymetrix microarrays to examine temporal gene expression patterns during chondrogenic differentiation in a mouse micromass culture system. Robust normalization of the data identified 3300 differentially expressed probe sets, which corresponds to 1772, 481, and 249 probe sets exhibiting minimum 2-, 5-, and 10-fold changes over the time period, respectively. GeneOntology annotations for molecular function show changes in the expression of molecules involved in transcriptional regulation and signal transduction among others. The expression of identified markers was confirmed by RT-PCR, and cluster analysis revealed groups of coexpressed transcripts. One gene that was up-regulated at later stages of chondrocyte differentiation was Rgs2. Overexpression of Rgs2 in the chondrogenic cell line ATDC5 resulted in accelerated hypertrophic differentiation, thus providing functional validation of microarray data. Collectively, these analyses provide novel information on the temporal expression of molecules regulating endochondral bone development.
Accumulating evidence suggests that mRNA degradation systems are crucial for various biological processes in eukaryotes. Here we provide evidence that an mRNA degradation system is associated with some plant hormones and stress responses in plants. We analysed a novel Arabidopsis abscisic acid (ABA)-hypersensitive mutant, ahg2-1, that showed ABA hypersensitivity not only in germination, but also at later developmental stages, and that displayed pleiotropic phenotypes. We found that ahg2-1 accumulated more endogenous ABA in seeds and mannitol-treated plants than did the wild type. Microarray experiments showed that the expressions of ABA-, salicylic acid- and stress-inducible genes were increased in normally grown ahg2-1 plants, suggesting that the ahg2-1 mutation somehow affects various stress responses as well as ABA responses. Map-based cloning of AHG2 revealed that this gene encodes a poly(A)-specific ribonuclease (AtPARN) that is presumed to function in mRNA degradation. Detailed analysis of the ahg2-1 mutation suggests that the mutation reduces AtPARN production. Interestingly, expression of AtPARN was induced by treatment with ABA, high salinity and osmotic stress. These results suggest that both upregulation and downregulation of gene expression by the mRNA-destabilizing activity of AtPARN are crucial for proper ABA, salicylic acid and stress responses.
BACKGROUND: Large-scale sequence comparison is a powerful tool for biological inference in modern molecular biology. Comparing new sequences to those in annotated databases is a useful source of functional and structural information about these sequences. Using software such as the basic local alignment search tool (BLAST) or HMMPFAM to identify statistically significant matches between newly sequenced segments of genetic material and those in databases is an important task for most molecular biologists. Searching algorithms are intrinsically slow and data-intensive, especially in light of the rapid growth of biological sequence databases due to the emergence of high throughput DNA sequencing techniques. Thus, traditional bioinformatics tools are impractical on PCs and even on dedicated UNIX servers. To take advantage of larger databases and more reliable methods, high performance computation becomes necessary. RESULTS: We describe the implementation of SS-Wrapper (Similarity Search Wrapper), a package of wrapper applications that can parallelize similarity search applications on a Linux cluster. Our wrapper utilizes a query segmentation-search (QS-search) approach to parallelize sequence database search applications. It takes into consideration load balancing between each node on the cluster to maximize resource usage. QS-search is designed to wrap many different search tools, such as BLAST and HMMPFAM using the same interface. This implementation does not alter the original program, so newly obtained programs and program updates should be accommodated easily. Benchmark experiments using QS-search to optimize BLAST and HMMPFAM showed that QS-search accelerated the performance of these programs almost linearly in proportion to the number of CPUs used. We have also implemented a wrapper that utilizes a database segmentation approach (DS-BLAST) that provides a complementary solution for BLAST searches when the database is too large to fit into the memory of a single node. CONCLUSIONS: Used together, QS-search and DS-BLAST provide a flexible solution to adapt sequential similarity searching applications in high performance computing environments. Their ease of use and their ability to wrap a variety of database search programs provide an analytical architecture to assist both the seasoned bioinformaticist and the wet-bench biologist.
Large-scale gene expression studies provide significant insight into genes differentially regulated in disease processes such as cancer. However, these investigations offer limited understanding of multisystem, multicellular diseases such as atherosclerosis. A systems biology approach that accounts for gene interactions, incorporates nontranscriptionally regulated genes, and integrates prior knowledge offers many advantages. We performed a comprehensive gene level assessment of coronary atherosclerosis using 51 coronary artery segments isolated from the explanted hearts of 22 cardiac transplant patients. After histological grading of vascular segments according to American Heart Association guidelines, isolated RNA was hybridized onto a customized 22-K oligonucleotide microarray, and significance analysis of microarrays and gene ontology analyses were performed to identify significant gene expression profiles. Our studies revealed that loss of differentiated smooth muscle cell gene expression is the primary expression signature of disease progression in atherosclerosis. Furthermore, we provide insight into the severe form of coronary artery disease associated with diabetes, reporting an overabundance of immune and inflammatory signals in diabetics. We present a novel approach to pathway development based on connectivity, determined by language parsing of the published literature, and ranking, determined by the significance of differentially regulated genes in the network. In doing this, we identify highly connected "nexus" genes that are attractive candidates for therapeutic targeting and followup studies. Our use of pathway techniques to study atherosclerosis as an integrated network of gene interactions expands on traditional microarray analysis methods and emphasizes the significant advantages of a systems-based approach to analyzing complex disease.
The successful completion of the Human Genome Project and the achievement of similar goals in other species have generated a huge amount of free available information about the genomic sequence of different organisms, opening the door to a postgenome era where new challenges arise. One of the most ambitious objectives of this new period, addressed by the emerging discipline of functional genomics, attempts to understand the genome and the products it encodes for, and how these gene products interact to produce complex living organisms. This new era is also characterized by the development of new technologies, which have produced genomic tools indispensable for understanding how gene products are regulated in normal and diseased conditions on a global genome scale. One of these technologies is DNA microarrays, turned into a very popular tool in the last years. Although the most common use of DNA microarrays is gene expression profiling, scientists have successfully used them for multiple applications, including genotyping, sequencing, DNA copy number analysis, and DNA-protein interactions, among others. In summary, DNA microarrays are changing the way biomedicine and other disciplines are addressing different biological questions and will allow the translation of genome research to the clinic.
Phenoloxidase (PO) is a major component of the insect immune system. The enzyme is involved in encapsulation and melanization processes as well as wound healing and cuticle sclerotization. PO is present as an inactive proenzyme, prophenoloxidase (PPO), which is activated via a protease cascade. In this study, we have cloned a full-length PPO1 cDNA and a partial PPO2 cDNA from the Indianmeal moth, Plodia interpunctella (Hubner) (Lepidoptera: Pyralidae) and documented changes in PO activity in larvae paralyzed and parasitized by the ectoparasitoid Habrobracon hebetor (Say) (Hymenoptera: Braconidae). The cDNA for PPO1 is 2,748 bp and encodes a protein of 681 amino acids with a calculated molecular weight of 78,328 and pI of 6.41 containing a conserved proteolytic cleavage site found in other PPOs. P. interpunctella PPO1 ranges from 71-78% identical to other known lepidopteran PPO-1 sequences. Percent identity decreases as comparisons are made to PPO-1 of more divergent species in the orders Diptera (Aa-48; As-49; and Sb-60%) and Coleoptera (Tm-58; Hd-50%). Paralyzation of host larvae of P. interpunctella by the idiobiont H. hebetor results in an increase in phenoloxidase activity in host hemolymph, a process that may protect the host from microbial infection during self-provisioning by this wasp. Subsequent parasitization by H. hebetor larvae causes a decrease in hemolymph PO activity, which suggests that the larval parasitoid may be secreting an immunosuppressant into the host larva during feeding.
BACKGROUND: We present a complete re-implementation of the segment-based approach to multiple protein alignment that contains a number of improvements compared to the previous version 2.2 of DIALIGN. This previous version is superior to Needleman-Wunsch-based multi-alignment programs on locally related sequence sets. However, it is often outperformed by these methods on data sets with global but weak similarity at the primary-sequence level. RESULTS: In the present paper, we discuss strengths and weaknesses of DIALIGN in view of the underlying objective function. Based on these results, we propose several heuristics to improve the segment-based alignment approach. For pairwise alignment, we implemented a fragment-chaining algorithm that favours chains of low-scoring local alignments over isolated high-scoring fragments. For multiple alignment, we use an improved greedy procedure that is less sensitive to spurious local sequence similarities. To evaluate our method on globally related protein families, we used the well-known database BAliBASE. For benchmarking tests on locally related sequences, we created a new reference database called IRMBASE which consists of simulated conserved motifs implanted into non-related random sequences. CONCLUSION: On BAliBASE, our new program performs significantly better than the previous version of DIALIGN and is comparable to the standard global aligner CLUSTAL W, though it is outperformed by some newly developed programs that focus on global alignment. On the locally related test sets in IRMBASE, our method outperforms all other programs that we evaluated.
Using massive cDNA sequencing, proteomics and customized computational biology approaches, we have isolated and identified the most abundant secreted proteins from the salivary glands of the sand fly Lutzomyia longipalpis. Out of 550 randomly isolated clones from a full-length salivary gland cDNA library, we found 143 clusters or families of related proteins. Out of these 143 families, 35 were predicted to be secreted proteins. We confirmed, by Edman degradation of Lu. longipalpis salivary proteins, the presence of 17 proteins from this group. Full-length sequence for 35 cDNA messages for secretory proteins is reported, including an RGD-containing peptide, three members of the yellow-related family of proteins, maxadilan, a PpSP15-related protein, six members of a family of putative anticoagulants, an antigen 5-related protein, a D7-related protein, a cDNA belonging to the Cimex apyrase family of proteins, a protein homologous to a silk protein with amino acid repeats resembling extracellular matrix proteins, a 5'-nucleotidase, a peptidase, a palmitoyl-hydrolase, an endonuclease, nine novel peptides and four different groups of proteins with no homologies to any protein deposited in accessible databases. Sixteen of these proteins appear to be unique to sand flies. With this approach, we have tripled the number of isolated secretory proteins from this sand fly. Because of the relationship between the vertebrate host immune response to salivary proteins and protection to parasite infection, these proteins are promising markers for vector exposure and attractive targets for vaccine development to control Leishmania chagasi infection.
Native bovine parathyroid hormone (bPTH) was found to be readily cleaved with human leukocyte elastase to yield the fragments bPTH(1-41) and bPTH(42-84). These were then isolated by reverse-phase HPLC and characterised by gas-phase sequencing and amino acid analysis. The biological activities of these fragments were assessed in an adenylate cyclase bioassay using the rat osteosarcoma cell line UMR106. bPTH(1-41) was found to have approximately twice the molar potency of the native hormone from which it was derived. bPTH(42-84) had no biological activity and did not modulate the adenylate cyclase response to these cells to the native hormone. The possible physiological significance of these observations is discussed.
The Coronaviridae family is characterized by a nucleocapsid that is composed of the genome RNA molecule in combination with the nucleoprotein (N protein) within a virion. The most striking physiochemical feature of the N protein of SARS-CoV is that it is a typical basic protein with a high predicted pI and high hydrophilicity, which is consistent with its function of binding to the ribophosphate backbone of the RNA molecule. The predicted high extent of phosphorylation of the N protein on multiple candidate phosphorylation sites demonstrates that it would be related to important functions, such as RNA-binding and localization to the nucleolus of host cells. Subsequent study shows that there is an SR-rich region in the N protein and this region might be involved in the protein-protein interaction. The abundant antigenic sites predicted in the N protein, as well as experimental evidence with synthesized polypeptides, indicate that the N protein is one of the major antigens of the SARS-CoV. Compared with other viral structural proteins, the low variation rate of the N protein with regards to its size suggests its importance to the survival of the virus.