PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

A computer method for finding common base paired helices in aligned sequences: application to the analysis of random sequences.

We describe a new computer program that identifies conserved secondary structures in aligned nucleotide sequences of related single-stranded RNAs. The program employs a series of hash tables to identify and sort common base paired helices that are located in identical positions in more than one sequence. The program gives information on the total number of base paired helices that are conserved between related sequences and provides detailed information about common helices that have a minimum of one or more compensating base changes. The program is useful in the analysis of large biological sequences. We have used it to examine the number and type of complementary segments (potential base paired helices) that can be found in common among related random sequences similar in base composition to 16S rRNA from Escherichia coli. Two types of random sequences were analyzed. One set consisted of sequences that were independent but they had the same mononucleotide composition as the 16S rRNA. The second set contained sequences that were 80% similar to one another. Different results were obtained in the analysis of these two types of random sequences. When 5 sequences that were 80% similar to one another were analyzed, significant numbers of potential helices with two or more independent base changes were observed. When 5 independent sequences were analyzed, no potential helices were found in common. The results of the analyses with random sequences were compared with the number and type of helices found in the phylogenetic model of the secondary structure of 16S ribosomal RNA. Many more helices are conserved among the ribosomal sequences than are found in common among similar random sequences. In addition, conserved helices in the 16S rRNAs are, on the average, longer than the complementary segments that are found in comparable random sequences. The significance of these results and their application in the analysis of long non-ribosomal nucleotide sequences is discussed.

Base Composition↗

Pathway logic: symbolic analysis of biological signaling.

The genomic sequencing of hundreds of organisms including homo sapiens, and the exponential growth in gene expression and proteomic data for many species has revolutionized research in biology. However, the computational analysis of these burgeoning datasets has been hampered by the sparse successes in combinations of data sources, representations, and algorithms. Here we propose the application of symbolic toolsets from the formal methods community to problems of biological interest, particularly signaling pathways, and more specifically mammalian mitogenic and stress responsive pathways. The results of formal symbolic analysis with extremely efficient representations of biological networks provide insights with potential biological impact. In particular, novel hypotheses may be generated which could lead to wet lab validation of new signaling possibilities. We demonstrate the graphic representation of the results of formal analysis of pathways, including navigational abilities, and describe the logical underpinnings of the approach. In summary, we propose and provide an initial description of an algebra and logic of signaling pathways and biologically plausible abstractions that provide the foundation for the application of high-powered tools such as model checkers to problems of biological interest.

Animals↗

Detailed protein sequence alignment based on Spectral Similarity Score (SSS).

BACKGROUND: The chemical property and biological function of a protein is a direct consequence of its primary structure. Several algorithms have been developed which determine alignment and similarity of primary protein sequences. However, character based similarity cannot provide insight into the structural aspects of a protein. We present a method based on spectral similarity to compare subsequences of amino acids that behave similarly but are not aligned well by considering amino acids as mere characters. This approach finds a similarity score between sequences based on any given attribute, like hydrophobicity of amino acids, on the basis of spectral information after partial conversion to the frequency domain. RESULTS: Distance matrices of various branches of the human kinome, that is the full complement of human kinases, were developed that matched the phylogenetic tree of the human kinome establishing the efficacy of the global alignment of the algorithm. PKCd and PKCe kinases share close biological properties and structural similarities but do not give high scores with character based alignments. Detailed comparison established close similarities between subsequences that do not have any significant character identity. We compared their known 3D structures to establish that the algorithm is able to pick subsequences that are not considered similar by character based matching algorithms but share structural similarities. Similarly many subsequences with low character identity were picked between xyna-theau and xyna-clotm F/10 xylanases. Comparison of 3D structures of the subsequences confirmed the claim of similarity in structure. CONCLUSION: An algorithm is developed which is inspired by successful application of spectral similarity applied to music sequences. The method captures subsequences that do not align by traditional character based alignment tools but give rise to similar secondary and tertiary structures. The Spectral Similarity Score (SSS) is an extension to the conventional similarity methods and results indicate that it holds a strong potential for analysis of various biological sequences and structural variations in proteins.

Algorithms↗

Bioinformatics in neurosurgery.

WITH THE COMPLETION of the Human Genome Project, the amount of molecular biological sequence data available in public databases has reached staggering proportions. Data continue to accumulate at an exponential rate in the postgenomic era. Compilation, storage, searching, sharing, studying, and transmitting of all these data present formidable challenges. To keep pace with this extant database, the science of bioinformatics (sometimes called computational biology) has evolved. Bioinformatics is the combination of biology and computers and usually involves the storage or analysis of molecular biological sequence data at either the deoxyribonucleic acid, ribonucleic acid, or protein (amino acid) level. Most bioinformatics tools are freely available on the Internet for use by investigators around the globe. The collective wisdom from bioinformatics databases worldwide will continue to spawn advances in the neurological sciences for generations to come. Neurosurgeons must be aware of the power and potential applications of bioinformatics for the analysis of neurosurgical diseases.

Animals↗

Peptide-based analysis of amino acid sequences important to the biological activity of eosinophil granule major basic protein.

Synthetic peptides corresponding to amino acid sequences in eosinophil granule major basic protein (MBP) were evaluated for cytotoxic activity toward K562 cells and for ability to stimulate basophil mediator release. Results obtained using 14 peptides spanning the 117-amino acid sequence of MBP in overlapping fashion indicated that the activities mapped to peptide sequences near the amino and carboxy termini of MBP. The activity of these regions was confirmed using two peptides corresponding to MBP residues 18-45 and 89-117. A 20-h incubation with 5 microM peptide 18-45 or peptide 89-117 caused approximately the same levels (>60%) of cytotoxicity in K562 cells as 5 microM MBP. Similarly, a 30-min incubation with peptides 18-44 and 89-117 stimulated basophil histamine release in a concentration-dependent manner over the range of 5-20 microM. The level of release stimulated by 20 microM peptide 89-117 approached that stimulated by 2 microM MBP. A 20 microM concentration of peptide 89-117 also stimulated leukotriene C4 (LTC4) production by the basophils. Neither peptide 18-45 nor peptide 89-117 was cytotoxic for basophils under the experimental conditions for histamine and LTC4 release, as determined by 51Cr release. These results indicate that two MBP peptide sequences, including one (89-117) that contains a unique carbohydrate-binding region, share the biologic activities of MBP.

Amino Acid Sequence↗

Using a comb filter to describe time-varying biological rhythmicities.

A very important problem in the analysis of biological data sequences is the detection of oscillations in the presence of random variations (noise). If the oscillations are not stationary, i.e., if they drift in frequency and amplitude, or occur in bursts, traditional analysis techniques utilizing the power spectrum or its time-domain equivalent, the autocorrelation function, can be both misleading and insensitive. Temporal filtering by a "comb" or se of band-pass filters is very effective for identifying and describing nonstationary oscillations. The basic procedures for interpreting the output of a comb filter are presented here, illustrated by examples using predefined input test sequences.

Computers↗

Biological applications of support vector machines.

One of the major tasks in bioinformatics is the classification and prediction of biological data. With the rapid increase in size of the biological databanks, it is essential to use computer programs to automate the classification process. At present, the computer programs that give the best prediction performance are support vector machines (SVMs). This is because SVMs are designed to maximise the margin to separate two classes so that the trained model generalises well on unseen data. Most other computer programs implement a classifier through the minimisation of error occurred in training, which leads to poorer generalisation. Because of this, SVMs have been widely applied to many areas of bioinformatics including protein function prediction, protease functional site recognition, transcription initiation site prediction and gene expression data classification. This paper will discuss the principles of SVMs and the applications of SVMs to the analysis of biological data, mainly protein and DNA sequences.

Algorithms↗

Genetic and metabolic diversity of Trichoderma: a case study on South-East Asian isolates.

We have used isolates of Trichoderma spp. collected in South-East Asia, including Taiwan and Western Indonesia, to assess the genetic and metabolic diversity of endemic species of Trichoderma. Ninety-six strains were isolated in total, and identified at the species level by analysis of morphological and biochemical characters (Biolog system), and by sequence analysis of their internal transcribed spacer regions 1 and 2 (ITS1 and 2) of the rDNA cluster, using ex-type strains and taxonomically established isolates of Trichoderma as reference. Seventy-eight isolates were positively identified as Trichoderma harzianum/Trichoderma inhamatum (37 strains) Trichoderma virens (16 strains), Trichoderma spirale (8 strains), Trichoderma koningii (3 strains), Trichoderma atroviride (3 strains), Trichoderma asperellum (4 strains), Hypocrea jecorina (anamorph: Trichoderma reesei; 2 strains), Trichoderma viride (2 strains), Trichoderma hamatum (1 strain), and Trichoderma ghanense (1 strain). Analysis of biochemical characters revealed that T. virens, T. spirale, T. asperellum, T. koningii, H. jecorina, and T. ghanense formed clearly defined clusters, thus exhibiting species-specific metabolic properties. In biochemical character analysis T. atroviride and T. viride formed partially overlapping clusters, indicating that these two species may share overlapping metabolic characteristics. This behavior was even more striking with T. harzianum/T. inhamatum where genotypes defined on the basis of ITS1 and 2 sequences overlapped significantly with adjacent genotypes in the biochemical character analysis, and four strains from the same location (Bali, Indonesia) even clustered with species from section Longibrachiatum. The data indicate that the T. harzianum/T. inhamatum group represents species with high metabolic diversity and partially unique metabolic characteristics. Nineteen strains yielded three different ITS1/2 sequence types which were not alignable with any known species. They were also uniquely characterized by morphological and biochemical characters and therefore represent three new taxa of Trichoderma.

Asia, Southeastern↗

A survey of nucleic acid services in core laboratories.

Core facility services related to DNA synthesis and sequencing were surveyed by the Association of Biomolecular Resource Facilities. Responses from 85 facilities offering DNA synthesis and 37 facilities offering DNA sequencing were obtained. Data on instrumentation, volume, number of users, cost, methodology and a number of other criteria were obtained. The volume of work performed by these centralized core facilities was quite substantial (combined synthesis output of 4 million bases per year and a combined sequencing output of 35 million bases per year). The large number of users supported by these facilities and the high sample throughput make these core resource facilities good indicators of technological trends.

Costs and Cost Analysis↗

Applying signal theory to the analysis of biomolecules.

MOTIVATION: The accumulation of sequence-related and other biological data for basic research and application purposes invites disaster. It appears very likely that neither traditional thinking nor current technologies (including their foreseeable evolutionary developments) will be able to cope with this ever intensifying situation. RESULTS: We present the detailed theoretical background for applying signal theory, as known from speech recognition and image analysis, to the analysis of biomolecules. The general scheme is as follows: biochemical and biophysical properties of biomolecules are used to model an n-dimensional signal which represents the entire information-bearing biomolecule. Such signals are used to search for biological principles, analogies or similarities between biomolecules. In a series of simple experiments (bacterial DNA, generation of real signals using melting enthalpies, detection filtering by convolution of signals) we have shown that the novel system for comparative analysis of the properties of information-bearing biomolecules works as in theory. SUPPLEMENTARY INFORMATION: http://genome.gbf.de/wavepaper.

Algorithms↗

A generic motif discovery algorithm for sequential data.

MOTIVATION: Motif discovery in sequential data is a problem of great interest and with many applications. However, previous methods have been unable to combine exhaustive search with complex motif representations and are each typically only applicable to a certain class of problems. RESULTS: Here we present a generic motif discovery algorithm (Gemoda) for sequential data. Gemoda can be applied to any dataset with a sequential character, including both categorical and real-valued data. As we show, Gemoda deterministically discovers motifs that are maximal in composition and length. As well, the algorithm allows any choice of similarity metric for finding motifs. Finally, Gemoda's output motifs are representation-agnostic: they can be represented using regular expressions, position weight matrices or any number of other models for any type of sequential data. We demonstrate a number of applications of the algorithm, including the discovery of motifs in amino acids sequences, a new solution to the (l,d)-motif problem in DNA sequences and the discovery of conserved protein substructures. AVAILABILITY: Gemoda is freely available at http://web.mit.edu/bamel/gemoda

Algorithms↗

CryoSCAPE: Scalable immune profiling using cryopreserved whole blood for multi-omic single cell and functional assays.

BACKGROUND: The field of single cell technologies has rapidly advanced our comprehension of the human immune system, offering unprecedented insights into cellular heterogeneity and immune function. While cryopreserved peripheral blood mononuclear cell (PBMC) samples enable deep characterization of immune cells, challenges in clinical isolation and preservation limit their application in underserved communities with limited access to research facilities. We present CryoSCAPE (Cryopreservation for Scalable Cellular And Proteomic Exploration), a scalable method for immune studies of human PBMC with multi-omic single cell assays using direct cryopreservation of whole blood. RESULTS: Comparative analyses of matched human PBMC from cryopreserved whole blood and density gradient isolation demonstrate the efficacy of this methodology in capturing cell proportions and molecular features. The method was then optimized and verified for high sample throughput using fixed single cell RNA sequencing and liquid handling automation with a single batch of 60 cryopreserved whole blood samples. Additionally, cryopreserved whole blood was demonstrated to be compatible with functional assays, enabling this sample preservation method for clinical research. CONCLUSIONS: The CryoSCAPE method, optimized for scalability and cost-effectiveness, allows for high-throughput single cell RNA sequencing and functional assays while minimizing sample handling challenges. Utilization of this method in the clinic has the potential to democratize access to single-cell assays and enhance our understanding of immune function across diverse populations.

Humans↗

Human hair histogenesis for the mitochondrial DNA forensic scientist.

Analysis of mitochondrial DNA (mtDNA) sequence from human hairs has proven to be a valuable complement to traditional hair comparison microscopy in forensic cases when nuclear DNA typing is not possible. However, while much is known about the specialties of hair biology and mtDNA sequence analysis, there has been little correlation of individual information. Hair microscopy and hair embryogenesis are subjects that are sometimes unfamiliar to the forensic DNA scientist. The continual growth and replacement of human hairs involves complex cellular transformation and regeneration events. In turn, the analysis of mtDNA sequence data can involve complex questions of interpretation (e.g., heteroplasmy and the sequence variation it may cause within an individual, or between related individuals. In this paper we review the details of hair developmental histology, including the migration of mitochondria in the growing hair, and the related interpretation issues regarding the analysis of mtDNA data in hair. Macroscopic and microscopic hair specimen classifications are provided as a possible guide to help forensic scientists better associate mtDNA sequence heteroplasmy data with the physical characteristics of a hair. These same hair specimen classifications may also be useful when evaluating the relative success in sequencing different types and/or forms of human hairs. The ultimate goal of this review is to bring the hair microscopist and forensic DNA scientist closer together, as the use of mtDNA sequence analysis continues to expand.

Cell Differentiation↗

Nonlinear analysis of biological systems using short M-sequences and sparse-stimulation techniques.

The m-sequence pseudorandom signal has been shown to be a more effective probing signal than traditional Gaussian white noise for studying nonlinear biological systems using cross-correlation techniques. The effectiveness is evidenced by the high signal-to-noise (S/N) ratio and speed of data acquisition. However, the "anomalies" that occur in the estimations of the cross-correlations represent an obstacle that prevents m-sequences from being more widely used for studying nonlinear systems. The sparse-stimulation method for measuring system kernels can help alleviate estimation errors caused by anomalies. In this paper, a "padded sparse-stimulation" method is evaluated, a modification of the "inserted sparse-stimulation" technique introduced by Sutter, along with a short m-sequence as a probing signal. Computer simulations show that both the "padded" and "inserted" methods can effectively eliminate the anomalies in the calculation of the second-order kernel, even when short m-sequences were used (length of 1023 for a binary m-sequence, and 728 for a ternary m-sequence). Preliminary experimental data from neuromagnetic studies of the human visual system are also presented, demonstrating that the system kernels can be measured with high signal-to-noise (S/N) ratios using short m-sequences.

Computer Simulation↗

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats↗

BGI-RIS: an integrated information resource and comparative analysis workbench for rice genomics.

Rice is a major food staple for the world's population and serves as a model species in cereal genome research. The Beijing Genomics Institute (BGI) has long been devoting itself to sequencing, information analysis and biological research of the rice and other crop genomes. In order to facilitate the application of the rice genomic information and to provide a foundation for functional and evolutionary studies of other important cereal crops, we implemented our Rice Information System (BGI-RIS), the most up-to-date integrated information resource as well as a workbench for comparative genomic analysis. In addition to comprehensive data from Oryza sativa L. ssp. indica sequenced by BGI, BGI-RIS also hosts carefully curated genome information from Oryza sativa L. ssp. japonica and EST sequences available from other cereal crops. In this resource, sequence contigs of indica (93-11) have been further assembled into Mbp-sized scaffolds and anchored onto the rice chromosomes referenced to physical/genetic markers, cDNAs and BAC-end sequences. We have annotated the rice genomes for gene content, repetitive elements, gene duplications (tandem and segmental) and single nucleotide polymorphisms between rice subspecies. Designed as a basic platform, BGI-RIS presents the sequenced genomes and related information in systematic and graphical ways for the convenience of in-depth comparative studies (http://rise.genomics.org.cn/).

China↗

The R protein of SARS-CoV: analyses of structure and function based on four complete genome sequences of isolates BJ01-BJ04.

The R (replicase) protein is the uniquely defined non-structural protein (NSP) responsible for RNA replication, mutation rate or fidelity, regulation of transcription in coronaviruses and many other ssRNA viruses. Based on our complete genome sequences of four isolates (BJ01-BJ04) of SARS-CoV from Beijing, China, we analyzed the structure and predicted functions of the R protein in comparison with 13 other isolates of SARS-CoV and 6 other coronaviruses. The entire ORF (open-reading frame) encodes for two major enzyme activities, RNA-dependent RNA polymerase (RdRp) and proteinase activities. The R polyprotein undergoes a complex proteolytic process to produce 15 function-related peptides. A hydrophobic domain (HOD) and a hydrophilic domain (HID) are newly identified within NSP1. The substitution rate of the R protein is close to the average of the SARS-CoV genome. The functional domains in all NSPs of the R protein give different phylogenetic results that suggest their different mutation rate under selective pressure. Eleven highly conserved regions in RdRp and twelve cleavage sites by 3CLP (chymotrypsin-like protein) have been identified as potential drug targets. Findings suggest that it is possible to obtain information about the phylogeny of SARS-CoV, as well as potential tools for drug design, genotyping and diagnostics of SARS.

Amino Acid Sequence↗

Comparative analysis of the Kekkon molecules, related members of the LIG superfamily.

Leucine-rich repeats (LRRs) and immunoglobulin (Ig) domains represent two of the most abundant sequence elements in metazoan proteomes. Despite this prevalence, comparatively few molecules containing both LRR and Ig (LIG) modules exist, and fewer still have been functionally defined. One LIG whose function has been investigated is the Drosophila protein Kekkon1 (Kek1). In vivo studies have demonstrated a role for Kek1 in Epidermal Growth Factor Receptor (EGFR) signaling and have suggested a role in neuronal pathfinding. Kek1 is the founding member of the Kek family, a group of six Drosophila transmembrane proteins that contain seven LRRs and a single Ig in their extracellular domains. While this arrangement of domains predicts a possible role as cell adhesion molecules (CAMs), to date little is known about the function or evolutionary relationship of these additional Kek molecules. Here we report that orthologs of Kek1, Kek2, Kek5, and Kek6 exist in the mosquito, Anopheles gambiae, and the honeybee, Apis mellifera, indicating that this family has been conserved for ~300 million years of evolutionary time. Comparative sequence analyses reveal remarkable identity among these orthologs, primarily in their extracellular regions. In contrast, the intracellular regions are more divergent, exhibiting only small pockets of conservation. In addition, we provide support for the general notion that these molecules may share common functions as CAMs, by demonstrating that Kek family members can form homotypic and heterotypic complexes.

Amino Acid Sequence↗