PubMed HealthSearch

SEARCH · PubMed Health

Results for “protein sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Emerging protein sequencing technologies: proteomics without mass spectrometry?

INTRODUCTION: Liquid chromatography-tandem mass spectrometry (LC-MS/MS) has been a leading method for proteomics for 30 years. Advantages provided by LC-MS/MS are offset by significant disadvantages, including cost. Recently, several non-mass spectrometric methods have emerged, but little information is available about their capacity to analyze the complex mixtures routine for mass spectrometry. AREAS COVERED: We review recent non-mass-spectrometric methods for sequencing proteins and peptides, including those using nanopores, sequencing by degradation, reverse translation, and short-epitope mapping, with comments on bioinformatics challenges, fundamental limitations, and areas where new technologies will be more or less competitive with LC-MS/MS. In addition to conventional literature searches, instrument vendor websites, patents, webinars, and preprints were also consulted to give a more up-to-date picture. EXPERT OPINION: Many new technologies are promising. However, demonstrations that they outperform mass spectrometry in terms of peptides and proteins identified have not yet been published, and astute observers note important disadvantages, especially relating to the dynamic range of single-molecule measurements of complex mixtures. Still, even if the performance of emerging methods proves inferior to LC-MS/MS, their low cost could create a different kind of revolution: a dramatic increase in the number of biology laboratories engaging in new forms of proteomics research.

Proteomics

GBSC: graph-based sequence clustering method for similar short tandem repeats in protein sequences.

MOTIVATION: Short tandem repeats (STRs) are abundant in protein sequences and play important role in determining their structures and functions. Strikingly, the unusual compositional characteristics of tandem repeats break classical sequence analysis tools. RESULTS: Here, we establish the first algorithm to effectively identify and cluster STRs: Graph-Based Sequence Clustering (GBSC) features linear time complexity, and clusters protein sequence fragments based on their STRs, while allowing for insertions and mutations and supporting the analysis of imperfect or cryptic repeats. Due to its computational efficacy, our algorithm can be used to systematically scan for patterns in large datasets. We compare our method both to state-of-the-art methods for identifying STRs in proteins and alternative clustering approaches. Unlike existing STR analysis methods, GBSC clusters repeat patterns rather than raw sequences, operating at the level of structural repeat identity, while tolerating biological variations and preventing erroneous merging of structurally and functionally distinct motifs. Whereas functional annotation is typically only available at the protein level, the functions of individual STRs and sequences of adjacent STRs remain largely unknown. On a challenging use case we here demonstrate and discuss how our method can be used to associate previously unannotated repetitive protein fragments with similar ones, allowing the transfer of annotation by similarity. For the first time, GBSC offers a tool that systematically extends this fundamental bioinformatics principle to low-complexity regions across large datasets. AVAILABILITY AND IMPLEMENTATION: GBSC is available at GitHub https://github.com/patryk-jarnot/GBSC and https://doi.org/10.5281/zenodo.18965247. The data and scripts to reproduce the analysis are available at https://doi.org/10.5281/zenodo.16906653.

Microsatellite Repeats

Alignment statistic for identifying related protein sequences.

Closely related proteins show an obvious kinship by having numerous matching amino acids in their aligned sequences. Kinship between anciently separated proteins requires a statistical evaluation to rule out fortuitous similarities. A simple statistic is developed which assumes equal probability for all codon pairs, and a table of critical values for amino acid sequence alignments of lengthnments of length 200 or less is presented. Applying this statistic to V and C regions of immunoglobulin chains, aligned on the basis of shared features of three-dimensional structure, provides evidence that the V and C sequences descended from a common ancestor. Similarly the distant evolutionary relationship of dehydrogenases, flavdoxin, and subtilisin, suggested by structural alignments, is verified. On the other hand, the statistic does not verify a common evolutionary origin for the heme binding pocket in globins and cytochrome bs. Empirical evidence from the distribution of MMD values of amino acid pairs in comparisons of misaligned polypeptide chains and from Monte Carlo trials of sequences aligned with arbitrary gaps supports the validity of the statistic.

Amino Acid Sequence

A comprehensive examination of protein sequences for evidence of internal gene duplication.

We have implemented a routine procedure for screening protein sequences for evidence of intragenic duplications. We tested 163 protein sequences representing 116 superfamilies of unrelated proteins. Twenty superfamilies contain proteins with internal gene duplications. The intragenic duplications detected can be divided into two major types. (1) One or more duplications of all or part of a gene produce a protein with two or several detectable regions of sequence homology. Sequences from 18 superfamilies contained this type of duplication. (2) Repeated reduplication of a small DNA segment can produce a protein that is repetitive over most of its length. Three superfamilies contain such repetitive sequences. We also investigated the limits of detection of ancient duplications using sequences derived by random mutation of a model sequence consisting of ten 10-residue repeats. The original repetitive nature of the sequence was usually detected after 250 point mutations even though the ancestral segment could not be accurately reconstructed.

Amino Acid Sequence

The physical and evolutionary energy landscapes of devolved protein sequences corresponding to pseudogenes.

Protein evolution is guided by structural, functional, and dynamical constraints ensuring organismal viability. Pseudogenes are genomic sequences identified in many eukaryotes that lack translational activity due to sequence degradation and thus over time have undergone "devolution." Previously pseudogenized genes sometimes regain their protein-coding function, suggesting they may still encode robust folding energy landscapes despite multiple mutations. We study both the physical folding landscapes of protein sequences corresponding to human pseudogenes using the Associative Memory, Water Mediated, Structure and Energy Model, and the evolutionary energy landscapes obtained using direct coupling analysis (DCA) on their parent protein families. We found that generally mutations that have occurred in pseudogene sequences have disrupted their native global network of stabilizing residue interactions, making it harder for them to fold if they were translated. In some cases, however, energetic frustration has apparently decreased when the functional constraints were removed. We analyzed this unexpected situation for Cyclophilin A, Profilin-1, and Small Ubiquitin-like Modifier 2 Protein. Our analysis reveals that when such mutations in the pseudogene ultimately stabilize folding, at the same time, they likely alter the pseudogenes' former biological activity, as estimated by DCA. We localize most of these stabilizing mutations generally to normally frustrated regions required for binding to other partners.

Cyclophilin A

A novel manual method for protein-sequence analysis.

A novel manual method for protein-sequence analysis is described. Three peptides, the hexapeptide (Leu-TRP-Met-Arg-Phe-Ala), insulin A chain and glucagon were used to test this technique. Peptides (1 or 2 nmol) were hydrolysed with acid and their qualitative amino acid compositions were confirmed by reacting with 4-NN-dimethylaminoazobenzene-4'-sulphonylchloride and 4-NN-dimethylaminoazobenzene 4'-isothiocyanate. Sequence determination of 20-200 nmol of peptide was then performed by the combined use of phenyl isothiocyanate and 4-NN-dimethylaminoazobenzene 4'-isothiocyanate, a new procedure that is analogous to the dansyl-Edman method with the replacement of dansyl chloride by 4-NN-dimethylaminoazobenzene 4'-isothiocyanate as the N-terminal residue determination reagent. On t.l.c. this new N-terminal reagent gave brightly coloured 4-NN-dimethylaminoazobenzene-4-thiohydantoins of amino acids and showed the following advantages: (1) the detection sensitivity is in the pmol range; (2) u.v. observation is not required; (3) there is no destruction of acid-labile amino acids; (4) two-dimensional t.l.c. separation is adequate to identify 24 amino acids, except leucine and isoleucine (this pair of amino acids can be resolved by using 4-NN-dimethylaminoazobenzene-4'-sulphonyl chloride); (5) the determination of a new N-terminal residue (from coupling to t.l.c. identification) takes only 3 h; (6) the colour difference beteen isothiocyanate, thiocarbamoyl and thiohydantoin derivatives facilitates the identifications.

Amino Acid Sequence

Evolution of homologous physiological mechanisms based on protein sequence data.

1. Genetic duplications can give rise to homologous physiological mechanisms that include structurally related protein components. There are many such examples of related proteins within the human body. 2. Evolutionary histories showing the origins and subsequent divergences of these distantly related proteins can be derived from the protein sequences and correlated with the functional characteristics of these proteins. 3. The hormones related to glucagon provide an example of homology of physiological mechanisms and emergence of new functions subsequent to gene duplications. 4. The proteins related to troponin C illustrate the participation of distantly related proteins in the same mechanism (muscle contraction), the relationship of proteins characteristic of a specialized tissue to proteins found in all eukaryote cells, and the correlation of genetic duplications with the evolutionary appearance of different types of muscle.

Amino Acid Sequence

Adenovirus-associated virus structural protein sequence homology.

Adenovirus-associated virus (AAV) structural proteins (VP1, VP2, and VP3) have been examined to determine if areas of sequence homology exist between these three virion proteins. Tryptic and chymotryptic maps have been produced which demonstrate extensive areas of sequence homology common to all three proteins. The amino acid compositions of the proteins were also determined and were found to be very similar. These data are consistent with the hypothesis that all three virion proteins arise either from a common precursor of similar transcripts.

Amino Acid Sequence

Silkmoth chorion proteins: sequence analysis of the products of a multigene family.

Five polypeptide components have been isolated from the eggshell (chorions) of a silkmoth. Two are homogeneous on sodium dodecyl sulfate and isoelectric focusing gels, and three contain predominantly two proteins each. Amino acid analyses show that all five components are similar to each other. These proteins have been sequenced from the amino terminus. Homogeneous components yielded single sequences; heterogeneous components yielded two residues at some positions, consistent with their containing two major electrophoretic components. Striking similarities are apparent among all these sequences. These similarities can be increased dramatically by separating each of the three protein mixtures into two sequences and introducing a small number of gaps or insertions. This is due in part to bringing into register a portion that contains short repeating subunits found in all sequences. All proteins are also characterized by a region of high cysteine content near the amino terminus followed by a longer low-cysteine region. The data suggest that these proteins share a common evolutionary origin and are encoded by a multigene family.

Amino Acid Sequence

Protein sequencing by computer graphics.

A computer graphics system has been used to fit a putative sequence to an electron density map of a sea snake neurotoxic protein at 2.2 A resolution. The complete sequence of this small protein could be determined from the map with very little ambiguity. In two places probable errors in the published chemical sequence were detected. This is the first instance in which a complete three-dimensional structure was solved with the use of computer graphics alone, without the construction of a physical model.

Amino Acid Sequence

Transfer ribonucleic acids from eleven immunoglobulin-secreting mouse plasmacytomas. Constant and variable chromatographic profiles compared with the myeloma protein sequences.

In order to test the concepts that aminoacyl-tRNAs in plasmacytomas may on the one hand modulate the protein synthesized or on the other hand reflect the structure of the synthesized protein, the RPC-5 chromatographic profiles of aminoacyl-tRNAs for all 20 amino acids were studied in tRNA prepared from normal mouse liver and 11 plasmacytomas. The patterns of isoaccepting tRNA were compared with the structure of the myeloma protein being synthesized. The elution profiles of aminoacyl-tRNAs for nine of the amino acids were constant, i.e. they were the same for liver and all plasmacytomas. Significant variability was observed in the profiles of the other 11 families of aminoacyl-tRNAs: asparagine, serine and tryptophan, had peaks of isoaccepting tRNAs found in tumors and not in liver; glutamic acid, histidine and lysine, had different patterns of aminoacyl-tRNAs in plasmacytomas which could be distinguished from the elution profile of liver; and isoleucine, proline, threonine and tyrosine, showed pattern variability in only a few of the tumors. Valyl-tRNA uniquely had one isoacceptor present in liver but absent in the tumors. This variability is thought to be associated with different posttranscriptional modification of the tRNAs rather than regulation of individual tRNA genes in response to particular amino acid sequences in secreted myeloma proteins. Similarily, the lack of correlation of isoacceptors with sequence differences makes the modulation of protein fine structure by tRNA availability unlikely.

Amino Acid Sequence

The evolution of protein sequences by repetitious gene duplication: clostridial flavodoxin.

Internal regularities of amino acid sequences of flavodoxins, FMN-containing, low molecular weight flavoproteins, were statistically examined using the minimum mutation method. The sequence of Clostridium pasteurianum flavodoxin shows statistically significant evidence of repetitious internal gene duplications at different levels of structure. Peptide pairs with a low chance probabilitiy of occurrence were frequently observed at a shift of 5 residues. The pairs with the lowest chance probabilities are a pair of heptapeptides at positions39--45 vs. 44--50, a 5 residue shift (p = 9 x 10(-6)). Most of the related pairs are consistent and could best be explained by the repeating pentapeptide sequence: (Lys-Gly-Ala-Asp-Val-)n and appropriate gaps. Internal repetitions with longer shifts were also suggested for other flavodoxins. Repetitious gene duplication is proposed for the early stages of flavodoxin evolution.

Amino Acid Sequence

Amino acid sequence of the basic trypsin inhibitor from bovine splenic capsule. A new fluorescent reagent for protein sequence analysis.

The amino acid sequences of tryptic peptides from the performic acid-oxidized trypsin inhibitor were determined. The degradation was performed with a new reagent, 3-isothiocyanato-4-methoxy-4'-nitrostilbene, and compared with the dansyl-Edman technique. The amino acid sequence of the trypsin inhibitor from bovine splenic capsule was found to be the same as that of the basic trypsin inhibitor from bovine pancreas, established by Kress & Laskowski (J. Biol. Chem. 1967, 242, 4925-4928).

Amino Acid Sequence

Evolution of two major chorion multigene families as inferred from cloned cDNA and protein sequences.

Complete or partial sequences are reported from six chorion cDNA clones of the silkmoth Antheraea polyphemus. The proteins encoded belong to the two major chorion protein classes, A and B, each of which is encoded by a multigene family. The sequence comparisons define some major features of the families and suggest how these genes may be evolving. Deletions and insertions might be involved in expanding or contracting internally repetitive regions. Sequence divergence is localized, thus defining sequence domains of distinct evolutionary properties and presumably distinct functions.

Amino Acid Sequence

4-NN-dimethylaminoazobenzene 4'-isothiocyanate, a new chromophoric reagent for protein sequence analysis.

4-NN-Dimethylaminoazobenzene 4'-isothiocyanate was synthesized for the purpose of improving the ease and sensitivity of peptide sequence analysis. The method of 4-NN-dimethylaminoazobenzene 4'-isothiocyanate synthesis, the preparation of 24 4-NN-dimethylaminoazobenzene-4'-thiohydantoins of amino acids and their t.l.c. separation are described. All the thiohydantoins, except those of leucine and isoleucine, could be satisfactorily separated by chromatography on a two-dimensional polyamide sheet. The sensitive azo group permits the detection of 4-NN-dimethylaminoazobenzene-4'-thiohydantoins of amino acids as red spots down to pmol amounts directly on the sheet. A simple sensitive method for sequencing dipeptides and the first two or three N-terminal amino acids of proteins is also reported. The colour change of the spots from purple to blue to red after being exposed to HCl vapour, corresponding to the chemical change from 4-NN-dimethylaminoazobenzene-4' isothiocyanate to the 4-NN-dimethylaminoazobenzene-4'-thiocarbamoyl amino acid derivative to the 4-NN-dimethylaminoazobenzene-4'-thiohydantoin amino acid derivative, reveals a very interesting and valuable feature of this reagent.

Amino Acid Sequence