PubMed HealthSearch

PubMed · 9367768

Do aligned sequences share the same fold?

Abstract

Sequence comparison remains a powerful tool to assess the structural relatedness of two proteins. To develop a sensitive sequence-based procedure for fold recognition, we performed an exhaustive global alignment (with zero end gap penalties) between sequences of protein domains with known three-dimensional folds. The subset of 1.3 million alignments between sequences of structurally unrelated domains was used to derive a set of analytical functions that represent the probability of structural significance for any sequence alignment at a given sequence identity, sequence similarity and alignment score. Analysis of overlap between structurally significant and insignificant alignments shows that sequence identity and sequence similarity measures are poor indicators of structural relatedness in the "twilight zone", while the alignment score allows much better discrimination between alignments of structurally related and unrelated sequences for a wide variety of alignment settings. A fold recognition benchmark was used to compare eight different substitution matrices with eight sets of gap penalties. The best performing matrices were Gonnet and Blosum50 with normalized gap penalties of 2.4/0.15 and 2.0/0.15, respectively, while the positive matrices were the worst performers. The derived functions and parameters can be used for fold recognition via a multilink chain of probability weighted pairwise sequence alignments.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

R A Abagyan, S Batalov. 1997-10-17. Do aligned sequences share the same fold?. https://doi.org/10.1006/jmbi.1997.1287

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

The Virginia generalist initiative: lessons learned in a statewide consortium.

In response to Virginia's need for an increased supply of generalist physicians, the state's three medical schools--Eastern Virginia Medical School, Virginia Commonwealth University School of Medicine, and the University of Virginia School of Medicine--have formed a partnership with key governmental stakeholders in the Virginia Generalist Initiative funded by The Robert Wood Johnson Foundation's Generalist Physician Initiative. These state-supported medical schools historically have functioned independently, with little cooperative effort. This paper describes the consortium, its activities, its successes, and its unmet objectives, and uses a series of cases in point to illustrate relevant lessons learned. Some of these lessons are that (1) stakeholders must be involved from the beginning of planning to identify mutual goals and establish consortium protocols; (2) all partners must share a philosophical commitment to the consortium's mission, as well as the time and resources needed; (3) an atmosphere that enables risk-taking behavior must be created; (4) stakeholders must be willing to revise goals and sustain an environment conductive to change; and (5) trust is essential and must be vigilantly maintained. The paper concludes that the Virginia Generalist Initiative has dramatically altered the goals, objectives and programs of the three schools and has succeeded in aligning the schools' strategic objectives with the state's priorities.

Databases as Topic

Regional diversity and breadth of the National Cancer Data Base.

BACKGROUND: The National Cancer Data Base (NCDB), a joint project of the Commission on Cancer of the American College of Surgeons and the American Cancer Society, is a cancer management and outcomes data base for health care organizations. It provides a comparative summary of patient care that is used by participating hospitals and communities for self-assessment. This article describes the most current (1995) data. METHODS: Since 1989, 7 calls for data have been issued, yielding a total of 5,558,389 cancer patient reports for the years 1985-1995. A total of 1849 hospital cancer registries have participated in at least 1 of the calls for data. RESULTS: One thousand one hundred and fourteen hospitals from 50 states and the District of Columbia reported 655,627 cases for the diagnosis year 1995. The hospitals represented a wide range of sizes (187 [16.8%] with 1000+ cases annually, 405 [36.4%] with 500-999 cases annually, 255 [22.9%] with 300-499 cases annually, 211 [18.9%] with 100-299 cases annually, and 56 [5%] with < 100 cases annually) and types (21 [1.9%] National Cancer Institute [NCI]-recognized cancer centers, 119 [10.7%] government hospitals, 102 [9.2%] teaching hospitals, 256 [23.0%] large community hospitals, 297 [26.7%] medium/small community hospitals, and 257 [23.1%] nongovernmental hospitals without approval status from the Commission on Cancer or NCI recognition). Remarkably similar distributions of cases by primary site and age were reported from each of six U.S. geographic regions. In addition, within each of these six regions, the cases were reported from a wide range of income strata and ethnicities. For several states, relatively few cancer cases were reported. For several examples of relatively rare patient and tumor groups, all reported cases between 1985-1995 included potentially useful quantities of patients in whom further study of such special groups was warranted. CONCLUSIONS: The authors conclude that the reported cases most likely are representative at the regional (but not state) level of cancer patients diagnosed and treated at U.S. hospitals with regard to types of cancer and ages of the patients. They conclude further that cancer reporting may be quite diverse within each region with regard to other known patient and reporting institution characteristics.

Databases as Topic

A comparison of statistical learning methods on the Gusto database.

We apply a battery of modern, adaptive non-linear learning methods to a large real database of cardiac patient data. We use each method to predict 30 day mortality from a large number of potential risk factors, and we compare their performances. We find that none of the methods could outperform a relatively simple logistic regression model previously developed for this problem.

Databases as Topic