References & Citations
Quantitative Biology > Molecular Networks
Title: Network approach integrates 3D structural and sequence data to improve protein classification
(Submitted on 24 May 2016 (this version), latest version 27 Feb 2017 (v2))
Abstract: Motivation: Early approaches for protein (structural) classification were sequence-based. Since amino acids that are distant in the sequence can be close in the 3-dimensional (3D) structure, 3D contact approaches can complement sequence approaches. Traditional 3D contact approaches study 3D structures directly. Instead, 3D structures can first be modeled as protein structure networks (PSNs). Then, network approaches can be used to classify the PSNs. Network approaches may improve upon traditional 3D contact approaches. We cannot use existing PSN approaches to test this, because: 1) They rely on naive measures of network topology that cannot capture the complexity of PSNs. 2) They are not robust to PSN size. They cannot integrate 3) multiple PSN measures or 4) PSN data with sequence data, although this could help because the different data types capture complementary biological knowledge.
Results: We address these limitations by: 1) exploiting well-established graphlet measures via a new network approach for protein classification, 2) introducing novel normalized graphlet measures to remove the bias of PSN size, 3) allowing for integrating multiple PSN measures, and 4) using ordered graphlets to combine the complementary ideas of PSN data and sequence data. We classify both synthetic networks and real-world PSNs more accurately and faster than existing network, 3D contact, or sequence approaches. Our approach finds PSN patterns that may be biochemically interesting.
Submission history
From: Khalique Newaz [view email][v1] Tue, 24 May 2016 00:49:26 GMT (575kb)
[v2] Mon, 27 Feb 2017 21:38:16 GMT (908kb)
Link back to: arXiv, form interface, contact.