Skip to search boxSkip to navigationSkip to main content

A quantitative sequence-aggregation relationship predictor applied as identification of self-assembled hexapeptides

  • Chen Chen
    ,
  • Yonglan Liu
    ,
  • Jin Zhang
    ,
  • Mingzhen Zhang
    ,
  • Jie Zheng
    ,
  • Yong Teng(corresponding author)
*Corresponding author for this work
  • Chongqing University
    ,
  • University of Akron
Scholary Output:
Contribution to journal
Article
Peer-review

Abstract

It is essential to predict aggregation-forming sequences for elucidation of protein misfolding mechanisms and the design of effective antiamyloid inhibitors. In this work, we predict and characterize self-assembled hexapeptides by a quantitative sequence-aggregation relationship (QSAR) model, which involves characterization of factor analysis scale of generalized amino acid information (FASGAI) and modeling of supporting vector machine (SVM) with radial basis function kernel. The QSAR model achieves maximum accuracy of 78.33% and area under the receiver operating characteristic curve of 0.83 with leave one out cross-validation on 180 training hexapeptides. We determine "hotspots" and key factors that largely contribute to the self-assembly of these hexapeptides by analyzing their sequence-aggregation relationships. We also explore the applications of the present model, e.g., the first is to identify the aggregation-forming sequences within both β-amyloid peptide (Aβ42) and human islet amyloid polypeptide (hIAPP) using a 6-residue slide window, and acquire good agreement with previous experimental observations, the second is to perform in silico design of potential aggregation-forming hexapeptides which are validated by all-atom molecular dynamics simulation and density functional theory calculations, and the third is to predict the potential self-assembled tri-, tetra- and pentapeptides, in which hydrophobic amino acids such as isoleucine, leucine, valine, phenylalanine, and methionine occur at higher frequencies. The present QSAR model is helpful for (i) predicting self-assembled behaviors of peptides, (ii) scanning and identifying aggregation-forming sequences within proteins, (iii) understanding action mechanisms of peptide/protein aggregation, and (iv) designing potential self-assembled sequences applied as drug discovery and nano-materials.

Publication Information

Output type

Scholary Output:
Contribution to journal
Article
Peer-review

Original language

English (US)

Pages from-to (Number of pages)

Pages 7-16 (10 pages)

Journal (Volume, Issue Number)

Chemometrics and Intelligent Laboratory Systems (Volume 145)

Publication milestones

  • Published - 07/05/2015

Publication status

Published - 07/05/2015

ISSN

0169-7439

Publication IDs

  • Scopus: 84928597737
  • ORCID: /0000-0002-1856-7289/work/62481108

Publication metrics

Metrics

Fractional count
1
Fractional count
0.14
Fractional count
6
Fractional count
0.86
Fractional count
1
Fractional count
1
SciVal
citations
7
Scopus
citations
SciVal
FWCI
0.52
SciVal
Author count
7
SciVal
Paper percentile
61

PlumX, opens in new tab

Captures
10
Citation count
9

Funding Details

J.Z. thanks for financial supports from the NSF CBET-1158447 and Alzheimer Association-New Investigator Research Grant ( 2015-NIRG-341372 ). G.L gratefully acknowledges supports of this research by the National Natural Science Foundation of China ( 10901169 ), and the Fundamental Research Funds for the Central Universities ( CQDXWL-2014-Z009 ).
FundersFunding numbers
NSF
2015-NIRG-341372, 1158447, CBET-1158447
NSFC
10901169
Fundamental Research Funds for Central Universities of the Central South University
CQDXWL-2014-Z009