Probabilistic variable-length segmentation of protein sequences for discriminative motif discovery (DiMotif) and sequence embedding (ProtVecX)

In this paper, we present peptide-pair encoding (PPE), a general-purpose probabilistic segmentation of protein sequences into commonly occurring variable-length sub-sequences. The idea of PPE segmentation is inspired by the byte-pair encoding (BPE) text compression algorithm, which has recently gain...

Full description

Saved in:
Bibliographic Details
Published in:Scientific reports Vol. 9; no. 1; p. 3577
Main Authors: Asgari, Ehsaneddin, McHardy, Alice C., Mofrad, Mohammad R. K.
Format: Journal Article
Language:English
Published: London Nature Publishing Group UK 05.03.2019
Nature Publishing Group
Subjects:
ISSN:2045-2322, 2045-2322
Online Access:Get full text
Tags: Add Tag
No Tags, Be the first to tag this record!
Be the first to leave a comment!
You must be logged in first