用于正类未标记和标签噪声学习的自适应采样方法及其生物信息学应用

AdaSampling for Positive-Unlabeled and Label Noise Learning With Bioinformatics Applications

IEEE Transactions on Cybernetics · 2018
被引 59
ABS 3

中文导读

提出自适应采样框架,同时处理正类未标记学习和标签噪声问题,通过迭代估计误标记概率降低训练风险,并在模拟和基准数据上验证,最后应用于激酶底物识别和转录因子靶基因预测。

Abstract

Class labels are required for supervised learning but may be corrupted or missing in various applications. In binary classification, for example, when only a subset of positive instances is labeled whereas the remaining are unlabeled, positive-unlabeled (PU) learning is required to model from both positive and unlabeled data. Similarly, when class labels are corrupted by mislabeled instances, methods are needed for learning in the presence of class label noise (LN). Here we propose adaptive sampling (AdaSampling), a framework for both PU learning and learning with class LN. By iteratively estimating the class mislabeling probability with an adaptive sampling procedure, the proposed method progressively reduces the risk of selecting mislabeled instances for model training and subsequently constructs highly generalizable models even when a large proportion of mislabeled instances is present in the data. We demonstrate the utilities of proposed methods using simulation and benchmark data, and compare them to alternative approaches that are commonly used for PU learning and/or learning with LN. We then introduce two novel bioinformatics applications where AdaSampling is used to: 1) identify kinase-substrates from mass spectrometry-based phosphoproteomics data and 2) predict transcription factor target genes by integrating various next-generation sequencing data.

机器学习生物信息学分类学习标签噪声