Co-training framework for enhancing survey accuracy while reducing respondent burden in travel data collection
提出一种基于协同训练的半监督学习方法,利用少量标注数据和大量未标注数据,在出行方式识别中同时提高预测精度并减少受访者手动标注负担。
A major bottleneck in travel behavior analysis is the need for a substantial amount of labeled data, which typically places a burden on survey respondents for collecting travel behavior data. Our study addresses this issue by leveraging semi-supervised learning, specifically utilizing the co-training algorithm, which effectively incorporates both labeled (active) and unlabeled (passive) data. We extend the semi-supervised learning concept to be a part of the survey scheme that involves both data collection process and enrichment process of travel attributes. Our experiments, focusing on travel mode identification using GPS data from Hiroshima, Japan, demonstrate that our proposed method outperforms existing conventional supervised learning methods such as neural networks, KNN, and SVM, particularly when incorporating an increased proportion of unlabeled data. This strategic use of unlabeled data achieves two apparently conflicting goals: (1) reduces the reliance on extensive manual labeling, thereby alleviating respondent burdens, and (2) increases the accuracy of the prediction. The results of our experiments also reveal that the delicate balance between labeled and unlabeled data proportions plays a pivotal role in co-training performance. Beyond serving as a mode identification tool, our findings underscore the transformative potential of co-training as a valuable data filtering method: By optimizing the interplay between labeled and unlabeled data, co-training efficiently filters noise and refines the dataset. This contributes to enhanced survey accuracy while minimizing labeling burdens. Our results provide useful information to design an adaptive scheme that dynamically tailors the information solicited from respondents to optimize the balance between data quality and respondent burden.