Learning to Classify with Missing and Corrupted Features (2008)

Authors

Abstract

After a classifier is trained using a machine learning algorithm and put to use in a real world system, it often faces noise which did not appear in the training data. Particularly, some subset of features may be missing or may become corrupted. We present two novel machine learning techniques that are robust to this type of classification-time noise. First, we solve an approximation to the learning problem using linear programming. We analyze the tightness of our approximation and prove statistical risk bounds for this approach. Second, we define the online-learning variant of our problem, address this variant using a modified Perceptron, and obtain a statistical learning algorithm using an online-to-batch technique. We conclude with a set of experiments that demonstrate the effectiveness of our algorithms.

Discussion

Enter your comment (wiki syntax is allowed):
GPNTJ
 
paper/2008/202.txt · Last modified: 2009/05/24 17:48 (external edit)
 
Driven by DokuWiki