A bagging SVM to learn from positive and unlabeled examples

Statistics – Machine Learning

Scientific paper

Rate now

  [ 0.00 ] – not rated yet Voters 0   Comments 0

Details

Scientific paper

We consider the problem of learning a binary classifier from a training set of positive and unlabeled examples, both in the inductive and in the transductive setting. This problem, often referred to as \emph{PU learning}, differs from the standard supervised classification problem by the lack of negative examples in the training set. It corresponds to an ubiquitous situation in many applications such as information retrieval or gene ranking, when we have identified a set of data of interest sharing a particular property, and we wish to automatically retrieve additional data sharing the same property among a large and easily available pool of unlabeled data. We propose a conceptually simple method, akin to bagging, to approach both inductive and transductive PU learning problems, by converting them into series of supervised binary classification problems discriminating the known positive examples from random subsamples of the unlabeled set. We empirically demonstrate the relevance of the method on simulated and real data, where it performs at least as well as existing methods while being faster.

No associations

LandOfFree

Say what you really think

Search LandOfFree.com for scientists and scientific papers. Rate them and share your experience with other people.

Rating

A bagging SVM to learn from positive and unlabeled examples does not yet have a rating. At this time, there are no reviews or comments for this scientific paper.

If you have personal experience with A bagging SVM to learn from positive and unlabeled examples, we encourage you to share that experience with our LandOfFree.com community. Your opinion is very important and A bagging SVM to learn from positive and unlabeled examples will most certainly appreciate the feedback.

Rate now

     

Profile ID: LFWR-SCP-O-86104

  Search
All data on this website is collected from public sources. Our data reflects the most accurate information available at the time of publication.