Demo D03 - OSD: An Online Web Spam Detection System
Web spam, which refers to any deliberate actions bringing to selected web pages an unjustifiable favorable relevance or importance, is one of the major obstacles for high quality information retrieval on the web. Most of the existing web spam detection methods are supervised that require a large and representative training set of web pages. Moreover, they often assume some global information such as a large web graph and snapshots of a large collection of web pages. However, in many situations such assumptions may not hold. Recently, we studied the problem of online web spam detection, and proposed the notion of spamicity to measure how likely a page is a spam web page [9, 7]. Spamicity is a more flexible and user-controllable measure than the traditional supervised classification methods. We developed e±cient online link spam and term spam detection methods using spamicity. In this paper, we present a demonstration of OSD, an Online Spam Detection system which can efficiently calculate a spamicity score online for any page on the web.