Papers
Topics
Authors
Recent
Search
2000 character limit reached

Automated System for Improving RSS Feeds Data Quality

Published 6 Apr 2015 in cs.IR | (1504.01433v1)

Abstract: Nowadays, the majority of RSS feeds provide incomplete information about their news items. The lack of information leads to engagement loss in users. We present a new automated system for improving the RSS feeds' data quality. RSS feeds provide a list of the latest news items ordered by date. Therefore, it makes it easy for a web crawler to precisely locate the item and extract its raw content. Then it identifies where the main content is located and extracts: main text corpus, relevant keywords, bigrams, best image and predicts the category of the item. The output of the system is an enhanced RSS feed. The proposed system showed an average item data quality improvement from 39.98% to 95.62%.

Citations (5)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Authors (1)

Collections

Sign up for free to add this paper to one or more collections.