We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.IR

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Information Retrieval

Title: How UMass-FSD Inadvertently Leverages Temporal Bias

Abstract: First Story Detection describes the task of identifying new events in a stream of documents. The UMass-FSD system is known for its strong performance in First Story Detection competitions. Recently, it has been frequently used as a high accuracy baseline in research publications. We are the first to discover that UMass-FSD inadvertently leverages temporal bias. Interestingly, the discovered bias contrasts previously known biases and performs significantly better. Our analysis reveals an increased contribution of temporally distant documents, resulting from an unusual way of handling incremental term statistics. We show that this form of temporal bias is also applicable to other well-known First Story Detection systems, where it improves the detection accuracy. To provide a more generalizable conclusion and demonstrate that the observed bias is not only an artefact of a particular implementation, we present a model that intentionally leverages a bias on temporal distance. Our model significantly improves the detection effectiveness of state-of-the-art First Story Detection systems.
Comments: Temporal Bias, First Story Detection, Topic Detection and Tracking, UMass-FSD, LSH-FSD
Subjects: Information Retrieval (cs.IR)
Journal reference: SIGIR 20, July 2020
DOI: 10.1145/3397271.3401306
Cite as: arXiv:2208.01347 [cs.IR]
  (or arXiv:2208.01347v1 [cs.IR] for this version)

Submission history

From: Dominik Wurzer [view email]
[v1] Tue, 2 Aug 2022 10:19:29 GMT (378kb,D)

Link back to: arXiv, form interface, contact.