We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.SI

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Social and Information Networks

Title: Media Cloud: Massive Open Source Collection of Global News on the Open Web

Abstract: We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access as well as user-facing tools. We also highlight the strengths and limitations of the Media Cloud collection strategy compared to relevant alternatives. We give an overview two sample datasets generated using Media Cloud and discuss how researchers can use the platform to create their own datasets.
Comments: 15 pages, 9 figures, accepted (minus the 3-page, 3-image appendix given here) for publication and forthcoming in Proceedings of the Fifteenth International AAAI Conference on Web and Social Media (ICWSM-2021)
Subjects: Social and Information Networks (cs.SI); Computers and Society (cs.CY)
ACM classes: J.4; H.3.5; J.7; J.5; K.4.1
Cite as: arXiv:2104.03702 [cs.SI]
  (or arXiv:2104.03702v3 [cs.SI] for this version)

Submission history

From: Momin M. Malik [view email]
[v1] Thu, 8 Apr 2021 11:51:13 GMT (1780kb,D)
[v2] Mon, 12 Apr 2021 13:22:30 GMT (1780kb,D)
[v3] Sat, 1 May 2021 23:01:20 GMT (1751kb,D)

Link back to: arXiv, form interface, contact.