We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.DL

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Digital Libraries

Title: Identifying the Development and Application of Artificial Intelligence in Scientific Text

Abstract: We describe a strategy for identifying the universe of research publications relevant to the application and development of artificial intelligence. The approach leverages the arXiv corpus of scientific preprints, in which authors choose subject tags for their papers from a set defined by editors. We compose a functional definition of AI relevance by learning these subjects from paper metadata, and then inferring the arXiv-subject labels of papers in larger corpora: Clarivate Web of Science, Digital Science Dimensions, and Microsoft Academic Graph. This yields predictive classification $F_1$ scores between .75 and .86 for Natural Language Processing (cs.CL), Computer Vision (cs.CV), and Robotics (cs.RO). For a single model that learns these and four other AI-relevant subjects (cs.AI, cs.LG, stat.ML, and cs.MA), we see precision of .83 and recall of .85. We evaluate the out-of-domain performance of our classifiers against other sources of topic information and predictions from alternative methods. We find that a supervised solution can generalize to identify publications that belong to the high-level fields of study represented on arXiv. This offers a method for identifying AI-relevant publications that updates at the pace of research output, without reliance on subject-matter experts for query development or labeling.
Comments: This revision expands our analysis in Section 5. We predict and evaluate labels for publications in Microsoft Academic Graph and Digital Science Dimensions, in addition to Clarivate Web of Science
Subjects: Digital Libraries (cs.DL); Information Retrieval (cs.IR)
Cite as: arXiv:2002.07143 [cs.DL]
  (or arXiv:2002.07143v2 [cs.DL] for this version)

Submission history

From: James Dunham [view email]
[v1] Mon, 17 Feb 2020 18:58:59 GMT (236kb,D)
[v2] Thu, 28 May 2020 15:35:17 GMT (51kb,D)

Link back to: arXiv, form interface, contact.