We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.IR

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Information Retrieval

Title: An experiment exploring the theoretical and methodological challenges in developing a semi-automated approach to analysis of small-N qualitative data

Authors: Sandro Tsang
Abstract: This paper experiments with designing a semi-automated qualitative data analysis (QDA) algorithm to analyse 20 transcripts by using freeware. Text-mining (TM) and QDA were guided by frequency and association measures, because these statistics remain robust when the sample size is small. The refined TM algorithm split the text into various sizes based on a manually revised dictionary. This lemmatisation approach may reflect the context of the text better than uniformly tokenising the text into one single size. TM results were used for initial coding. Code repacking was guided by association measures and external data to implement a general inductive QDA approach. The information retrieved by TM and QDA was depicted in subgraphs for comparisons. The analyses were completed in 6-7 days. Both algorithms retrieved contextually consistent and relevant information. However, the QDA algorithm retrieved more specific information than TM alone. The QDA algorithm does not strictly comply with the convention of TM or of QDA, but becomes a more efficient, systematic and transparent text analysis approach than a conventional QDA approach. Scaling up QDA to reliably discover knowledge from text was exactly the research purpose. This paper also sheds light on understanding the relations between information technologies, theory and methodologies.
Comments: Page 2: "qualitative" research (QR); Page 3 and Appendix: Cited the second and third authors of a paper; Page 4: Redundant citation of author names; Page 8: Replaced wordclouds with higher resolution ones (cited facilities on page 3); Page 13: TM is a big-data/"quantitative" method and replaced "supplementary material" with "appendix"; Showed author names after see or cf. on pages 4, 5, 6 and 15
Subjects: Information Retrieval (cs.IR); Computation and Language (cs.CL)
ACM classes: I.1.2; I.1.4; I.2.7
Cite as: arXiv:2002.04513 [cs.IR]
  (or arXiv:2002.04513v2 [cs.IR] for this version)

Submission history

From: Sandro Tsang Dr [view email]
[v1] Mon, 3 Feb 2020 17:55:19 GMT (2138kb,D)
[v2] Sat, 15 Feb 2020 19:10:29 GMT (2693kb,D)

Link back to: arXiv, form interface, contact.