We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CL

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Computation and Language

Title: Clustering and Network Analysis for the Embedding Spaces of Sentences and Sub-Sentences

Abstract: Sentence embedding methods offer a powerful approach for working with short textual constructs or sequences of words. By representing sentences as dense numerical vectors, many natural language processing (NLP) applications have improved their performance. However, relatively little is understood about the latent structure of sentence embeddings. Specifically, research has not addressed whether the length and structure of sentences impact the sentence embedding space and topology. This paper reports research on a set of comprehensive clustering and network analyses targeting sentence and sub-sentence embedding spaces. Results show that one method generates the most clusterable embeddings. In general, the embeddings of span sub-sentences have better clustering properties than the original sentences. The results have implications for future sentence embedding models and applications.
Comments: Accepted by The International Conference on Intelligent Data Science Technologies and Applications (IDSTA2021)
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2110.00697 [cs.CL]
  (or arXiv:2110.00697v1 [cs.CL] for this version)

Submission history

From: Yuan An [view email]
[v1] Sat, 2 Oct 2021 00:47:35 GMT (373kb,D)

Link back to: arXiv, form interface, contact.