Discovering Salient Neurons in Deep NLP Models

Durrani, Nadir; Dalvi, Fahim; Sajjad, Hassan

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2206

Change to browse by:

Computer Science > Computation and Language

Title: Discovering Salient Neurons in Deep NLP Models

Authors: Nadir Durrani, Fahim Dalvi, Hassan Sajjad

(Submitted on 27 Jun 2022 (v1), last revised 14 Jan 2024 (this version, v2))

Abstract: While a lot of work has been done in understanding representations learned within deep NLP models and what knowledge they capture, little attention has been paid towards individual neurons. We present a technique called as Linguistic Correlation Analysis to extract salient neurons in the model, with respect to any extrinsic property - with the goal of understanding how such a knowledge is preserved within neurons. We carry out a fine-grained analysis to answer the following questions: (i) can we identify subsets of neurons in the network that capture specific linguistic properties? (ii) how localized or distributed neurons are across the network? iii) how redundantly is the information preserved? iv) how fine-tuning pre-trained models towards downstream NLP tasks, impacts the learned linguistic knowledge? iv) how do architectures vary in learning different linguistic properties? Our data-driven, quantitative analysis illuminates interesting findings: (i) we found small subsets of neurons that can predict different linguistic tasks, ii) with neurons capturing basic lexical information (such as suffixation) localized in lower most layers, iii) while those learning complex concepts (such as syntactic role) predominantly in middle and higher layers, iii) that salient linguistic neurons are relocated from higher to lower layers during transfer learning, as the network preserve the higher layers for task specific information, iv) we found interesting differences across pre-trained models, with respect to how linguistic information is preserved within, and v) we found that concept exhibit similar neuron distribution across different languages in the multilingual transformer models. Our code is publicly available as part of the NeuroX toolkit.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2206.13288 [cs.CL]
	(or arXiv:2206.13288v2 [cs.CL] for this version)

Submission history

From: Fahim Dalvi [view email]
[v1] Mon, 27 Jun 2022 13:31:49 GMT (2333kb,D)
[v2] Sun, 14 Jan 2024 13:25:01 GMT (689kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2206.13288

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Discovering Salient Neurons in Deep NLP Models

Submission history