Block-Sparse Recurrent Neural Networks

Narang, Sharan; Undersander, Eric; Diamos, Gregory

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 1711

Computer Science > Machine Learning

Title: Block-Sparse Recurrent Neural Networks

Authors: Sharan Narang, Eric Undersander, Gregory Diamos

(Submitted on 8 Nov 2017)

Abstract: Recurrent Neural Networks (RNNs) are used in state-of-the-art models in domains such as speech recognition, machine translation, and language modelling. Sparsity is a technique to reduce compute and memory requirements of deep learning models. Sparse RNNs are easier to deploy on devices and high-end server processors. Even though sparse operations need less compute and memory relative to their dense counterparts, the speed-up observed by using sparse operations is less than expected on different hardware platforms. In order to address this issue, we investigate two different approaches to induce block sparsity in RNNs: pruning blocks of weights in a layer and using group lasso regularization to create blocks of weights with zeros. Using these techniques, we demonstrate that we can create block-sparse RNNs with sparsity ranging from 80% to 90% with small loss in accuracy. This allows us to reduce the model size by roughly 10x. Additionally, we can prune a larger dense network to recover this loss in accuracy while maintaining high block sparsity and reducing the overall parameter count. Our technique works with a variety of block sizes up to 32x32. Block-sparse RNNs eliminate overheads related to data storage and irregular memory accesses while increasing hardware efficiency compared to unstructured sparsity.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:1711.02782 [cs.LG]
	(or arXiv:1711.02782v1 [cs.LG] for this version)

Submission history

From: Sharan Narang [view email]
[v1] Wed, 8 Nov 2017 00:57:54 GMT (180kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1711.02782

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Block-Sparse Recurrent Neural Networks

Submission history