L3 Fusion: Fast Transformed Convolutions on CPUs

Gelashvili, Rati; Shavit, Nir; Zlateski, Aleksandar

Full-text links:

Download:

Current browse context:

cs.DC

< prev | next >

new | recent | 1912

Computer Science > Distributed, Parallel, and Cluster Computing

Title: L3 Fusion: Fast Transformed Convolutions on CPUs

Authors: Rati Gelashvili, Nir Shavit, Aleksandar Zlateski

(Submitted on 4 Dec 2019)

Abstract: Fast convolutions via transforms, either Winograd or FFT, had emerged as a preferred way of performing the computation of convolutional layers, as it greatly reduces the number of required operations. Recent work shows that, for many layer structures, a well--designed implementation of fast convolutions can greatly utilize modern CPUs, significantly reducing the compute time. However, the generous amount of shared L3 cache present on modern CPUs is often neglected, and the algorithms are optimized solely for the private L2 cache. In this paper we propose an efficient `L3 Fusion` algorithm that is specifically designed for CPUs with significant amount of shared L3 cache. Using the hierarchical roofline model, we show that in many cases, especially for layers with fewer channels, the `L3 fused` approach can greatly outperform standard 3 stage one provided by big vendors such as Intel. We validate our theoretical findings, by benchmarking our `L3 fused` implementation against publicly available state of the art.

Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
Cite as:	arXiv:1912.02165 [cs.DC]
	(or arXiv:1912.02165v1 [cs.DC] for this version)

Submission history

From: Rati Gelashvili [view email]
[v1] Wed, 4 Dec 2019 18:34:58 GMT (72kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1912.02165

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Distributed, Parallel, and Cluster Computing

Title: L3 Fusion: Fast Transformed Convolutions on CPUs

Submission history