High-level Modeling of Manufacturing Faults in Deep Neural Network Accelerators

Kundu, Shamik; Soyyiğit, Ahmet; Hoque, Khaza Anuarul; Basu, Kanad

doi:10.1109/IOLTS50870.2020.9159704

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2006

Computer Science > Machine Learning

Title: High-level Modeling of Manufacturing Faults in Deep Neural Network Accelerators

Authors: Shamik Kundu, Ahmet Soyyiğit, Khaza Anuarul Hoque, Kanad Basu

(Submitted on 5 Jun 2020 (v1), last revised 26 Oct 2020 (this version, v2))

Abstract: The advent of data-driven real-time applications requires the implementation of Deep Neural Networks (DNNs) on Machine Learning accelerators. Google's Tensor Processing Unit (TPU) is one such neural network accelerator that uses systolic array-based matrix multiplication hardware for computation in its crux. Manufacturing faults at any state element of the matrix multiplication unit can cause unexpected errors in these inference networks. In this paper, we propose a formal model of permanent faults and their propagation in a TPU using the Discrete-Time Markov Chain (DTMC) formalism. The proposed model is analyzed using the probabilistic model checking technique to reason about the likelihood of faulty outputs. The obtained quantitative results show that the classification accuracy is sensitive to the type of permanent faults as well as their location, bit position and the number of layers in the neural network. The conclusions from our theoretical model have been validated using experiments on a digit recognition-based DNN.

Comments:	4 pages, 2 figures
Subjects:	Machine Learning (cs.LG); Performance (cs.PF); Machine Learning (stat.ML)
DOI:	10.1109/IOLTS50870.2020.9159704
Cite as:	arXiv:2006.03616 [cs.LG]
	(or arXiv:2006.03616v2 [cs.LG] for this version)

Submission history

From: Khaza Anuarul Hoque [view email]
[v1] Fri, 5 Jun 2020 18:11:14 GMT (488kb,D)
[v2] Mon, 26 Oct 2020 15:31:56 GMT (489kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2006.03616v2

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: High-level Modeling of Manufacturing Faults in Deep Neural Network Accelerators

Submission history