Complementary Language Model and Parallel Bi-LRNN for False Trigger Mitigation

Agarwal, Rishika; Niu, Xiaochuan; Dighe, Pranay; Vishnubhotla, Srikanth; Badaskar, Sameer; Naik, Devang

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2008

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Complementary Language Model and Parallel Bi-LRNN for False Trigger Mitigation

Authors: Rishika Agarwal, Xiaochuan Niu, Pranay Dighe, Srikanth Vishnubhotla, Sameer Badaskar, Devang Naik

(Submitted on 18 Aug 2020)

Abstract: False triggers in voice assistants are unintended invocations of the assistant, which not only degrade the user experience but may also compromise privacy. False trigger mitigation (FTM) is a process to detect the false trigger events and respond appropriately to the user. In this paper, we propose a novel solution to the FTM problem by introducing a parallel ASR decoding process with a special language model trained from "out-of-domain" data sources. Such language model is complementary to the existing language model optimized for the assistant task. A bidirectional lattice RNN (Bi-LRNN) classifier trained from the lattices generated by the complementary language model shows a $38.34\%$ relative reduction of the false trigger (FT) rate at the fixed rate of $0.4\%$ false suppression (FS) of correct invocations, compared to the current Bi-LRNN model. In addition, we propose to train a parallel Bi-LRNN model based on the decoding lattices from both language models, and examine various ways of implementation. The resulting model leads to further reduction in the false trigger rate by $10.8\%$.

Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:2008.08113 [eess.AS]
	(or arXiv:2008.08113v1 [eess.AS] for this version)

Submission history

From: Rishika Agarwal [view email]
[v1] Tue, 18 Aug 2020 18:21:33 GMT (835kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2008.08113

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Complementary Language Model and Parallel Bi-LRNN for False Trigger Mitigation

Submission history