We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CV

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Computer Vision and Pattern Recognition

Title: TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition

Abstract: Few-shot action recognition aims to recognize novel action classes (query) using just a few samples (support). The majority of current approaches follow the metric learning paradigm, which learns to compare the similarity between videos. Recently, it has been observed that directly measuring this similarity is not ideal since different action instances may show distinctive temporal distribution, resulting in severe misalignment issues across query and support videos. In this paper, we arrest this problem from two distinct aspects -- action duration misalignment and action evolution misalignment. We address them sequentially through a Two-stage Action Alignment Network (TA2N). The first stage locates the action by learning a temporal affine transform, which warps each video feature to its action duration while dismissing the action-irrelevant feature (e.g. background). Next, the second stage coordinates query feature to match the spatial-temporal action evolution of support by performing temporally rearrange and spatially offset prediction. Extensive experiments on benchmark datasets show the potential of the proposed method in achieving state-of-the-art performance for few-shot action recognition.
Comments: manuscripts
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2107.04782 [cs.CV]
  (or arXiv:2107.04782v2 [cs.CV] for this version)

Submission history

From: Shuyuan Li [view email]
[v1] Sat, 10 Jul 2021 07:22:49 GMT (9178kb,D)
[v2] Wed, 22 Sep 2021 04:40:53 GMT (18978kb,D)

Link back to: arXiv, form interface, contact.