We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CV

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Computer Vision and Pattern Recognition

Title: Motion Representation Using Residual Frames with 3D CNN

Abstract: Recently, 3D convolutional networks (3D ConvNets) yield good performance in action recognition. However, optical flow stream is still needed to ensure better performance, the cost of which is very high. In this paper, we propose a fast but effective way to extract motion features from videos utilizing residual frames as the input data in 3D ConvNets. By replacing traditional stacked RGB frames with residual ones, 35.6% and 26.6% points improvements over top-1 accuracy can be obtained on the UCF101 and HMDB51 datasets when ResNet-18 models are trained from scratch. And we achieved the state-of-the-art results in this training mode. Analysis shows that better motion features can be extracted using residual frames compared to RGB counterpart. By combining with a simple appearance path, our proposal can be even better than some methods using optical flow streams.
Comments: Accepted in IEEE ICIP 2020. arXiv admin note: substantial text overlap with arXiv:2001.05661
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2006.13017 [cs.CV]
  (or arXiv:2006.13017v1 [cs.CV] for this version)

Submission history

From: Li Tao [view email]
[v1] Sun, 21 Jun 2020 07:35:41 GMT (6524kb)

Link back to: arXiv, form interface, contact.