Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation

Dadashzadeh, Amirhossein; Whone, Alan; Mirmehdi, Majid

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2112

Change to browse by:

Computer Science > Computer Vision and Pattern Recognition

Title: Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation

Authors: Amirhossein Dadashzadeh, Alan Whone, Majid Mirmehdi

(Submitted on 7 Dec 2021 (v1), last revised 25 Apr 2022 (this version, v3))

Abstract: Despite the outstanding success of self-supervised pretraining methods for video representation learning, they generalise poorly when the unlabeled dataset for pretraining is small or the domain difference between unlabelled data in source task (pretraining) and labeled data in target task (finetuning) is significant. To mitigate these issues, we propose a novel approach to complement self-supervised pretraining via an auxiliary pretraining phase, based on knowledge similarity distillation, auxSKD, for better generalisation with a significantly smaller amount of video data, e.g. Kinetics-100 rather than Kinetics-400. Our method deploys a teacher network that iteratively distills its knowledge to the student model by capturing the similarity information between segments of unlabelled video data. The student model meanwhile solves a pretext task by exploiting this prior knowledge. We also introduce a novel pretext task, Video Segment Pace Prediction or VSPP, which requires our model to predict the playback speed of a randomly selected segment of the input video to provide more reliable self-supervised representations. Our experimental results show superior results to the state of the art on both UCF101 and HMDB51 datasets when pretraining on K100 in apple-to-apple comparisons. Additionally, we show that our auxiliary pretraining, auxSKD, when added as an extra pretraining phase to recent state of the art self-supervised methods (i.e. VCOP, VideoPace, and RSPNet), improves their results on UCF101 and HMDB51. Our code is available at this https URL

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2112.04011 [cs.CV]
	(or arXiv:2112.04011v3 [cs.CV] for this version)

Submission history

From: Amirhossein Dadashzadeh [view email]
[v1] Tue, 7 Dec 2021 21:50:40 GMT (2203kb,D)
[v2] Mon, 14 Mar 2022 20:39:41 GMT (2487kb,D)
[v3] Mon, 25 Apr 2022 14:25:55 GMT (2487kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2112.04011

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: Auxiliary Learning for Self-Supervised Video Representation via Similarity-based Knowledge Distillation

Submission history