We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:


References & Citations

DBLP - CS Bibliography


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Computer Vision and Pattern Recognition

Title: TransMOT: Spatial-Temporal Graph Transformer for Multiple Object Tracking

Abstract: Tracking multiple objects in videos relies on modeling the spatial-temporal interactions of the objects. In this paper, we propose a solution named TransMOT, which leverages powerful graph transformers to efficiently model the spatial and temporal interactions among the objects. TransMOT effectively models the interactions of a large number of objects by arranging the trajectories of the tracked objects as a set of sparse weighted graphs, and constructing a spatial graph transformer encoder layer, a temporal transformer encoder layer, and a spatial graph transformer decoder layer based on the graphs. TransMOT is not only more computationally efficient than the traditional Transformer, but it also achieves better tracking accuracy. To further improve the tracking speed and accuracy, we propose a cascade association framework to handle low-score detections and long-term occlusions that require large computational resources to model in TransMOT. The proposed method is evaluated on multiple benchmark datasets including MOT15, MOT16, MOT17, and MOT20, and it achieves state-of-the-art performance on all the datasets.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2104.00194 [cs.CV]
  (or arXiv:2104.00194v2 [cs.CV] for this version)

Submission history

From: Peng Chu [view email]
[v1] Thu, 1 Apr 2021 01:49:05 GMT (1949kb,D)
[v2] Sat, 3 Apr 2021 05:12:03 GMT (1950kb,D)

Link back to: arXiv, form interface, contact.