SF-Net: Single-Frame Supervision for Temporal Action Localization

Ma, F; Zhu, L; Yang, Y; Zha, S; Kundu, G; Feiszli, M; Shou, Z

SF-Net: Single-Frame Supervision for Temporal Action Localization

Ma, F Zhu, L

Yang, Y Zha, S Kundu, G Feiszli, M Shou, Z

Permalink

Publisher:: Springer International Publishing
Publication Type:: Conference Proceeding
Citation:: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2020, 12349 LNCS, pp. 420-437
Issue Date:: 2020-01-01

Open Access

Copyright Clearance Process

Recently Added
In Progress
Open Access

This item is open access.

Download Accepted ManuscriptAdobe PDF (1.1 MB)

View on publisher's site

View statistics

Full metadata record

Field	Value	Language
dc.contributor.advisor	This is a post-peer-review, pre-copyedit version of an article published in [Lecture Notes in Computer Science (including subse]. The final authenticated version is available online at: https://link.springer.com/chapter/10.1007%2F978-3-030-58548-8_25]”
dc.contributor.author	Ma, F
dc.contributor.author	Zhu, L https://orcid.org/0000-0002-4093-7557
dc.contributor.author	Yang, Y
dc.contributor.author	Zha, S
dc.contributor.author	Kundu, G
dc.contributor.author	Feiszli, M
dc.contributor.author	Shou, Z
dc.date.accessioned	2021-04-21T06:43:52Z
dc.date.available	2021-04-21T06:43:52Z
dc.date.issued	2020-01-01
dc.identifier.citation	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2020, 12349 LNCS, pp. 420-437
dc.identifier.isbn	9783030585471
dc.identifier.issn	0302-9743
dc.identifier.issn	1611-3349
dc.identifier.uri	http://hdl.handle.net/10453/148249
dc.description.abstract	In this paper, we study an intermediate form of supervision, i.e., single-frame supervision, for temporal action localization (TAL). To obtain the single-frame supervision, the annotators are asked to identify only a single frame within the temporal window of an action. This can significantly reduce the labor cost of obtaining full supervision which requires annotating the action boundary. Compared to the weak supervision that only annotates the video-level label, the single-frame supervision introduces extra temporal action signals while maintaining low annotation overhead. To make full use of such single-frame supervision, we propose a unified system called SF-Net. First, we propose to predict an actionness score for each video frame. Along with a typical category score, the actionness score can provide comprehensive information about the occurrence of a potential action and aid the temporal boundary refinement during inference. Second, we mine pseudo action and background frames based on the single-frame annotations. We identify pseudo action frames by adaptively expanding each annotated single frame to its nearby, contextual frames and we mine pseudo background frames from all the unannotated frames across multiple videos. Together with the ground-truth labeled frames, these pseudo-labeled frames are further used for training the classifier. In extensive experiments on THUMOS14, GTEA, and BEOID, SF-Net significantly improves upon state-of-the-art weakly-supervised methods in terms of both segment localization and single-frame localization. Notably, SF-Net achieves comparable results to its fully-supervised counterpart which requires much more resource intensive annotations. The code is available at https://github.com/Flowerfan/SF-Net.
dc.language	en
dc.publisher	Springer International Publishing
dc.relation	http://purl.org/au-research/grants/arc/DP200100938
dc.relation.ispartof	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
dc.relation.isbasedon	10.1007/978-3-030-58548-8_25
dc.rights	info:eu-repo/semantics/openAccess
dc.rights	This is a post-peer-review, pre-copyedit version of an article published in [Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2020, 12349 LNCS, pp. 420-437]. The final authenticated version is available online at: https://link.springer.com/chapter/10.1007%2F978-3-030-58548-8_25]”
dc.subject.classification	Artificial Intelligence & Image Processing
dc.title	SF-Net: Single-Frame Supervision for Temporal Action Localization
dc.type	Conference Proceeding
utslib.citation.volume	12349 LNCS
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Strength - AAII - Australian Artificial Intelligence Institute
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology/School of Computer Science
utslib.copyright.status	open_access	*
dc.date.updated	2021-04-21T06:43:47Z
pubs.publication-status	Published
pubs.volume	12349 LNCS

Abstract:

In this paper, we study an intermediate form of supervision, i.e., single-frame supervision, for temporal action localization (TAL). To obtain the single-frame supervision, the annotators are asked to identify only a single frame within the temporal window of an action. This can significantly reduce the labor cost of obtaining full supervision which requires annotating the action boundary. Compared to the weak supervision that only annotates the video-level label, the single-frame supervision introduces extra temporal action signals while maintaining low annotation overhead. To make full use of such single-frame supervision, we propose a unified system called SF-Net. First, we propose to predict an actionness score for each video frame. Along with a typical category score, the actionness score can provide comprehensive information about the occurrence of a potential action and aid the temporal boundary refinement during inference. Second, we mine pseudo action and background frames based on the single-frame annotations. We identify pseudo action frames by adaptively expanding each annotated single frame to its nearby, contextual frames and we mine pseudo background frames from all the unannotated frames across multiple videos. Together with the ground-truth labeled frames, these pseudo-labeled frames are further used for training the classifier. In extensive experiments on THUMOS14, GTEA, and BEOID, SF-Net significantly improves upon state-of-the-art weakly-supervised methods in terms of both segment localization and single-frame localization. Notably, SF-Net achieves comparable results to its fully-supervised counterpart which requires much more resource intensive annotations. The code is available at https://github.com/Flowerfan/SF-Net.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/148249