Learning common and specific features for RGB-D semantic segmentation with deconvolutional networks

Wang, J; Wang, Z; Tao, D; See, S; Wang, G

Learning common and specific features for RGB-D semantic segmentation with deconvolutional networks

Wang, J Wang, Z Tao, D

See, S Wang, G

Permalink

Publication Type:: Conference Proceeding
Citation:: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2016, 9909 LNCS pp. 664 - 679
Issue Date:: 2016-01-01

Closed Access

	Filename	Description	Size
	LEARNING COMMON AND SPECIFIC FEATURES FOR RGB-D SEMANTIC SEGMENTATION WITH DECONVOLUTIONAL NETWORKS.pdf	Published version	1.66 MB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Wang, J	en_US
dc.contributor.author	Wang, Z	en_US
dc.contributor.author	Tao, D https://orcid.org/0000-0001-7225-5449	en_US
dc.contributor.author	See, S	en_US
dc.contributor.author	Wang, G	en_US
dc.date.issued	2016-01-01	en_US
dc.identifier.citation	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2016, 9909 LNCS pp. 664 - 679	en_US
dc.identifier.isbn	9783319464534	en_US
dc.identifier.issn	0302-9743	en_US
dc.identifier.uri	http://hdl.handle.net/10453/104417
dc.description.abstract	© Springer International Publishing AG 2016. In this paper, we tackle the problem of RGB-D semantic segmentation of indoor images. We take advantage of deconvolutional networks which can predict pixel-wise class labels, and develop a new structure for deconvolution of multiple modalities. We propose a novel feature transformation network to bridge the convolutional networks and deconvolutional networks. In the feature transformation network, we correlate the two modalities by discovering common features between them, as well as characterize each modality by discovering modality specific features. With the common features, we not only closely correlate the two modalities, but also allow them to borrow features from each other to enhance the representation of shared information. With specific features, we capture the visual patterns that are only visible in one modality. The proposed network achieves competitive segmentation accuracy on NYU depth dataset V1 and V2.	en_US
dc.relation	http://purl.org/au-research/grants/arc/DP140102164
dc.relation	http://purl.org/au-research/grants/arc/FT130101457
dc.relation.ispartof	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)	en_US
dc.relation.isbasedon	10.1007/978-3-319-46454-1_40	en_US
dc.subject.classification	Artificial Intelligence & Image Processing	en_US
dc.title	Learning common and specific features for RGB-D semantic segmentation with deconvolutional networks	en_US
dc.type	Conference Proceeding
utslib.citation.volume	9909 LNCS	en_US
utslib.for	0801 Artificial Intelligence and Image Processing	en_US
pubs.embargo.period	Not known	en_US
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
utslib.copyright.status	closed_access
pubs.publication-status	Published	en_US
pubs.volume	9909 LNCS	en_US

Abstract:

© Springer International Publishing AG 2016. In this paper, we tackle the problem of RGB-D semantic segmentation of indoor images. We take advantage of deconvolutional networks which can predict pixel-wise class labels, and develop a new structure for deconvolution of multiple modalities. We propose a novel feature transformation network to bridge the convolutional networks and deconvolutional networks. In the feature transformation network, we correlate the two modalities by discovering common features between them, as well as characterize each modality by discovering modality specific features. With the common features, we not only closely correlate the two modalities, but also allow them to borrow features from each other to enhance the representation of shared information. With specific features, we capture the visual patterns that are only visible in one modality. The proposed network achieves competitive segmentation accuracy on NYU depth dataset V1 and V2.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/104417