On designing socially acceptable reward shaping

Raza, SA; Clark, J; Williams, MA

On designing socially acceptable reward shaping

Raza, SA

Clark, J

Williams, MA

Permalink

Publication Type:: Conference Proceeding
Citation:: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2016, 9979 LNAI pp. 860 - 869
Issue Date:: 2016-01-01

Closed Access

	Filename	Description	Size
	Conference paper.pdf	Published version	505.13 kB	Adobe PDF	View/Open

Copyright Clearance Process

Recently Added
In Progress
Closed Access

This item is closed access and not available.

Full metadata record

Field	Value	Language
dc.contributor.author	Raza, SA https://orcid.org/0000-0001-6570-4808	en_US
dc.contributor.author	Clark, J https://orcid.org/0000-0002-1920-341X	en_US
dc.contributor.author	Williams, MA https://orcid.org/0000-0002-1047-0503	en_US
dc.date.issued	2016-01-01	en_US
dc.identifier.citation	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2016, 9979 LNAI pp. 860 - 869	en_US
dc.identifier.isbn	9783319474366	en_US
dc.identifier.issn	0302-9743	en_US
dc.identifier.uri	http://hdl.handle.net/10453/98936
dc.description.abstract	© Springer International Publishing AG 2016. For social robots, learning from an ordinary user should be socially appealing. Unfortunately, machine learning demands an enormous amount of human data, and a prolonged interactive teaching session becomes anti-social. We have addressed this problem in the context of reward shaping for reinforcement learning. For efficient reward shaping, a continuous stream of rewards is expected from the teacher. We present a simple framework which seeks rewards for a small number of steps from each of a large number of human teachers. Therefore, it simplifies the job of an individual teacher. The framework was tested with online crowd workers on a transport puzzle. We thoroughly analyzed the quality of the learned policies and crowd’s teaching behavior. Our results showed that nearly perfect policies can be learned using this framework. The framework was generally acceptable in the crowd’s opinion.	en_US
dc.relation.ispartof	Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)	en_US
dc.relation.isbasedon	10.1007/978-3-319-47437-3_84	en_US
dc.subject.classification	Artificial Intelligence & Image Processing	en_US
dc.title	On designing socially acceptable reward shaping	en_US
dc.type	Conference Proceeding
utslib.citation.volume	9979 LNAI	en_US
utslib.for	0806 Information Systems	en_US
utslib.for	080101 Adaptive Agents and Intelligent Robotics	en_US
pubs.embargo.period	Not known	en_US
pubs.organisational-group	/University of Technology Sydney
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology
pubs.organisational-group	/University of Technology Sydney/Faculty of Engineering and Information Technology/School of Computer Science
pubs.organisational-group	/University of Technology Sydney/Strength - CAI - Centre for Artificial Intelligence
utslib.copyright.status	closed_access
pubs.publication-status	Published	en_US
pubs.volume	9979 LNAI	en_US

Abstract:

© Springer International Publishing AG 2016. For social robots, learning from an ordinary user should be socially appealing. Unfortunately, machine learning demands an enormous amount of human data, and a prolonged interactive teaching session becomes anti-social. We have addressed this problem in the context of reward shaping for reinforcement learning. For efficient reward shaping, a continuous stream of rewards is expected from the teacher. We present a simple framework which seeks rewards for a small number of steps from each of a large number of human teachers. Therefore, it simplifies the job of an individual teacher. The framework was tested with online crowd workers on a transport puzzle. We thoroughly analyzed the quality of the learned policies and crowd’s teaching behavior. Our results showed that nearly perfect policies can be learned using this framework. The framework was generally acceptable in the crowd’s opinion.

Please use this identifier to cite or link to this item:

http://hdl.handle.net/10453/98936