<oai_dc:dc xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
  <dc:contributor>Schmidhuber, Jürgen</dc:contributor>
  <dc:contributor>Förster, Alexander</dc:contributor>
  <dc:creator>Frank, Mikhail Alexander</dc:creator>
  <dc:date>2014-10-21</dc:date>
  <dc:description xmlns:ns0="xml" ns0:lang="en">The next generation of intelligent robots will need to be able to plan reaches. Not just  ballistic point to point reaches, but reaches around things such as the edge of a table,  a nearby human, or any other known object in the robot’s workspace. Planning  reaches may seem easy to us humans, because we do it so intuitively, but it has  proven to be a challenging problem, which continues to limit the versatility of what  robots can do today. In this document, I propose a novel intrinsically motivated RL  system that draws on both Path/Motion Planning and Reactive Control. Through  Reinforcement Learning, it tightly integrates these two previously disparate  approaches to robotics. The RL system is evaluated on a task, which is as yet  unsolved by roboticists in practice. That is to put the palm of the iCub humanoid robot  on arbitrary target objects in its workspace, start- ing from arbitrary initial  configurations. Such motions can be generated by planning, or searching the  configuration space, but this typically results in some kind of trajectory, which must  then be tracked by a separate controller, and such an approach offers a brit- tle  runtime solution because it is inflexible. Purely reactive systems are robust to many  problems that render a planned trajectory infeasible, but lacking the capacity to search,  they tend to get stuck behind constraints, and therefore do not replace motion  planners. The planner/controller proposed here is novel in that it deliberately plans  reaches without the need to track trajectories. Instead, reaches are composed of  sequences of reactive motion primitives, implemented by my Modular Behavioral  Environment (MoBeE), which provides (fictitious) force control with reactive collision  avoidance by way of a realtime kinematic/geometric model of the robot and its  workspace. Thus, to the best of my knowledge, mine is the first reach planning  approach to simultaneously offer the best of both the Path/Motion Planning and  Reactive Control approaches. By controlling the real, physical robot directly, and  feeling the influence of the con- straints imposed by MoBeE, the proposed system  learns a stochastic model of the iCub’s configuration space. Then, the model is  exploited as a multiple query path planner to find sensible pre-reach poses, from which  to initiate reaching actions. Experiments show that the system can autonomously find  practical reaches to target objects in workspace and offers excellent robustness to  changes in the workspace configuration as well as noise in the robot’s sensory-motor  apparatus.</dc:description>
  <dc:format>application/pdf</dc:format>
  <dc:identifier>https://susi.usi.ch/global/documents/318488</dc:identifier>
  <dc:identifier>https://n2t.net/ark:/12658/srd1318488</dc:identifier>
  <dc:identifier>https://susi.usi.ch/documents/318488/files/2014INFO011.pdf</dc:identifier>
  <dc:language>eng</dc:language>
  <dc:relation>info:eu-repo/semantics/altIdentifier/urn/urn:nbn:ch:rero-006-113712</dc:relation>
  <dc:relation>info:eu-repo/semantics/altIdentifier/ark/12658/srd1318488</dc:relation>
  <dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
  <dc:rights>License undefined</dc:rights>
  <dc:subject xmlns:ns1="xml" ns1:lang="en">Humanoid robotics</dc:subject>
  <dc:subject xmlns:ns2="xml" ns2:lang="en">Developmental robotics</dc:subject>
  <dc:subject xmlns:ns3="xml" ns3:lang="en">Motion planning</dc:subject>
  <dc:subject xmlns:ns4="xml" ns4:lang="en">Control systems</dc:subject>
  <dc:subject xmlns:ns5="xml" ns5:lang="en">Reinforcement learning</dc:subject>
  <dc:subject>info:eu-repo/classification/udc/004</dc:subject>
  <dc:title xmlns:ns6="xml" ns6:lang="en">Learning to reach and reaching to learn : a unified approach to path planning and reactive control through reinforcement learning</dc:title>
  <dc:type>http://purl.org/coar/resource_type/c_db06</dc:type>
</oai_dc:dc>
