<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-19T10:47:52Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/101086" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/101086</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_10761</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_10755</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Lazebnik, Svetlana</dc:contributor>
          <dc:creator>Ge, Victor</dc:creator>
          <dc:date>2018-09-04T20:32:01Z</dc:date>
          <dc:date>2018-09-04T20:32:01Z</dc:date>
          <dc:date>2018-04-26</dc:date>
          <dc:date>2018-05</dc:date>
          <dc:description>Deep reinforcement learning methods are capable of learning complex heuristics starting with no prior knowledge, but struggle in environments where the learning signal is sparse. In contrast, planning methods can discover the optimal path to a goal in the absence of external rewards, but often require a hand-crafted heuristic function to be effective. In this thesis, we describe a model-based reinforcement learning method that bridges the middle ground between these two approaches. When evaluated on the complex domain of Sokoban, the model-based method was found to be more performant, stable and sample-efficient than a model-free baseline.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms</dc:description>
          <dc:description>The student, Victor Ge, accepted the attached license on 2018-04-25 at 18:20.</dc:description>
          <dc:description>The student, Victor Ge, submitted this Thesis for approval on 2018-04-25 at 18:28.</dc:description>
          <dc:description>This Thesis was approved for publication on 2018-04-26 at 15:04.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #12504 on 2018-08-31 at 17:15:02</dc:description>
          <dc:description>Made available in DSpace on 2018-09-04T20:32:01Z (GMT). No. of bitstreams: 2
GE-THESIS-2018.pdf: 671819 bytes, checksum: 161b3332ca985b02b599f786084fffd2 (MD5)
LICENSE.txt: 4206 bytes, checksum: aa9b2beecc67250f9c9c384d5880b2d6 (MD5)
  Previous issue date: 2018-04-26</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/101086</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2018 Victor Ge</dc:rights>
          <dc:subject>reinforcement learning</dc:subject>
          <dc:subject>mcts</dc:subject>
          <dc:subject>sokoban</dc:subject>
          <dc:subject>a*</dc:subject>
          <dc:subject>heuristic</dc:subject>
          <dc:title>Solving planning problems with deep reinforcement learning and tree search</dc:title>
          <dc:type>text</dc:type>
          <dc:type>text</dc:type>
          <degree>
            <department>Computer Science</department>
            <discipline>Computer Science</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Thesis</level>
            <name>M.S.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
