<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-21T00:04:17Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/113048" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/113048</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_10761</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_10755</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Jiang, Nan</dc:contributor>
          <dc:creator>Agrawal, Priyank</dc:creator>
          <dc:date>2022-01-12T21:45:44Z</dc:date>
          <dc:date>2022-01-12T21:45:44Z</dc:date>
          <dc:date>2021-07-15</dc:date>
          <dc:date>2021-08</dc:date>
          <dc:description>This work studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping variant of one classical Thompson Sampling (TS)-like algorithm, randomized least-squares value iteration (RLSVI). Our $\tilde{\mathrm{O}}(H^2S\sqrt{AT})$ high-probability worst-case regret bound improves the previous sharpest worst-case regret bounds for RLSVI and matches the existing state-of-the-art worst-case TS-based regret bounds.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms</dc:description>
          <dc:description>The student, Priyank Agrawal, accepted the attached license on 2021-07-14 at 20:46.</dc:description>
          <dc:description>The student, Priyank Agrawal, submitted this Thesis for approval on 2021-07-14 at 20:59.</dc:description>
          <dc:description>This Thesis was approved for publication on 2021-07-15 at 15:44.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #16946 on 2022-01-12 at 12:45:32</dc:description>
          <dc:description>Made available in DSpace on 2022-01-12T21:45:44Z (GMT). No. of bitstreams: 2
AGRAWAL-THESIS-2021.pdf: 427442 bytes, checksum: bbc5e57adc02afd39317656f59724340 (MD5)
LICENSE.txt: 4212 bytes, checksum: 330f6b569f4b36aad71e7f0c0f652d33 (MD5)
  Previous issue date: 2021-07-15</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/113048</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2021 Priyank Agrawal</dc:rights>
          <dc:subject>Reinforcement Learning</dc:subject>
          <dc:subject>Exploration-Exploitation</dc:subject>
          <dc:title>Improved worst-case regret bounds for randomized least-squares value iteration</dc:title>
          <dc:type>text</dc:type>
          <dc:type>Thesis</dc:type>
          <degree>
            <department>Computer Science</department>
            <discipline>Computer Science</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Thesis</level>
            <name>M.S.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
