<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-21T14:08:45Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/129852" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/129852</identifier>
        <datestamp>2025-10-25</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_8888</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_8887</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:format>application/pdf</dc:format>
          <dc:language>en</dc:language>
          <dc:type>text</dc:type>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms</dc:description>
          <dc:description>The student, Yashaswini Murthy, accepted the attached license on 2025-07-09 at 18:14.</dc:description>
          <dc:description>The student, Yashaswini Murthy, submitted this Dissertation for approval on 2025-07-10 at 18:49.</dc:description>
          <dc:description>This Dissertation was approved for publication on 2025-07-11 at 15:46.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #22474 on 2025-10-20 at 16:57:39</dc:description>
          <dc:title>Policy-based average-reward and robust Markov decision processes and reinforcement learning</dc:title>
          <dc:creator>Murthy, Yashaswini</dc:creator>
          <dc:date>2025-07-11</dc:date>
          <dc:contributor>Srikant, Rayadurgam</dc:contributor>
          <dc:contributor>Srikant, Rayadurgam</dc:contributor>
          <dc:contributor>Hajek, Bruce</dc:contributor>
          <dc:contributor>Stolyar, Aleksandr</dc:contributor>
          <dc:contributor>Hu, Bin</dc:contributor>
          <dc:subject>Reinforcement Learning</dc:subject>
          <dc:subject>Markov Decision Processes</dc:subject>
          <dc:language>eng</dc:language>
          <dc:description>This thesis addresses critical challenges in applying Reinforcement Learning (RL) to complex, real-world control problems by developing and analyzing policy-based algorithms for Markov Decision Processes (MDPs) under average-reward, robust (risk-sensitive), and countable-space settings. Traditional RL methods often fall short due to assumptions of finite state/action spaces, bounded costs, or reliance on discounted reward criteria, which may not align with long-run performance objectives in dynamic systems. The research presented herein makes several key contributions. First, for average-reward MDPs, we establish rigorous finite-time performance bounds for Approximate Policy Iteration (API) and prove global convergence with $O(\log(T))$ regret for Projected Policy Gradient (PPG) methods, notably by proving the smoothness of the average-reward function. We also demonstrate global convergence for Natural Policy Gradient (NPG)/Mirror Descent Methods (MDM) in the tabular average-reward setting. Second, to tackle problems with countably infinite state spaces and unbounded costs, common in queuing systems, we develop an NPG algorithm with state-dependent step sizes derived from novel policy-independent bounds on the relative value function, achieving non-trivial $O(\sqrt{T})$ regret bounds and relaxing learning error assumptions. Third, for robust decision-making under model uncertainty and cost variability, we introduce Modified Policy Iteration (MPI) for risk-sensitive exponential cost average-cost MDPs, proving its finite-time convergence. This is extended to an Approximate MPI (AMPI) framework and an Approximate PI framework, providing the first convergence guarantees for approximate methods in this risk-sensitive, robust setting and deriving a unified stability condition for error propagation across risk-neutral and risk-sensitive regimes. Collectively, this work advances the theoretical understanding and practical applicability of policy-based RL, providing new algorithms, convergence guarantees, and analytical tools for optimizing long-run performance in large-scale, uncertain, and risk-aware environments.</dc:description>
          <dc:date>2025-08</dc:date>
          <dc:type>Text</dc:type>
          <dc:identifier>https://hdl.handle.net/2142/129852</dc:identifier>
          <dc:rights>Copyright 2025 Yashaswini Murthy</dc:rights>
          <degree>
            <department>Electrical &amp; Computer Eng</department>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <grantor>University of Illinois Urbana-Champaign</grantor>
            <name>Ph.D.</name>
            <level>Dissertation</level>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
