<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-21T05:00:42Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/18352" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/18352</identifier>
        <datestamp>2023-07-10</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>com_2142_5130</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:subject>machine learning</dc:subject>
          <dc:title>Parametrized Stochastic Multi-armed Bandits with Binary Rewards</dc:title>
          <dc:contributor>Srikant, Rayadurgam</dc:contributor>
          <dc:creator>Jiang, Chong</dc:creator>
          <dc:date>2011-01-14T22:47:13Z</dc:date>
          <dc:date>2011-01-14T22:47:13Z</dc:date>
          <dc:date>2011-01-14T22:47:13Z</dc:date>
          <dc:description>In this thesis, we consider the problem of multi-armed bandits with a 
large number of correlated arms. We assume that the arms have Bernoulli distributed rewards, independent across arms 
and across time, where the probabilities of success are parametrized by known 
attribute vectors for each arm, as well as an unknown preference vector. 
For this model, we seek an algorithm with a total regret that 
is sub-linear in time and independent of the number of arms. We present 
such an algorithm, which we call the Three-phase Algorithm, and analyze 
its performance. We show an upper bound on the total regret which applies uniformly in time.
The asymptotics of this bound show that for any $f \in \omega(\log(T))$, the total 
regret can be made to be $O(f(T))$, independent of the number of arms.</dc:description>
          <dc:description>Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-12-09T18:39:12Z
Item was in collections:
University of Illinois Theses &amp; Dissertations (ID: 1)
No. of bitstreams: 1
Jiang_Chong.pdf: 325450 bytes, checksum: 9f6372630df4d279f19fca565c89d472 (MD5)</dc:description>
          <dc:description>Made available in DSpace on 2011-01-14T22:47:13Z (GMT). No. of bitstreams: 2
Jiang_Chong.pdf: 325450 bytes, checksum: 9f6372630df4d279f19fca565c89d472 (MD5)
license.txt: 4060 bytes, checksum: 9148706d902e93953c049052831ab198 (MD5)</dc:description>
          <dc:identifier>http://hdl.handle.net/2142/18352</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2010 Chong Jiang</dc:rights>
          <dc:subject>multi-armed bandits</dc:subject>
          <dc:date>2010-12</dc:date>
          <degree>
            <department>Electrical &amp; Computer Eng</department>
            <departmentCode>1933</departmentCode>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <disciplineCode>1200</disciplineCode>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Thesis</level>
            <name>M.S.</name>
            <program>PHD:Electr &amp; Computer Eng-UIUC</program>
            <programCode>10KS1200PHD</programCode>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
