<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-22T11:18:25Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/100977" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/100977</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_10761</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_10755</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Lazebnik, Svetlana</dc:contributor>
          <dc:contributor>Lazebnik, Svetlana</dc:contributor>
          <dc:contributor>Hockenmaier, Julia</dc:contributor>
          <dc:contributor>Hoiem, Derek</dc:contributor>
          <dc:contributor>Brown, Matthew</dc:contributor>
          <dc:creator>Plummer, Bryan A.</dc:creator>
          <dc:date>2018-09-04T20:27:05Z</dc:date>
          <dc:date>2018-09-04T20:27:05Z</dc:date>
          <dc:date>2018-04-16</dc:date>
          <dc:date>2018-05</dc:date>
          <dc:description>Grounding language in images has shown it can help improve performance on many image-language tasks. To spur research on this topic, this dissertation introduces a new dataset which provides the ground truth annotations of the location of noun phrase chunks in image captions.  I begin by introducing a constituent task termed phrase localization, where the goal is to localize an entity known to exist in an image when provided with a natural language query.  To address this task, I introduce a model which learns a set of models, each of which capture a different concept which is useful in our task.  These concepts can be predefined, such as attributes gleamed from the adjectives, as well as those which are automatically learned in a single-end-to-end neural network.  I also address the more challenging detection style task, where the goal is to localize a phrase and determine if it is associated with an image.  Multiple applications of the models presented in this work demonstrate their value beyond the phrase localization task.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2018-08-31 without embargo terms</dc:description>
          <dc:description>The student, Bryan Plummer, accepted the attached license on 2018-04-14 at 05:42.</dc:description>
          <dc:description>The student, Bryan Plummer, submitted this Dissertation for approval on 2018-04-14 at 05:57.</dc:description>
          <dc:description>This Dissertation was approved for publication on 2018-04-16 at 11:21.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #12248 on 2018-08-31 at 17:12:23</dc:description>
          <dc:description>Made available in DSpace on 2018-09-04T20:27:05Z (GMT). No. of bitstreams: 3
PLUMMER-DISSERTATION-2018.pdf: 9214167 bytes, checksum: 7c4ec584f104360a0a99ea9a290b5a09 (MD5)
grounding-natural-language.zip: 8908155 bytes, checksum: 79b1f7c90322e047fc2559c7486ab3c1 (MD5)
LICENSE.txt: 4210 bytes, checksum: 398aac9d4dcb281c202797929dc1603f (MD5)
  Previous issue date: 2018-04-16</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/100977</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2018 Bryan A. Plummer</dc:rights>
          <dc:subject>Computer Vision, Natural Language Processing, Phrase Grounding</dc:subject>
          <dc:title>Grounding natural language phrases in images and video</dc:title>
          <dc:type>text</dc:type>
          <dc:type>text</dc:type>
          <degree>
            <department>Computer Science</department>
            <discipline>Computer Science</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Dissertation</level>
            <name>Ph.D.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
