<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-22T12:45:11Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/108055" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/108055</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_8888</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_8887</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:subject>machine translation</dc:subject>
          <dc:title>A translation framework for discovering word-like units from visual scenes and spoken descriptions</dc:title>
          <dc:type>text</dc:type>
          <dc:type>Thesis</dc:type>
          <dc:contributor>Hasegawa-Johnson, Mark A</dc:contributor>
          <dc:creator>Wang, Liming</dc:creator>
          <dc:date>2020-08-26T21:58:07Z</dc:date>
          <dc:date>2020-08-26T21:58:07Z</dc:date>
          <dc:date>2020-05-14</dc:date>
          <dc:date>2020-05</dc:date>
          <dc:description>In the absence of dictionaries, translators, or grammars, it is still possible to learn some of the words of a new language by listening to spoken descriptions of images. If several images, each containing a particular visually salient object, each co-occur with a particular sequence of speech sounds, we can infer that those speech sounds are a word whose definition is the visible object. A multimodal word discovery system accepts, as input, a database of spoken descriptions of images (or a set of corresponding phone transcriptions) and learns a mapping from waveform segments (or phone strings) to their associated image concepts. In this thesis, we propose a novel framework for multimodal word discovery systems based on statistical machine translation (SMT) and neural machine translation (NMT). We extend the existing theoretical frameworks on unsupervised word discovery and demonstrate a class of effective models for end-to-end word discovery from image regions and spoken descriptions. Finally, we provide a careful ablation study on components of my system and present some of the challenges in multimodal spoken word discovery.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2020-08-25 without embargo terms</dc:description>
          <dc:description>The student, Liming Wang, accepted the attached license on 2020-05-13 at 10:26.</dc:description>
          <dc:description>The student, Liming Wang, submitted this Thesis for approval on 2020-05-13 at 10:28.</dc:description>
          <dc:description>This Thesis was approved for publication on 2020-05-14 at 09:24.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #15375 on 2020-08-25 at 17:14:33</dc:description>
          <dc:description>Made available in DSpace on 2020-08-26T21:58:07Z (GMT). No. of bitstreams: 2
WANG-THESIS-2020.pdf: 1955236 bytes, checksum: d466468eb8c76cf7cef7185d035e1cb1 (MD5)
LICENSE.txt: 4208 bytes, checksum: 404bb44630b880d96166085d5cde918c (MD5)
  Previous issue date: 2020-05-14</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/108055</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2020 Liming Wang</dc:rights>
          <dc:subject>multimodal learning</dc:subject>
          <dc:subject>low-resource speech technology</dc:subject>
          <degree>
            <department>Electrical &amp; Computer Eng</department>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Thesis</level>
            <name>M.S.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
