<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-24T00:30:03Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/127369" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/127369</identifier>
        <datestamp>2026-02-03</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_8888</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_8887</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:title>Audiovisual processing for generation and enhancement</dc:title>
          <dc:creator>Fan, Xulin</dc:creator>
          <dc:date>2024-11-27</dc:date>
          <dc:contributor>Hasegawa-Johnson, Mark Allan</dc:contributor>
          <dc:subject>Multimodal Signal Processing</dc:subject>
          <dc:subject>Speech Processing</dc:subject>
          <dc:subject>Speech Enhancement</dc:subject>
          <dc:subject>Signal Processing</dc:subject>
          <dc:language>eng</dc:language>
          <dc:format>application/pdf</dc:format>
          <dc:language>en</dc:language>
          <dc:type>text</dc:type>
          <dc:description>Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01</dc:description>
          <dc:description>The student, Xulin Fan, accepted the attached license on 2024-11-26 at 14:26.</dc:description>
          <dc:description>The student, Xulin Fan, submitted this Thesis for approval on 2024-11-26 at 14:39.</dc:description>
          <dc:description>This Thesis was approved for publication on 2024-11-27 at 09:26.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #21403 on 2025-03-28 at 14:43:27</dc:description>
          <dc:description>With advances in deep neural networks and the increasing availability of computational power, machine learning researchers working on unimodal tasks, such as text, audio, or vision, have begun to explore methods that address multimodal problems, which involve inputs from multiple modalities. Each modality carries distinct types of information represented in different formats. For example, audiovisual tasks typically involve an audio-visual aligned video, where the video channel is essentially a sequence of images represented as a 4D tensor (number of frames, RGB channels, height, width), while the audio channel is a 1D signal sampled at a much higher rate. Popular multimodal neural architectures generally consist of three stages: modality-specific encoders, modality fusion, and one or more task-specific decoders. To handle inputs of varying formats, a common design choice is to employ modality-specific encoders, which project each modality into a learnable embedding space. This embedding space is structured to facilitate cross-modal similarity and ease the subsequent fusion process. The modality fusion step, however, is tailored to the requirements of the specific downstream task. For instance, an audiovisual task that outputs an audio signal, such as audiovisual speech enhancement, may require high temporal resolution, whereas tasks that produce textual outputs, such as audiovisual automatic speech recognition, may prioritize different fusion strategies. In this thesis, we explore various design choices for two audiovisual tasks: audio-driven talking head synthesis and audiovisual target speaker extraction. Through extensive experimentation, we identify key considerations for adapting transformer-based and diffusion-based methods to multimodal scenarios.</dc:description>
          <dc:date>2024-12</dc:date>
          <dc:type>Thesis</dc:type>
          <dc:identifier>https://hdl.handle.net/2142/127369</dc:identifier>
          <dc:rights>Copyright 2024 Xulin Fan</dc:rights>
          <degree>
            <department>Electrical &amp; Computer Eng</department>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <name>M.S.</name>
            <level>Thesis</level>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
