<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-18T17:04:08Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/127396" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/127396</identifier>
        <datestamp>2026-02-03</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_10761</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_10755</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:format>application/pdf</dc:format>
          <dc:language>en</dc:language>
          <dc:type>text</dc:type>
          <dc:description>Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2026-12-01</dc:description>
          <dc:description>The student, Xiaoyang Wang, accepted the attached license on 2024-12-04 at 02:22.</dc:description>
          <dc:description>The student, Xiaoyang Wang, submitted this Dissertation for approval on 2024-12-04 at 02:30.</dc:description>
          <dc:description>This Dissertation was approved for publication on 2024-12-04 at 13:32.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #21491 on 2025-03-28 at 14:44:40</dc:description>
          <dc:title>Invariant learning for learning in the wild</dc:title>
          <dc:creator>Wang, Xiaoyang</dc:creator>
          <dc:date>2024-12-04</dc:date>
          <dc:contributor>Koyejo, Oluwasanmi</dc:contributor>
          <dc:contributor>Nahrstedt, Klara</dc:contributor>
          <dc:contributor>Tong, Hanghang</dc:contributor>
          <dc:contributor>Dimitriadis, Dimitrios</dc:contributor>
          <dc:subject>Machine Learning</dc:subject>
          <dc:subject>Invariance</dc:subject>
          <dc:subject>Robustness</dc:subject>
          <dc:language>eng</dc:language>
          <dc:description>Machine learning models are increasingly deployed in production (i.e., the wild) but may fail for various reasons. For example, fraud detection models can protect numerous users against phishing emails but are subject to intentional poisoning and may fail to identify novel types of phishing. Similarly, a language model provides timely answers to user questions. However, the answer quality can decrease significantly or be harmful even if minor changes apply to the questions. Common failures of machine learning models in production environments fall into two categories: (1) data quality and (2) data shift. Data quality problems can be caused by malicious adversaries that aim to corrupt machine learning models, uncurated crowdsourced data from the web, etc. Meanwhile, data shift problems often occur due to the mismatch between the offline training data and the continuously evolving data in online production environments. Tackling the data quality and shift problems requires methods that help machine learning models continuously learn generally useful patterns from the data without entangling the harmful ones. In this dissertation, we introduce invariant learning as a paradigm to meet the aforementioned requirement and address the data quality and shift problems in the wild. In particular, we first study a data quality problem with multiple data sources with mixed data qualities. Our main contribution to this problem is a novel algorithm that helps machine learning models learn invariant patterns from multiple data sources and selectively filter out the contribution of low-quality data. Then, we further study a setting that requires machine learning models to be fine-tuned (i.e., customized) to a particular data source with improved performance but does not sacrifice the invariance benefit. The last part of this dissertation applies invariant learning to an active fine-tuning problem, which requires machine learning models to continuously learn new data with improved data efficiency. Our invariance-aware approach selects subsets of data samples that invariantly benefit the full dataset with minimal neglect of unselected data samples and helps machine learning models adapt to shifting data more effectively.</dc:description>
          <dc:date>2024-12</dc:date>
          <dc:type>Thesis</dc:type>
          <dc:identifier>https://hdl.handle.net/2142/127396</dc:identifier>
          <dc:rights>Copyright 2024 Xiaoyang Wang</dc:rights>
          <degree>
            <department>Siebel School Comp &amp; Data Sci</department>
            <discipline>Computer Science</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <name>Ph.D.</name>
            <level>Dissertation</level>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
