<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-20T16:34:35Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/132689" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/132689</identifier>
        <datestamp>2026-03-24</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_8951</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_8950</setSpec>
        <setSpec>com_2142_150</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:format>application/pdf</dc:format>
          <dc:language>en</dc:language>
          <dc:type>text</dc:type>
          <dc:description>Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2027-12-01</dc:description>
          <dc:description>The student, Xitong (Jacqueline) Zhang, accepted the attached license on 2025-12-04 at 12:24.</dc:description>
          <dc:description>The student, Xitong (Jacqueline) Zhang, submitted this Thesis for approval on 2025-12-04 at 12:35.</dc:description>
          <dc:description>This Thesis was approved for publication on 2025-12-08 at 16:07.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #23060 on 2026-02-19 at 18:46:43</dc:description>
          <dc:title>Enhancing knowledge distillation in large language models via domain adaptation</dc:title>
          <dc:creator>Zhang, Xitong (Jacqueline)</dc:creator>
          <dc:date>2025-12-08</dc:date>
          <dc:contributor>He, Jingrui</dc:contributor>
          <dc:contributor>Ma, Jiaqi</dc:contributor>
          <dc:subject>Knowledge Distillation</dc:subject>
          <dc:subject>Large Language Models</dc:subject>
          <dc:subject>Deep Learning</dc:subject>
          <dc:subject>Machine Learning</dc:subject>
          <dc:subject>Artificial Intelligence</dc:subject>
          <dc:language>eng</dc:language>
          <dc:description>Domain-Adaptive Pre-Training (DAPT) is widely used to improve Large Language Models on specialized domains, yet its interaction with knowledge distillation (KD) remains poorly understood. In particular, intermediate DAPT checkpoints are rarely analyzed, and the evolution of teacher uncertainty across such checkpoints has not been systematically studied. This thesis develops a unified framework to examine how DAPT reshapes teacher confidence and how these shifts influence KD performance, downstream performance and calibration. Using LLaMA-2-7B teachers adapted for 2,000, 5,000, 7,500, and 10,000 DAPT steps, together with Sheared-LLaMA-1.3B students distilled under four KD variants, we evaluate two biomedical QA benchmarks: PubMedQA and BioASQ. We analyze teacher entropy, entropy–performance correlations, and student Expected Calibration Error (ECE) across checkpoints.

Our findings reveal three key insights: (1) teacher entropy shifts moderately with deeper DAPT but redistributes most strongly over semantically informative tokens; (2) moderate entropy reduction yields the strongest KD gains for abstractive reasoning tasks such as PubMedQA, whereas extractive QA tasks benefit more from heavily domain-adapted teachers whose predictions are sharper and more concentrated, and (3) student calibration closely tracks teacher entropy, with sharper teachers generally producing better-calibrated models, though excessively low entropy can introduce calibration trade-offs.</dc:description>
          <dc:date>2025-12</dc:date>
          <dc:type>Thesis</dc:type>
          <dc:identifier>https://hdl.handle.net/2142/132689</dc:identifier>
          <dc:rights>Copyright 2025 Xitong (Jacqueline) Zhang</dc:rights>
          <degree>
            <department>Information Sciences</department>
            <discipline>Bioinformatics</discipline>
            <grantor>University of Illinois Urbana-Champaign</grantor>
            <name>M.S.</name>
            <level>Thesis</level>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
