<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-19T09:02:47Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/117774" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/117774</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_8888</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_8887</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Kim, Nam Sung</dc:contributor>
          <dc:contributor>Kim, Nam Sung</dc:contributor>
          <dc:contributor>Hwu, Wen-mei</dc:contributor>
          <dc:contributor>Torrellas, Josep</dc:contributor>
          <dc:contributor>Chen, Deming</dc:contributor>
          <dc:date>2022-12</dc:date>
          <dc:format>application/pdf</dc:format>
          <dc:language>en</dc:language>
          <dc:type>text</dc:type>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2023-04-12 without embargo terms</dc:description>
          <dc:description>The student, Youjie Li, accepted the attached license on 2022-11-22 at 20:06.</dc:description>
          <dc:description>The student, Youjie Li, submitted this Dissertation for approval on 2022-11-22 at 21:09.</dc:description>
          <dc:description>This Dissertation was approved for publication on 2022-11-23 at 13:35.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #18624 on 2023-04-12 at 07:31:54</dc:description>
          <dc:title>Communication-centric cross-stack acceleration for distributed machine learning</dc:title>
          <dc:creator>Li, Youjie</dc:creator>
          <dc:date>2022-11-23</dc:date>
          <dc:subject>Distributed Machine Learning</dc:subject>
          <dc:subject>Distributed Training</dc:subject>
          <dc:subject>Parallel Computing</dc:subject>
          <dc:subject>In-network Computing</dc:subject>
          <dc:subject>Smart Nics</dc:subject>
          <dc:subject>Programmable Switches</dc:subject>
          <dc:subject>Pipelined Sgd</dc:subject>
          <dc:subject>Machine Learning Framework</dc:subject>
          <dc:description>Distributed training has been the Holy Grail in machine learning systems, as it is an indispensable technique for addressing the ever-growing computation demands and memory requirements due to the unprecedented scaling of model sizes and data volumes. However, even distributed training takes inordinate time, of which a large fraction is paid for communication overhead in either inter-server networks or intra-server interconnects. In this dissertation, we propose cross-stack solutions that span hardware, software, and algorithms to accelerate and scale distributed training systems. First, we propose in-network computing by leveraging novel network hardware in modern data-centers, such as programmable network interface cards and switches, for not only compressing traffic volume in real time but also reducing network hops on the fly. Second, we present new algorithms surrounding efficient communication, such as gradient compression and pipelined computation with communication, for further shrinking the network overhead while maintaining the training convergence and model accuracy. Third, we develop a next-generation machine learning framework by novel schemes of task decomposition and late binding, to train massive models even without sufficient memory while drastically reducing communication overhead within a server's interconnects.</dc:description>
          <dc:type>Thesis</dc:type>
          <dc:language>eng</dc:language>
          <dc:identifier>https://hdl.handle.net/2142/117774</dc:identifier>
          <dc:rights>In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of [university/educational entity's name goes here]'s products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink.</dc:rights>
          <degree>
            <name>Ph.D.</name>
            <level>Dissertation</level>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <department>Electrical &amp; Computer Eng</department>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
