<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-19T09:37:55Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/102465" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/102465</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>col_2142_10761</setSpec>
        <setSpec>com_2142_5130</setSpec>
        <setSpec>com_2142_10755</setSpec>
        <setSpec>com_2142_234</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Han, Jiawei</dc:contributor>
          <dc:contributor>Han, Jiawei</dc:contributor>
          <dc:contributor>Zhai, ChengXiang</dc:contributor>
          <dc:contributor>Abdelzaher, Tarek</dc:contributor>
          <dc:contributor>Mei, Qiaozhu</dc:contributor>
          <dc:creator>Zhang, Chao</dc:creator>
          <dc:date>2019-02-06T19:36:25Z</dc:date>
          <dc:date>2019-02-06T19:36:25Z</dc:date>
          <dc:date>2018-12-03</dc:date>
          <dc:date>2018-12</dc:date>
          <dc:description>As one of the most important data forms, unstructured text data plays a crucial role in data-driven decision making in domains ranging from social networking and information retrieval to healthcare and scientific research.  In many emerging applications, people's information needs from text data are becoming multi-dimensional---they demand useful insights for multiple aspects from the given text corpus.  However, turning massive text data into multi-dimensional knowledge remains a challenge that cannot be readily addressed by existing data mining techniques.
In this thesis, we propose algorithms that turn unstructured text data into multi-dimensional knowledge with limited supervision. We investigate two core questions: 
1. How to identify task-relevant data with declarative queries in multiple dimensions?
2. How to distill knowledge from data in a multi-dimensional space?
To address the above questions, we propose an integrated cube construction and exploitation framework.  First, we develop a cube construction module that organizes unstructured data into a cube structure, by discovering latent multi-dimensional and multi-granular structure from the unstructured text corpus and allocating documents into the structure.  Second, we develop a cube exploitation module that models multiple dimensions in the cube space, thereby distilling multi-dimensional knowledge from data to provide insights along multiple dimensions.  Together, these two modules constitute an integrated pipeline: leveraging the cube structure, users can perform multi-dimensional, multi-granular data selection with declarative queries; and with cube exploitation algorithms, users can make accurate cross-dimension predictions or extract multi-dimensional patterns for decision making.
The proposed framework has two distinctive advantages when turning text data into multi-dimensional knowledge: flexibility and label-efficiency. First, it enables acquiring multi-dimensional knowledge flexibly, as the cube structure allows users to easily identify task-relevant data along multiple dimensions at varied granularities and further distill multi-dimensional knowledge. Second, the algorithms for cube construction and exploitation require little supervision; this makes the framework appealing for many applications where labeled data are expensive to obtain.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-02-05 without embargo terms</dc:description>
          <dc:description>The student, Chao Zhang, accepted the attached license on 2018-12-02 at 13:48.</dc:description>
          <dc:description>The student, Chao Zhang, submitted this Dissertation for approval on 2018-12-02 at 14:03.</dc:description>
          <dc:description>This Dissertation was approved for publication on 2018-12-03 at 10:17.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #13173 on 2019-02-05 at 11:13:29</dc:description>
          <dc:description>Made available in DSpace on 2019-02-06T19:36:25Z (GMT). No. of bitstreams: 3
ZHANG-DISSERTATION-2018.pdf: 16001723 bytes, checksum: 33dbcc1eba134a7f9ddb45de5832d8c2 (MD5)
LICENSE.txt: 4207 bytes, checksum: 02f937f556cae3ea0735c1ce842d291b (MD5)
PROQUEST_LICENSE.txt: 4553 bytes, checksum: 84435e94bc2e715ebadcbe89164b6023 (MD5)
  Previous issue date: 2018-12-03</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/102465</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2018 Chao Zhang</dc:rights>
          <dc:subject>data mining</dc:subject>
          <dc:subject>multi-dimensional analysis</dc:subject>
          <dc:subject>less supervision</dc:subject>
          <dc:title>Multi-dimensional mining of unstructured data with limited supervision</dc:title>
          <dc:type>text</dc:type>
          <dc:type>text</dc:type>
          <degree>
            <department>Computer Science</department>
            <discipline>Computer Science</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Dissertation</level>
            <name>Ph.D.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
