<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" href="/oai-pmh.xsl"?>
<OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd">
  <responseDate>2026-09-19T22:28:52Z</responseDate>
  <request identifier="oai:www.ideals.illinois.edu:2142/92851" metadataPrefix="etdms" verb="GetRecord">https://www.ideals.illinois.edu/oai-pmh</request>
  <GetRecord>
    <record>
      <header>
        <identifier>oai:www.ideals.illinois.edu:2142/92851</identifier>
        <datestamp>2023-07-11</datestamp>
        <setSpec>col_2142_5131</setSpec>
        <setSpec>com_2142_5130</setSpec>
      </header>
      <metadata>
        <thesis xmlns="http://www.ndltd.org/standards/metadata/etdms/1.1/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.ndltd.org/standards/metadata/etdms/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdms11.xsd http://purl.org/dc/elements/1.1/ http://www.ndltd.org/standards/metadata/etdms/1.1/etdmsdc.xsd">
          <dc:contributor>Viswanath, Pramod</dc:contributor>
          <dc:contributor>Bhat, Suma P.</dc:contributor>
          <dc:creator>Mu, Jiaqi</dc:creator>
          <dc:date>2016-11-10T17:55:14Z</dc:date>
          <dc:date>2016-11-10T17:55:14Z</dc:date>
          <dc:date>2016-07-15</dc:date>
          <dc:date>2016-08</dc:date>
          <dc:description>Knowledge bases (KB) store relational facts and constitute a significant resource for a variety of natural language processing (NLP) tasks. Improving their coverage and refining the relations is a basic and pressing research effort. In this thesis we propose a novel approach towards this canonical task by using the unstructured Wikipedia corpus: we extract low-dimensional embeddings for title pages of the Wikipedia corpus and show that they can be used to significantly outperform state-of-the-art approaches on a variety of metrics in three concrete tasks: measuring semantic relatedness,  solving semantic analogies, and  KB completion and refinement. A central feature of our work is a new log-linear discriminative model for the annotations inside a Wikipedia document that we name IBOE (isotropic bag-of-entities): we hypothesize that the parameters of the model satisfy a geometric symmetry property (isotropy). We show that the isotropy property leads to self-normalization allowing for the design of an efficient parameter estimation algorithm that we christen wiki2vec. The self-normalization property of IBOE is validated empirically on the Wikipedia corpus and is also of independent mathematical interest.</dc:description>
          <dc:description>Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms</dc:description>
          <dc:description>The student, Jiaqi Mu, accepted the attached license on 2016-07-15 at 10:25.</dc:description>
          <dc:description>The student, Jiaqi Mu, submitted this Thesis for approval on 2016-07-15 at 10:39.</dc:description>
          <dc:description>This Thesis was approved for publication on 2016-07-15 at 13:08.</dc:description>
          <dc:description>DSpace SAF Submission Ingestion Package generated from Vireo submission #9961 on 2016-11-09 at 10:25:14</dc:description>
          <dc:description>Made available in DSpace on 2016-11-10T17:55:14Z (GMT). No. of bitstreams: 2
MU-THESIS-2016.pdf: 1259589 bytes, checksum: 0c2de48d9460151f7e468008a704635e (MD5)
LICENSE.txt: 4205 bytes, checksum: 1db9778bc0832fa860692456466d1b5a (MD5)
  Previous issue date: 2016-07-15</dc:description>
          <dc:format>application/pdf</dc:format>
          <dc:identifier>http://hdl.handle.net/2142/92851</dc:identifier>
          <dc:language>en</dc:language>
          <dc:rights>Copyright 2016 Jiaqi Mu</dc:rights>
          <dc:subject>entity embedding</dc:subject>
          <dc:subject>knowledge base completion</dc:subject>
          <dc:title>Semantic modeling of the natural language of Wikipedia annotations</dc:title>
          <dc:type>text</dc:type>
          <dc:type>text</dc:type>
          <degree>
            <department>Electrical &amp; Computer Eng</department>
            <discipline>Electrical &amp; Computer Engr</discipline>
            <grantor>University of Illinois at Urbana-Champaign</grantor>
            <level>Thesis</level>
            <name>M.S.</name>
          </degree>
        </thesis>
      </metadata>
    </record>
  </GetRecord>
</OAI-PMH>
