<!DOCTYPE art SYSTEM 'http://www.biomedcentral.com/xml/article.dtd'>
<art>
   <ui>1471-2105-8-264</ui>
   <ji>1471-2105</ji>
   <fm>
      <dochead>Research article</dochead>
      <bibl>
         <title>
            <p>Using contextual and lexical features to restructure and validate the classification of biomedical concepts</p>
         </title>
         <aug>
            <au id="A1">
               <snm>Fan</snm>
               <fnm>Jung-Wei</fnm>
               <insr iid="I1"/>
               <email>jung-wei.fan@dbmi.columbia.edu</email>
            </au>
            <au id="A2">
               <snm>Xu</snm>
               <fnm>Hua</fnm>
               <insr iid="I1"/>
               <email>hua.xu@dbmi.columbia.edu</email>
            </au>
            <au id="A3" ca="yes">
               <snm>Friedman</snm>
               <fnm>Carol</fnm>
               <insr iid="I1"/>
               <email>carol.friedman@dbmi.columbia.edu</email>
            </au>
         </aug>
         <insg>
            <ins id="I1">
               <p>Department of Biomedical Informatics, Columbia University Vanderbilt Clinic, 5th Floor, 622 West 168th Street, New York, NY 10032, USA</p>
            </ins>
         </insg>
         <source>BMC Bioinformatics</source>
         <issn>1471-2105</issn>
         <pubdate>2007</pubdate>
         <volume>8</volume>
         <issue>1</issue>
         <fpage>264</fpage>
         <url>http://www.biomedcentral.com/1471-2105/8/264</url>
         <xrefbib>
            <pubidlist>
               <pubid idtype="pmpid">17650333</pubid>
               <pubid idtype="doi">10.1186/1471-2105-8-264</pubid>
            </pubidlist>
         </xrefbib>
      </bibl>
      <history>
         <rec>
            <date>
               <day>06</day>
               <month>3</month>
               <year>2007</year>
            </date>
         </rec>
         <acc>
            <date>
               <day>24</day>
               <month>7</month>
               <year>2007</year>
            </date>
         </acc>
         <pub>
            <date>
               <day>24</day>
               <month>7</month>
               <year>2007</year>
            </date>
         </pub>
      </history>
      <cpyrt>
         <year>2007</year>
         <collab>Fan et al; licensee BioMed Central Ltd.</collab>
         <note>This is an Open Access article distributed under the terms of the Creative Commons Attribution License (<url>http://creativecommons.org/licenses/by/2.0</url>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</note>
      </cpyrt>
      <abs>
         <sec>
            <st>
               <p>Abstract</p>
            </st>
            <sec>
               <st>
                  <p>Background</p>
               </st>
               <p>Biomedical ontologies are critical for integration of data from diverse sources and for use by knowledge-based biomedical applications, especially natural language processing as well as associated mining and reasoning systems. The effectiveness of these systems is heavily dependent on the quality of the ontological terms and their classifications. To assist in developing and maintaining the ontologies objectively, we propose automatic approaches to classify and/or validate their semantic categories. In previous work, we developed an approach using contextual syntactic features obtained from a large domain corpus to reclassify and validate concepts of the Unified Medical Language System (UMLS), a comprehensive resource of biomedical terminology. In this paper, we introduce another classification approach based on words of the concept strings and compare it to the contextual syntactic approach.</p>
            </sec>
            <sec>
               <st>
                  <p>Results</p>
               </st>
               <p>The string-based approach achieved an error rate of 0.143, with a mean reciprocal rank of 0.907. The context-based and string-based approaches were found to be complementary, and the error rate was reduced further by applying a linear combination of the two classifiers. The advantage of combining the two approaches was especially manifested on test data with sufficient contextual features, achieving the lowest error rate of 0.055 and a mean reciprocal rank of 0.969.</p>
            </sec>
            <sec>
               <st>
                  <p>Conclusion</p>
               </st>
               <p>The lexical features provide another semantic dimension in addition to syntactic contextual features that support the classification of ontological concepts. The classification errors of each dimension can be further reduced through appropriate combination of the complementary classifiers.</p>
            </sec>
         </sec>
      </abs>
   </fm>
   <bdy>
      <sec>
         <st>
            <p>Background</p>
         </st>
         <sec>
            <st>
               <p>Introduction</p>
            </st>
            <p>Biomedical ontologies such as Gene Ontology (GO) <abbrgrp><abbr bid="B1">1</abbr></abbrgrp>, the Foundational Model of Anatomy (FMA) <abbrgrp><abbr bid="B2">2</abbr></abbrgrp>, and the Unified Medical Language System (UMLS) <abbrgrp><abbr bid="B3">3</abbr><abbr bid="B4">4</abbr></abbrgrp> are important for terminology management, data sharing/integration, and decision support <abbrgrp><abbr bid="B5">5</abbr></abbrgrp>. The ontologies specify not only the definitions of biomedical terms but also associate them with normalized concepts and semantic categories within the ontological structures. Therefore, they provide abundant lexical and semantic knowledge that is especially valuable to Natural Language Processing (NLP) systems. The overhead involved in the costly and time-consuming system development process could be substantially reduced with the aid of these knowledge sources. NLP techniques have been playing an increasingly critical role in bioinformatics research <abbrgrp><abbr bid="B6">6</abbr><abbr bid="B7">7</abbr><abbr bid="B8">8</abbr></abbrgrp>. Spasic et al. <abbrgrp><abbr bid="B9">9</abbr></abbrgrp> summarized various approaches that applied ontologies in biomedical NLP tasks.</p>
            <p>High-quality semantic classification is crucial for NLP and for other ontology-based applications that take advantage of conceptualization and reasoning, because the semantic accuracy can affect the correctness and/or the flexibility of the applications. We have been investigating automated methods to assist in developing and maintaining the semantic classification of biomedical ontologies for NLP purposes. The goal is two-fold: to directly improve the ontologies and to indirectly improve the NLP applications built upon the ontologies. Our work currently focuses on semantic classification within the UMLS because it has evolved <abbrgrp><abbr bid="B10">10</abbr></abbrgrp> to include vocabularies such as NCBI taxonomy <abbrgrp><abbr bid="B11">11</abbr></abbrgrp>, GO, and OMIM <abbrgrp><abbr bid="B12">12</abbr></abbrgrp>, making it valuable not only for clinical applications but for bioinformatics applications as well. In addition to currently being the most comprehensive resource of biomedical terms, the UMLS has a broad user population within the biomedical domain, and is continually maintained by the National Library of Medicine (NLM).</p>
            <p>In a previous work <abbrgrp><abbr bid="B13">13</abbr></abbrgrp> we demonstrated the feasibility of using a corpus-based, distributional similarity approach to semantic classification of UMLS concepts that appear in text. An evaluation of the method demonstrated that it achieved a lowest error rate of 0.198. The distributional approach was based on Harris' sublanguage theory stating that the syntactic dependence of words on other words exhibits unequal likelihood constraints especially in specialized domain <abbrgrp><abbr bid="B14">14</abbr></abbrgrp>. For example, in the biomedical domain the head noun of the adjective "homologous" is more likely to be a <it>gene or protein </it>than a <it>disorder</it>, and such unequal distribution of syntactic dependences can be used to characterize different semantic categories. Subsequently, we recognized that in the biomedical domain the terms used to name the concepts are generally descriptive, and can be as important as the contextual information for characterizing semantic categories. For example, words such as "syndrome", "deficiency", and "malignant" occur frequently in concepts of the <it>disorder </it>class and thus should be statistically representative of the class. Therefore, in this paper we developed a string-based classification approach and compared it with the distributional approach. The string-based approach uses words constituting the concepts. We hypothesized that it is feasible to characterize a semantic class through the lexical usage of terms associated with concepts in that class.</p>
            <p>Both the distributional and the string-based approaches are automated and are based on empirical language usage, so they differ from manually classifying the concepts and should perform more consistently and objectively than humans. We expect that the automated methods will assist human experts by proposing classifications in a high-throughput manner. Additionally, the semantic classes used by the two approaches can be defined with varying coverage or granularity, preserving the flexibility to customize the classifiers for different applications. To the best of our knowledge, this is the first time the string-based approach has been used and compared with a distributional approach for semantically classifying UMLS concepts.</p>
            <p>In the following subsections we introduce background material concerning the methods and resources related to our study, as well as the issues involved, and then describe the experiments, followed by the results, discussion, conclusion, and details of the methods.</p>
         </sec>
         <sec>
            <st>
               <p>The UMLS Semantic Network</p>
            </st>
            <p>The current release (2007AA) of UMLS integrates 139 source vocabularies into a concept-centered terminology, with each concept assigned a Concept Unique Identifier (CUI). Each CUI is assigned to one or more semantic types in the Semantic Network (SN) <abbrgrp><abbr bid="B15">15</abbr></abbrgrp>, which specifies semantic relations between the semantic types. The biomedical knowledge in the SN is preserved in both the semantic classification and the semantic relations. For example, type-relation-type triplets such as T195:("Antibiotic") <b>interrupts </b>T043:("Cell Function") encode propositions that can be used for various purposes. The UMLS concepts and their semantic types have been used by some NLP applications to determine relations among the extracted terms through designated semantic patterns <abbrgrp><abbr bid="B16">16</abbr><abbr bid="B17">17</abbr></abbrgrp>. Although semantic classification is crucial to the NLP applications, Cimino et al. have reported inconsistencies <abbrgrp><abbr bid="B18">18</abbr></abbrgrp> in the SN classification. The SN is also known to associate numerous concepts with questionable semantic types. For example, the current release assigns "Cavitated nodule", "Calcified nodule", and "Ossified nodule" to T169 "Functional Concept", but "Functional Concept" is very general and these terms are semantically closer to T190 "Anatomical Abnormality".</p>
            <p>Another issue involved in the SN classification is that the granularity may be appropriate for certain knowledge-based applications but may be too fine-grained for general biomedical NLP systems. For example, in the previous work we grouped "Fungus", "Virus", "Rickettsia or Chlamydia", "Bacterium", and "Archaeon" into a single <it>microorganism </it>class because, for that group, the textual patterns and textual relations with other biomedical entities are generally similar. From an ontological perspective, Burgun et al. <abbrgrp><abbr bid="B19">19</abbr></abbrgrp> also suggested that the SN types should be further simplified to form basic-level semantic categories. There has been research on simplifying <abbrgrp><abbr bid="B20">20</abbr><abbr bid="B21">21</abbr><abbr bid="B22">22</abbr></abbrgrp> and auditing <abbrgrp><abbr bid="B23">23</abbr><abbr bid="B24">24</abbr><abbr bid="B25">25</abbr></abbrgrp> the UMLS semantic classification. However, the approaches relied on existing semantic types to form broader categories as well as audit questionable semantic assignments, so they could not avoid the limitations of the existing SN structure. For example, if a gene has been assigned to only T025 "Cell" from the very beginning, it will be still incorrectly grouped into an anatomy-related broad class without any contradictions detected. In other words, re-organizing the semantic types <it>in situ </it>does not help remove the erroneous semantic assignments, and none of the auditing methods provides automated classification(s) for individual concepts. Therefore, in this paper we propose approaches different from the above related work in that we reclassify the concepts directly and automatically.</p>
         </sec>
         <sec>
            <st>
               <p>Corpus-based semantic similarity</p>
            </st>
            <p>The corpus-based approach derived from Harris' distributional hypothesis <abbrgrp><abbr bid="B26">26</abbr></abbrgrp> states that terms can be semantically characterized and classified by the distribution of frequently co-occurring terms associated with specific syntactic relations. For example, a syntactic dependency of "cellular defense response" in the sample sentence "Cellular defense response is affected by temperature" would be <b>passive_verb</b>(affected), denoting that "cellular defense response" is the object of "affected". If we zoom in a little bit, we can also specify the syntactic dependency <b>modifier_adjective</b>(cellular) for the term "defense response". In general, the syntactic dependencies serve as more informative features for semantic classification, and it has been reported in related work <abbrgrp><abbr bid="B27">27</abbr><abbr bid="B28">28</abbr></abbrgrp> that they outperformed simpler approaches that used only co-occurring terms without considering syntax. From a parsed corpus, we can extract sets of such syntactic dependencies for individual concepts (e.g. "cellular defense response") and for semantic classes (e.g. <it>biologic function</it>). The syntactic dependencies of each concept are collected, assigned weights, and normalized into a probability distribution profile for the concept. In order to create a distributional profile for each class, a similar procedure is followed. In this case, the syntactic dependencies of all the concepts associated with the class are aggregated, assigned weights, and normalized to form its distributional profile. Through the distributional profiles, we can perform classification by computing the semantic similarities between the profile of an individual concept and that of a set of semantic classes. Similarity measures (e.g. Lee's &#945;-skew divergence <abbrgrp><abbr bid="B29">29</abbr></abbrgrp>) that quantify the overlapping between probability distributions can be used, and a concept will be classified as the semantic class with which it has the highest similarity score. Readers may refer to our previous work <abbrgrp><abbr bid="B13">13</abbr></abbrgrp> for more details about the methods.</p>
            <p>In the biomedical domain, Sibanda et al. <abbrgrp><abbr bid="B30">30</abbr></abbrgrp> aimed to show that syntactic dependencies are useful for determining the semantic categories of terms and clauses in discharge summaries, though a Support Vector Machine was used instead of a canonical distributional model. They defined eight semantic categories, including complex categories such as <it>results</it>. Some of the categories can be defined directly as a subset of SN types, but some (e.g. <it>dosages</it>) have no SN type equivalents. Using syntactic features along with orthographic and lexical features, <it>F</it>-measures of >90% were achieved for most categories. Weeds et al. <abbrgrp><abbr bid="B31">31</abbr></abbrgrp> applied their co-occurrence retrieval distributional similarity method based on syntactic dependencies and a nearest neighbor voting process to classify terms into the 36 semantic categories of the GENIA corpus <abbrgrp><abbr bid="B32">32</abbr></abbrgrp>, achieving a best accuracy of 0.768. In our previous work <abbrgrp><abbr bid="B13">13</abbr></abbrgrp> of reclassifying UMLS concepts into seven broad semantic categories, we used syntactic dependencies that were initially extracted from shallow-parsed PubMed abstracts and then organized into concept-based distributional profiles for each of the categories. Lee's &#945;-skew divergence was used as the similarity metric, and a best accuracy of 0.802 was achieved. For this paper we applied the same method to build our distributional classifiers but built them using a larger training corpus and incorporated some minor implementation refinements.</p>
         </sec>
         <sec>
            <st>
               <p>String-based approach to semantic similarity</p>
            </st>
            <p>Lexical features (e.g. content words or phrases) are usually used in text categorization tasks, with wide applications in general NLP such as news topic classification <abbrgrp><abbr bid="B33">33</abbr></abbrgrp> and authorship verification <abbrgrp><abbr bid="B34">34</abbr></abbrgrp> (see Sebastiani's review <abbrgrp><abbr bid="B35">35</abbr></abbrgrp> on text categorization). Regardless of the text size, tasks dealing with small pieces of text can still be considered as a special case of text categorization. For example, once a topic classifier is built, it can be used to classify headlines consisting of even only a few words. In the biomedical domain, words as well as other processed lexical tokens were shown to be useful for named entity recognition/classification. For example, Lee et al. <abbrgrp><abbr bid="B36">36</abbr></abbrgrp> built SVM classifiers based on both contextual (surrounding words and bigrams) and lexical features to classify terms into GENIA semantic categories. Working also on GENIA semantic classification, Torii et al. <abbrgrp><abbr bid="B37">37</abbr></abbrgrp> applied a decision list algorithm and found lexical features more helpful than contextual features (adjacent words). In Zhang and colleague's work of partitioning the Semantic Network <abbrgrp><abbr bid="B22">22</abbr></abbrgrp>, words of the SN type definitions were used to compute semantic similarity for grouping the types.</p>
            <p>The Na&#239;ve Bayesian model <abbrgrp><abbr bid="B38">38</abbr></abbrgrp> is widely used by text categorization systems. In Na&#239;ve Bayesian classification the target posterior probability for a specific class <it>c</it><sub>i </sub>given a set of features <it>F </it>is computed as:</p>
            <p>
               <display-formula id="M1">P(<it>c</it><sub>i</sub>|F) = P(<it>F</it>|<it>c</it><sub>i</sub>)&#183;P(<it>c</it><sub>i</sub>)/P(<it>F</it>)</display-formula>
            </p>
            <p>
               <display-formula id="M2">&#8776;P(<it>f</it><sub>1</sub>, <it>f</it><sub>2</sub>, ..., <it>f</it><sub>n</sub>|<it>c</it><sub>i</sub>)&#183;P(<it>c</it><sub>i</sub>) &#160;&#160;&#160; (omit the common denominator)</display-formula>
            </p>
            <p>
               <display-formula id="M3">&#8776;P(<it>f</it><sub>1</sub>|<it>c</it><sub>i</sub>)&#183;P(<it>f</it><sub>2</sub>|<it>c</it><sub>i</sub>)&#183;...&#183;P(<it>f</it><sub>n</sub>|<it>c</it><sub>i</sub>)&#183;P(<it>c</it><sub>i</sub>) &#160;&#160;&#160; (assume conditional independence)</display-formula>
            </p>
            <p>The correct class is predicted as the <it>c</it><sub>i </sub>which maximizes the product of formula <b>3</b>. For this paper we classify UMLS concepts into high-level semantic categories such as <it>biologic function </it>and <it>disorder</it>, i.e. the <it>c</it><sub>i </sub>in the above formulas. The features by our approach are words of the strings associated with each CUI. For example, the string "cellular defense response" of the concept C1155076 provides "cellular", "defense", and "response" as the features of the concept. The prior probabilities P(<it>c</it><sub>i</sub>) and the conditional probabilities P(<it>f</it><sub>j</sub>|<it>c</it><sub>i</sub>) can both be estimated from the training data. By pooling the words of multiple CUIs under certain SN type or arbitrarily defined high-level semantic class, a larger lexical profile can be formed for Na&#239;ve Bayesian classification. The details of the implementation are described in Methods.</p>
         </sec>
         <sec>
            <st>
               <p>The UMLS for building the string-based classifier</p>
            </st>
            <p>The UMLS has well-defined SN types and some less well-defined, semantically general or vague SN types. SN types such as T047 "Disease or Syndrome" are well-defined, and the concepts under them are semantically homogeneous and could reliably be used for training, whereas some types are not appropriate for building classifiers. For example, T033 "Finding", is very general. It includes concepts corresponding to biological functions such as "Basal gastric acid output" and "Nitrogen balance", but also includes many disorder-related concepts such as "Hyperlactatemia", "Progressive renal failure", and some very general findings such as "Unemployment" and "Birth place". By pooling lexical features of well-defined SN types into semantic classes we can use the UMLS to train the Na&#239;ve Bayesian model introduced earlier, while excluding those less well-defined SN types during the training process.</p>
            <p>A subset of the UMLS source vocabularies is known for well-established ontological structure. In addition, some have semantic annotations embedded in the UMLS strings. For example, some GO terms contain a parenthesized "sensu" note for specifying taxonomical information, e.g. "cell wall polysaccharide anabolism (sensu Fungi)". Similar but not strictly ontological, the HUGO Nomenclature <abbrgrp><abbr bid="B39">39</abbr></abbrgrp> has some taxonomical notes parenthesized as in "mindbomb homolog 2 (Drosophila)" or some semantic qualifiers parenthesized as in "endothelial cell growth factor 1 (platelet-derived)". Cohen and colleagues have reported that parenthesized information contributes to naming variants within and between genes <abbrgrp><abbr bid="B40">40</abbr></abbrgrp>. Another valuable source ontology is SNOMED-CT <abbrgrp><abbr bid="B41">41</abbr></abbrgrp> because it has sound ontological properties and a broad coverage of biomedicine. SNOMED semantic classes such as "Disease", "Body structure", "Function", and "Substance" occasionally occur in the UMLS string terms as parenthesized annotations, e.g. "Oxidative phosphorylation (function)". These parenthesized annotations provide extra information of the ontological views from different source vocabularies. Therefore, in building the Na&#239;ve Bayesian classifiers we experimented with features consisting of only the pure strings (i.e. discarding parenthesized annotations) and those consisting of strings including the annotations. To reduce training noise, we also excluded strings that were marked as suppressible synonyms because they were considered obsolescent by the corresponding source vocabularies or by the NLM.</p>
         </sec>
         <sec>
            <st>
               <p>Summary of experiments in the current study</p>
            </st>
            <p>We evaluated the string-based approach and compared it with the distributional approach for classifying UMLS concepts into eight broad semantic classes: <it>anatomy </it>(above the molecular level), <it>behavior</it>, <it>biologic function</it>, <it>disorder</it>, <it>gene or protein</it>, <it>microorganism</it>, <it>procedure</it>, and <it>substance</it>. We chose these specific classes because according to the recent reviews of the field <abbrgrp><abbr bid="B6">6</abbr><abbr bid="B7">7</abbr><abbr bid="B8">8</abbr></abbrgrp> they comprised the most relevant ones for biomedical text mining applications. For each approach, the eight classes were trained using 64 well-defined SN types ' [see Additional file <supplr sid="S1">1</supplr>]' out of the total 135, while noisy types such as T033 discussed in the previous section were not used. For the distributional approach we varied the number of available syntactic dependencies that were used, and for the string-based approach we used strings either with or without parenthesized annotations. The two approaches were compared on the same test sets. The gold standard was automatically generated based on the CUIs that had their SN types modified within the 2005AA~ 2006AA interval (by comparing the UMLS MRSTY.RRF tables). The gold standard had been evaluated by a biomedical expert with an M.D. degree using a subset of CUIs that were randomly sampled from the testing data. The expert's agreement with the gold standard's classification was more than 86% (43/50), and therefore the gold standard was determined to be reliable. For methodological details concerning the gold standard creation and evaluation, please refer to our previous work <abbrgrp><abbr bid="B13">13</abbr></abbrgrp>.</p>
            <suppl id="S1">
               <title>
                  <p>Additional file 1</p>
               </title>
               <text>
                  <p>The 8 broad semantic classes and their constituent SN types. A detailed list of the Semantic Types selected to build each of our eight broad classes.</p>
               </text>
               <file name="1471-2105-8-264-S1.pdf">
                  <p>Click here for file</p>
               </file>
            </suppl>
            <p>The main quantitative evaluation reported in this paper consists of the error rates of the different automated classifiers. Each error rate was computed by the formula:</p>
            <p>
               <display-formula id="M4">
                  <m:math name="1471-2105-8-264-i1" xmlns:m="http://www.w3.org/1998/Math/MathML">
                     <m:semantics>
                        <m:mrow>
                           <m:mfrac>
                              <m:mrow>
                                 <m:mi>M</m:mi>
                                 <m:mo>+</m:mo>
                                 <m:mn>0.5</m:mn>
                                 <m:mo>&#8901;</m:mo>
                                 <m:mi>T</m:mi>
                              </m:mrow>
                              <m:mi>N</m:mi>
                           </m:mfrac>
                        </m:mrow>
                        <m:annotation encoding="MathType-MTEF">
 MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaadaWcaaqaaiabd2eanjabgUcaRiabicdaWiabc6caUiabiwda1iabgwSixlabdsfaubqaaiabd6eaobaaaaa@362B@</m:annotation>
                     </m:semantics>
                  </m:math>
               </display-formula>
            </p>
            <p>where <it>M </it>is the number of misclassifications, <it>T </it>is the number of ties, and <it>N </it>is the total number of CUIs tested. A tie is defined when the correct or incorrect predictions receive equal similarity scores from the classifier. The correct/incorrect classification counting for a single test CUI was divided by the number of classes it was associated with in the gold standard. For example, if in the gold standard a CUI belonged to both <it>gene or protein </it>and <it>substance</it>, and in the top 2 predictions the classifier only got <it>substance</it>, it was counted as a 0.5 correct classification and a 0.5 misclassification. The error rate can be understood as a measure complementary to classification accuracy, with tied cases penalized as a half misclassification. We also applied linear combinations varying the relative weights of the two approaches and evaluated the error rates of the hybrid classifiers.</p>
            <p>To evaluate the efficacy of the classifiers' ranking function (i.e. the ability of bringing the correct class up to the topmost prediction), we calculated the mean reciprocal rank (MRR) for each experiment:</p>
            <p>
               <display-formula id="M5">
                  <m:math name="1471-2105-8-264-i2" xmlns:m="http://www.w3.org/1998/Math/MathML">
                     <m:semantics>
                        <m:mrow>
                           <m:mfrac>
                              <m:mrow>
                                 <m:mstyle displaystyle="true">
                                    <m:mo>&#8721;</m:mo>
                                    <m:mrow>
                                       <m:mrow>
                                          <m:mo>(</m:mo>
                                          <m:mrow>
                                             <m:mn>1</m:mn>
                                             <m:mo>/</m:mo>
                                             <m:mtext>R</m:mtext>
                                             <m:mi>t</m:mi>
                                          </m:mrow>
                                          <m:mo>)</m:mo>
                                       </m:mrow>
                                    </m:mrow>
                                 </m:mstyle>
                              </m:mrow>
                              <m:mi>T</m:mi>
                           </m:mfrac>
                        </m:mrow>
                        <m:annotation encoding="MathType-MTEF">
 MathType@MTEF@5@5@+=feaafiart1ev1aaatCvAUfKttLearuWrP9MDH5MBPbIqV92AaeXatLxBI9gBaebbnrfifHhDYfgasaacH8akY=wiFfYdH8Gipec8Eeeu0xXdbba9frFj0=OqFfea0dXdd9vqai=hGuQ8kuc9pgc9s8qqaq=dirpe0xb9q8qiLsFr0=vr0=vr0dc8meaabaqaciaacaGaaeqabaqabeGadaaakeaadaWcaaqaamaaqaeabaWaaeWaaeaacqaIXaqmcqGGVaWlcqqGsbGucqWG0baDaiaawIcacaGLPaaaaSqabeqaniabggHiLdaakeaacqWGubavaaaaaa@3606@</m:annotation>
                     </m:semantics>
                  </m:math>
               </display-formula>
            </p>
            <p>where <it>T </it>is the total number of targets to be predicted and R<it>t </it>is the rank of a specific target <it>t </it>among the classifier's predictions. MRR has been used in the TREC Genomics Track <abbrgrp><abbr bid="B42">42</abbr></abbrgrp> for an information retrieval task similar to question-answering. To evaluate the ranking performance of our classifiers we adapted the MRR by assuming that the gold standard class was the target to be retrieved. The MRR value lies in the range of (0, 1), and a value of 1 is considered the best the performance.</p>
         </sec>
      </sec>
      <sec>
         <st>
            <p>Results</p>
         </st>
         <sec>
            <st>
               <p>Error rates of individual approaches</p>
            </st>
            <p>The error rates computed by formula <b>4 </b>are presented in Table <tblr tid="T1">1</tblr>, where <it>N </it>represents the number of CUIs tested. The columns show results on test sets with varying number of syntactic dependencies (the complete test set disregarding syntactic dependencies, &#8805; 1, and &#8805; 10). The table also shows the MRR values of the individual experiments. The majority of the MRR values were above 0.85 (except for the distributional approach experiment where there were insufficient features), indicating that the ranking mechanism of the classifiers functioned well. The string-based approach achieved a highest MRR of 0.907 (when including parenthesized annotations), and the distributional approach achieved that of 0.887.</p>
            <tbl id="T1">
               <title>
                  <p>Table 1</p>
               </title>
               <caption>
                  <p>Error rates of the experiments</p>
               </caption>
               <tblbdy cols="7">
                  <r>
                     <c>
                        <p/>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>The entire test set </b>
                        </p>
                        <p>(<it>N </it>= 223)</p>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>&#8805; 1 syntactic dependencies </b>
                        </p>
                        <p>(<it>N </it>= 192)</p>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>&#8805; 10 syntactic dependencies </b>
                        </p>
                        <p>(<it>N </it>= 91)</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="7">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>
                           <b>Type of features</b>
                        </p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="7">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>Syntactic dependencies</p>
                     </c>
                     <c ca="center">
                        <p>--</p>
                     </c>
                     <c ca="center">
                        <p>--</p>
                     </c>
                     <c ca="center">
                        <p>0.315</p>
                     </c>
                     <c ca="center">
                        <p>0.792</p>
                     </c>
                     <c ca="center">
                        <p>0.187</p>
                     </c>
                     <c ca="center">
                        <p>
                           <b>0.887</b>
                        </p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>Strings without annotations*</p>
                     </c>
                     <c ca="center">
                        <p>0.191</p>
                     </c>
                     <c ca="center">
                        <p>0.881</p>
                     </c>
                     <c ca="center">
                        <p>0.221<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.862<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.247<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.853<sup>#</sup></p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>Strings with annotations*</p>
                     </c>
                     <c ca="center">
                        <p>0.143</p>
                     </c>
                     <c ca="center">
                        <p>
                           <b>0.907</b>
                        </p>
                     </c>
                     <c ca="center">
                        <p>0.164<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.895<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.192<sup>#</sup></p>
                     </c>
                     <c ca="center">
                        <p>0.886<sup>#</sup></p>
                     </c>
                  </r>
               </tblbdy>
               <tblfn>
                  <p>* Note that the string-based approach is independent of the number of available syntactic features used by the distributional approach, and the alignment (<sup>#</sup>) is just for comparison purpose.</p>
               </tblfn>
            </tbl>
            <p>The advantage of the string-based approach over the distributional approach was stronger on test sets where the latter obtained fewer features (e.g. the error rate 0.164 versus the 0.315 on the test CUIs with &#8805; 1 syntactic dependencies). However, it should be clarified that the string-based approach is independent of the number of available syntactic features used by the distributional approach, and the alignment (the entires marked with <sup>#</sup> in the table) is just to allow for comparison on the same test sets. Therefore, the first column (entire test set) actually presents a general evaluation of the string-based approach, and the 223 test CUIs form the superset of all the other test sets in the table. It can also be estimated that about 86% (192/223) of the test CUIs had &#8805; 1 syntactic dependencies for the eight selected broad classes, suggesting the proportion which the distributional approach influenced.</p>
         </sec>
         <sec>
            <st>
               <p>Summary of the misclassifications</p>
            </st>
            <p>We manually analyzed the misclassifications in the column with &#8805; 10 syntactic dependencies in Table <tblr tid="T1">1</tblr>. The distributional classifiers made 17 misclassifications (see Table <tblr tid="T2">2</tblr> for the confusion matrix), of which the top 3 misclassified classes were: <it>biologic function </it>(6), <it>disorder </it>(3), and <it>procedure </it>(3). For example, "Down-regulation" was misclassified as <it>procedure</it>, and "Acetylation" was misclassified as <it>gene or protein</it>. The string-based classifier without using parenthesized annotations made 22.5 misclassifications, of which the top 3 misclassified classes were: <it>anatomy </it>(7), <it>substance </it>(6), and <it>biologic function </it>(5). For example, "Shoulder" and "Foot" were misclassified as <it>disorder</it>. The <it>substance </it>"Alkylating agents" was misclassified as <it>disorder</it>, and the <it>biologic function </it>"Anergy" was misclassified as <it>procedure</it>. There were 7 overlapping misclassifications between the 17 and 22.5, but 4 of them were misclassified as different incorrect classes.</p>
            <tbl id="T2">
               <title>
                  <p>Table 2</p>
               </title>
               <caption>
                  <p>Confusion matrix for the 17 misclassifications by the distributional approach on the <it>N </it>= 91 test set</p>
               </caption>
               <tblbdy cols="9">
                  <r>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1)</p>
                     </c>
                     <c ca="center">
                        <p>2)</p>
                     </c>
                     <c ca="center">
                        <p>3)</p>
                     </c>
                     <c ca="center">
                        <p>4)</p>
                     </c>
                     <c ca="center">
                        <p>5)</p>
                     </c>
                     <c ca="center">
                        <p>6)</p>
                     </c>
                     <c ca="center">
                        <p>7)</p>
                     </c>
                     <c ca="center">
                        <p>8)</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="9">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>1) anatomy</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>2</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>2) behavior</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>3) biologic function</p>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>2</p>
                     </c>
                     <c ca="center">
                        <p>2</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>4) disorder</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>3</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>5) gene or protein</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>6) microorganism</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>7) procedure</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>8) substance</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
               </tblbdy>
            </tbl>
            <p>The string-based classifier using parenthesized annotations made 17.5 misclassifications (see Table <tblr tid="T3">3</tblr> for the confusion matrix), of which the top 3 misclassified classes were also: <it>anatomy </it>(5), <it>substance </it>(4), and <it>biologic function </it>(4), but each had fewer misclassifications as noted in parentheses. For example, the anatomical modifiers "Peritoneal" and "Popliteal" were both misclassified as <it>procedure</it>. The <it>substance </it>"Quinolones" was misclassified as <it>disorder</it>, and the <it>biologic function </it>"Hemolysis" was misclassified as <it>behavior</it>. By including parenthesized annotations, 6 of the 22.5 misclassifications by the pure string-based approach were corrected. However, the <it>procedure </it>"Acoustic Evoked Brain Stem Potentials" was misclassified as <it>biologic function </it>only after including the parenthesized annotations. Please ' [see Additional file <supplr sid="S2">2</supplr>, <supplr sid="S3">3</supplr>, <supplr sid="S4">4</supplr>]' for complete lists of the misclassifications in the three experiments summarized above.</p>
            <suppl id="S2">
               <title>
                  <p>Additional file 2</p>
               </title>
               <text>
                  <p>Summary of the 17 misclassifications by the distributional approach. A more detailed list of the misclassified concepts by the distributional approach.</p>
               </text>
               <file name="1471-2105-8-264-S2.pdf">
                  <p>Click here for file</p>
               </file>
            </suppl>
            <suppl id="S3">
               <title>
                  <p>Additional file 3</p>
               </title>
               <text>
                  <p>Summary of the 22.5 misclassifications by the string-based approach without parenthesized annotations. A more detailed list of the misclassified concepts by the string-based approach without using the parenthesized annotations.</p>
               </text>
               <file name="1471-2105-8-264-S3.pdf">
                  <p>Click here for file</p>
               </file>
            </suppl>
            <suppl id="S4">
               <title>
                  <p>Additional file 4</p>
               </title>
               <text>
                  <p>Summary of the 17.5 misclassifications by the string-based approach with parenthesized annotations. A more detailed list of the misclassified concepts by the string-based approach using the parenthesized annotations.</p>
               </text>
               <file name="1471-2105-8-264-S4.pdf">
                  <p>Click here for file</p>
               </file>
            </suppl>
            <tbl id="T3">
               <title>
                  <p>Table 3</p>
               </title>
               <caption>
                  <p>Confusion matrix for the 17.5 misclassifications by the string-based approach (with parenthesized annotations) on the <it>N </it>= 91 test set</p>
               </caption>
               <tblbdy cols="9">
                  <r>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1)</p>
                     </c>
                     <c ca="center">
                        <p>2)</p>
                     </c>
                     <c ca="center">
                        <p>3)</p>
                     </c>
                     <c ca="center">
                        <p>4)</p>
                     </c>
                     <c ca="center">
                        <p>5)</p>
                     </c>
                     <c ca="center">
                        <p>6)</p>
                     </c>
                     <c ca="center">
                        <p>7)</p>
                     </c>
                     <c ca="center">
                        <p>8)</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="9">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>1) anatomy</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>3</p>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>2) behavior</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>3) biologic function</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>3</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>4) disorder</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>5) gene or protein</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>6) microorganism</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>7) procedure</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>1</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>8) substance</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c ca="center">
                        <p>3.5</p>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                     <c>
                        <p/>
                     </c>
                  </r>
               </tblbdy>
            </tbl>
         </sec>
         <sec>
            <st>
               <p>Complementariness study</p>
            </st>
            <p>The three experiments concerning misclassification analysis summarized in the preceding section used the same test set (91 CUIs), and represented three approaches respectively: distributional, string-based, and string-based plus parenthesized annotations. We computed how many misclassifications by one approach were correct using the other approach. In Table <tblr tid="T4">4</tblr> the values in each cell represent the proportion of how many misclassifications (denominator) by the approach represented by the row were correctly classified (numerator) by the approach represented by the column. For example, the top right value 12/17 means 12 out of 17 misclassifications by the distributional approach were correctly classified by the string-based approach with parenthesized annotations.</p>
            <tbl id="T4">
               <title>
                  <p>Table 4</p>
               </title>
               <caption>
                  <p>Misclassifications that are complemented to be correct by different approaches</p>
               </caption>
               <tblbdy cols="4">
                  <r>
                     <c ca="center">
                        <p>
                           <b>Features used</b>
                        </p>
                        <p>
                           <b>+ Similarity measure</b>
                        </p>
                     </c>
                     <c ca="center">
                        <p>Distributional dependencies</p>
                        <p>+ &#945;-skew divergence</p>
                     </c>
                     <c ca="center">
                        <p>Strings without annotations</p>
                        <p>+ Na&#239;ve Bayesian</p>
                     </c>
                     <c ca="center">
                        <p>Strings with annotations</p>
                        <p>+ Na&#239;ve Bayesian</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="4">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="center">
                        <p>Distributional dependencies</p>
                        <p>+ &#945;-skew divergence</p>
                     </c>
                     <c ca="center">
                        <p>--</p>
                     </c>
                     <c ca="center">
                        <p>10/17</p>
                     </c>
                     <c ca="center">
                        <p>12/17</p>
                     </c>
                  </r>
                  <r>
                     <c ca="center">
                        <p>Strings without annotations</p>
                        <p>+ Na&#239;ve Bayesian</p>
                     </c>
                     <c ca="center">
                        <p>15.5/22.5</p>
                     </c>
                     <c ca="center">
                        <p>--</p>
                     </c>
                     <c ca="center">
                        <p>6/22.5</p>
                     </c>
                  </r>
                  <r>
                     <c ca="center">
                        <p>Strings with annotations</p>
                        <p>+ Na&#239;ve Bayesian</p>
                     </c>
                     <c ca="center">
                        <p>12.5/17.5</p>
                     </c>
                     <c ca="center">
                        <p>1/17.5</p>
                     </c>
                     <c ca="center">
                        <p>--</p>
                     </c>
                  </r>
               </tblbdy>
            </tbl>
            <p>The first row indicates that more than half (10/17) of the misclassifications by the distributional approach were correct when using the string-based approach. The string-based approach was strengthened by the parenthesized annotations, recovering over two thirds (12/17) of the misclassifications by the distributional approach. The second row shows that more than two thirds (15.5/22.5) of the misclassifications by the pure string-based approach were correct using the distributional approach. The third row shows that more than two thirds (12.5/17.5) of the misclassifications by the string-based approach including annotations were correct using the distributional approach.</p>
         </sec>
         <sec>
            <st>
               <p>Error rates of the combined approaches</p>
            </st>
            <p>Table <tblr tid="T5">5</tblr> shows the error rates for linearly combined classifiers. The coefficient <it>w </it>in the table denotes the manually optimized weight for the distributional classifier, while 1-<it>w </it>was applied to weigh the string-based classifier complementarily. The lowest error rate of 0.055 was obtained by combining the distributional approach with &#8805; 10 syntactic dependencies and the string-based approach that included parenthesized annotations. The advantage of combining was apparent especially when the distributional classifier was built with more contextual features because the error rate decreased from 0.143 to 0.055 as the number of dependencies increased from &#8805; 1 to &#8805; 10. The increase in MRR values shows that the ranking function also benefited from the hybrid effect. Implementation details of the linear combination are described in the Methods section.</p>
            <tbl id="T5">
               <title>
                  <p>Table 5</p>
               </title>
               <caption>
                  <p>Error rates of the combined classifier</p>
               </caption>
               <tblbdy cols="7">
                  <r>
                     <c>
                        <p/>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>The entire test set </b>
                        </p>
                        <p>(<it>N </it>= 223)</p>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>&#8805; 1 syntactic dependencies </b>
                        </p>
                        <p>(<it>N </it>= 192)</p>
                     </c>
                     <c cspan="2" ca="center">
                        <p>
                           <b>&#8805; 10 syntactic dependencies </b>
                        </p>
                        <p>(<it>N </it>= 91)</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="7">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>
                           <b>Type of features</b>
                        </p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                     <c ca="center">
                        <p>Error rate</p>
                     </c>
                     <c ca="center">
                        <p>MRR</p>
                     </c>
                  </r>
                  <r>
                     <c cspan="7">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>Syntactic dependencies</p>
                        <p>+</p>
                        <p>Strings without annotations</p>
                     </c>
                     <c ca="center">
                        <p>0.191</p>
                        <p><it>w </it>= 0</p>
                     </c>
                     <c ca="center">
                        <p>0.881</p>
                     </c>
                     <c ca="center">
                        <p>0.167</p>
                        <p><it>w </it>= 0.2</p>
                     </c>
                     <c ca="center">
                        <p>0.900</p>
                     </c>
                     <c ca="center">
                        <p>0.110</p>
                        <p><it>w </it>= 0.3</p>
                     </c>
                     <c ca="center">
                        <p>0.935</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>Syntactic dependencies</p>
                        <p>+</p>
                        <p>Strings with annotations</p>
                     </c>
                     <c ca="center">
                        <p>0.143</p>
                        <p><it>w </it>= 0</p>
                     </c>
                     <c ca="center">
                        <p>0.907</p>
                     </c>
                     <c ca="center">
                        <p>0.143</p>
                        <p><it>w </it>= 0.3</p>
                     </c>
                     <c ca="center">
                        <p>0.912</p>
                     </c>
                     <c ca="center">
                        <p>0.055</p>
                        <p><it>w </it>= 0.3</p>
                     </c>
                     <c ca="center">
                        <p>0.969</p>
                     </c>
                  </r>
               </tblbdy>
               <tblfn>
                  <p>* Note that in the first column we set <it>w </it>= 0 because the distributional approach is not applicable.</p>
               </tblfn>
            </tbl>
         </sec>
      </sec>
      <sec>
         <st>
            <p>Discussion</p>
         </st>
         <p>In this paper we classify ontological concepts for purposes of curation or customization for specific types of applications, whereas the related work <abbrgrp><abbr bid="B36">36</abbr><abbr bid="B37">37</abbr></abbrgrp> dealt with classification of terms occurring in text. One implication was that they could exploit positional cues for lexical features (e.g. head nouns usually reside around the rightmost) but we were prohibited by the prevalent permutation of ontological strings (e.g. "transcription, RNA-dependent") which rarely appear in real text. Therefore, we assumed a simpler order-insensitive, independent bag-of-words model. In Torii et al. <abbrgrp><abbr bid="B37">37</abbr></abbrgrp>, informative suffixes were also included as lexical features. For contextual features, we incorporated syntactic relations, while the related work used only surrounding terms. Our classifiers were built using the resources publicly offered by the NLM, and our semantic classes included higher-level entities than the specific ones of molecular biology tested in the related work. Our eight semantic classes covered the major entities the bioinformatics community is interested in. For example, <it>gene or protein </it>and <it>biologic function </it>are apparently relevant, while cellular components are covered under <it>anatomy</it>, phenotypes are covered by <it>behavior </it>and <it>disorder</it>, and <it>substance </it>covers chemicals and drugs. Please ' [see Additional file <supplr sid="S1">1</supplr>]' for detailed contents of the eight classes.</p>
         <p>The results demonstrated that the performance was very good, especially when combining the two complementary approaches. The methods proposed in this paper can help validate classification of the CUIs used in training (e.g. the CUIs of the 64 SN types used in training our eight classes) and help reclassify CUIs in noisy but potentially useful SN types (e.g. T033 "Finding" and T169 "Functional Concepts"). New concepts added to an existing source vocabulary or that in a newly added source vocabulary can also be classified with the assistance of our methods. However, there are concepts of some SN types that are not specifically useful for biomedical applications (at least 30 of the 135), and we do not plan to reclassify them (e.g. "Urban Plannings" in T064 "Governmental or Regulatory Activity", "Democracy" in T078 "Idea or Concept", and "Algorithms" in T170 "Intellectual Product"). The methods are generalizable to any ontology wherever strings of the concepts are available and when the concepts can be identified in a training corpus. The training classes can be formed based on an existing ontological structure or by any arbitrary selection criteria, depending on the intended application. In the following subsections we discuss issues related to our results and methods.</p>
         <sec>
            <st>
               <p>Analysis of the string-based approach</p>
            </st>
            <p>When using the string-based approaches, many <it>anatomy </it>concepts were misclassified as <it>disorder</it>, for example, C0016504 "Foot" and C0037004 "Shoulder". We observed that the words of these basic body parts also appear frequently in the strings of <it>disorder </it>concepts involving these body parts, e.g. "Mycetoma of foot" and "Congenital valgus deformity of foot". Similarly, two <it>anatomy </it>modifier adjectives "peritoneal" and "popliteal" were misclassified as <it>procedure </it>owing to the presence of numerous procedure terms like "Continuous ambulatory <ul>peritoneal</ul> dialysis" and "Incision of <ul>popliteal</ul> space". However, some concepts had overlapping classes, and could be considered correctly classified as both classes. For example, C0431085 "Unspecified tumor cell NOS" is semantically a <it>disorder </it>and an <it>anatomy </it>(as an abnormal type of cell). Several <it>substance </it>concepts were misclassified as <it>disorder</it>. We found the unique term "Macrolides" of C0282563 appear frequently in concepts such as "poisoning by <ul>macrolides</ul>" or "<ul>macrolides</ul> causing adverse effects in therapeutic use", which belong to the SN type "Injury or Poisoning" (selected in training our <it>disorder </it>class). As a result, the representative string of the concept prevails more in the <it>disorder </it>class rather than in the <it>substance </it>class. The same situation happened to C0034428 "Quinolones" and C0002073 "Alkylating agents", e.g. the latter has strings frequently involved in a type of therapy-related acute myeloid leukaemia (<it>disorder</it>).</p>
            <p>Overall, the string-based approach was an unexpected success, since it is based on a very simplified compositional view concerning meaning: concepts can be semantically characterized by their lexical components (e.g. "single simple ovarian cyst") regardless of the order and syntactic roles of the components. Possibly this is due to characteristics relating to the descriptive way in which biomedical concepts are named, which basically corresponds to noun phrases, in which case the syntactic roles within the noun phrases may not be as important as syntactic roles within complete sentences. For example, in "cellular defense response", the fact that "cellular" and "defense" are modifiers of "response" is similar to the fact that they are words comprising the term. The results demonstrated that the simple approach did not compromise performance much while it avoided execution of a more complex and difficult task involving syntactic analysis. In the error analysis, however, we observed a typical disturbance to the Na&#239;ve Bayesian model: features (words) that are not semantically representative of certain classes occur frequently in those classes because of compositionality. For example, many words of the <it>anatomy </it>class occur frequently in <it>disorder </it>or <it>procedure </it>concepts, as discussed earlier. This implied that for test CUIs with only simple strings (and usually consisting of single words), the conditional probabilities that had been overestimated in certain classes might adversely affect the classification. If a test CUI consists of only a few non-discriminative words (with close conditional probabilities over every class), the classification could also be determined more by the priors of the classes. We anticipate the problem would still exist when applied to a more specific ontology such as GO, in which prevalent compositional structures have been reported by Ogren et al <abbrgrp><abbr bid="B43">43</abbr></abbrgrp>. For example, many substance terms such as "abscisic acid" are embedded in both Biological Process (e.g. "abscisic acid mediated signaling") and Molecular Function (e.g. "abscisic acid binding activity") concepts, but the presence of "abscisic acid" does not help much in determining which of the classes the concept should be assigned. A possible solution would be to perform syntactic analysis in order to determine the head nouns (e.g. the "signaling" and "activity" above) and to assign them higher weights when building the string-based classifier. However, this is left for future work.</p>
            <p>The result that string-based models trained with parenthesized annotations performed the best implies that many of the annotations from the source vocabularies agree with the SN types, and therefore the annotations provided straightforward support in making the correct classification. For example, by the augmented string-based approach "Foot" and "Shoulder" were correctly classified owing to the class annotation <it>body structure </it>from SNOMED-CT and <it>anatomy </it>from the Psychological Index Terms. Although these annotations are useful for obtaining the best reclassification, they could be based on manual curation and thus are not completely objective and are prone to human inconsistencies. However, if the manual re-assignment process is independent of the source classification, the parenthesized annotations could be thought as the surrogates of certain high-level intensional (i.e. to philosophically define a concept by stating its properties) meanings from different source vocabularies and in the string-based approach they contributed a type of highly weighted feature. Along with the strings, the annotations (whenever available) were aggregated under each CUI to serve as an extensional (i.e. to philosophically define a concept by enumerating its instances) meaning of the concept, as interpreted by Campbell et al. <abbrgrp><abbr bid="B44">44</abbr></abbrgrp>. Therefore, when the annotations of more than one or two source vocabularies consistently favor a semantic class incompatible to the SN type, we would doubt the appropriateness of the SN type assignment. To summarize, although the annotations are useful, further study is required to determine how objective they are.</p>
         </sec>
         <sec>
            <st>
               <p>Issues in comparing and combining the two approaches</p>
            </st>
            <p>In Table <tblr tid="T4">4</tblr> we have shown how the two approaches are quantitatively complementary. Although the string-based approach seemed generally more robust, the distributional approach with sufficient features could make a decisive contribution to the combined model. For example, our results showed that about 41% (91/223) of the test CUIs could be classified with an error rate of 0.055 (see Table <tblr tid="T5">5</tblr>). If the training corpus for the distributional approach were made larger, it is likely that the error rate would be even lower. The confusion matrices also revealed qualitative complementariness of the two approaches in terms of their relative strength on different classes. For example, a cluster of <it>anatomy </it>concepts misclassified as <it>procedure </it>(with two examples discussed earlier) were only observed in the string-based approach. This sheds light on solutions of applying class-dependent combination for the two approaches in the future, and we will need larger testing data to plot richer and more stable confusion matrices for analysis.</p>
            <p>Since the distributional approach uses contextual features and different similarity metrics, it is less susceptible to the weakness of the Na&#239;ve Bayesian model. An interesting example was C0066928 "Mullerian-inhibiting hormone", assigned to <it>substance </it>in the gold standard. The string-based approaches misclassified it as <it>biologic function</it>, probably because it contained the substring "inhibiting" while the distributional approach correctly classified it as <it>gene or protein </it>(top 1) and <it>substance </it>(top 2). We verified that it is a member of the transforming growth factor-beta gene family (e.g. in <it>Homo sapiens</it>, Entrez GeneID: 268) by searching the NCBI database. This also demonstrates that the approach could capture a more specific sense to augment the existing SN classification. However, we observed that the distributional approach incurred different types of misclassifications. The typical misclassifications by the distributional approach resulted from inadequate discriminating power of the syntactic dependencies. For example, the molecular function C0001038 "Acetylation" was misclassified as <it>gene or protein</it>, indicating that the function was not differentiated from the molecule itself through the contextual information.</p>
            <p>The string-based approach we used was based on CUIs and the strings associated with the CUIs. Many of the UMLS concepts are complex multi-word phrases, which are compositional. In such a phrase the meaning is based on a combination of the meanings of the parts constituting the phrase (e.g. "ribosomal protein S6 kinase") as a single concept. When using the distributional approach, however, the entire phrase representing a concept is treated as an atomic entity and therefore the modifiers within that phrase (e.g. "ribosomal") are not counted as syntactic dependencies. As a result, when using the distributional approach based on concepts many syntactic dependencies were lost that would have been obtained if we were using just words. In contrast, when using the string-based approach the loss of contextual features associated with multi-word concept terms are inadvertently compensated for because the words of the strings are utilized in building the classifier. For example, the word "ribosomal", which is characteristic of the <it>gene or protein </it>class, actually occurred 583 times P(ribosomal|<it>gene or protein</it>) = 0.00158 in the string-based profile but only 33 times P(<b>modifier_adjective</b>(ribosomal)|<it>gene or protein</it>) = 0.00018) as a modifier adjective in the distributional profile, demonstrating in this case, that the distributional information was more strongly captured by the string-based model than by the distributional model. This phenomenon suggests why the string-based approach performed better than the distributional approach because when informative adjuncts, such as "ribosomal", were absorbed into the UMLS concepts, the syntactic dependencies of some <it>gene or protein </it>concepts might have become indistinguishable from that of the more general <it>substance </it>class.</p>
            <p>There were also cases where the two approaches agreed with each other, but disagreed with the gold standard. It seemed that some of them represent borderline categories where the concepts could be semantically interpreted in more than one way. For example, C0079319 "Acoustic Evoked Brain Stem Potentials", which according to the UMLS definition means electrical waves generated in the brain when responding to auditory click stimuli. Therefore this concept refers to the functional condition of the brain stem as determined by a diagnostic procedure. It was classified by the distributional approach as <it>biologic function</it>, but the gold standard class was <it>procedure</it>. The string-based approach using parenthesize annotations agreed with the distributional approach, since the concept was annotated as <it>function </it>in SNOMED-CT. The examples discussed above also indicate that the classification accuracy might be underestimated a little by using the automatically generated gold standard.</p>
         </sec>
         <sec>
            <st>
               <p>Limitations</p>
            </st>
            <p>It is possible that the SN updates from the UMLS, which we used to automatically obtain the gold standard, are not completely objective for testing because they might represent more difficult cases that had been misclassified even by human curators in earlier releases. In addition, they may not be completely error-free. Although we believe that the semantic classes selected in this study have covered the major entities that the bioinformatics community is interested in, the constituent SN types included/excluded based on our own judgment according to recent reviews <abbrgrp><abbr bid="B6">6</abbr><abbr bid="B7">7</abbr><abbr bid="B8">8</abbr></abbrgrp> may not be free of bias. The performance of the distributional approach in this paper was bound to the randomly sampled training corpus, so that both the optimal size and representativeness could not be guaranteed.</p>
         </sec>
         <sec>
            <st>
               <p>Future work</p>
            </st>
            <p>We plan to 1) use syntactic analysis of the phrasal structures to improve the feature weighting of the string-based approach, 2) seek other gold standard sources that would offer greater randomness and testing size, 3) develop an automatic method to determine how to optimally combine the two approaches, 4) use the hybrid classifier to reclassify potentially useful concepts under general and vague SN types that were not used for training, e.g. T033 "Finding", 5) compare our results with other machine learning algorithms based on the same existing features, 6) propose a revised classification scheme for the UMLS, and 7) integrate the new classification into NLP applications.</p>
         </sec>
      </sec>
      <sec>
         <st>
            <p>Conclusion</p>
         </st>
         <p>We developed a string-based approach to classify UMLS concepts into eight broad semantic classes, and compared its performance with that of a method based on a distributional approach, which we previously developed. The performance of the string-based approach appeared to be comparable to the distributional approach, with an error rate of 0.143 and a mean reciprocal rank of 0.907. By comparing misclassifications associated with the distributional and string-based approaches, the two were found to be complementary. Classification performance was further enhanced after applying linear combinations of the two approaches, achieving the lowest error rate of 0.055 and a mean reciprocal rank of 0.969. Based on our results, we conclude that it is highly feasible to use the hybrid classifier to assist human experts in validating/reclassifying UMLS concepts.</p>
      </sec>
      <sec>
         <st>
            <p>Methods</p>
         </st>
         <sec>
            <st>
               <p>Determining the training classes and their CUIs</p>
            </st>
            <p>In the previous work, we formed seven broad semantic classes that we considered to be most significant for biomedical applications. In this paper, we added the <it>behavior </it>class in order to cover concepts related to behavioral phenotypes (e.g. C0242659 "Female homosexuality). We then selected 64 well-defined (out of a total of 135) SN types and mapped them into eight broad classes to make each of them semantically coherent. For example, the <it>biologic function </it>class included T038 "Biologic Function", T039 "Physiologic Function", T040 "Organism Function", T042 "Organ or Tissue Function", T043 "Cell Function", T044 "Molecular Function", and T045 "Genetic Function". Please ' [see Additional file <supplr sid="S1">1</supplr>]' for the complete list of SN types under each class. The CUIs belonging to each class were automatically identified, so that each broad class contained multiple SN types and included all their underlying CUIs (as illustrated by Figure <figr fid="F1">1</figr>). All the CUIs in each class were then used to train the classifier, as elaborated below.</p>
            <fig id="F1">
               <title>
                  <p>Figure 1</p>
               </title>
               <caption>
                  <p>Hierarchical illustration of the eight semantic classes (with some of the constituent SN types and CUIs omitted for conciseness)</p>
               </caption>
               <text>
                  <p>Hierarchical illustration of the eight semantic classes (with some of the constituent SN types and CUIs omitted for conciseness).</p>
               </text>
               <graphic file="1471-2105-8-264-1"/>
            </fig>
         </sec>
         <sec>
            <st>
               <p>Training the string-based Na&#239;ve Bayesian classifier</p>
            </st>
            <p>The Na&#239;ve Bayesian classifier was built based on words in the UMLS concept strings. Suppressible strings that have been obsolesced (flagged as 'O') by the source or by the NLM were excluded, e.g. string variations simply with an additional note <it>NOS </it>(Not Otherwise Specified) were excluded in this process. From the 2006AA MRCONSO.RRF table, we selected the unique English strings of each CUI. By unique strings, we mean the strings corresponding to the non-redundant String Unique Identifier (SUI) pertaining to each CUI. The selected strings were broken into words and transformed to lowercase. All punctuations were removed, but lexical variants (e.g. inflections or derivational suffixes) were not normalized for simplicity of implementation. Words consisting of only one character (e.g. "1" or "i") or those belonging to the PubMed stoplist <abbrgrp><abbr bid="B45">45</abbr></abbrgrp> were discarded; these words were considered non-informative for semantic classification and constituted about 0.05% of all the distinct words extracted. For example, from the string texts of Table <tblr tid="T6">6</tblr> we extracted the individual words and their frequencies as displayed in Table <tblr tid="T7">7</tblr>. Then via the string words from all the CUIs associated with each class, we calculated the prior probabilities of the classes and the conditional probabilities of the words given each class (for use in formula <b>3 </b>to train the classifier). Two versions of the classifier were trained, using strings either with or without parenthesized annotations. For words not seen in the trained models, their prior probabilities were used for smoothing.</p>
            <tbl id="T6">
               <title>
                  <p>Table 6</p>
               </title>
               <caption>
                  <p>Example strings of the concept C0079380</p>
               </caption>
               <tblbdy cols="2">
                  <r>
                     <c ca="left">
                        <p>
                           <b>String Unique Identifier</b>
                        </p>
                     </c>
                     <c ca="left">
                        <p>
                           <b>String text</b>
                        </p>
                     </c>
                  </r>
                  <r>
                     <c cspan="2">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>S6441909</p>
                     </c>
                     <c ca="left">
                        <p>Frameshift Mutation function</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>S0042786</p>
                     </c>
                     <c ca="left">
                        <p>Frameshift Mutation</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>S6148830</p>
                     </c>
                     <c ca="left">
                        <p>Reading Frame Shift Mutation</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>S6364685</p>
                     </c>
                     <c ca="left">
                        <p>Out-of-Frame Mutation</p>
                     </c>
                  </r>
               </tblbdy>
            </tbl>
            <tbl id="T7">
               <title>
                  <p>Table 7</p>
               </title>
               <caption>
                  <p>Words and frequencies obtained from the string texts of Table 6</p>
               </caption>
               <tblbdy cols="2">
                  <r>
                     <c ca="left">
                        <p>
                           <b>Word</b>
                        </p>
                     </c>
                     <c ca="left">
                        <p>
                           <b>Frequency</b>
                        </p>
                     </c>
                  </r>
                  <r>
                     <c cspan="2">
                        <hr/>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>frameshift</p>
                     </c>
                     <c ca="left">
                        <p>2</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>mutation</p>
                     </c>
                     <c ca="left">
                        <p>4</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>function</p>
                     </c>
                     <c ca="left">
                        <p>1</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>reading</p>
                     </c>
                     <c ca="left">
                        <p>1</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>frame</p>
                     </c>
                     <c ca="left">
                        <p>2</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>shift</p>
                     </c>
                     <c ca="left">
                        <p>1</p>
                     </c>
                  </r>
                  <r>
                     <c ca="left">
                        <p>out</p>
                     </c>
                     <c ca="left">
                        <p>1</p>
                     </c>
                  </r>
               </tblbdy>
            </tbl>
         </sec>
         <sec>
            <st>
               <p>Evaluating the two classification approaches</p>
            </st>
            <p>Based on our previous method of automatically generating test CUIs and the corresponding gold standard, we obtained 223 CUIs that belonged to the eight broad classes and were unambiguously identified from the training corpus (199,313 processed PubMed abstracts) used in our previous work. The 223 CUIs with gold standard classification served as the testing data and had been excluded from training the classifiers. We did not consider CUIs that never occurred in the corpus since there are many CUIs that would never be found in text, and our objective in semantic classification is primarily to improve text mining tasks. Using only the CUIs that occur in biomedical text is practical for these tasks, as McCray et al. <abbrgrp><abbr bid="B46">46</abbr></abbrgrp> have reported that only about 10% of the UMLS strings actually occur in MEDLINE. The string-based approach was evaluated on the entire test set of 223 CUIs. Then we varied the number of available contextual features (&#8805; 1 and &#8805; 10 syntactic dependencies) for the distributional approach and obtained two other subsets for testing (192 and 91 respectively out of the entire set of 223), and the two approaches were both evaluated on the two subsets for comparison. The gold standard was derived according to the 2006AA SN classification, and based on that the error rate (formula <b>4</b>) and mean reciprocal rank (formula <b>5</b>) were computed. We manually analyzed and summarized misclassifications by each approach on the test set (91 CUIs) where the distributional approach had &#8805; 10 features. Confusion matrices for the misclassifications were also generated. To evaluate whether the two approaches were complementary, a program was created to automatically compute the number of misclassifications by one approach that were correctly classified when using the other.</p>
         </sec>
         <sec>
            <st>
               <p>Linear combination of the two approaches</p>
            </st>
            <p>Before applying a linear combination, the raw similarity scores of each approach were rescaled and normalized into pseudo-probabilities by the following formulas respectively:</p>
            <p>
               <display-formula id="M6">[log(min(<it>X</it>))/log(<it>x</it>)] - 1 &#160;&#160;&#160; (for Nat&#239;ve Bayesian)</display-formula>
            </p>
            <p>
               <display-formula id="M7">[max(<it>X</it>)/<it>x</it>] - 1 &#160;&#160;&#160; (for &#945;-skew divergence)</display-formula>
            </p>
            <p>where <it>X </it>represents the similarity scores between the test CUI and all the eight classes and <it>x </it>is the similarity score of the specific class being considered. Then the following formula was applied to combine the adjusted scores of the two approaches:</p>
            <p>
               <display-formula id="M8"><it>w</it>&#183;SD + (1 - <it>w</it>)&#183;NB</display-formula>
            </p>
            <p>where SD and NB stand respectively for the adjusted scores of the distributional and the string-based approach, and <it>w </it>was varied from 0.1 to 0.9 with an interval of 0.1. The combined similarity scores for each CUI were computed so that a score was obtained for each of the eight semantic classes, and then the eight scores were sorted in descending order, so that the class with the highest combined score was taken the top prediction.</p>
         </sec>
      </sec>
      <sec>
         <st>
            <p>Authors' contributions</p>
         </st>
         <p>JF developed the programs, conducted the experiments, and prepared the manuscript. JF and HX both conducted the data analysis. The research was performed under the advice and supervision of CF. All authors read and approved the final manuscript.</p>
      </sec>
   </bdy>
   <bm>
      <ack>
         <sec>
            <st>
               <p>Acknowledgements</p>
            </st>
            <p>We would like to thank Dr. Alan Aronson, Guy Divita and James Mork from the NLM for help with using the UMLS-associated databases and tools. We also would like to thank Dr. George Hripcsak for performing the expert evaluation of the gold standard, Lyudmila Shagina for technical assistance, and Dr. Amy Chused for explaining some UMLS concepts in our analysis of misclassifications. This work was supported by Grants R01 LM7659 and R01 LM8635 from the National Library of Medicine.</p>
         </sec>
      </ack>
      <refgrp>
         <bibl id="B1">
            <title>
               <p>Gene ontology: tool for the unification of biology. The Gene Ontology Consortium</p>
            </title>
            <aug>
               <au>
                  <snm>Ashburner</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Ball</snm>
                  <fnm>CA</fnm>
               </au>
               <au>
                  <snm>Blake</snm>
                  <fnm>JA</fnm>
               </au>
               <au>
                  <snm>Botstein</snm>
                  <fnm>D</fnm>
               </au>
               <au>
                  <snm>Butler</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Cherry</snm>
                  <fnm>JM</fnm>
               </au>
               <au>
                  <snm>Davis</snm>
                  <fnm>AP</fnm>
               </au>
               <au>
                  <snm>Dolinski</snm>
                  <fnm>K</fnm>
               </au>
               <au>
                  <snm>Dwight</snm>
                  <fnm>SS</fnm>
               </au>
               <au>
                  <snm>Eppig</snm>
                  <fnm>JT</fnm>
               </au>
               <etal/>
            </aug>
            <source>Nat Genet</source>
            <pubdate>2000</pubdate>
            <volume>25</volume>
            <issue>1</issue>
            <fpage>25</fpage>
            <lpage>29</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/75556</pubid>
                  <pubid idtype="pmpid" link="fulltext">10802651</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B2">
            <title>
               <p>A reference ontology for biomedical informatics: the Foundational Model of Anatomy</p>
            </title>
            <aug>
               <au>
                  <snm>Rosse</snm>
                  <fnm>C</fnm>
               </au>
               <au>
                  <snm>Mejino</snm>
                  <fnm>JL</fnm>
                  <suf>Jr</suf>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2003</pubdate>
            <volume>36</volume>
            <issue>6</issue>
            <fpage>478</fpage>
            <lpage>500</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2003.11.007</pubid>
                  <pubid idtype="pmpid" link="fulltext">14759820</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B3">
            <title>
               <p>The Unified Medical Language System</p>
            </title>
            <aug>
               <au>
                  <snm>Lindberg</snm>
                  <fnm>DA</fnm>
               </au>
               <au>
                  <snm>Humphreys</snm>
                  <fnm>BL</fnm>
               </au>
               <au>
                  <snm>McCray</snm>
                  <fnm>AT</fnm>
               </au>
            </aug>
            <source>Methods Inf Med</source>
            <pubdate>1993</pubdate>
            <volume>32</volume>
            <issue>4</issue>
            <fpage>281</fpage>
            <lpage>291</lpage>
            <xrefbib>
               <pubid idtype="pmpid">8412823</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B4">
            <title>
               <p>The Unified Medical Language System: toward a collaborative approach for solving terminologic problems</p>
            </title>
            <aug>
               <au>
                  <snm>Campbell</snm>
                  <fnm>KE</fnm>
               </au>
               <au>
                  <snm>Oliver</snm>
                  <fnm>DE</fnm>
               </au>
               <au>
                  <snm>Shortliffe</snm>
                  <fnm>EH</fnm>
               </au>
            </aug>
            <source>J Am Med Inform Assoc</source>
            <pubdate>1998</pubdate>
            <volume>5</volume>
            <issue>1</issue>
            <fpage>12</fpage>
            <lpage>16</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">61272</pubid>
                  <pubid idtype="pmpid" link="fulltext">9452982</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B5">
            <title>
               <p>Methods in biomedical ontology</p>
            </title>
            <aug>
               <au>
                  <snm>Yu</snm>
                  <fnm>AC</fnm>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2006</pubdate>
            <volume>39</volume>
            <issue>3</issue>
            <fpage>252</fpage>
            <lpage>266</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2005.11.006</pubid>
                  <pubid idtype="pmpid" link="fulltext">16387553</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B6">
            <title>
               <p>A survey of current work in biomedical text mining</p>
            </title>
            <aug>
               <au>
                  <snm>Cohen</snm>
                  <fnm>AM</fnm>
               </au>
               <au>
                  <snm>Hersh</snm>
                  <fnm>WR</fnm>
               </au>
            </aug>
            <source>Brief Bioinform</source>
            <pubdate>2005</pubdate>
            <volume>6</volume>
            <issue>1</issue>
            <fpage>57</fpage>
            <lpage>71</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1093/bib/6.1.57</pubid>
                  <pubid idtype="pmpid" link="fulltext">15826357</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B7">
            <title>
               <p>Knowledge discovery in biology and biotechnology texts: a review of techniques, evaluation strategies, and applications</p>
            </title>
            <aug>
               <au>
                  <snm>Natarajan</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Berrar</snm>
                  <fnm>D</fnm>
               </au>
               <au>
                  <snm>Hack</snm>
                  <fnm>CJ</fnm>
               </au>
               <au>
                  <snm>Dubitzky</snm>
                  <fnm>W</fnm>
               </au>
            </aug>
            <source>Crit Rev Biotechnol</source>
            <pubdate>2005</pubdate>
            <volume>25</volume>
            <issue>1&#8211;2</issue>
            <fpage>31</fpage>
            <lpage>52</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1080/07388550590935571</pubid>
                  <pubid idtype="pmpid">15999851</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B8">
            <title>
               <p>Literature mining for the biologist: from information retrieval to biological discovery</p>
            </title>
            <aug>
               <au>
                  <snm>Jensen</snm>
                  <fnm>LJ</fnm>
               </au>
               <au>
                  <snm>Saric</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Bork</snm>
                  <fnm>P</fnm>
               </au>
            </aug>
            <source>Nat Rev Genet</source>
            <pubdate>2006</pubdate>
            <volume>7</volume>
            <issue>2</issue>
            <fpage>119</fpage>
            <lpage>129</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/nrg1768</pubid>
                  <pubid idtype="pmpid" link="fulltext">16418747</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B9">
            <title>
               <p>Text mining and ontologies in biomedicine: making sense of raw text</p>
            </title>
            <aug>
               <au>
                  <snm>Spasic</snm>
                  <fnm>I</fnm>
               </au>
               <au>
                  <snm>Ananiadou</snm>
                  <fnm>S</fnm>
               </au>
               <au>
                  <snm>McNaught</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Kumar</snm>
                  <fnm>A</fnm>
               </au>
            </aug>
            <source>Brief Bioinform</source>
            <pubdate>2005</pubdate>
            <volume>6</volume>
            <issue>3</issue>
            <fpage>239</fpage>
            <lpage>251</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1093/bib/6.3.239</pubid>
                  <pubid idtype="pmpid" link="fulltext">16212772</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B10">
            <title>
               <p>The Unified Medical Language System (UMLS): integrating biomedical terminology</p>
            </title>
            <aug>
               <au>
                  <snm>Bodenreider</snm>
                  <fnm>O</fnm>
               </au>
            </aug>
            <source>Nucleic Acids Res</source>
            <pubdate>2004</pubdate>
            <issue>32 Database</issue>
            <fpage>D267</fpage>
            <lpage>270</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">308795</pubid>
                  <pubid idtype="pmpid" link="fulltext">14681409</pubid>
                  <pubid idtype="doi">10.1093/nar/gkh061</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B11">
            <title>
               <p>The NCBI Entrez Taxonomy</p>
            </title>
            <url>http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=Taxonomy</url>
         </bibl>
         <bibl id="B12">
            <title>
               <p>Online Mendelian Inheritance in Man, OMIM (TM)</p>
            </title>
            <url>http://www.ncbi.nlm.nih.gov/sites/entrez?db=OMIM</url>
         </bibl>
         <bibl id="B13">
            <title>
               <p>Semantic classification of biomedical concepts using distributional similarity</p>
            </title>
            <aug>
               <au>
                  <snm>Fan</snm>
                  <fnm>JW</fnm>
               </au>
               <au>
                  <snm>Friedman</snm>
                  <fnm>C</fnm>
               </au>
            </aug>
            <source> J Am Med Inform Assoc</source>
            <pubdate>2007</pubdate>
            <volume>14</volume>
            <issue>4</issue>
            <fpage>467</fpage>
            <lpage>77</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmpid" link="fulltext">17460124</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B14">
            <title>
               <p>A theory of language and information: a mathematical approach</p>
            </title>
            <aug>
               <au>
                  <snm>Harris</snm>
                  <fnm>ZS</fnm>
               </au>
            </aug>
            <publisher>Oxford [England], New York: Clarendon Press; Oxford University Press</publisher>
            <pubdate>1991</pubdate>
         </bibl>
         <bibl id="B15">
            <title>
               <p>An upper level ontology for the biomedical domain</p>
            </title>
            <aug>
               <au>
                  <snm>McCray</snm>
                  <fnm>AT</fnm>
               </au>
            </aug>
            <source>Comp Funct Genom</source>
            <pubdate>2003</pubdate>
            <volume>4</volume>
            <fpage>80</fpage>
            <lpage>84</lpage>
            <xrefbib>
               <pubid idtype="doi">10.1002/cfg.255</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B16">
            <title>
               <p>Semantic relations asserting the etiology of genetic diseases</p>
            </title>
            <aug>
               <au>
                  <snm>Rindflesch</snm>
                  <fnm>TC</fnm>
               </au>
               <au>
                  <snm>Libbus</snm>
                  <fnm>B</fnm>
               </au>
               <au>
                  <snm>Hristovski</snm>
                  <fnm>D</fnm>
               </au>
               <au>
                  <snm>Aronson</snm>
                  <fnm>AR</fnm>
               </au>
               <au>
                  <snm>Kilicoglu</snm>
                  <fnm>H</fnm>
               </au>
            </aug>
            <source>AMIA Annu Symp Proc</source>
            <pubdate>2003</pubdate>
            <fpage>554</fpage>
            <lpage>558</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">1480275</pubid>
                  <pubid idtype="pmpid">14728234</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B17">
            <title>
               <p>The interaction of domain knowledge and linguistic structure in natural language processing: interpreting hypernymic propositions in biomedical text</p>
            </title>
            <aug>
               <au>
                  <snm>Rindflesch</snm>
                  <fnm>TC</fnm>
               </au>
               <au>
                  <snm>Fiszman</snm>
                  <fnm>M</fnm>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2003</pubdate>
            <volume>36</volume>
            <issue>6</issue>
            <fpage>462</fpage>
            <lpage>477</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2003.11.003</pubid>
                  <pubid idtype="pmpid" link="fulltext">14759819</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B18">
            <title>
               <p>Consistency across the hierarchies of the UMLS Semantic Network and Metathesaurus</p>
            </title>
            <aug>
               <au>
                  <snm>Cimino</snm>
                  <fnm>JJ</fnm>
               </au>
               <au>
                  <snm>Min</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Perl</snm>
                  <fnm>Y</fnm>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2003</pubdate>
            <volume>36</volume>
            <issue>6</issue>
            <fpage>450</fpage>
            <lpage>461</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2003.11.001</pubid>
                  <pubid idtype="pmpid" link="fulltext">14759818</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B19">
            <title>
               <p>Sharing knowledge in medicine: semantic and ontologic facets of medical concepts</p>
            </title>
            <aug>
               <au>
                  <snm>Burgun</snm>
                  <fnm>A</fnm>
               </au>
               <au>
                  <snm>Botti</snm>
                  <fnm>G</fnm>
               </au>
               <au>
                  <snm>Fieschi</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Le Beux</snm>
                  <fnm>P</fnm>
               </au>
            </aug>
            <source>IEEE Proc Syst Mans Cybern</source>
            <pubdate>1999</pubdate>
            <fpage>300</fpage>
            <lpage>305</lpage>
         </bibl>
         <bibl id="B20">
            <title>
               <p>Aggregating UMLS semantic types for reducing conceptual complexity</p>
            </title>
            <aug>
               <au>
                  <snm>McCray</snm>
                  <fnm>AT</fnm>
               </au>
               <au>
                  <snm>Burgun</snm>
                  <fnm>A</fnm>
               </au>
               <au>
                  <snm>Bodenreider</snm>
                  <fnm>O</fnm>
               </au>
            </aug>
            <source>Medinfo</source>
            <pubdate>2001</pubdate>
            <fpage>216</fpage>
            <lpage>220</lpage>
            <xrefbib>
               <pubid idtype="pmpid">11604736</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B21">
            <title>
               <p>Partitioning the UMLS semantic network</p>
            </title>
            <aug>
               <au>
                  <snm>Chen</snm>
                  <fnm>Z</fnm>
               </au>
               <au>
                  <snm>Perl</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Halper</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Geller</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Gu</snm>
                  <fnm>H</fnm>
               </au>
            </aug>
            <source>IEEE Trans Inf Technol Biomed</source>
            <pubdate>2002</pubdate>
            <volume>6</volume>
            <issue>2</issue>
            <fpage>102</fpage>
            <lpage>108</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1109/TITB.2002.1006296</pubid>
                  <pubid idtype="pmpid">12075663</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B22">
            <title>
               <p>A lexical metaschema for the UMLS semantic network</p>
            </title>
            <aug>
               <au>
                  <snm>Zhang</snm>
                  <fnm>L</fnm>
               </au>
               <au>
                  <snm>Perl</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Halper</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Geller</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Hripcsak</snm>
                  <fnm>G</fnm>
               </au>
            </aug>
            <source>Artif Intell Med</source>
            <pubdate>2005</pubdate>
            <volume>33</volume>
            <issue>1</issue>
            <fpage>41</fpage>
            <lpage>59</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.artmed.2004.06.002</pubid>
                  <pubid idtype="pmpid" link="fulltext">15617981</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B23">
            <title>
               <p>Auditing concept categorizations in the UMLS</p>
            </title>
            <aug>
               <au>
                  <snm>Gu</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Perl</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Elhanan</snm>
                  <fnm>G</fnm>
               </au>
               <au>
                  <snm>Min</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Zhang</snm>
                  <fnm>L</fnm>
               </au>
               <au>
                  <snm>Peng</snm>
                  <fnm>Y</fnm>
               </au>
            </aug>
            <source>Artif Intell Med</source>
            <pubdate>2004</pubdate>
            <volume>31</volume>
            <issue>1</issue>
            <fpage>29</fpage>
            <lpage>44</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.artmed.2004.02.002</pubid>
                  <pubid idtype="pmpid" link="fulltext">15182845</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B24">
            <title>
               <p>Auditing the Unified Medical Language System with semantic methods</p>
            </title>
            <aug>
               <au>
                  <snm>Cimino</snm>
                  <fnm>JJ</fnm>
               </au>
            </aug>
            <source>J Am Med Inform Assoc</source>
            <pubdate>1998</pubdate>
            <volume>5</volume>
            <issue>1</issue>
            <fpage>41</fpage>
            <lpage>51</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">61274</pubid>
                  <pubid idtype="pmpid" link="fulltext">9452984</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B25">
            <title>
               <p>Auditing the UMLS for redundant classifications</p>
            </title>
            <aug>
               <au>
                  <snm>Peng</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Halper</snm>
                  <fnm>MH</fnm>
               </au>
               <au>
                  <snm>Perl</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Geller</snm>
                  <fnm>J</fnm>
               </au>
            </aug>
            <source>Proc AMIA Symp</source>
            <pubdate>2002</pubdate>
            <fpage>612</fpage>
            <lpage>616</lpage>
            <xrefbib>
               <pubid idtype="pmpid">12463896</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B26">
            <title>
               <p>Mathematical structures of language</p>
            </title>
            <aug>
               <au>
                  <snm>Harris</snm>
                  <fnm>ZS</fnm>
               </au>
            </aug>
            <publisher>New York: Interscience Publishers</publisher>
            <pubdate>1968</pubdate>
         </bibl>
         <bibl id="B27">
            <title>
               <p>Towards large-scale, open-domain and ontology-based named entity classification</p>
            </title>
            <aug>
               <au>
                  <snm>Cimiano</snm>
                  <fnm>P</fnm>
               </au>
               <au>
                  <snm>V&#246;lker</snm>
                  <fnm>J</fnm>
               </au>
            </aug>
            <source>Proc Intl Conf Recent Adv Nat Lang Process</source>
            <pubdate>2005</pubdate>
            <fpage>166</fpage>
            <lpage>172</lpage>
         </bibl>
         <bibl id="B28">
            <title>
               <p>Do we need linguistics when we have statistics? A comparative analysis of the contributions of linguistic cues to a statistical word grouping system</p>
            </title>
            <aug>
               <au>
                  <snm>Hatzivassiloglou</snm>
                  <fnm>V</fnm>
               </au>
            </aug>
            <source>The balancing act: combining symbolic and statistical approaches to language</source>
            <publisher>Cambridge (MA): MIT Press</publisher>
            <editor>Klavans JL, Resnik P</editor>
            <pubdate>1996</pubdate>
            <fpage>67</fpage>
            <lpage>94</lpage>
         </bibl>
         <bibl id="B29">
            <title>
               <p>Measures of distributional similarity</p>
            </title>
            <aug>
               <au>
                  <snm>Lee</snm>
                  <fnm>L</fnm>
               </au>
            </aug>
            <source>Proc Annu Meet Assoc Comput Linguist</source>
            <pubdate>1999</pubdate>
            <fpage>25</fpage>
            <lpage>32</lpage>
         </bibl>
         <bibl id="B30">
            <title>
               <p>Syntactically-informed semantic category recognition in discharge summaries</p>
            </title>
            <aug>
               <au>
                  <snm>Sibanda</snm>
                  <fnm>T</fnm>
               </au>
               <au>
                  <snm>He</snm>
                  <fnm>T</fnm>
               </au>
               <au>
                  <snm>Szolovits</snm>
                  <fnm>P</fnm>
               </au>
               <au>
                  <snm>Uzuner</snm>
                  <fnm>O</fnm>
               </au>
            </aug>
            <source>AMIA Annu Symp Proc</source>
            <pubdate>2006</pubdate>
            <fpage>714</fpage>
            <lpage>718</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">1839398</pubid>
                  <pubid idtype="pmpid">17238434</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B31">
            <title>
               <p>Using distributional similarity to organize biomedical terminology</p>
            </title>
            <aug>
               <au>
                  <snm>Weeds</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Dowdall</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Schneider</snm>
                  <fnm>G</fnm>
               </au>
               <au>
                  <snm>Keller</snm>
                  <fnm>B</fnm>
               </au>
               <au>
                  <snm>Weir</snm>
                  <fnm>D</fnm>
               </au>
            </aug>
            <source>Terminology</source>
            <pubdate>2005</pubdate>
            <volume>11</volume>
            <issue>1</issue>
            <fpage>107</fpage>
            <lpage>141</lpage>
         </bibl>
         <bibl id="B32">
            <title>
               <p>GENIA corpus &#8211; semantically annotated corpus for bio-textmining</p>
            </title>
            <aug>
               <au>
                  <snm>Kim</snm>
                  <fnm>JD</fnm>
               </au>
               <au>
                  <snm>Ohta</snm>
                  <fnm>T</fnm>
               </au>
               <au>
                  <snm>Tateisi</snm>
                  <fnm>Y</fnm>
               </au>
               <au>
                  <snm>Tsujii</snm>
                  <fnm>J</fnm>
               </au>
            </aug>
            <source>Bioinformatics</source>
            <pubdate>2003</pubdate>
            <volume>19</volume>
            <issue>Suppl 1</issue>
            <fpage>i180</fpage>
            <lpage>182</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1093/bioinformatics/btg1023</pubid>
                  <pubid idtype="pmpid" link="fulltext">12855455</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B33">
            <title>
               <p>Managing content with automatic document classification</p>
            </title>
            <aug>
               <au>
                  <snm>Calvo</snm>
                  <fnm>RA</fnm>
               </au>
               <au>
                  <snm>Lee</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Li</snm>
                  <fnm>X</fnm>
               </au>
            </aug>
            <source>J Digit Inf</source>
            <pubdate>2004</pubdate>
            <volume>5</volume>
            <issue>2</issue>
            <fpage>282</fpage>
         </bibl>
         <bibl id="B34">
            <title>
               <p>Authorship attribution with support vector machines</p>
            </title>
            <aug>
               <au>
                  <snm>Diederich</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Kindermann</snm>
                  <fnm>O</fnm>
               </au>
               <au>
                  <snm>Leopold</snm>
                  <fnm>E</fnm>
               </au>
               <au>
                  <snm>Paass</snm>
                  <fnm>G</fnm>
               </au>
            </aug>
            <source>APPL INTELL</source>
            <pubdate>2003</pubdate>
            <volume>19</volume>
            <issue>1&#8211;2</issue>
            <fpage>109</fpage>
            <lpage>123</lpage>
            <xrefbib>
               <pubid idtype="doi">10.1023/A:1023824908771</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B35">
            <title>
               <p>Text categorization</p>
            </title>
            <aug>
               <au>
                  <snm>Sebastiani</snm>
                  <fnm>F</fnm>
               </au>
            </aug>
            <source>Text mining and its applications</source>
            <publisher>Southampton, UK: WIT Press</publisher>
            <editor>Zanasi A</editor>
            <pubdate>2005</pubdate>
            <fpage>109</fpage>
            <lpage>129</lpage>
         </bibl>
         <bibl id="B36">
            <title>
               <p>Biomedical named entity recognition using two-phase model based on SVMs</p>
            </title>
            <aug>
               <au>
                  <snm>Lee</snm>
                  <fnm>KJ</fnm>
               </au>
               <au>
                  <snm>Hwang</snm>
                  <fnm>YS</fnm>
               </au>
               <au>
                  <snm>Kim</snm>
                  <fnm>S</fnm>
               </au>
               <au>
                  <snm>Rim</snm>
                  <fnm>HC</fnm>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2004</pubdate>
            <volume>37</volume>
            <issue>6</issue>
            <fpage>436</fpage>
            <lpage>447</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2004.08.012</pubid>
                  <pubid idtype="pmpid" link="fulltext">15542017</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B37">
            <title>
               <p>Using name-internal and contextual features to classify biological terms</p>
            </title>
            <aug>
               <au>
                  <snm>Torii</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Kamboj</snm>
                  <fnm>S</fnm>
               </au>
               <au>
                  <snm>Vijay-Shanker</snm>
                  <fnm>K</fnm>
               </au>
            </aug>
            <source>J Biomed Inform</source>
            <pubdate>2004</pubdate>
            <volume>37</volume>
            <issue>6</issue>
            <fpage>498</fpage>
            <lpage>511</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1016/j.jbi.2004.08.007</pubid>
                  <pubid idtype="pmpid" link="fulltext">15542022</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B38">
            <title>
               <p>Machine Learning</p>
            </title>
            <aug>
               <au>
                  <snm>Mitchell</snm>
                  <fnm>TM</fnm>
               </au>
            </aug>
            <publisher>New York: McGraw-Hill</publisher>
            <pubdate>1997</pubdate>
         </bibl>
         <bibl id="B39">
            <title>
               <p>Genew: the Human Gene Nomenclature Database, 2004 updates</p>
            </title>
            <aug>
               <au>
                  <snm>Wain</snm>
                  <fnm>HM</fnm>
               </au>
               <au>
                  <snm>Lush</snm>
                  <fnm>MJ</fnm>
               </au>
               <au>
                  <snm>Ducluzeau</snm>
                  <fnm>F</fnm>
               </au>
               <au>
                  <snm>Khodiyar</snm>
                  <fnm>VK</fnm>
               </au>
               <au>
                  <snm>Povey</snm>
                  <fnm>S</fnm>
               </au>
            </aug>
            <source>Nucleic Acids Res</source>
            <pubdate>2004</pubdate>
            <issue>32 Database</issue>
            <fpage>D255</fpage>
            <lpage>257</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">308806</pubid>
                  <pubid idtype="pmpid" link="fulltext">14681406</pubid>
                  <pubid idtype="doi">10.1093/nar/gkh072</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B40">
            <title>
               <p>Contrast and variability in gene names</p>
            </title>
            <aug>
               <au>
                  <snm>Cohen</snm>
                  <fnm>KB</fnm>
               </au>
               <au>
                  <snm>Acquaah-Mensah</snm>
                  <fnm>GK</fnm>
               </au>
               <au>
                  <snm>Dolbey</snm>
                  <fnm>AE</fnm>
               </au>
               <au>
                  <snm>Hunter</snm>
                  <fnm>L</fnm>
               </au>
            </aug>
            <source>Proceedings of Workshop on NLP in the Biomedical Domain, ACL 2002; Philadelphia</source>
            <fpage>14</fpage>
            <lpage>20</lpage>
         </bibl>
         <bibl id="B41">
            <title>
               <p>SNOMED: SNOMED CT</p>
            </title>
            <url>http://www.ihtsdo.org/our-standards/snomed-ct/</url>
         </bibl>
         <bibl id="B42">
            <title>
               <p>TREC Genomics Track &#8211; Roadmap</p>
            </title>
            <url>http://ir.ohsu.edu/genomics/roadmap.html</url>
         </bibl>
         <bibl id="B43">
            <title>
               <p>The compositional structure of Gene Ontology terms</p>
            </title>
            <aug>
               <au>
                  <snm>Ogren</snm>
                  <fnm>PV</fnm>
               </au>
               <au>
                  <snm>Cohen</snm>
                  <fnm>KB</fnm>
               </au>
               <au>
                  <snm>Acquaah-Mensah</snm>
                  <fnm>GK</fnm>
               </au>
               <au>
                  <snm>Eberlein</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Hunter</snm>
                  <fnm>L</fnm>
               </au>
            </aug>
            <source>Pac Symp Biocomput</source>
            <pubdate>2004</pubdate>
            <fpage>214</fpage>
            <lpage>225</lpage>
            <xrefbib>
               <pubid idtype="pmpid">14992505</pubid>
            </xrefbib>
         </bibl>
         <bibl id="B44">
            <title>
               <p>Representing thoughts, words, and things in the UMLS</p>
            </title>
            <aug>
               <au>
                  <snm>Campbell</snm>
                  <fnm>KE</fnm>
               </au>
               <au>
                  <snm>Oliver</snm>
                  <fnm>DE</fnm>
               </au>
               <au>
                  <snm>Spackman</snm>
                  <fnm>KA</fnm>
               </au>
               <au>
                  <snm>Shortliffe</snm>
                  <fnm>EH</fnm>
               </au>
            </aug>
            <source>J Am Med Inform Assoc</source>
            <pubdate>1998</pubdate>
            <volume>5</volume>
            <issue>5</issue>
            <fpage>421</fpage>
            <lpage>431</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">61323</pubid>
                  <pubid idtype="pmpid" link="fulltext">9760390</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B45">
            <title>
               <p>PubMed stopwords</p>
            </title>
            <url>http://www.ncbi.nlm.nih.gov/books/bv.fcgi?rid=helppubmed.table.pubmedhelp.T43</url>
         </bibl>
         <bibl id="B46">
            <title>
               <p>Evaluating UMLS strings for natural language processing</p>
            </title>
            <aug>
               <au>
                  <snm>McCray</snm>
                  <fnm>AT</fnm>
               </au>
               <au>
                  <snm>Bodenreider</snm>
                  <fnm>O</fnm>
               </au>
               <au>
                  <snm>Malley</snm>
                  <fnm>JD</fnm>
               </au>
               <au>
                  <snm>Browne</snm>
                  <fnm>AC</fnm>
               </au>
            </aug>
            <source>Proc AMIA Symp</source>
            <pubdate>2001</pubdate>
            <fpage>448</fpage>
            <lpage>452</lpage>
            <xrefbib>
               <pubid idtype="pmpid">11825228</pubid>
            </xrefbib>
         </bibl>
      </refgrp>
   </bm>
</art>
