• ISO 24614-2:2011

    Current The latest, up-to-date edition.

    Language resource management — Word segmentation of written texts — Part 2: Word segmentation for Chinese, Japanese and Korean

    Available format(s):  Hardcopy, PDF, PDF 3 Users, PDF 5 Users, PDF 9 Users

    Language(s):  English

    Published date:  25-08-2011

    Publisher:  International Organization for Standardization

    Add To Cart

    Abstract - (Show below) - (Hide below)

    The basic concepts and general principles of word segmentation as defined in ISO 24614-1 apply to Chinese, Japanese and Korean. Text needs to be segmented into tokens, words, phrases or some other types of smaller textual units in order to perform certain computational applications on language resources, such as natural language processing, information retrieval and machine translation. ISO 24614-2:2011 is restricted to the segmentation of a text into words or other word segmentation units (WSUs). This task is distinct from morphological or syntactic analysis per se, although it greatly depends on morphosyntactic analysis. It is also different from the task of laying out a framework for constructing a lexicon and identifying its lexical entries, namely lemmas and lexemes. The frameworks for the latter tasks are provided by ISO 24611, ISO 24613 and ISO 24615.

    ISO 24614-2:2011 specifies rules for delineating WSUs for Chinese, Japanese and Korean. Some rules are common to all three languages, though each language also has its own distinct rules for identifying WSUs. The common features are discussed, then the distinct rules are laid out for Chinese, for Japanese and for Korean.

    General Product Information - (Show below) - (Hide below)

    Committee ISO/TC 37/SC 4
    Development Note Supersedes ISO/DIS 24614-2. (08/2011)
    Document Type Standard
    Publisher International Organization for Standardization
    Status Current

    Standards Referenced By This Book - (Show below) - (Hide below)

    14/30266984 DC : 0 BS ISO 7098 - INFORMATION AND DOCUMENTATION - ROMANIZATION OF CHINESE
    BS ISO 7098:2015 Information and documentation. Romanization of Chinese
    ISO 7098:2015 Information and documentation Romanization of Chinese

    Standards Referencing This Book - (Show below) - (Hide below)

    ISO 24613:2008 Language resource management - Lexical markup framework (LMF)
    ISO 24612:2012 Language resource management — Linguistic annotation framework (LAF)
    ISO 24611:2012 Language resource management — Morpho-syntactic annotation framework (MAF)
    ISO 24614-1:2010 Language resource management Word segmentation of written texts Part 1: Basic concepts and general principles
    ISO 24615:2010 Language resource management Syntactic annotation framework (SynAF)
    • Access your standards online with a subscription

      Features

      • Simple online access to standards, technical information and regulations
      • Critical updates of standards and customisable alerts and notifications
      • Multi - user online standards collection: secure, flexibile and cost effective