Tagasi

ISO 24614-2:2011

Language resource management -- Word segmentation of written texts -- Part 2: Word segmentation for Chinese, Japanese and Korean

Üldinfo
Kehtiv alates 25.08.2011
Direktiivid või määrused
puuduvad

Standardi ajalugu

Staatus
Kuupäev
Tüüp
Nimetus
25.08.2011
Põhitekst
The basic concepts and general principles of word segmentation as defined in ISO 24614-1 apply to Chinese, Japanese and Korean. Text needs to be segmented into tokens, words, phrases or some other types of smaller textual units in order to perform certain computational applications on language resources, such as natural language processing, information retrieval and machine translation. ISO 24614-2:2011 is restricted to the segmentation of a text into words or other word segmentation units (WSUs). This task is distinct from morphological or syntactic analysis per se, although it greatly depends on morphosyntactic analysis. It is also different from the task of laying out a framework for constructing a lexicon and identifying its lexical entries, namely lemmas and lexemes. The frameworks for the latter tasks are provided by ISO 24611, ISO 24613 and ISO 24615.
ISO 24614-2:2011 specifies rules for delineating WSUs for Chinese, Japanese and Korean. Some rules are common to all three languages, though each language also has its own distinct rules for identifying WSUs. The common features are discussed, then the distinct rules are laid out for Chinese, for Japanese and for Korean.
*
*
*
PDF
226,82 € koos KM-ga
Paber
226,82 € koos KM-ga
Standardi monitooring

Teised on ostnud veel

Põhitekst

ISO 24615-1:2014

Language resource management -- Syntactic annotation framework (SynAF) -- Part 1: Syntactic model
Uusim versioon Kehtiv alates 05.02.2014
Põhitekst

ISO/TS 24617-5:2014

Language resource management -- Semantic annotation framework (SemAF) -- Part 5: Discourse structure (SemAF-DS)
Uusim versioon Kehtiv alates 05.03.2014
Põhitekst

ISO 24610-1:2006

Language resource management -- Feature structures -- Part 1: Feature structure representation
Uusim versioon Kehtiv alates 10.04.2006
Põhitekst

ISO 24612:2012

Language resource management -- Linguistic annotation framework (LAF)
Uusim versioon Kehtiv alates 15.06.2012