Chinese Word Segmentation

July 21, 2021 · View on GitHub

Overview

Chinese is written using characters (hanzi), where each character represents a syllable. A word is usually taken to consist of one or more character tokens. There are no spaces between words. Less than 3500 distinct characters are normally encountered. Word segmentation (or tokenization) is the process of dividing up a sequence of characters into a sequence of words.

Example

Input:

亲 请问有什么可以帮您的吗?

Output:

亲 请问 有 什么 可以 帮 您 的 吗 ?

Metrics

Word F1 score:

Gold: 共同 创造 美好 的 新 世纪 —— 二○○一年 新年 贺词

Hypothesis: 共同 创造 美 好 的 新 世纪 —— 二○○一年 新年 贺词

Precision = 9 / 11 = 0.818

Recall = 9 / 10 = 0.9

F1 = 0.857

The Second International Chinese Word Segmentation Bakeoff in SIGHAN 2005 Workshop (Emerson, 2005).

CorpusAbbrev.EncodingTest Size (Tokens/Types)
Traditional Chinese
Academia Sinica(Taipei)ASUnicode/Big Five Plus122K / 19K
City University of Hong KongCityUHKSCS Unicode/Big Five104K / 13K
Simplified Chinese
Peking UniversityPKCP936/Unicode41K / 9K
Microsoft ResearchMSRACP936/Unicode107K / 13K

Results

ModelASCITYUMSRAPKU
Ke et al. (2021)97.098.298.596.9
Qiu, Pei, Yan, Huang (2020)96.496.998.196.4
Tian, Song, Xia, Zhang, Wang (2020)96.697.998.496.5
Meng et al. (2019)96.7*97.9*98.396.7
Huang et al. (2019)96.697.697.996.6
Ma et al. (2018)96.297.297.496.1
Yang et al. (2017)95.796.997.596.3
Zhou et al. (2017)97.896.0

* Unlike others, Meng et al. (2019) do not report converting traditional Chinese to simplified Chinese.

Resources

Train setTraining Size(Words)
AS5.45M
CityU1.46M
MSRA2.37M
PKU1.1M

Chinese Penn Treebank.

  • Website
  • Includes 3 datasets:
    • CTB6: consisting of 780,000 words (over 1.28 million Chinese characters)
    • CTB7: consists of 2,448 text files, 51,447 sentences, 1,196,329 words and 1,931,381 hanzi (Chinese characters)
    • CTB9: consists of 3,726 text files, 132,076 sentences, 2,084,387 words, 3,247,331 characters (hanzi or foreign)
Data setTest set (Tokens)
CTB682K
CTB7245K
CTB9242K

Results

ModelCTB6CTB7CTB9
Ke et al. (2021)97.9
Tian, Song, Ao, Xia, Quan, Zhang, Wang (2020)97.597.397.8
Tian, Song, Xia, Zhang, Wang (2020)97.3
Yan et al. (2020)97.197.6
Huang et al. (2019)97.6
Ma et al. (2018)96.796.6**
Yang et al. (2017)96.2
Zhou et al. (2017)96.2

** Ma et al. (2018) report different statistics for their CTB7 split (950K/60K/82K), so the results might not be comparable.

Resources

Train setTraining Size (Words)
CTB6641K
CTB7718K
CTB91,696K

Chinese Universal Treebank (UD).

Data setTest set(Tokens)
UD12,012

Results

ModelUD
Ke et al. (2021)98.6
Tian, Song, Ao, Xia, Quan, Zhang, Wang (2020)98.3
Huang et al. (2019)97.3
Ma et al. (2018)96.9

Resources

Train setTraining Size(Words)
UD98,608

NLPCC2016 WordSeg Weibo.

# Sentences# Words# Characters
Weibo8,592-315,857

Results

ModelWeibo
Yang et al. (2017)95.5

Resources

# Sentences# Words# Characters
Train20,135421,166688,734
Dev2,05243,69773,244

Other Resources


Suggestions? Changes? Please send email to chinesenlp.xyz@gmail.com