计算机科学
杠杆(统计)
分割
边距(机器学习)
水准点(测量)
人工智能
性格(数学)
冲程(发动机)
编码(集合论)
机器学习
代表(政治)
自然语言处理
模式识别(心理学)
几何学
政治
大地测量学
集合(抽象数据类型)
地理
程序设计语言
政治学
机械工程
数学
工程类
法学
作者
Lizhao Liu,Kunyang Lin,Shangxin Huang,Zhongli Li,Chao Li,Yunbo Cao,Qingyu Zhou
标识
DOI:10.48550/arxiv.2210.13826
摘要
Stroke is the basic element of Chinese character and stroke extraction has been an important and long-standing endeavor. Existing stroke extraction methods are often handcrafted and highly depend on domain expertise due to the limited training data. Moreover, there are no standardized benchmarks to provide a fair comparison between different stroke extraction methods, which, we believe, is a major impediment to the development of Chinese character stroke understanding and related tasks. In this work, we present the first public available Chinese Character Stroke Extraction (CCSE) benchmark, with two new large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW). With the large-scale datasets, we hope to leverage the representation power of deep models such as CNNs to solve the stroke extraction task, which, however, remains an open question. To this end, we turn the stroke extraction problem into a stroke instance segmentation problem. Using the proposed datasets to train a stroke instance segmentation model, we surpass previous methods by a large margin. Moreover, the models trained with the proposed datasets benefit the downstream font generation and handwritten aesthetic assessment tasks. We hope these benchmark results can facilitate further research. The source code and datasets are publicly available at: https://github.com/lizhaoliu-Lec/CCSE.
科研通智能强力驱动
Strongly Powered by AbleSci AI