Command Palette
Search for a command to run...
LCCC Large Clean Chinese Conversational Corpus
Date
Size
Publish URL
Paper URL
LCCC (Large-scale Cleaned Chinese Conversation corpus) was released by Tsinghua University and Samsung China Research Institute in 2020. The dataset mainly consists of two parts: LCCC-base (6.8 million dialogues) and LCCC-large (12 million dialogues). The research team designed a strict data filtering process to ensure the quality of the dialogue data in the dataset. The process is based on a set of rules and a classifier trained on 110K manually annotated dialogue pairs. The noise filtered by the research team includes: dirty words, special characters, emoticons, grammatically incorrect sentences, and irrelevant dialogues in the context. The cleaned dataset and pre-trained model will promote the research of short text dialogue modeling.
Citation
Please kindly cite our paper if you use the datasets or models in your research: @inproceedings{wang2020chinese, title={A Large-Scale Chinese Short-Text Conversation Dataset}, author={Wang, Yida and Ke, Pei and Zheng, Yinhe and Huang, Kaili and Jiang, Yong and Zhu, Xiaoyan and Huang, Minlie} booktitle={NLPCC}, year={2020}, url={https://arxiv.org/abs/2008.03946} }
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.