sophonplus / chinesenlpcorpus Goto Github PK

View Code? Open in Web Editor NEW

5.5K 5.5K 1.4K 11.48 MB

搜集、整理、发布中文自然语言处理语料/数据集，与有志之士共同促进中文自然语言处理的发展。

Jupyter Notebook 100.00%

chinesenlpcorpus's People

Contributors

Stargazers

Watchers

Forkers

lymcurry xianglizuel simmoncn cyy0523xc coeasy infinityfuture chenghuige chenjun0210 xhappy obaby liangsheng yndu13 amarry garrett-cn jiaqi18 ninedaywang bingtel nonva charlottesean wushicanasl yorkchu1995 mingleili moolighty fendaq iamzn havesupper dream1202 lhmzll artist100 1780041410 micahxie 1jasonzhang tianyikenan amoliu ericxsun dream5788 yulinlin0828 shannonyu allensmile mocha-pudding drganghe qiwsir robingong little1tow richardsun-voyager rolinston4 limengmingx liujiabing derekgrant haishuofang yifeixian lionxu feiwofeifeixiaowo zombie366 zouxiaoyuonly gongqingyi-github awesome-archive hulalazz yyht huguanglong yiyinianhua qiansi ycsuperlife weiyunfei zpeng1989 fword wybert studylzpvit joe2hpimn wishchen lplping skywindy zenghang cuitxubin colionx chenbing-ml demonsong greengrass2015 songxianjin x-hacker haiyunj1234 ahappycutedog changaolin lyjsz mrrabbit0o0 thelisq yutaoxxx excelsimon baizhengbiao wangjing0128 wibruce ikuangye benjamesbabala liudan8 mowayao someone-xsl freshzy zdsicecoco liu4lin jiniaoxu

chinesenlpcorpus's Issues

你们百度云盘还能下载吗，我在大陆一直提示“无法获取分享文件”

谢谢回复

豆瓣评论限制问题

豆瓣的影评现在只能最多查看500页的信息，请问是怎么做到爬取数万条影评信息的呢？

waimai_10k.csv的编码问题

您好，我使用pd.read_csv()读取waimai_10k.csv数据，却发现始终存在编码问题，尝试了各种编码均不行；请问您使用了什么预处理手段达到您intro.ipynb那样没有编码错误的展示吗

simplifyweibo_4_moods百度网盘无法下载

simplifyweibo_4_moods百度网盘无法下载，请问有新的下载链接吗？

分享100G中文数据

资源公布在https://github.com/pass-lin/misaka-writer/blob/main
其中100G中文数据里百分之九十以上都是小说
除了小说还有名人传记、百科等

能否提供faq未脱敏原始数据，感谢大神！[email protected]

yf_dianping 有老哥有比较好的使用思路么

刚开始学推荐，想用yf_dianping的数据集来试下，但是没有使用思路；

本来想用来推荐每种品类的餐厅，但是好像没有分类；很难搞

论文中可否引用该数据？

本科毕业论文想引用豆瓣影评中的一些数据。想请问本数据可否引用，若引用参考文献中怎么写才好？谢谢(#^.^#)

could I use it in bert and how I should do the preprocessing for the data? are emoticons out of vocabulary?

Weibo senti 100k is very likely labelled by the emoticons

I downloaded this dataset(ChineseNlpCorpus/datasets/weibo_senti_100k) to train a model for chinese sentiment analysis. Upon treating this dataset I observed that 100% of the posts contain emoticons. Here is the distribution of the top10 emoticons according to the positive and negative polarity:

1013 emoticons in total. They are: [('泪', 44489), ('哈哈', 40510), ('嘻嘻', 22370), ('抓狂', 17262), ('鼓掌', 15923), ('爱你', 12685), ('怒', 12011), ('衰', 10466), ('晕', 9440), ('偷笑', 8375)]

710 emoticons in the positive set. They are: [('哈哈', 35764), ('嘻嘻', 20115), ('鼓掌', 14836), ('爱你', 11349), ('偷笑', 5223), ('太开心', 3820), ('可爱', 3809), ('心', 2122), ('赞', 1991), ('给力', 1976)]

695 emoticons in the negative set. They are: [('泪', 43248), ('抓狂', 16643), ('怒', 11830), ('衰', 10202), ('晕', 9022), ('哈哈', 4746), ('偷笑', 3152), ('蜡烛', 2887), ('汗', 2456), ('嘻嘻', 2255)]

I trained a very simple model to classify and I obtained 98% of accuracy in 2 epochs. Therefore, the emoticons have a strong bias in the classification. It led me to conclude that this dataset is not manually annotated. Probably whoever annotated the dataset manually classified some frequent emoticons and use them to tag the posts. Just saying for anyone who want to gather this data, you'd probably like to clean the emoticons out of it to avoid bias.

Peace!