sophonplus / chinesenlpcorpus Goto Github PK
View Code? Open in Web Editor NEW搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。
搜集、整理、发布 中文 自然语言处理 语料/数据集,与 有志之士 共同 促进 中文 自然语言处理 的 发展。
谢谢回复
豆瓣的影评现在只能最多查看500页的信息,请问是怎么做到爬取数万条影评信息的呢?
您好,我使用pd.read_csv()读取waimai_10k.csv数据,却发现始终存在编码问题,尝试了各种编码均不行;请问您使用了什么预处理手段达到您intro.ipynb那样没有编码错误的展示吗
simplifyweibo_4_moods百度网盘无法下载,请问有新的下载链接吗?
项目看来没人维护?
资源公布在https://github.com/pass-lin/misaka-writer/blob/main
其中100G中文数据里百分之九十以上都是小说
除了小说还有名人传记、百科等
能否提供faq未脱敏原始数据,感谢大神![email protected]
刚开始学推荐,想用yf_dianping的数据集来试下,但是没有使用思路;
本来想用来推荐每种品类的餐厅,但是好像没有分类;很难搞
本科毕业论文想引用豆瓣影评中的一些数据。想请问本数据可否引用,若引用参考文献中怎么写才好?谢谢(#^.^#)
百度網盤無法存取
online_shopping_10_cats这个,下载不了,请问怎么解决呀
跳转链接与前面那个是同一个,只有8000条数据
请问一下保险的数据集可以用来做其他用途吗?开源协议是什么?
I downloaded this dataset(ChineseNlpCorpus/datasets/weibo_senti_100k) to train a model for chinese sentiment analysis. Upon treating this dataset I observed that 100% of the posts contain emoticons. Here is the distribution of the top10 emoticons according to the positive and negative polarity:
1013 emoticons in total. They are: [('泪', 44489), ('哈哈', 40510), ('嘻嘻', 22370), ('抓狂', 17262), ('鼓掌', 15923), ('爱你', 12685), ('怒', 12011), ('衰', 10466), ('晕', 9440), ('偷笑', 8375)]
710 emoticons in the positive set. They are: [('哈哈', 35764), ('嘻嘻', 20115), ('鼓掌', 14836), ('爱你', 11349), ('偷笑', 5223), ('太开心', 3820), ('可爱', 3809), ('心', 2122), ('赞', 1991), ('给力', 1976)]
695 emoticons in the negative set. They are: [('泪', 43248), ('抓狂', 16643), ('怒', 11830), ('衰', 10202), ('晕', 9022), ('哈哈', 4746), ('偷笑', 3152), ('蜡烛', 2887), ('汗', 2456), ('嘻嘻', 2255)]
I trained a very simple model to classify and I obtained 98% of accuracy in 2 epochs. Therefore, the emoticons have a strong bias in the classification. It led me to conclude that this dataset is not manually annotated. Probably whoever annotated the dataset manually classified some frequent emoticons and use them to tag the posts. Just saying for anyone who want to gather this data, you'd probably like to clean the emoticons out of it to avoid bias.
Peace!
您好,看到您在对yf_amazon数据集进行介绍时,描述的是“数据来源:亚马逊”,但是在原数据集地方写的是“JD.com E-Commerce Data,Yongfeng Zhang 教授为 WWW 2015 会议论文而搜集的数据”,在具体的数据集评论中发现确实存在“亚马逊”等字眼,所以想跟您确认一下yf_amazon数据集的来源,谢谢!
标的不准确,感觉是基于情感词直接打的标签
Is there any dataset for multi-label semantic analysis?
Such as happiness, anger, etc.
为什么金融数据集里面的reply一项很多后面都带有省略号,这不是不全吗
A declarative, efficient, and flexible JavaScript library for building user interfaces.
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
An Open Source Machine Learning Framework for Everyone
The Web framework for perfectionists with deadlines.
A PHP framework for web artisans
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
Some thing interesting about web. New door for the world.
A server is a program made to process requests and deliver data to clients.
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
Some thing interesting about visualization, use data art
Some thing interesting about game, make everyone happy.
We are working to build community through open source technology. NB: members must have two-factor auth.
Open source projects and samples from Microsoft.
Google ❤️ Open Source for everyone.
Alibaba Open Source for everyone
Data-Driven Documents codes.
China tencent open source team.