Comments (4)
Of course. Thanks for the offer.
I'll do a code clean up after the Mercari competition closes this week. This includes fixing bugs, and removing differences to sklearn API, for example how .fit() and .transform() behave. It would be a good time to document the code at that point.
from wordbatch.
Sure thanks. I would also start working on this after the competition ends. Besides, Are you interested in adding BM25 module? i.e. similar to WordBag but the score of each word for each document is BM25 score, it seems better than simple TF-IDF.
from wordbatch.
We can do this, it won't be difficult to add. Average document lengths need to be tracked for the BM25 document length normalization, but that's all. BM25 IDF is applied the same as in TF-IDF, it's only the term+length normalization that works different.
Argument options for TF-IDF/BM25 configurations can get messy, so those need to be sorted out. One option would be refactoring the normalizations into a separate class, that could be used on any set of sparse matrix features.
from wordbatch.
That's great. By the way how can I contact you except through Github? I will be better if we can contact more instantly. My email is [email protected], and if you use Facebook or something else it will also be okay for me.
from wordbatch.
Related Issues (20)
- WordVec extractor failing due to decode error HOT 1
- cannot install on windows 8.1 HOT 4
- "Illegal operation" when importing wordbatch.extractors HOT 2
- Licensing for commercial use without open source? HOT 1
- Tried to pickle the fitted wordbatch model, but bumped into this Error: AttributeError: 'function' object has no attribute 'im_self' HOT 3
- Import FTRL fails HOT 1
- Error on trying to import FM_FTRL HOT 1
- predict() takes a very long time HOT 1
- from wordbatch.data_utils import * HOT 3
- IndexError: too many indices for array HOT 1
- Illegal instruction (core dumped) HOT 1
- TypeError: only size-1 arrays can be converted to Python scalars (Windows, Python 3.5) HOT 1
- Multiprocessing Hanging in Python 3.6+ HOT 7
- are this times normal? HOT 2
- AttributeError: Can't get attribute 'normalize_text' on <module '__main__'> HOT 1
- About Wordbatch HOT 2
- pip install wordbatch on macos---error: command 'gcc-7' failed with exit status 1
- 'tuple' object has no attribute 'transform' HOT 3
- cross validation and grid search HOT 3
- will it work for Windows ?
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from wordbatch.