Comments (2)
3-4 seconds for 300 words sounds too high.
Your script could be loading or copying in memory large objects with each call, which would cause this. If you use multiprocessing as the backend, each subprocess will copy all of the global variables, and variables referred in the passed data. If per-document transforms with "serial" as the method are fast, then this subprocess memory copying is likely the issue. If you want to use multiprocessing, then make sure your global variables don't include big objects, and you don't pass references to them.
from wordbatch.
thanks a lot!
At the end I just use the multiprocessing batch call to all the tf-idf transform that I need to do and the time for all the thousands of transformation is the same !!( so practically all the time is loading my model to the memory).
Thanks again for your amazing tool, greetings from Chile.
from wordbatch.
Related Issues (20)
- WordVec extractor failing due to decode error HOT 1
- cannot install on windows 8.1 HOT 4
- "Illegal operation" when importing wordbatch.extractors HOT 2
- Licensing for commercial use without open source? HOT 1
- Tried to pickle the fitted wordbatch model, but bumped into this Error: AttributeError: 'function' object has no attribute 'im_self' HOT 3
- Import FTRL fails HOT 1
- Error on trying to import FM_FTRL HOT 1
- predict() takes a very long time HOT 1
- from wordbatch.data_utils import * HOT 3
- IndexError: too many indices for array HOT 1
- Illegal instruction (core dumped) HOT 1
- TypeError: only size-1 arrays can be converted to Python scalars (Windows, Python 3.5) HOT 1
- Multiprocessing Hanging in Python 3.6+ HOT 7
- AttributeError: Can't get attribute 'normalize_text' on <module '__main__'> HOT 1
- About Wordbatch HOT 2
- pip install wordbatch on macos---error: command 'gcc-7' failed with exit status 1
- 'tuple' object has no attribute 'transform' HOT 3
- cross validation and grid search HOT 3
- will it work for Windows ?
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from wordbatch.