error runing 03_glove_build_counts.py
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 32/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start with 03_glove_build_counts.py and the logged command that pipes data\S2vcorpusMODELV05\corpusMODELV05-1.s2v into vocab_count. Check how the script resolves data/glove.6B.200d.txt/vocab_count on Windows and compare that with the downloaded GloVe files. Done means the command completes and creates data\S2VVocabMODELV05\vocab.txt.
Written by the indexing model from the issue text.
Description
I followed your step to train my own S2V for my corpus on my customized NER model, thill step 2 everything is fine,.
corpusMODELV05.spacy is made and also corpusMODELV05-1.s2v
but in step 3 I faced with this error
ℹ Using 1 input files
✔ Created output directory data/S2VVocabMODELV05
ℹ Creating vocabulary counts
cat data\S2vcorpusMODELV05\corpusMODELV05-1.s2v | data/glove.6B.200d.txt/vocab_count -min-count 5 -verbose 2 > data\S2VVocabMODELV05\vocab.txt
✘ Failed creating vocab counts
I am working on Win 10 machine and have used this version of the glove
Wikipedia 2014 + Gigaword 5 (6B tokens, 400K vocab, uncased, 50d, 100d, 200d, & 300d vectors, 822 MB download): glove.6B.zip
https://nlp.stanford.edu/projects/glove/
it seems the number of VOC in
glove.6B.200d.txt/vocab_count is not in line with something
can someone help me ?
many thanks in advance
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 236
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from explosion/sense2vec
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
-
import errorOpenusage
Difficulty 4/5 3-5 days Newbie friendliness 20/100
-
usage
Difficulty 3/5 1-2 days Newbie friendliness 30/100
-
provide citationOpen
Difficulty 1/5 Under an hour Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
All issues in explosion/sense2vec
Similar issues
-
docs pydanty:is-working
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
pydantic/pydantic-ai#9800 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
-
[Bug]: With --api-server-count > 1, gauges such as vllm:num_requests_running have no samples until the first requestPossibly taken @roy6n23 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
vllm-project/vllm#59988 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
pymc-devs/pymc-examples#897 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 91/100