Add dataset: clmet_3-1
@davanstrien ci sta già lavorando.
Dal 18/7/2022.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
A URL for this dataset
http://fedora.clarin-d.uni-saarland.de/clmet/clmet.html
Dataset description
The Corpus of Late Modern English Texts, version 3.1 (CLMET3.1) is a principled collection of public domain texts drawn from various online archiving projects. In total, the corpus contains some 34 million words of running text. It incorporates CLMET, CLMETEV, and CLMET3.0, and has been compiled following roughly the same principles, that is:
The corpus covers the period 1710–1920, divided into three 70-year sub-periods.
The texts making up the corpus have all been written by British and Irish authors who are native speakers of English.
The corpus never contains more than three texts by the same author.
The texts within each sub-period have been written by authors born within a correspondingly restricted sub-period.
Size: 34 million words
Annotation: PoS-tagged; genre.
Dataset modality
Text
Dataset licence
Creative Commons Attribution Non Commercial Share Alike 4.0 International
Other licence
No response
How can you access this data
As a download from a repository/website
Confirm the dataset has an open licence
- To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian
No response
- Lingua principale
- Nessun dato sulla lingua
- Stelle
- 91
- Fork
- 8
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di bigscience-workshop/lam
-
dataset good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
bigscience-workshop/lam#86 · 1 commento ·
-
dataset
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
bigscience-workshop/lam#65 · 1 commento ·
-
candidate-dataset
Difficoltà 2/5 1-3 ore Idoneità per principianti 35/100
bigscience-workshop/lam#92 · 1 commento ·
-
dataset
Difficoltà 2/5 1-3 ore Idoneità per principianti 35/100
bigscience-workshop/lam#87 ·
-
dataset
Difficoltà 4/5 3-5 giorni Idoneità per principianti 42/100
bigscience-workshop/lam#85 · 1 reazione ·