Add dataset: clmet_3-1
@davanstrien 已经在做这个了。
开始于 2022年7月18日。
评估
这个 Issue 还没有评估数据。
描述
A URL for this dataset
http://fedora.clarin-d.uni-saarland.de/clmet/clmet.html
Dataset description
The Corpus of Late Modern English Texts, version 3.1 (CLMET3.1) is a principled collection of public domain texts drawn from various online archiving projects. In total, the corpus contains some 34 million words of running text. It incorporates CLMET, CLMETEV, and CLMET3.0, and has been compiled following roughly the same principles, that is:
The corpus covers the period 1710–1920, divided into three 70-year sub-periods.
The texts making up the corpus have all been written by British and Irish authors who are native speakers of English.
The corpus never contains more than three texts by the same author.
The texts within each sub-period have been written by authors born within a correspondingly restricted sub-period.
Size: 34 million words
Annotation: PoS-tagged; genre.
Dataset modality
Text
Dataset licence
Creative Commons Attribution Non Commercial Share Alike 4.0 International
Other licence
No response
How can you access this data
As a download from a repository/website
Confirm the dataset has an open licence
- To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian
No response
- 主要语言
- 没有语言数据
- 星标
- 91
- 派生
- 8
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
bigscience-workshop/lam 的其他 Issue
-
dataset good first issue
难度 2/5 1-3 小时 新手友好度 68/100
bigscience-workshop/lam#86 · 1 条评论 ·
-
dataset
难度 2/5 1-3 小时 新手友好度 68/100
bigscience-workshop/lam#65 · 1 条评论 ·
-
candidate-dataset
难度 2/5 1-3 小时 新手友好度 35/100
bigscience-workshop/lam#92 · 1 条评论 ·
-
dataset
难度 2/5 1-3 小时 新手友好度 35/100
bigscience-workshop/lam#87 ·
-
dataset
难度 4/5 3-5 天 新手友好度 42/100
bigscience-workshop/lam#85 · 1 个 reaction ·