Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Add dataset: clmet_3-1

未关闭
#58 5 条评论 0 个 reaction 已指派 2 人 在 GitHub 查看

@davanstrien 已经在做这个了。

开始于 2022年7月18日。

评估

这个 Issue 还没有评估数据。

描述

dataset ready for review
A URL for this dataset

http://fedora.clarin-d.uni-saarland.de/clmet/clmet.html

Dataset description

The Corpus of Late Modern English Texts, version 3.1 (CLMET3.1) is a principled collection of public domain texts drawn from various online archiving projects. In total, the corpus contains some 34 million words of running text. It incorporates CLMET, CLMETEV, and CLMET3.0, and has been compiled following roughly the same principles, that is:

The corpus covers the period 1710–1920, divided into three 70-year sub-periods.
The texts making up the corpus have all been written by British and Irish authors who are native speakers of English.
The corpus never contains more than three texts by the same author.
The texts within each sub-period have been written by authors born within a correspondingly restricted sub-period.

Size: 34 million words

Annotation: PoS-tagged; genre.

Dataset modality

Text

Dataset licence

Creative Commons Attribution Non Commercial Share Alike 4.0 International

Other licence

No response

How can you access this data

As a download from a repository/website

Confirm the dataset has an open licence
  • To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian

No response

主要语言
没有语言数据
星标
91
派生
8
PR 合并指标
30 天内没有已合并 PR

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

bigscience-workshop/lam 的其他 Issue

查看 bigscience-workshop/lam 的全部 Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。