Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Add dataset: chronicling_america

未关闭
#85 0 条评论 1 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
42/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
停滞

调研方向

首先定位数据集加载脚本,然后检查 https://chroniclingamerica.loc.gov/newspapers.json 上的 Chronicling America 报纸 API,以及其中链接的标题、期刊、OCR 和图像端点。查看加载所需的日期和其他筛选条件;当可以通过开放 API 添加数据集,并且其访问和许可详细信息已记录时,即表示完成。

由索引模型根据 Issue 内容生成。

描述

dataset
A URL for this dataset

https://chroniclingamerica.loc.gov/about/api/#bulk-data

Dataset description

Chronicling America is a Library of Congress project to digitise historic newspapers. The collection contains mostly English but also contains other languages. Breakdown by language: https://public.tableau.com/app/profile/chronicling.america#!/vizhome/ChroniclingAmericaLanguageCoverageBubble/All_Lang

Various ways of accessing this data include bulk downloads and an API. The API may be the most helpful way of accessing this dataset (via dataset loading script) because this dataset is not static (more titles are digitised and added on a rolling basis).

The 'newspapers' API (https://chroniclingamerica.loc.gov/newspapers.json) is probably the best starting point. This starts instead from a list of Newspaper titles for which digital content is held. A title, i.e. https://chroniclingamerica.loc.gov/lccn/sn86072192.json, contains a bunch of metadata.

Screenshot 2022-09-27 at 16 32 26.

This API also contains all the issues for that title. For each issue, you get a set of pages. Each page contains the plain text generated from the OCR for that page, e.g. https://chroniclingamerica.loc.gov/lccn/sn82014726/1888-04-07/ed-1/seq-1/ocr.txt and a link to the image of that page, e.g. https://chroniclingamerica.loc.gov/lccn/sn82014726/1888-04-07/ed-1/seq-1.jp2.

My suggested approach to loading this dataset would be to call https://chroniclingamerica.loc.gov/newspapers.json at the start of the script and, depending on some filters defined in the loading script, i.e. start/end date of interest, build up a list of relevant URLs for the text/images for each page.

If you want to work on this dataset, please cc @davanstrien and @albertvillanova!

Dataset modality

Mixed

Dataset licence

Other license

Other licence

https://chroniclingamerica.loc.gov/about/#rights

How can you access this data

Via an open API

size of dataset

10GB

Confirm the dataset has an open licence
  • To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian

No response

主要语言
没有语言数据
星标
91
派生
8
PR 合并指标
30 天内没有已合并 PR

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

bigscience-workshop/lam 的其他 Issue

查看 bigscience-workshop/lam 的全部 Issue

相似的 Issue

更多 Data Engineering Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。