Add dataset: chronicling_america
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 42/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 停滞
调研方向
首先定位数据集加载脚本,然后检查 https://chroniclingamerica.loc.gov/newspapers.json 上的 Chronicling America 报纸 API,以及其中链接的标题、期刊、OCR 和图像端点。查看加载所需的日期和其他筛选条件;当可以通过开放 API 添加数据集,并且其访问和许可详细信息已记录时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
A URL for this dataset
https://chroniclingamerica.loc.gov/about/api/#bulk-data
Dataset description
Chronicling America is a Library of Congress project to digitise historic newspapers. The collection contains mostly English but also contains other languages. Breakdown by language: https://public.tableau.com/app/profile/chronicling.america#!/vizhome/ChroniclingAmericaLanguageCoverageBubble/All_Lang
Various ways of accessing this data include bulk downloads and an API. The API may be the most helpful way of accessing this dataset (via dataset loading script) because this dataset is not static (more titles are digitised and added on a rolling basis).
The 'newspapers' API (https://chroniclingamerica.loc.gov/newspapers.json) is probably the best starting point. This starts instead from a list of Newspaper titles for which digital content is held. A title, i.e. https://chroniclingamerica.loc.gov/lccn/sn86072192.json, contains a bunch of metadata.
.
This API also contains all the issues for that title. For each issue, you get a set of pages. Each page contains the plain text generated from the OCR for that page, e.g. https://chroniclingamerica.loc.gov/lccn/sn82014726/1888-04-07/ed-1/seq-1/ocr.txt and a link to the image of that page, e.g. https://chroniclingamerica.loc.gov/lccn/sn82014726/1888-04-07/ed-1/seq-1.jp2.
My suggested approach to loading this dataset would be to call https://chroniclingamerica.loc.gov/newspapers.json at the start of the script and, depending on some filters defined in the loading script, i.e. start/end date of interest, build up a list of relevant URLs for the text/images for each page.
If you want to work on this dataset, please cc @davanstrien and @albertvillanova!
Dataset modality
Mixed
Dataset licence
Other license
Other licence
https://chroniclingamerica.loc.gov/about/#rights
How can you access this data
Via an open API
size of dataset
10GB
Confirm the dataset has an open licence
- To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian
No response
- 主要语言
- 没有语言数据
- 星标
- 91
- 派生
- 8
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
bigscience-workshop/lam 的其他 Issue
-
dataset good first issue
难度 2/5 1-3 小时 新手友好度 68/100
bigscience-workshop/lam#86 · 1 条评论 ·
-
dataset
难度 2/5 1-3 小时 新手友好度 68/100
bigscience-workshop/lam#65 · 1 条评论 ·
-
candidate-dataset
难度 2/5 1-3 小时 新手友好度 35/100
bigscience-workshop/lam#92 · 1 条评论 ·
-
dataset
难度 2/5 1-3 小时 新手友好度 35/100
bigscience-workshop/lam#87 ·
-
dataset
难度 3/5 1-2 天 新手友好度 35/100
bigscience-workshop/lam#84 ·
查看 bigscience-workshop/lam 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
FireDynamics/fdsreader#123 ·
-
Missing: Aircall for Startups可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 1/5 1 小时以内 新手友好度 85/100
sourcey/startup-credits#1633 · 1 条评论 ·
维护者通常 10 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
Sienna-Platform/PowerSystemCaseBuilder.jl#239 ·
维护者通常 1 天内回复
-
area:resolver diagnostic:error diagnostic:unresolved-replacement reason:lesson-not-found schedule-diagnostic scope:actual shift:first
难度 1/5 1 小时以内 新手友好度 90/100
xTCry/ygk-schedule#262 ·
-
难度 1/5 1 小时以内 新手友好度 85/100
UKGovernmentBEIS/inspect_ai#5712 · 1 条评论 ·
维护者通常 2 天内回复