Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Add dataset: beyond_words_coco

オープン
#37 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞
領域
data

調査の方向性

まず、リンク先の beyond_words_data URL にあるデータセットとメタデータを確認します。アクセスとライセンスの詳細も含めて確認してください。次に、リポジトリのデータセット登録またはカタログのエントリーポイントを特定します。beyond_words_coco データセットが、その説明、モダリティ、アクセス URL、ライセンス情報、アノテーション数とともに追加されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

dataset
A URL for this dataset

https://github.com/LibraryOfCongress/newspaper-navigator/tree/master/beyond_words_data

Dataset description

This is a dataset containing crowdsourced annotastions of the bounding boxes of different types of 'visual content in historic US newspapers. It can be used to train object detection models to extract visual content from historic newspaper collections. Breakdown of annotations:

Category # in Full Dataset
Photograph 4,254
Illustration 1,048
Map 215
Comics/Cartoon 1,150
Editorial Cartoon 293
Headline 27,868
Advertisement 13,581
Total 48,409
Dataset modality

Image

Dataset licence

Other license

Other licence

https://chroniclingamerica.loc.gov/about/

How can you access this data

As a download from a repository/website

Confirm the dataset has an open licence
  • To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian

No response

主要言語
言語のデータがありません
スター
91
フォーク
8
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

bigscience-workshop/lam のほかの issue

bigscience-workshop/lam の issue をすべて見る

似ている issue

Data Engineering の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。