Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Add dataset: TexBiG

オープン
#84 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞

調査の方向性

この issue には Zenodo URL とデータセットのメタデータしか記載されておらず、対象ファイル、テスト、エントリーポイントは指定されていません。まずリポジトリを調査して、データセットカタログまたはコントリビューション形式を確認し、その後、TexBiG のレコードとライセンスの詳細を正しい情報源として使用してください。データセットがプロジェクトで想定されている形式で追加され、関連する検証に合格すれば完了です。

索引モデルが issue の本文から書いたものです。

説明

dataset
A URL for this dataset

https://zenodo.org/record/6885144

Dataset description

TexBiG (from the German Text-Bild-Gefüge, meaning Text-Image-Structure) is a document layout analysis dataset for historical documents in the late 19th and early 20th century. The dataset provides instance segmentation (bounding boxes and polygons/masks) annotations for 19 different classes with more then 52.000 instances. Annotations are manually annotated by experts and evaluated with Krippendorff's Alpha, for each document image are least two different annotators have labeled the document. Further details can be found in the Paper.

Dataset modality

Mixed

Dataset licence

Creative Commons Attribution 4.0 International

Other licence

No response

How can you access this data

As a download from a repository/website

size of dataset

10GB

Confirm the dataset has an open licence
  • To the best of my knowledge, this dataset is accessible via an open licence
Contact details for data custodian

No response

主要言語
言語のデータがありません
スター
91
フォーク
8
PR マージ指標
30日以内にマージされた PR はありません

環境構築

このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

bigscience-workshop/lam のほかの issue

bigscience-workshop/lam の issue をすべて見る

似ている issue

Data Engineering の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。