Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Brainstorming replacing QA-Tiles

オープン
#38 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
20/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
停滞
技術スタック
python

調査の方向性

これは実装タスクではなくブレインストーミングの提案であり、ファイルやテストは指定していません。まず現在の label-maker アーキテクチャと Python モジュールのエントリーポイントを確認し、次に提案されている GeoJSON、Mapbox imagery、augmentation、caching、COCO の出力と比較してください。実装を開始する前にスコープと設計について合意できれば、Done です。

索引モデルが issue の本文から書いたものです。

説明

I need to rework my https://github.com/jremillard/images-to-osm project to use Mapbox tiles. The problem that label-maker is attempting to solve is right at the center of the planned rework. I just wanted to communicate what label-maker would look like if it was a perfect fit for my needs.

The input data (training ) to label-maker should be a set of geojson files. There is a rich and mature existing infrastructure of generating them from OSM and other data sources. They are easy to write code against in any language. Let other tools deal with it.

Label maker config would be

  1. output zoom level OR a metric output (.5 m/pixel).
  2. output image size for the training network (say 800x800), not an even tile boundary.
  3. data augmentation options (center object, randomly slide object around, up/down, left/right flips, % scale change, edge buffer zone, allow clipped features, etc).
  4. How many sample images to make.
  5. training/validation split %.
  6. Sat image TMS URL (someday support Bing when they can change the license).
  7. Max sat image cache size, directory, also need max ago of sat image cache (mapbox is 30 days).
  8. % of images to create that are negative samples (no objects in them).

The final output would be intermediate files (training, and validating), not the training images.

When the network is training, the intermediate files can be opened up, and single images can be generated on the fly from a python module. The python module would handle either fetching and forming the training images or getting them from the sat image cache. It would stitch the sat images together, crop them correctly, and output bounding boxes, segmentation masks, and instance masks. The one image at a time would allow data sets that don't fit into memory to be used, keep performance good, and not violate sat image caching licensing restrictions.

If you want to be really nice to people, have an option to write out MS COCO files, since basically everyone is using that data set right now for benchmarking.

主要言語
Python
スター
472
フォーク
106
PR マージ指標
30日以内にマージされた PR はありません

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

developmentseed/label-maker のほかの issue

developmentseed/label-maker の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。