Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Brainstorming replacing QA-Tiles

Đang mở
#38 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
20/100
Loại issue
Tính năng
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
python

Hướng nghiên cứu

Đây là một đề xuất brainstorming thay vì một nhiệm vụ triển khai, và không nêu tên tệp hay bài kiểm thử nào. Hãy bắt đầu bằng việc xem xét kiến trúc label-maker hiện tại và các entry point của module Python, sau đó so sánh chúng với các đầu ra GeoJSON, Mapbox imagery, augmentation, caching và COCO được đề xuất. Done yêu cầu phạm vi và thiết kế được thống nhất trước khi có thể bắt đầu triển khai.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

I need to rework my https://github.com/jremillard/images-to-osm project to use Mapbox tiles. The problem that label-maker is attempting to solve is right at the center of the planned rework. I just wanted to communicate what label-maker would look like if it was a perfect fit for my needs.

The input data (training ) to label-maker should be a set of geojson files. There is a rich and mature existing infrastructure of generating them from OSM and other data sources. They are easy to write code against in any language. Let other tools deal with it.

Label maker config would be

  1. output zoom level OR a metric output (.5 m/pixel).
  2. output image size for the training network (say 800x800), not an even tile boundary.
  3. data augmentation options (center object, randomly slide object around, up/down, left/right flips, % scale change, edge buffer zone, allow clipped features, etc).
  4. How many sample images to make.
  5. training/validation split %.
  6. Sat image TMS URL (someday support Bing when they can change the license).
  7. Max sat image cache size, directory, also need max ago of sat image cache (mapbox is 30 days).
  8. % of images to create that are negative samples (no objects in them).

The final output would be intermediate files (training, and validating), not the training images.

When the network is training, the intermediate files can be opened up, and single images can be generated on the fly from a python module. The python module would handle either fetching and forming the training images or getting them from the sat image cache. It would stitch the sat images together, crop them correctly, and output bounding boxes, segmentation masks, and instance masks. The one image at a time would allow data sets that don't fit into memory to be used, keep performance good, and not violate sat image caching licensing restrictions.

If you want to be really nice to people, have an option to write out MS COCO files, since basically everyone is using that data set right now for benchmarking.

Ngôn ngữ chính
Python
Star
472
Fork
106
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của developmentseed/label-maker

Tất cả issue của developmentseed/label-maker

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.