Brainstorming replacing QA-Tiles
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 20/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- python
- Área
- computer-vision, data, machine-learning
Línea de trabajo
Esta es una propuesta de brainstorming, no una tarea de implementación, y no menciona archivos ni pruebas. Empieza revisando la arquitectura actual de label-maker y los puntos de entrada de los módulos Python; después, compáralos con las salidas propuestas de GeoJSON, imágenes de Mapbox, augmentation, caching y COCO. Se considerará Done cuando se hayan acordado el alcance y el diseño antes de que pueda comenzar la implementación.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
I need to rework my https://github.com/jremillard/images-to-osm project to use Mapbox tiles. The problem that label-maker is attempting to solve is right at the center of the planned rework. I just wanted to communicate what label-maker would look like if it was a perfect fit for my needs.
The input data (training ) to label-maker should be a set of geojson files. There is a rich and mature existing infrastructure of generating them from OSM and other data sources. They are easy to write code against in any language. Let other tools deal with it.
Label maker config would be
- output zoom level OR a metric output (.5 m/pixel).
- output image size for the training network (say 800x800), not an even tile boundary.
- data augmentation options (center object, randomly slide object around, up/down, left/right flips, % scale change, edge buffer zone, allow clipped features, etc).
- How many sample images to make.
- training/validation split %.
- Sat image TMS URL (someday support Bing when they can change the license).
- Max sat image cache size, directory, also need max ago of sat image cache (mapbox is 30 days).
- % of images to create that are negative samples (no objects in them).
The final output would be intermediate files (training, and validating), not the training images.
When the network is training, the intermediate files can be opened up, and single images can be generated on the fly from a python module. The python module would handle either fetching and forming the training images or getting them from the sat image cache. It would stitch the sat images together, crop them correctly, and output bounding boxes, segmentation masks, and instance masks. The one image at a time would allow data sets that don't fit into memory to be used, keep performance good, and not violate sat image caching licensing restrictions.
If you want to be really nice to people, have an option to write out MS COCO files, since basically everyone is using that data set right now for benchmarking.
- Lenguaje dominante
- Python
- Estrellas
- 472
- Forks
- 106
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de developmentseed/label-maker
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 35/100
developmentseed/label-maker#195 ·
-
Use logging or drop log optionAbierto
Dificultad 3/5 1-2 días Aptitud para principiantes 35/100
developmentseed/label-maker#186 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 35/100
developmentseed/label-maker#185 ·
-
Generalize Final data.npz outputQuizá libre de nuevo @martham93 la tomó hace 2223 días y no hay ningún pull request abierto. Abiertoenhancement
developmentseed/label-maker#181 · 1 comentario · 1 asignado ·
-
No support for Tensorflow 2Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 25/100
developmentseed/label-maker#180 · 1 comentario ·
Todos los issues de developmentseed/label-maker
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-2 días Aptitud para principiantes 70/100
-
FingerprintSplitter raises ZeroDivisionError when int(frac_train * len(dataset)) floors to zeroAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 7 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
lmstudio-ai/mlx-engine#376 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
pyiron/bagofholding#166 ·