Issue 1
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- aws, jupyter-notebook, machine-learning
- Domain
- cloud, data-engineering, machine-learning
Research direction
Start by reviewing the linked labs issue and the listed Phase 1 milestones. The payload identifies no repository files, tests, or entry points. Completion is not defined beyond planning pipelines, generating embeddings, storing them as Zarr on Source Cooperative, and providing downstream examples.
Written by the indexing model from the issue text.
Description
Goals
-
Help increase ease of access to various embeddings in cloud native formats
-
Provide an example of what cloud native easily accessible, interoperable embeddings from Geospatial Foundation Models could look like by storing embeddings derived from 3 different models on Source Cooperative for 8 years (2017-2024) for all a sample geographic area (ie all of Kenya)
-
Make it easier to use embeddings in downstream modeling classes
Partner and Stakeholders
Taylor Geospatial Engine, Fields of the World, Open Source Developers: Issac Corley, Jeff Albrecht, Source Cooperative
Team
Dev Seed Gaia + other labs contributors
Milestones
Phase 1 November - December 2025
-
Simple + scalable engineering AWS pipeline for S2 and S1 download needed for embedding generation
-
simple + scalable engineering AWS pipeline for embeddings generations for models we want to run inference with
-
Experiments to guide initial decisions around: zarr file chunking, projections
-
Generate (when necessary) embeddings, store as Zarr - determine chunking scheme host on source cooperative
Phase 2 (Application Based) December/Jan 2025
- tools + examples to facilitate comparison for downstream use cases (ie field boundary delineation) with the embeddings that are stored on source cooperative
Future Work:
-
Enable easier visualization + visual comparison of different model's embeddings
-
Connect Embeddings to datasets/stac collections that they are derived from
-
Going beyond country level scale - do our initial decisions hold out at areas larger than a country
Project Organization
- slack channel tbd
- project board - coming soon 🚧
Documents
https://github.com/developmentseed/labs/issues/458
cc @zacdezgeo
- Dominant language
- Jupyter Notebook
- Stars
- 32
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from developmentseed/pixelverse
-
documentation
Difficulty 2/5 1-2 days Newbie friendliness 68/100
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
developmentseed/pixelverse#44 · 2 comments ·
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Use geo-zarr toolkit Open
developmentseed/pixelverse#33 · 1 assignee ·
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
developmentseed/pixelverse#32 · 1 comment ·
All issues in developmentseed/pixelverse
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
level/task module/gcp type/bug
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
-
bug needs-triage service/elbv2
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
hashicorp/terraform-provider-aws#50100 · 1 comment ·