Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Proposal: New implementation of ECSCluster

Aperta
#330 3 commenti 5 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Funzionalità
Chiarezza
Da chiarire
Stato di attività
Ferma
Stack tecnologico
aws, python

Direzione di ricerca

Il punto di ingresso è l’implementazione ECSCluster esistente; leggila prima insieme alle issue correlate #313, #121 e #262 per mappare i limiti segnalati. La issue propone una nuova implementazione e API ECSCluster, quindi conferma l’ambito e il design prima di scrivere il codice; per considerare il lavoro completato sarà necessario risolvere i problemi elencati relativi al ciclo di vita dei task, al ridimensionamento, alla configurazione, alle autorizzazioni e alla pulizia.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

provider/aws/ecs

I use ECSCluster heavily, but I find it has lots of small implementation issues that make it hard to reliably run large Dask clusters. Here's a laundry list of some of the problems I've run into.

  • Task names are always derived from the cluster name. In a shared ECS cluster, task definitions overlap between multiple ECSCluster instances, making it impossible to tell what tasks belong to which Dask cluster.
  • API rate limits are not handled properly. Combined with the log parsing for addresses (relates to #313, #121), large clusters are hard to reliably instantiate because the worker IP addresses can't be found.
  • ECSCluster directly instantiates tasks without using a service. There's no good way to do placement strategies like binpack.
  • Exited tasks are not handled and rescheduled. When workers run on spot instances, the Dask cluster can gradually lose workers as spot instances come and go.
  • There's no way to configure capacity providers for tasks.
  • There's no way to configure different subnets, environment variables, etc. for schedulers vs workers.
  • There's no way to configure driver to something other than awslogs.
  • Scaling a cluster while a previous scale is still in progress sometimes fails.
  • Too many IAM permissions are required, even when using pre-existing ECS clusters and resources.
  • Deprovisioning of tasks for both workers and schedulers is not clean (relates to #262).
  • Closing the client and cluster objects results in dangling hooks.

Rather than trying to morph the existing ECSCluster class, would this project be open to a completely new implementation (ECSCluster2?). I anticipate API changes are required (i.e., the arguments to ECSCluster). I'm willing to tackle this myself.

Lingua principale
Python
Stelle
147
Fork
119
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di dask/dask-cloudprovider

Tutte le issue di dask/dask-cloudprovider

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.