a direct way to specify the worker spec
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- python
- Lĩnh vực
- distributed-systems, hpc
Hướng nghiên cứu
Không có tệp, bài kiểm thử hoặc entry point nào được nêu tên. Hãy bắt đầu bằng cách xem xét các API tài nguyên jobqueue và cấu hình worker hiện có, sau đó so sánh cách các giới hạn của scheduler được cung cấp cho các queue được hỗ trợ. Công việc được xem là hoàn tất khi có một API được định nghĩa có thể tiếp nhận đặc tả worker mong muốn, phân phối worker giữa các job và thất bại sớm khi đặc tả được yêu cầu không thể đáp ứng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
I'm frequently confused by the way the API requires me to specify resources for the jobqueue: I need to specify the job size and the number of workers per job (there's quite a few more knobs, of course), and it would evenly distribute the resources to each worker. I can then choose how many jobs to submit.
However, as a user (with admittedly a limited knowledge of how HPC work, so what I'm describing may be naive), my view is usually something like:
- my local jobqueue allows individual jobs to request up to 115GB of memory, 28 cores and a certain walltime, and for some queues there's also minimum resource limits
- I want to have about 14 workers, with about 15 GB and 2 threads each, where the concrete worker specs are often somewhat arbitrary and depend on my knowledge of the problem I'm trying to compute
This usually leads to me trying to group the workers manually to optimally fit the resource limits (so I don't get de-prioritized by submitting too many jobs).
Instead, I ideally would like an API allows me to specify (or retrieve) the resource limits per job of the jobqueue and the desired worker spec. It would then try to optimally distribute the workers and submit the jobs for me (and fail early if the resource limits don't allow the worker spec I requested).
Does something like this exist already? If not, would you be open to adding something like that? Is there anything I'm missing that would inhibit something like this?
- Ngôn ngữ chính
- Python
- Star
- 256
- Fork
- 149
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của dask/dask-jobqueue
-
bug LSF
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
dask/dask-jobqueue#703 · 1 bình luận ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
dask/dask-jobqueue#701 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 45/100
dask/dask-jobqueue#699 · 2 bình luận ·
-
bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 38/100
dask/dask-jobqueue#692 · 1 bình luận ·
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
dask/dask-jobqueue#691 · 7 bình luận ·
Tất cả issue của dask/dask-jobqueue
Issue tương tự
-
ACK_WAITING HELP_WANTED UPDATE_CS
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
OWASP/CheatSheetSeries#2458 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 90/100
BasedHardware/omi#19711 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Qwen3_5MoeModel no longer returns router_logits, breaking aux loss with output_router_logits=TrueĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
huggingface/transformers#49172 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
vllm-project/vllm-metal#885 ·
Maintainer thường phản hồi trong vòng 1 ngày