Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Consider using object pools for internal state keeping

Đang mở
#8,901 3 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
20/100
Loại issue
Tái cấu trúc
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Đình trệ
Công nghệ
python
Lĩnh vực
performance

Hướng nghiên cứu

Issue không nêu tệp nguồn, điểm vào hoặc các bài kiểm thử nào; hãy bắt đầu bằng cách lập hồ sơ việc cấp phát set trong scheduler và đo tác động của nó đến bộ nhớ và thời gian chạy. So sánh các phép đo đó với khái niệm SetPool/PooledSet được đề xuất, và chỉ xem công việc là hoàn tất khi lợi ích đo được cùng hành vi tái sử dụng an toàn đã được chứng minh trong các bài kiểm thử của dự án.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

discussion

Dask is creating many small objects and not just for very large graphs but this is pretty much a built in thing. This is one of not the most prominent reason why dask was originally built as a "tuple of tuple of tuple ..." machinery. Even the scheduler internally only adopted the usage of custom classes only a couple of years back because instantiation can be very costly. However, modern python versions have gotten much better at managing overhead so this is negligible for most classes.

The one type of objects we're still affected by, both in terms of memory but also in terms of runtime are sets. Yes, sets! (I might share profiles but this is mostly an issue to preseve an idea). One way to work around instantiation cost of sets is to use an object pool design. Effectively, we'd "disable" garbage collection and would resurrect objects on finalization in a way that would allow us to reuse them.

A minimal version of this would look like that

import sys

class SetPool:
    def __init__(self):
        self._sets = []

    def add(self, obj):
        obj.clear()
        self._sets.append(obj)

    def get(self):
        try:
            new = self._sets.pop()
            return new
        except:
            return None

    def stored_size(self):
        return sum(map(sys.getsizeof, self._sets))
    
globalpool = SetPool()

class PooledSet(set):

    def __new__(cls, *iterables):
        obj = globalpool.get()
        if obj is not None:
            return obj
        return super().__new__(cls, *iterables)
    
    def __del__(self):
        print(f"Resurrecting {id(self)}")
        globalpool.add(self)

image

This could be expanded on need, e.g. by hinting towards whether this would be an empty set, a very large set or a somewhat normal one.

Haven't tried out what the actual impact would be but it is a fun concept that could help if we actually want/need to optimize for this

Ngôn ngữ chính
Python
Star
1.7k
Fork
778
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của dask/distributed

Tất cả issue của dask/distributed

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.