Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Documentation] Explain performance improvements

Open
#670 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
35/100
Issue type
Documentation
Clarity
Mostly clear
Activity status
Stale
Tech stack
numpy, pandas, python

Research direction

The issue does not name a documentation file; start by locating the existing SingleNodeExecutor usage documentation and compare it with the provided benchmark examples. Document the progression from repeated data loading to block allocation and init_function-based loading, including the stated performance results and usage patterns.

Written by the indexing model from the issue text.

Description

documentation

Generate data:

import numpy as np
import pandas as pd

N = 1_000_000
data = pd.DataFrame({
    "c": np.random.choice(["a", "b", "c"], size=N),
    "x": np.random.uniform(size=N),
    "y": np.random.normal(size=N)
})

data.to_csv("blob.csv")  # File is about 45 Mb

Slow execution: 24.1 s ± 115 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

import numpy as np
import pandas as pd
from executorlib import SingleNodeExecutor

def get_sum(i, df):
    return i, df["x"].sum(), df["y"].sum()

with SingleNodeExecutor(max_workers=10) as exe:
    future_lst = [exe.submit(get_sum, df=pd.read_csv("blob.csv"), i=i) for i in range(100)]
    result_lst = [f.result() for f in future_lst]

Reduce the startup time for the processes: 19.5 s ± 31.9 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

import numpy as np
import pandas as pd
from executorlib import SingleNodeExecutor

def get_sum(i, df):
    return i, df["x"].sum(), df["y"].sum()

with SingleNodeExecutor(max_workers=10, block_allocation=True) as exe:
    future_lst = [exe.submit(get_sum, df=pd.read_csv("blob.csv"), i=i) for i in range(100)]
    result_lst = [f.result() for f in future_lst]

Load the data only once for each process: 946 ms ± 24.3 ms per loop (mean ± std. dev. of 7 runs, 1 loop each)

import numpy as np
import pandas as pd
from executorlib import SingleNodeExecutor

def get_sum(i, df):
    return i, df["x"].sum(), df["y"].sum()

def init_funct():
    return {"df": pd.read_csv("blob.csv")}

with SingleNodeExecutor(max_workers=10, block_allocation=True, init_function=init_funct) as exe:
    future_lst = [exe.submit(get_sum, i=i) for i in range(100)]
    result_lst = [f.result() for f in future_lst]
Dominant language
Python
Stars
77
Forks
7
Avg merge
10h 32m
Merged PRs (30d)
12

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from pyiron/executorlib

All issues in pyiron/executorlib

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.