Possible reference leak of the argument tuple in `FunctionCall()`
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 76/100
Research direction
Start in src/map.cpp at FunctionCall and review the CPython ownership rules for PyTuple_Pack and PyObject_CallObject. Verify the argument tuple is released on both successful and failed calls, while df_obj retains its existing ownership handling. Done means repeated DuckDBPyRelation.map() calls no longer leak the tuple or retain the input DataFrame.
Written by the indexing model from the issue text.
Description
What happens?
FunctionCall() passes a newly created tuple directly to PyObject_CallObject():
File: src/map.cpp
Function: FunctionCall
auto *df_obj = PyObject_CallObject(function, PyTuple_Pack(1, in_df.ptr()));
PyTuple_Pack() returns a new reference, while PyObject_CallObject() does
not steal its args reference. Because the tuple is not stored in a local
variable, it is never passed to Py_DECREF().
As a result, every invocation leaks one tuple. The tuple also owns a reference
to in_df, so the input pandas DataFrame remains alive after FunctionCall()
returns. This occurs on both successful and failed calls.
The function is used during bind-time schema inference and query execution, so
the leak is reachable through ordinary DuckDBPyRelation.map() operations.
The handling of df_obj is unrelated and correct:
auto df = py::reinterpret_steal<py::object>(df_obj);
PyObject_CallObject() returns a new reference on success, which
reinterpret_steal() adopts.
To Reproduce
This issue can be confirmed directly from the reference ownership in
src/map.cpp.
In FunctionCall(), the argument tuple is created inline:
auto *df_obj = PyObject_CallObject(function, PyTuple_Pack(1, in_df.ptr()));
According to the CPython C API reference ownership rules:
PyTuple_Pack()returns a new reference.PyObject_CallObject()does not steal the reference passed asargs.- The tuple pointer is not stored, so there is no subsequent
Py_DECREF()for that new reference. - The tuple therefore leaks on every call and retains its reference to
in_df.
This issue is specific to the Python API and is not reproducible through plain
SQL in the DuckDB CLI.
OS:
x86_64
DuckDB Package Version:
latest version
Python Version:
3.12
Full Name:
Ksx
Affiliation:
SMU
What is the latest build you tested with? If possible, we recommend testing with the latest nightly build.
I have not tested with any build
Did you include all relevant data sets for reproducing the issue?
No - Other reason (please specify in the issue body)
Did you include all code required to reproduce the issue?
- Yes, I have
Did you include all relevant configuration to reproduce the issue?
- Yes, I have
- Dominant language
- Python
- Stars
- 189
- Forks
- 117
- Avg merge
- 1d 58m
- Merged PRs (30d)
- 15
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from duckdb/duckdb-python
-
`overwrite` parameter in `to_parquet` does nothingPossibly taken @00200200 claimed this 6 days ago. Openneeds triage
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
duckdb/duckdb-python#633 · 1 comment ·
Maintainers usually reply within 1 day
-
Type stubs omit the `connection` keyword on module-level `read_json` and `write_csv`Possibly taken @tojacob03 claimed this 6 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
duckdb/duckdb-python#627 · 1 comment ·
Maintainers usually reply within 1 day
-
`read_json` accepts a list of paths at runtime but the type stub only allows a single path.Possibly taken @aleks-drozy claimed this 64 days ago. Openneeds triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
duckdb/duckdb-python#576 · 3 comments ·
Maintainers usually reply within 1 day
-
Missing ROW_GROUPS_PER_FILE argument in write_parquet functionPossibly taken A pull request linked to this issue is open or already merged. Openneeds triage
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
duckdb/duckdb-python#386 ·
Maintainers usually reply within 1 day
-
needs triage
Difficulty 4/5 3-5 days Newbie friendliness 45/100
duckdb/duckdb-python#645 ·
Maintainers usually reply within 1 day
All issues in duckdb/duckdb-python
Similar issues
-
needs-human needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
gke-labs/kube-agents#2400 · 1 comment ·
Maintainers usually reply within 1 day
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedPossibly taken @sylvesterkaczmarek claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
google-deepmind/bsuite#56 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
LearningCircuit/local-deep-research#7206 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Maintainers usually reply within 1 day