Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Loading and saving back from file with netcdf fails

Open Beginner friendly
#11,672 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@mokashang is already working on this.

Since Oct 5, 2026.

  • #11673 by @mokashang — open

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
74/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
data

Research direction

Start with the second loop in DataTree.load() around datatree.py#L2478-L2482, then check the existing DataTree tests for a place to add a regression case. Reproduce the issue with a variable named like its group, then verify loading brings that variable into memory and saving back to the same netCDF file succeeds without losing it.

Written by the indexing model from the issue text.

Description

bug topic-backends
What happened?

I load a DataTree from a netCDF file with xr.load_datatree and save it back to the same path. The save fails with RuntimeError: NetCDF: Write to read only. The failed save also leaves the file damaged: the group is still there, but the variable is gone.

This only happens when a variable has the same name as its group (e.g. variable a in group /a). If you rename the variable, the save works.

What did you expect to happen?

load_datatree loads everything into memory, so saving back to the same file should work, as it does with xr.load_dataset

Minimal Complete Verifiable Example
# /// script
# requires-python = ">=3.11"
# dependencies = [
#   "xarray[complete]@git+https://github.com/pydata/xarray.git@main",
# ]
# ///

import xarray as xr

xr.DataTree.from_dict({"/a": xr.Dataset({"a": 0})}).to_netcdf("test.nc")
xr.load_datatree("test.nc").to_netcdf("test.nc")
Steps to reproduce

Execute the script above with uv run

MVCE confirmation
  • Minimal example — the example is as focused as reasonably possible to demonstrate the underlying issue in xarray.
  • Complete example — the example is self-contained, including all data and the text of any traceback.
  • Verifiable example — the example copy & pastes into an IPython prompt or Binder notebook, returning the result.
  • New issue — a search of GitHub Issues suggests this is not a duplicate.
  • Recent environment — the issue occurs with the latest version of xarray and its dependencies.
Relevant log output
...
RuntimeError: NetCDF: Write to read only
Raised while encoding variable 'a' with value <xarray.Variable ()> Size: 8B
[1 values with dtype=int64]
Anything else we need to know?

The cause seems to be in DataTree.load(). Its second loop checks if k not in lazy_data, but lazy_data is keyed by node path, not by variable name (datatree.py#L2478-L2482). So when a variable has the same name as a node path, it isn't loaded and stays lazy. When it's written, xarray reads it from the file that is being overwritten at that moment.

Environment
INSTALLED VERSIONS ------------------ commit: None python: 3.14.4 (main, Apr 14 2026, 14:46:33) [Clang 22.1.3 ] python-bits: 64 OS: Darwin OS-release: 27.0.0 machine: arm64 processor: arm byteorder: little LC_ALL: en_US.utf-8 LANG: en_US.utf-8 LOCALE: ('en_US', 'UTF-8') libhdf5: 1.14.6 libnetcdf: 4.9.3

xarray: 2026.9.1.dev21+g78b8e9f96
pandas: 3.0.6
numpy: 2.5.3
scipy: 1.18.1
netCDF4: 1.7.4
pydap: 3.5.11
h5netcdf: 1.8.1
h5py: 3.16.0
zarr: 3.4.0
cftime: 1.6.6
nc_time_axis: 1.4.1
iris: None
bottleneck: 1.6.0
dask: 2026.8.0
distributed: 2026.8.0
matplotlib: 3.11.2
cartopy: 0.26.0
seaborn: 0.13.2
numbagg: 0.9.6
fsspec: 2026.9.0
cupy: None
pint: None
sparse: 0.19.2
flox: 0.11.2
numpy_groupies: 0.12.3
setuptools: None
pip: None
conda: None
pytest: None
mypy: None
IPython: None
sphinx: None

Dominant language
Python
Stars
4.2k
Forks
1.4k
Avg merge
2d 19h
Merged PRs (30d)
48

Getting set up

Open in Codespaces

Starts the project's dev container in your browser, under your own GitHub account.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from pydata/xarray

All issues in pydata/xarray

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.