Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[BUG]: px.sunburst / px.treemap / px.icicle with path give a different sector order on every run for Polars DataFrames

Closed
#5,765 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
35/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start in plotly/express/_core.py at process_dataframe_hierarchy and trace the group_by calls used for path-based charts. Compare the Polars, pandas, and PyArrow hierarchy ordering, then inspect the regression test included or proposed by the reporter. Done means repeated Polars runs produce stable, first-appearance ordering consistent with the other supported inputs.

Written by the indexing model from the issue text.

Description

Description

When path= is used with a Polars DataFrame, px.sunburst, px.treemap and px.icicle build their ids / labels / parents / values arrays in a different order every time the script is run. The same data as a pandas DataFrame or a PyArrow table always gives the same order: the order in which the sectors first appear in the data.

The cause is in process_dataframe_hierarchy (plotly/express/_core.py). Each level of the hierarchy is built with df.group_by(path[i:]).agg(...). With pandas (narwhals uses sort=False) and PyArrow the groups come back in order of first appearance, but Polars' group_by does not guarantee any order, so the order of the output changes from run to run.

Consequences:

  • fig.to_json() / fig.write_html() output is not reproducible with Polars input (snapshot tests, caching, diffs of generated HTML).
  • With sort=False, or when sectors have equal values, the chart itself is laid out differently on each run.
  • Polars results differ from pandas / PyArrow results for identical data.
Screenshots/Video

N/A: the difference is in the figure data; see the output below.

Steps to reproduce
import plotly
import plotly.express as px
import polars as pl

df = pl.DataFrame(
    {
        "region": ["South", "North", "South", "West", "North", "West"],
        "sector": ["Tech", "Finance", "Finance", "Tech", "Tech", "Finance"],
        "sales": [1, 2, 3, 4, 5, 6],
    }
)
fig = px.sunburst(df, path=["region", "sector"], values="sales")
print(plotly.__version__, pl.__version__, list(fig.data[0].ids))

Running the script three times (plotly 7.1.0, polars 1.44.2):

7.1.0 1.44.2 ['West/Tech', 'West/Finance', 'North/Finance', 'South/Finance', 'North/Tech', 'South/Tech', 'South', 'North', 'West']
7.1.0 1.44.2 ['South/Tech', 'West/Tech', 'North/Tech', 'South/Finance', 'North/Finance', 'West/Finance', 'West', 'North', 'South']
7.1.0 1.44.2 ['South/Tech', 'West/Tech', 'North/Tech', 'North/Finance', 'South/Finance', 'West/Finance', 'West', 'South', 'North']

With pd.DataFrame(...) instead, every run prints:

['South/Tech', 'North/Finance', 'South/Finance', 'West/Tech', 'North/Tech', 'West/Finance', 'South', 'North', 'West']
Notes

I have a small fix with a regression test and will open a PR for it.

Dominant language
Python
Stars
18.8k
Forks
2.8k
Avg merge
13h 41m
Merged PRs (30d)
21

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from plotly/plotly.py

All issues in plotly/plotly.py

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.