map_sync with pandas operation function does not finish.
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 25/100
- Issue 類型
- 缺陷
- 描述清晰度
- 需要釐清
- 活躍度
- 停滯
- 技術堆疊
- jupyter, numpy, pandas, python
研究方向
從提供的 Windows 重現開始,使用 ipyparallel 的 map_sync、40 個 sub-dataframes 以及 pandas 的 groupby/apply 操作;當處理在大約 10 個目標 dataframe 後停止時,檢查 client 和 worker 的輸出。確定不完整 map_sync 的原因,並記錄可重現的修正或已確認的限制後,該 issue 即完成。
由索引模型根據 Issue 內容生成。
描述
Map_sync with pandas operation function does not finish.
I have very long dataframe. So I split the dataframe into 40 sub-dataframes, and apply pandas operation to 40 sub-dataframes parallelly by using map_sync. The pandas operation is just about groupby and apply.
My code is like this:
PEN = 40
dfs = np.array_split(target_df, PEN)
c = ipp.Cluster(n=PEN)
with c as rc:
e_all = rc[:]
results = e_all.map_sync(FUCTION, dfs)
results
I have 30 target_dfs. For the first 10 target dfs map_sync worked fine. But after that map_sync didn't complete.
I have found that without parallelism, the pandas job applied to target_df completes in under 2 hours.
I use window os and Ipyparallel version is the lastest.
- 主要語言
- Jupyter Notebook
- 星號
- 2.6k
- 分支
- 1k
- PR 合併指標
- 30 天內沒有已合併 PR
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
ipython/ipyparallel 的其他 Issue
-
難度 4/5 3-5 天 新手友好度 48/100
ipython/ipyparallel#1029 · 9 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 35/100
ipython/ipyparallel#984 ·
-
難度 4/5 3-5 天 新手友好度 35/100
ipython/ipyparallel#937 · 4 則留言 ·
-
question
難度 5/5 一週以上 新手友好度 25/100
ipython/ipyparallel#897 · 12 則留言 ·
-
bug
難度 4/5 3-5 天 新手友好度 35/100
ipython/ipyparallel#882 ·
查看 ipython/ipyparallel 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 75/100
-
channels:add check:passed
難度 1/5 1 小時以內 新手友好度 90/100
-
難度 2/5 1-3 小時 新手友好度 75/100
avniproject/avni-webapp#1811 ·
-
inceleme-kuyrugu
難度 2/5 1-3 小時 新手友好度 70/100
Greater-Turkiye/platform#114 ·
-
bug triage
難度 2/5 1-3 小時 新手友好度 70/100
ultralytics/ultralytics#26305 · 1 則留言 ·