map_sync with pandas operation function does not finish.
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 25/100
- Issue-Typ
- Bug
- Klarheit
- Muss geklärt werden
- Aktivitätsstatus
- Veraltet
- Tech-Stack
- jupyter, numpy, pandas, python
- Bereich
- data, distributed-systems
Rechercherichtung
Beginnen Sie mit der bereitgestellten Windows-Reproduktion unter Verwendung von ipyparallel's map_sync, 40 sub-dataframes und der pandas groupby/apply-Operation; untersuchen Sie die Client- und Worker-Ausgabe, wenn die Verarbeitung nach ungefähr 10 Zieldatenframes stoppt. Das Issue ist abgeschlossen, sobald die Ursache für das unvollständige map_sync identifiziert und entweder ein reproduzierbarer Fix oder eine bestätigte Einschränkung dokumentiert wurde.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Map_sync with pandas operation function does not finish.
I have very long dataframe. So I split the dataframe into 40 sub-dataframes, and apply pandas operation to 40 sub-dataframes parallelly by using map_sync. The pandas operation is just about groupby and apply.
My code is like this:
PEN = 40
dfs = np.array_split(target_df, PEN)
c = ipp.Cluster(n=PEN)
with c as rc:
e_all = rc[:]
results = e_all.map_sync(FUCTION, dfs)
results
I have 30 target_dfs. For the first 10 target dfs map_sync worked fine. But after that map_sync didn't complete.
I have found that without parallelism, the pandas job applied to target_df completes in under 2 hours.
I use window os and Ipyparallel version is the lastest.
- Vorherrschende Sprache
- Jupyter Notebook
- Sterne
- 2.6k
- Forks
- 1k
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Keine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus ipython/ipyparallel
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
ipython/ipyparallel#1029 · 9 Kommentare ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 35/100
ipython/ipyparallel#984 ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 35/100
ipython/ipyparallel#937 · 4 Kommentare ·
-
question
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 25/100
ipython/ipyparallel#897 · 12 Kommentare ·
-
bug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 35/100
ipython/ipyparallel#882 ·
Alle Issues in ipython/ipyparallel
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
Yuliang-Liu/MultimodalOCR#103 ·
-
blocklist:remove check:passed
Schwierigkeit 2/5 Unter einer Stunde Anfängerfreundlichkeit 66/100
iptv-org/database#36968 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 9 Tagen
-
Add: Prima Comedy SDOffencheck:passed streams:add
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 78/100
Maintainer antworten meist innerhalb von 1 Tag
-
Component: Ruby Type: bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 74/100
Maintainer antworten meist innerhalb von 1 Tag
-
enhancement
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 82/100
Maintainer antworten meist innerhalb von 4 Tagen