Using partition "breaks" program logic
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 28/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- python
- Ambito
- stream-processing
Direzione di ricerca
Start by tracing stream.partition, stream.emit, and accumulate, then read the documented async def process_file example in “Processing Time and Back Pressure.” Determine how a partition flushes when input reaches EOF and how callers can wait for pending processing; done means the final count includes all lines before the concluding print runs.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
I am struggling to use partition in a pipeline because it "breaks" the logic of my program; presumably because it introduces asynchronous processing.
As a simplified example, I have something that works along the lines of this:
import streamz
def main():
state = {
"cnt": 0,
}
stream = streamz.Stream()
cntd = stream.accumulate(cnt,
returns_state=True,
start=state)
cntd.sink(print)
with open("many_lines.txt", "r") as fh:
for line in fh:
stream.emit(line)
print(f"found {state.get('cnt')} lines")
def cnt(state, itm):
state["cnt"] += 1
return state, itm
if __name__ == "__main__":
main()
This basically runs through all the lines in the file many_lines.txt, counts and prints them and then reports
found 10000 lines
So far so good.
When I introduce partition now, like this:
import streamz
def main():
state = {
"cnt": 0,
}
stream = streamz.Stream()
parted = stream.partition(10001, timeout=2) # <= PARTITION HERE
cntd = parted.accumulate(cnt,
returns_state=True,
start=state)
cntd.sink(print)
with open("many_lines.txt", "r") as fh:
for line in fh:
stream.emit(line)
print(f"found {state.get('cnt')} lines")
def cnt(state, itm):
state["cnt"] += 1
return state, itm
if __name__ == "__main__":
main()
I would want to see basically the same result. But I see nothing for some time and then
found 0 lines
I know, there are only 10'000 lines in many_lines.txt so the partition will never fill up, but it should hit the timeout at some point and "release" the data, no?
I suspect that the program terminates before the partition hits the timeout, so I tried (many variations of) awaiting stream.emit(line). That was inspired by the async def process_file(fn): function in Processing Time and Back Pressure.
For example like this:
import streamz
def main():
state = {
"cnt": 0,
}
stream = streamz.Stream()
parted = stream.partition(10001, timeout=2)
cntd = parted.accumulate(cnt,
returns_state=True,
start=state)
cntd.sink(print)
with open("many_lines.txt", "r") as fh:
for line in fh:
await stream.emit(line) # <= USE AWAIT HERE
print(f"found {state.get('cnt')} lines")
def cnt(state, itm):
state["cnt"] += 1
return state, itm
if __name__ == "__main__":
main()
But this (obviously) does not work (SyntaxError: 'await' outside async function). And I also did not find a way to make it work.
(How) Can I make sure the for loop terminates before the print statement (or any remaining code, for that matter) is executed? Or am I getting this completely wrong?
My use case is to read (all) lines in pretty big files (I cannot load into memory at once), send them through a streamz pipeline and then continue with my program. "Then" meaning, after all lines are processed (also those that might be "stuck" in a partition when no more lines are emitted because we reached EOF; this is why I need the timeout, I believe).
- Lingua principale
- Python
- Stelle
- 1.3k
- Fork
- 149
- Merge medio
- 17h 39m
- PR unite (30g)
- 1
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di python-streamz/streamz
-
pkg_resources warning Aperta
Difficoltà 3/5 1-2 giorni Idoneità per principianti 35/100
python-streamz/streamz#481 · 4 commenti ·
-
Combining the streamz.Stream.filenames() and streamz.Stream.from_textfile() using dask scatter? Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
python-streamz/streamz#480 · 2 commenti ·
-
Compile the code into c++ Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
python-streamz/streamz#479 · 1 commento ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
python-streamz/streamz#478 · 6 commenti ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 10/100
python-streamz/streamz#476 · 17 commenti · 2 reazioni ·
Tutte le issue di python-streamz/streamz
Issue simili
-
bug confirmed issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
open-webui/open-webui#30750 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
OpenwaterHealth/openmotion-bloodflow-app#604 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
good first issue
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100