bug: Compaction does not work and the chat continues where it left of
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
Research direction
Start by reproducing the compaction flow described in the issue and compare the request that returns the summary with the following chat request. Trace where conversation history is selected between those requests; done means the post-compaction request uses the summary instead of continuing from the pre-compaction history, with the behavior verified by a regression test.
Written by the indexing model from the issue text.
Description
Hello,
I discovered that the compaction feature does not work correctly. Eventhough the compaction gets the model to return something the next message continues right where the last one left of. My assumption is that the compaction correctly gets the model to summaries but subsequent request dont cut off the chat history before that.
0177.20.719.593 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 29.54.311.332 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.996 (> 0.100 thold), f_keep = 1.000
[55467] 29.54.311.928 I slot launch_slot_: id 0 | task 387 | processing task, is_child = 0
[55467] 30.02.697.707 I slot print_timing: id 0 | task 387 | n_gen = 101, tg = 23.37 t/s, tg_3s = 23.60 t/s
[55467] 30.05.736.972 I slot print_timing: id 0 | task 387 | n_gen = 194, tg = 26.37 t/s, tg_3s = 30.60 t/s
[55467] 30.08.774.428 I slot print_timing: id 0 | task 387 | n_gen = 265, tg = 25.49 t/s, tg_3s = 23.37 t/s
[55467] 30.11.814.724 I slot print_timing: id 0 | task 387 | n_gen = 316, tg = 23.51 t/s, tg_3s = 16.77 t/s
[55467] 30.12.301.236 I slot print_timing: id 0 | task 387 | prompt eval time = 4106.47 ms / 640 tokens ( 6.42 ms per token, 155.85 tokens per second)
[55467] 30.12.301.238 I slot print_timing: id 0 | task 387 | eval time = 13882.63 ms / 324 tokens ( 42.98 ms per token, 23.27 tokens per second)
[55467] 30.12.301.239 I slot print_timing: id 0 | task 387 | total time = 17989.10 ms / 964 tokens
[55467] 30.12.301.239 I slot print_timing: id 0 | task 387 | graphs reused = 383
[55467] 30.12.301.243 I slot print_timing: id 0 | task 387 | draft acceptance = 0.61696 ( 211 accepted / 342 generated), mean len = 2.85
// long context i want to compact
[55467] 30.12.304.163 I slot release: id 0 | task 387 | stop processing: n_tokens = 176122, truncated = 0
10178.12.623.354 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 30.45.948.746 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 611151107915
[55467] 30.49.168.967 I slot launch_slot_: id 0 | task 504 | processing task, is_child = 0
[55467] 30.53.320.942 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 4171, progress = 0.56, t = 3.68 s / 1133.63 tokens per second
[55467] 30.55.275.116 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 6219, progress = 0.84, t = 5.61 s / 1107.64 tokens per second
[55467] 30.55.548.147 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 6391, progress = 0.86, t = 6.28 s / 1018.43 tokens per second
[55467] 30.56.840.228 I slot print_timing: id 0 | task 504 | prompt processing, n_tokens = 7415, progress = 1.00, t = 7.19 s / 1031.60 tokens per second
[55467] 30.59.959.252 I slot print_timing: id 0 | task 504 | n_gen = 162, tg = 53.53 t/s, tg_3s = 53.86 t/s
[55467] 31.02.962.961 I slot print_timing: id 0 | task 504 | n_gen = 292, tg = 48.41 t/s, tg_3s = 43.28 t/s
[55467] 31.05.972.238 I slot print_timing: id 0 | task 504 | n_gen = 447, tg = 49.44 t/s, tg_3s = 51.51 t/s
[55467] 31.09.000.928 I slot print_timing: id 0 | task 504 | n_gen = 598, tg = 49.55 t/s, tg_3s = 49.86 t/s
[55467] 31.12.049.402 I slot print_timing: id 0 | task 504 | n_gen = 750, tg = 49.61 t/s, tg_3s = 49.86 t/s
[55467] 31.13.253.884 I slot print_timing: id 0 | task 504 | prompt eval time = 7782.71 ms / 7419 tokens ( 1.05 ms per token, 953.27 tokens per second)
[55467] 31.13.253.886 I slot print_timing: id 0 | task 504 | eval time = 16301.98 ms / 794 tokens ( 20.56 ms per token, 48.64 tokens per second)
[55467] 31.13.253.886 I slot print_timing: id 0 | task 504 | total time = 24084.69 ms / 8213 tokens
[55467] 31.13.253.887 I slot print_timing: id 0 | task 504 | graphs reused = 649
[55467] 31.13.253.891 I slot print_timing: id 0 | task 504 | draft acceptance = 0.64453 ( 524 accepted / 813 generated), mean len = 2.93
[55467] 31.13.254.085 I slot release: id 0 | task 504 | stop processing: n_tokens = 8214, truncated = 0
// returned presumably the summary
10178.39.960.509 E srv operator(): http client error: Connection handling canceled
10180.54.269.523 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 33.27.873.571 I slot get_availabl: id 0 | task -1 | selected slot by LRU, t_last = 611212057838
[55467] 33.28.965.816 I slot launch_slot_: id 0 | task 783 | processing task, is_child = 0
[55467] 33.33.605.682 I slot print_timing: id 0 | task 783 | n_gen = 103, tg = 26.94 t/s, tg_3s = 27.20 t/s
[55467] 33.36.643.610 I slot print_timing: id 0 | task 783 | n_gen = 197, tg = 28.72 t/s, tg_3s = 30.94 t/s
[55467] 33.39.698.431 I slot print_timing: id 0 | task 783 | n_gen = 285, tg = 28.75 t/s, tg_3s = 28.81 t/s
[55467] 33.40.674.123 I slot print_timing: id 0 | task 783 | prompt eval time = 852.88 ms / 25 tokens ( 34.12 ms per token, 29.31 tokens per second)
[55467] 33.40.674.124 I slot print_timing: id 0 | task 783 | eval time = 10855.21 ms / 316 tokens ( 34.46 ms per token, 29.02 tokens per second)
[55467] 33.40.674.125 I slot print_timing: id 0 | task 783 | total time = 11708.09 ms / 341 tokens
[55467] 33.40.674.125 I slot print_timing: id 0 | task 783 | graphs reused = 736
[55467] 33.40.674.127 I slot print_timing: id 0 | task 783 | draft acceptance = 0.85019 ( 227 accepted / 267 generated), mean len = 3.55
// continues before the compaction?
[55467] 33.40.677.165 I slot release: id 0 | task 783 | stop processing: n_tokens = 176463, truncated = 0
10181.08.589.202 I srv proxy_reques: proxying request to model Qwen3.8 27B (medium) on port 55467
[55467] 33.42.179.481 I slot get_availabl: id 0 | task -1 | selected slot by LCP similarity, f_sim_best = 0.998 (> 0.100 thold), f_keep = 1.000
[55467] 33.42.180.025 I slot launch_slot_: id 0 | task 875 | processing task, is_child = 0
[55467] 33.48.601.836 I slot print_timing: id 0 | task 875 | n_gen = 100, tg = 25.31 t/s, tg_3s = 25.56 t/s
[55467] 33.51.663.108 I slot print_timing: id 0 | task 875 | n_gen = 192, tg = 27.39 t/s, tg_3s = 30.05 t/s
[55467] 33.54.717.522 I slot print_timing: id 0 | task 875 | n_gen = 291, tg = 28.92 t/s, tg_3s = 32.41 t/s
Thank you for your time and consideration.
Best regards,
Joshua Krimmer
- Dominant language
- Python
- Stars
- 620
- Forks
- 85
- PR merge metrics
- No merged PRs in 30d
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from open-webui/computer
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
open-webui/computer#293 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
open-webui/computer#286 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
open-webui/computer#276 · 1 comment · 1 reaction ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
open-webui/computer#271 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
open-webui/computer#269 · 1 comment ·
All issues in open-webui/computer
Similar issues
-
needs-human needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
gke-labs/kube-agents#2400 · 1 comment ·
Maintainers usually reply within 1 day
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedPossibly taken @sylvesterkaczmarek claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
google-deepmind/bsuite#56 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
LearningCircuit/local-deep-research#7206 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Maintainers usually reply within 1 day