Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Windows: generation slows when the server window is minimized (EcoQoS moves threads to E-cores) — fix included

Open Beginner friendly
#691 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
68/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
performance

Research direction

Read serve/winjob.py and the included change, focusing on no_throttle and how contain() handles child processes. Run python -m unittest serve.test_winjob and check the Windows process-throttling behavior described in the issue. Done means the server and contained children are opted out of throttling and the test still passes.

Written by the indexing model from the issue text.

Description

Problem

On Windows 11 with a hybrid CPU (here an i9-13980HX, P+E cores), generation slows down (about 20%) when the server's console window is minimized. Having another window focused makes no difference. Windows power throttling (EcoQoS) treats the server and its children (strata.exe, strata-vision.exe) as background work and moves their threads to the E-cores.

Fix

Opt the server and every contained child out of power throttling with SetProcessInformation(ProcessPowerThrottling). This is done in serve/winjob.py, since contain() already handles every child process:

diff --git a/serve/winjob.py b/serve/winjob.py
index 008642a..c5db5f3 100644
--- a/serve/winjob.py
+++ b/serve/winjob.py
@@ -4,7 +4,8 @@ Closing the console window, Task Manager or a crash end the server without runni
 the vision encoder and the MCP servers kept running on their own. `contain(proc)` puts a child in a job object that
 is set to kill everything in it when its last handle closes - and the only handle is this process's, which the OS
 closes however the server ends. Processes the child starts later join the same job. The server itself stays out of
-the job, so a browser that `--open` starts is not tied to it. Elsewhere, and if the job cannot be made, a no-op.
+the job, so a browser that `--open` starts is not tied to it. Each child (and the server) is also opted out of
+Windows power throttling, see `no_throttle`. Elsewhere, and if the job cannot be made, a no-op.
 """
 from __future__ import annotations
 
@@ -57,6 +58,21 @@ if os.name == "nt":
             return None
         return job
 
+    class _PowerThrottling(ctypes.Structure):
+        _fields_ = [("Version", wintypes.ULONG), ("ControlMask", wintypes.ULONG), ("StateMask", wintypes.ULONG)]
+
+    _k32.SetProcessInformation.argtypes = (wintypes.HANDLE, ctypes.c_int, wintypes.LPVOID, wintypes.DWORD)
+
+    def no_throttle(handle) -> bool:
+        """Opt a process out of Windows power throttling (EcoQoS). Without this, Windows 11 treats a process whose
+        console window is minimized as background and, on hybrid CPUs, moves its threads to the E-cores:
+        generation slowed down by about 20% while the server window was minimized."""
+        info = _PowerThrottling(1, 0x1 | 0x4, 0)    # control EXECUTION_SPEED and IGNORE_TIMER_RESOLUTION, both off
+        return bool(_k32.SetProcessInformation(handle, 4, ctypes.byref(info), ctypes.sizeof(info)))  # ProcessPowerThrottling
+
+    _k32.GetCurrentProcess.restype = wintypes.HANDLE
+    no_throttle(_k32.GetCurrentProcess())           # the server itself (tokenizing, streaming) too
+
 
 def contain(proc) -> bool:
     """Put a started subprocess.Popen in the kill-on-close job. True if it is in; False (harmlessly) otherwise."""
@@ -69,6 +85,7 @@ def contain(proc) -> bool:
         if not _job:
             return False
     try:
+        no_throttle(int(proc._handle))
         return bool(_k32.AssignProcessToJobObject(_job, int(proc._handle)))
     except (AttributeError, OSError, ValueError):
         return False

python -m unittest serve.test_winjob still passes.

Measured with the patch

Setup: RTX 4090 Laptop 16 GB, i9-13980HX, 96 GB RAM, Qwen3.8-Flash-Next IQ3_S, AC power, Balanced plan. Same prompt at temperature 0 with 400 max tokens, 3 rounds each:

Server window Median tok/s Runs
Focused 44.3 42.1 / 46.8 / 44.3
Another window focused 43.0 43.0 / 40.4 / 44.3
Minimized 45.0 41.9 / 45.9 / 45.0

With the patch, minimizing no longer makes a difference beyond run-to-run noise. Numbers without the patch are in the comment below. Not yet tested on battery, where Windows throttles more aggressively.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.