Testing, performance and upstreaming for fizzy's SDL patches and backend
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 12/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- c, zig
- Domain
- desktop-dev, performance, testing-qa
Research direction
This is a multi-part tracker, not a single task. Pick one unchecked step, such as A6 (scripts/live-resize/analyze.py threshold and exit code, plus a "Checking a rebase" section in docs/DEPENDENCIES.md) or B1's demo test in tests/integration.zig. Start by reading the named file and running the existing check. Done means that step's checkbox is met and the PR says Part of #280.
Written by the indexing model from the issue text.
Description
This plan has two tracks of tests. Track A tests fizzy's SDL patches in isolation, inside the fork that carries them. Track B tests fizzy running on patched SDL, under the kind of use a person gives it, with demo tapes as the driver. Both are needed because the tape player adds its events to dvui right after Window.begin (app/automation/Input.zig), above SDL. A tape never passes through SDL's input translation or the OS's modal resize loop, and it cannot see whether alpha reached DWM. Track A covers that. What tapes can do is drive the real backend through the paths the patches exist for (transparent windows, glass, floats popped out into OS windows, full screen, resize). Track B turns that into checks that fail when something breaks.
Today the patches are checked by hand at each rebase (docs/DEPENDENCIES.md, "Where a rebase conflicts"). The one measured check is scripts/live-resize/run.sh, and it prints numbers rather than a pass or fail. The bundled demos, which fizzyed.it embeds, are never played in a test. The demo tests (tests/integration.zig, "demo: …") use a synthetic stage and a tape built in the test.
Goals
- Test everything we do, rigorously. Every SDL patch, every backend path it serves, and every bundled demo gets a check that fails when it breaks, ideally in CI, otherwise as a documented local check with a pass/fail result.
- Make fizzy as fast as it can be, and keep it there. The same runs that check correctness also measure: frame times, presents, swapchain recreations, native calls per frame. Each measurement gets a baseline, so a regression shows up as a number, not as a feeling.
- Distill the patches into changes SDL upstream would take. A patch is ready to propose when it has a test that fails on upstream, measurements on real hardware, and no fizzy-specific assumptions. When upstream merges it, the fork's stack gets shorter.
Each step is its own PR saying Part of #280. Track A's PRs land in fizzyedit/SDL and fizzyedit/sdl_zig and link back here.
Track A: the SDL fork tests its own patches
Where the tests live. Suites go in fizzyedit/SDL as test/testautomation_fizzy*.c, on SDL's own test harness (src/test/SDL_test_harness.c, SDLTest_AssertCheck, --filter). A test then moves through every rebase with the patch it checks, and it can go upstream in that patch's PR. fizzyedit/sdl_zig gets a zig build test-fizzy step that builds SDL_test and the suites against the library it already builds. It uses zig cc, so binaries for every target cross-compile from one machine.
Every test must fail on upstream. The fork's CI runs each suite against release-3.4.N and expects it to fail there, then against fizzy-3.4 and expects it to pass. If a test passes on upstream, it doesn't test its patch. Once a patch is merged upstream, its test passes on the new tag, and that is the signal to drop the patch at the next rebase.
Patch (DEPENDENCIES.md) |
What the test asserts | How it's observed | Where it runs |
|---|---|---|---|
| 1. Transparent claim | SDL_ClaimWindowForGPUDevice succeeds for a SDL_WINDOW_TRANSPARENT window on Metal and Vulkan |
API result | Linux CI (lavapipe), a Mac |
| 2. D3D12 DirectComposition | Claim, clear to 50% red, present: the composited desktop shows the blend over a known backdrop. Resizing logs no errors | BitBlt from the desktop DC, which captures DWM's output, then a pixel probe |
windows-latest on WARP if its session allows screen captures, otherwise the UTM VM |
| 3. Metal present with transaction | With presentsWithTransaction set, N presents inside an explicit CATransaction finish within a timeout, and each drawable gets a presentedTime |
Timeout plus drawable callbacks | A Mac |
| 4+5. macOS live-resize sync | One app frame per size step, driven from displayLayer: |
Callbacks counted per size change, with a synthetic drag (scripts/live-resize/record.swift) |
A local Mac only (Accessibility and Screen Recording) |
| 6. Wayland frame insets | set_window_geometry is the frame without the insets. The published inset props match. Maximizing zeroes them. The input region is the frame plus at most 8 |
Exact protocol requests in a WAYLAND_DEBUG=1 trace |
Linux CI under weston --backend=headless |
| 7. Windows live-resize sync | One app frame per client-size change in WM_WINDOWPOSCHANGED |
Event-watch callbacks counted against size changes, during a SendInput drag |
The UTM VM; the CI runner if SendInput works there |
GitHub's hosted macOS runners have Metal (found in A1). macos-14, macos-15 and macos-latest (macOS 26) each expose one "Apple Paravirtual device", and SDL's Metal driver claims a window and presents on all three. On macos-14, MTLCreateSystemDefaultDevice() returns nil from a command-line tool even though MTLCopyAllDevices() lists the device, so probe with the device list. CI can therefore run the Metal suites for patches 3–5, as far as having a device goes. Whether a live resize behaves there is still unknown.
Linux suites run on the oldest supported distribution as well as a current one. A1 found that SDL built this way can't load its Wayland driver where libxkbcommon is older than 1.10, which includes Ubuntu 24.04 (#289). A test job on ubuntu:24.04 would have caught it.
Steps:
- A1.
sdl_zig: atest-fizzystep building SDL_test plus a suite for patch 1, with Linux CI on lavapipe: fizzyedit/SDL#1 →fizzy-3.4.16-6, fizzyedit/sdl_zig#1 →fizzy-1.0.3+3.4.16-6; fizzy repins in #294. Passes on the fork on a Mac, Linux (X11 and Wayland) and all three macOS runners; fails on upstream 3.4.16 (checked on Metal) - A2.
SDL: the patch 6 suite on the Wayland debug trace, under headless weston: fizzyedit/SDL#2 →fizzy-3.4.16-8, fizzyedit/sdl_zig#2 →fizzy-1.0.3+3.4.16-8; fizzy repins in #294. Passes at scale 1 and 2, fails on upstream. Headless weston can't cover a compositor-chosen floating size, tiled states, fractional scale or xdg-decoration - A3.
SDL: the patch 2 suite on WARP, first claim and present, then the pixel probe if the runner allows captures - A4.
SDLCI: run every suite against upstream and against the fork, and fail if any passes on upstream - A4b.
SDL: patch 8,xkb_keymap_mod_get_maskoptional at runtime, with a suite that fails onubuntu:24.04without it (#289): fizzyedit/SDL#3 →fizzy-3.4.16-7, fizzyedit/sdl_zig#3 →fizzy-1.0.3+3.4.16-7; fizzy repins in #294. The upstream PR text is drafted, not yet opened - A5.
SDL: the suites for patches 3, 4+5 and 7, as local rebase checks rather than CI - A6. fizzy:
scripts/live-resize/analyze.pygets a threshold and an exit code, anddocs/DEPENDENCIES.mdgets a "Checking a rebase" section naming every check above
Track B: tapes drive fizzy, and the run returns a verdict
- B1. Every bundled demo plays to the end in
test-integration. #293 (draft): fizzy's realDemostage runs under dvui-testing, in its own test binary; the test found and fixed five seek bugs. Still open: with seed0x6, the explorer tree's two-way open-folder sync overrides a restored snapshot. Each catalog demo (src/editor/demos/) plays through fizzy's realDemostage. The test asserts that no wait gives up and thatPlayer.mismatches == 0. It then seeks to seeded random times and asserts that the fingerprint at each matches what straight play reached. That is a property test of every document owner'scaptureDocumentStateand restore, which nothing tests directly now. It also catches an anchor renamed under a demo before the demo breaks on fizzyed.it. First, find out whetherEditor.Demoruns under dvui-testing. If it doesn't, the missing piece is a seam in the stage, not a second copy of the demos. Merged as #293. - B2. Backend health counters. A small struct the native backend keeps up to date: presents, swapchain recreations, OS windows alive, SDL errors (captured with
SDL_SetLogOutputFunction) and frame-time percentiles. Read through a named seam, it replaces the one-offFIZZY_LIVE_RESIZE_TRACEplumbing, which becomes one reader of it. Merged as #309:SDLBackend.health()returns aHealth.Snapshot, about 10 ns a frame in ReleaseFast. - B3. A headless run of the real app that ends in a verdict. Something like
fizzy --profile <dir> --tape <file> --exit-when-done. It plays a live tape (app/automation/LiveDriver.zig) or a demo, then exits non-zero if the tape lost its place, a replay mismatched, SDL logged an error,DebugAllocatorfound leaks, or the counters break the tape's expectations. It also writes the counters as ZON for CI to keep. CI runs it on Linux under headless weston with lavapipe, and onwindows-latestwith WARP. Merged as #313:FIZZY_VERDICT=out.zon FIZZY_DEMO=<demo or tape> fizzy --profile <dir>, plus a requiredVerdict run (Linux)CI job. It found #315 (merged), #316, #317 and #318. - B4. Soak tapes. Tapes that repeat the risky things: pop a float out into an OS window and back 50 times and expect one window alive, toggle full screen, switch the glass commands (
fizzy.window.glass.*), open and close dialogs over a frost. These are the bugs found by eye so far: a closed popped-out window kept alive for lack of an autorelease pool, ghost windows from animations that never finished. They complement #243 step 5, which tests the windows' plan against a fake backend. These test the real backend reconciling it. - B5. Window operations in
Stage. AStage.window(op)seam for resize, move and scale change, so a tape reaches swapchain resize and the transparent claim at every step. It is not a live resize, since there is no modal loop, so Track A's live-resize checks stay. Butscripts/live-resize/run.shcan then play a tape first, so its drag is measured over a real scene (glass, floats, a long document) rather than an empty window. - B6. A seeded monkey driver. Each frame it picks a random visible anchor (anchors are published per frame through
core.anchor.want) and clicks, drags, scrolls or types, in a safety-checked build with B2's counters on. A failure saves its seed and the binary.tapeit produced, which are the repro. Delta-debugging over the ops shrinks the tape to the fewest steps that still fail. - B7. Reference images at chapter marks, a few. Demo time is deterministic, so a tape's frame N is reproducible. Compare a handful of frames on the software rasterizers, with a tolerance, to catch regressions in glass, frost and premultiplied alpha. The frame-dump names on dvui-dev are the likely hook. Keep the set small, because reference images are brittle.
Deferred, with its seam: once the recorder in docs/AUTOMATION_PLAN.md lands, a person's bug report can be a tape, and B3 plays it as a regression test with no new machinery.
Performance: the same runs, measured
Correctness runs double as benchmarks, so a fast path is never untested and a test never hides a slowdown.
- P1. Benchmark tapes on real hardware. B3's runner with B2's counters, playing a fixed scene (glass, two floats, a long document, a markdown preview) for a fixed demo time. Output: frame-time p50/p95/p99, presents, Objective-C messages and native calls per frame. Run locally on a Mac, the Windows VM and Linux. Hosted CI is too noisy to fail a PR on timing.
- P2. Baselines in the repo, compared on demand. A
zig build bench-tapestep that writes results as ZON and compares them against a checked-in baseline per machine class, flagging regressions past a threshold. Thebench-replayandbench-tapesteps that already exist are the model. - P3. Counts in CI, not times. Things that are deterministic per frame (draw calls, render-target switches, native calls, swapchain recreations, allocations) are asserted in CI with B3's runner, because a count doesn't wobble the way a time does. #243 step 4's per-kind budgets for Objective-C messages are one such count.
- P4. SDL-side measurements for patches whose job is timing (3, 4+5, 7). Track A's local suites print the same numbers that
scripts/live-resizemeasures, so a rebase is checked for regressions as well as breakage.
Track C: from fork patches to upstream SDL
On hold (2026-10-09): nothing goes upstream for now. SDL's PR template requires certifying a contribution contains no code generated by a Large Language Model, so agent-written patches can't be proposed as PRs. When this resumes, the options are a bug report (an issue carries no code) or a fix the maintainer writes. C1 and C2/C2b still apply: they keep the fork's stack clean and tested. Don't open anything on libsdl-org/SDL.
Track A's tests are what make a patch proposable. This track takes each patch the rest of the way.
-
C1. Tidy the stacks. Squash the fix-ups that
DEPENDENCIES.mdalready lists (SDL 5 into 4; sdl_zig 3–6 into 1) at the next rebase, so each SDL patch is a single change with a single test. -
C2. A readiness checklist per patch in
DEPENDENCIES.md: a test that fails on upstream (Track A), measurements on real hardware (P4), no fizzy-specific names or assumptions, and SDL's own code style and docs (hint docs inSDL_hints.h, property docs in headers). Each patch shows where it stands against that checklist. -
C2b. Patch 6 (Wayland frame insets) before proposing. A2 found these: Fixed on the forks (
fizzy-3.4.16-9), plus a sixth bug (framing lost after hide and show). fizzy repins in #314. Upstreammainhas its own insets API now; see DEPENDENCIES.md.- Recreating the window loses the insets.
SDL_RecreateWindowcallsCreateSDLWindow(_this, window, 0)(SDL_video.c), so the creation props are gone, e.g. after an OpenGL renderer is attached to a window made withoutSDL_WINDOW_OPENGL. Found by reading the code; untested. Fix: a setter, or the insets kept onSDL_Window. Upstream will want a setter anyway. - An opaque window's opaque region covers its shadow.
set_opaque_regiongets the whole surface (wl_region.add(0,0,400,300)in the trace) instead of the frame. Fizzy's window is transparent, so it isn't affected. - Toplevel bounds are compared against the wrong size. In the floating 0x0 configure path, the compositor's bounds (frame units) are compared with the surface size, without subtracting the insets.
- The 8-unit input band is hard-coded. A public API may want it configurable. Fizzy's 6-point resize edge fits inside it.
- The props appear only once the window is shown. They're published on the first
ApplyFrameInsets, so they're absent before show.
Each fix gets an assertion in A2's suite where headless weston can show it (1, 2, 5), and a note where it can't (3 needs a compositor-chosen floating size).
- Recreating the window loses the insets.
-
C3. Propose in order of readiness. First patch 1 (one check moved into the D3D12 driver; its test is A1). Then 6 (Wayland insets; upstream will want a setter as well as the create props). Then 2 (DirectComposition). Then 3, 4+5 and 7 together as one live-resize proposal, with the measurements from macOS and Windows. Each upstream PR links its fork test and evidence.
-
C4. When upstream merges a patch, its suite passes on the new upstream tag (A4 flags it). The patch is dropped at the next rebase, its suite moves into the upstream PR's own tests or is deleted, and
DEPENDENCIES.mdrecords the change.
Order
Do A1 and B1 first. Both are small and need nothing new. A2 is the best return in Track A. B2 → B3 → B4 is where bugs between SDL and fizzy actually show up. B5 and B6 build on B3, and B7 comes last. P1–P3 follow B2/B3, since they use the same runner. C1 and C2 can start at any time; C3 for each patch waits on that patch's Track A suite.
🤖 Generated with Claude Code
- Dominant language
- Zig
- Stars
- 1.5k
- Forks
- 40
- Avg merge
- 5h 35m
- Merged PRs (30d)
- 151
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from fizzyedit/fizzy
-
replay: bench-replay no longer compiles, and nothing in CI builds itPossibly taken A pull request linked to this issue is open or already merged. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
fizzyedit/fizzy#306 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 8/100
Maintainers usually reply within 1 day
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 25/100
Maintainers usually reply within 1 day
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 12/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
windows 再現済み 要トリアージ 誤判定
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
yksr-melt/Meltype#419 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
Facepunch/sbox-public#12063 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NousResearch/hermes-agent#136483 ·
Maintainers usually reply within 1 day
-
effort:S priority:P2
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day