Design: table-driven CodecStrategy to shrink generated code
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- rust
- Domain
- backend-api-design, performance
Research direction
Start by reading the existing Config API, generated Message implementation, and the buffa::table entry point described in the issue, then review the proof-of-concept measurements. Done means adding the opt-in CodecStrategy configuration with Unrolled as the default, defining the documented fallback and warning behavior, and covering supported and unsupported message shapes with differential tests.
Written by the indexing model from the issue text.
Description
@jlucaso1 — this builds on your code-size investigation in https://github.com/jlucaso1/buffa/pull/1, and on the binary-size discussion in #270.
A table-driven codec strategy would roughly halve the size of a large generated crate at opt-level = "z", at a run-time cost that depends on the message shape and on which crate the shared interpreter is compiled in. I propose adding it as a configuration option, CodecStrategy, with two values. Unrolled is today's per-message specialised code and stays the default. Table emits one static descriptor per message plus a small shared interpreter, and can be selected globally or for specific message types.
This issue records the measurements from a proof-of-concept and the design decisions that follow from them. The proof-of-concept is a measurement tool and not a candidate implementation, and it has not been pushed.
The per-field code is 731 K of a 1,644 K binary
The generated Message impl for each type contains three per-field code sequences, compute_size, write_to, and merge_field, plus merge_to_limit and merge_length_delimited wrappers. I measured the whatsapp waproto schema (334 top-level messages, 3,477 fields, owned types only) built at opt-level = "z". These five functions total 731 K of a 1,644 K binary: merge_field 211 K, merge_to_limit 174 K, write_to 147 K, compute_size 131 K, and merge_length_delimited 67 K. Across 3,477 fields, that is about 210 bytes per field. The changes in jlucaso1/buffa#1 reduce the duplication inside each of these sequences and in the wrappers around them. The table strategy removes the per-field sequences themselves.
The table strategy replaces per-field code with 12 bytes per field and one shared interpreter
Each message gets a static MessageTable holding a sorted array of 12-byte entries {tag, offset, kind, tag_len, aux}, a dense number-to-index lookup for small field numbers, and the offset of the unknown-fields slot. kind is one of 80 values (field type crossed with cardinality), so the interpreter dispatches through a single jump table per field. Field offsets come from core::mem::offset_of!. Message-typed, repeated, and enum fields have a small per-field-type accessor (place, push, get, set), referenced from aux. The generated Message impl forwards compute_size, write_to, merge_field, merge_to_limit, and merge_length_delimited to interpreters in buffa::table. Decode runs over a contiguous &[u8], gathering a non-contiguous Buf once first.
At z the five per-field functions shrink from 731 K to about 100 K: 42 K of tables (12 bytes per field), 13 K of shared interpreter, and 45 K of per-type accessors, which comes to about 29 bytes per field. Drop glue (37 K) and Default construction stay per message and set the size that remains after this change.
Table roughly halves the binary at z
Whatsapp waproto, x86-64, fat LTO, codegen-units = 1, panic = "abort", static relocation model, text + data in units of 1,000 bytes. "Main" is v0.9.2 (main was within 16 bytes of it). "PR minus outlined writers" is jlucaso1/buffa#1 with the outlined-writer change omitted, because outlining the writers cost 8–35% of encode throughput at O3 in my measurements.
| opt-level | main | PR minus outlined writers | Table |
|---|---|---|---|
| z | 1,644 K | 1,328 K | 817 K |
| s | 2,479 K | 2,135 K | 874 K |
| 3 | 3,408 K | 3,159 K | 960 K |
The size runs flatten the schema's roughly 150 oneof members into optional fields and its 3 map fields into repeated entry messages, because the proof-of-concept does not yet support either. Both transformations preserve the wire format, and they are applied to every column so the comparison is like for like.
Table costs up to 3.5x on dense small messages and nothing on bulk data
The interpreter runs about 10 instructions of dispatch per field where the unrolled code runs about 2, so the cost concentrates in messages made of many small scalar fields and disappears in messages dominated by bulk data. On google_message1 encode, perf stat shows an IPC of 4.1 and 72 K branch misses in 8.7 G branches, so the loss is instruction count and not misprediction.
These were measured on an AWS c7i.metal-24xl spot instance with a pinned core, each benchmark built under three code-alignment layouts. As a control, I also built the unrolled binary under the other two layouts. It differs from the first layout by at most 2% on 35 of 54 measurements and at most 7% on 48 of 54, and the extremes are 0.91 and 1.16 (log_record and analytics_event decode). Ratios of 1.1 or below are within layout noise. Ratios below are time with Table divided by time with Unrolled, both built at O3, averaged over the three layouts. The three shapes marked * were built with oneofs and maps flattened as described under the size table, so hash-map costs are not measured.
| shape | decode | encode | compute_size |
|---|---|---|---|
| google_message1 (42 fields) | 1.40 | 3.49 | 3.34 |
| api_response | 1.07 | 1.50 | 1.76 |
| packed_tile | 1.06 | 1.10 | 1.11 |
| mesh (tiny nested messages) | 1.60 (median) | 2.18 | 2.72 |
| column_batch | 0.98 | 0.95 | 1.01 |
| analytics_event* | 1.14 | 1.68 | 2.11 |
| log_record* | 1.10 | 1.29 | 0.85 |
| media_frame* | 1.01 | 1.13 | 1.16 |
| whatsapp corpus (2,657 records, average 55 bytes) | 1.00 | 1.78 | 2.27 |
Compiling the interpreter at O3 recovers most of the speed for 42 K
The interpreter's encode path is generic over the sink, and merge_to_limit is generic over the input Buf, so these functions are compiled in the crate that calls them, at that crate's opt-level. A size-constrained user builds at z, so the interpreter is also compiled at z, and that is where Table looked worst: on the whatsapp corpus it took 1.22x for decode, 2.52x for encode, and 2.52x for compute_size against unrolled at O3.
Setting [profile.release.package.buffa] opt-level = 3 while the generated crate stays at z recovers most of it, and the setting survives fat LTO. Encode needed one more change, a non-generic Vec<u8> entry point inside buffa, because the generic write path would otherwise be instantiated back in the user's crate at z. The proof-of-concept detects the sink with a type_name comparison; a real implementation would override encode_to_vec in Table messages to call a non-generic function in buffa. Times are relative to unrolled at O3:
| whatsapp corpus | size | decode | encode | compute_size |
|---|---|---|---|---|
| Unrolled, everything at z | 1,644 K | 1.12 | 1.80 | 1.18 |
| Table, everything at z | 817 K | 1.22 | 2.52 | 2.52 |
Table, buffa at O3, generated crate at z |
859 K | 1.03 | 1.82 | 2.28 |
| Table, everything at O3 | 960 K | 1.00 | 1.78 | 2.27 |
The same configuration on the other shapes, with unrolled at z beside it for reference (one layout, so ±5–10%):
| shape | decode: Unrolled z / Table + O3 runtime | encode: Unrolled z / Table + O3 runtime | compute_size: Unrolled z / Table + O3 runtime |
|---|---|---|---|
| google_message1 | 1.94 / 1.50 | 5.85 / 3.60 | 2.15 / 3.29 |
| api_response | 1.89 / 1.32 | 5.59 / 1.59 | 2.99 / 1.68 |
| packed_tile | 1.36 / 1.10 | 5.20 / 1.10 | 4.47 / 1.10 |
| mesh | 2.68 / 2.76 | 5.73 / 2.12 | 2.40 / 2.67 |
| column_batch | 1.81 / 0.99 | 5.65 / 0.96 | 3.86 / 1.02 |
Unrolled code at z is 1.1–5.9x slower than at O3, so a project that builds everything at z today already pays more than Table with an O3 runtime would cost. The table build is faster on encode for every shape and on decode for every shape except mesh, where the two are equal. It is slower only on compute_size for the small-message shapes (google_message1 1.5x, mesh 1.1x, and the whatsapp corpus 1.9x). The consequence for the configuration is that an Unrolled message that must stay fast needs to be compiled at O3, which under a z build means placing it in a separate crate, because Cargo sets opt-level per crate and not per module.
Decision: a configuration option, Unrolled by default
Unrolled remains the default, so existing users see no change. Table is opt-in, either for every message or for named messages.
Config::new()
.codec_strategy(CodecStrategy::Table)
.codec_strategy_in(CodecStrategy::Unrolled, &[".wa.Message", ".wa.Receipt"])
- The plugin equivalents are
codec_strategy=tableand a repeatablecodec_strategy_in=<path>=unrolled, spelled likeoverride_feature_in, sobuf.gen.yamlopts work. - Paths use the existing prefix grammar (
matches_proto_prefix), nested messages included, and the last matching rule wins. A rule that matches no generated message produces a warning, aspreserve_unknown_fields_indoes. - The option selects only the binary
Messageimplementation. It never changes the encoding, and JSON, text, and view code are generated per message either way. The name iscodec_strategyand notwire_codecso that it cannot be read as changing the wire format. - A message that
Tablecannot handle follows theunbox_oneof_inprecedent. The unsupported cases are oneofs, maps, and groups in the proof-of-concept, and custom pointer, string, or bytes representations. A prefix rule keeps such a messageUnrolledand reports aCodeGenWarning, and a rule that specifies the message by its exact path is an error. - The interpreter requires a
buffafeature because it usesunsafe. A crate with#![forbid(unsafe_code)]cannot opt in. - An
Unrolledchild inside aTableparent needs a small thunk set (size, merge over a slice, write). A static table cannot hold a function that is generic over the sink, so the write thunk would take&mut Vec<u8>and other sinks would encode into a scratchVecandput_sliceit, at the cost of one copy forRopeandBytesMut. ATablechild inside anUnrolledparent needs nothing beyond the existingMessagetrait. The bridge is not built.
Keeping a hot Unrolled message costs about 700 bytes per field at O3 (the difference between the unrolled and table builds divided by the field count), so 20 hot messages of 10 fields cost about 140 K.
The tests so far cover five benchmark schemas and all 334 whatsapp top-level messages
Differential tests against the Unrolled codec on five of the benchmark schemas cover round trips, merging twice, non-contiguous input, unknown fields, truncation, and 3,000 corrupted inputs with identical Ok and Err results. On the whatsapp schema, 16,032 arbitrary-generated instances across all 334 top-level messages match on round trip, encoded_len, double merge, and error result. Re-encoding the flattened benchmark datasets produces identical bytes. Miri, a 32-bit target, and any platform other than x86-64 have not been tried.
What is not yet built or decided
- The generated code uses
offset_of!, stable since Rust 1.77, so the option needsrustversionchecks or a higher MSRV than the current 1.75. - Table messages implement
clear()by assigningDefault, soclear()no longer keeps allocated capacity. - Decode of a non-contiguous
Bufgathers it into one buffer first. - Oneof, map, and group support is not implemented. Oneofs need a per-variant accessor pair. Maps need generic key and value handling that avoids the sink-generic problem above.
- A separate decode and encode setting is possible, since decode is the cheaper half (1.0–1.5x against up to 3.6x for encode) and merge is about 60% of the per-field code in main. I would start with one option.
@jlucaso1: what opt-level and target are you building for, and would setting buffa to O3 with the generated crate at z fit your build? Your workload's mix of oneof, map, and dense scalar messages also decides how much of the measured encode cost you would see.
- Dominant language
- Rust
- Stars
- 900
- Forks
- 95
- Avg merge
- 5d 10h
- Merged PRs (30d)
- 57
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from anthropics/buffa
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
anthropics/buffa#487 ·
Maintainers usually reply within 2 days
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
anthropics/buffa#490 ·
Maintainers usually reply within 2 days
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 55/100
anthropics/buffa#482 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
anthropics/buffa#450 ·
Maintainers usually reply within 2 days
-
buffa-build: skip_debug to suppress the generated `impl Debug` for selected messages (prost parity)Open
Difficulty 4/5 3-5 days Newbie friendliness 68/100
anthropics/buffa#448 ·
Maintainers usually reply within 2 days
All issues in anthropics/buffa
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
area: cli bug priority: P2 ready-for-agent
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day