Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Design: table-driven CodecStrategy to shrink generated code

Open
#463 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
rust

Research direction

Start by reading the existing Config API, generated Message implementation, and the buffa::table entry point described in the issue, then review the proof-of-concept measurements. Done means adding the opt-in CodecStrategy configuration with Unrolled as the default, defining the documented fallback and warning behavior, and covering supported and unsupported message shapes with differential tests.

Written by the indexing model from the issue text.

Description

enhancement

@jlucaso1 — this builds on your code-size investigation in https://github.com/jlucaso1/buffa/pull/1, and on the binary-size discussion in #270.

A table-driven codec strategy would roughly halve the size of a large generated crate at opt-level = "z", at a run-time cost that depends on the message shape and on which crate the shared interpreter is compiled in. I propose adding it as a configuration option, CodecStrategy, with two values. Unrolled is today's per-message specialised code and stays the default. Table emits one static descriptor per message plus a small shared interpreter, and can be selected globally or for specific message types.

This issue records the measurements from a proof-of-concept and the design decisions that follow from them. The proof-of-concept is a measurement tool and not a candidate implementation, and it has not been pushed.

The per-field code is 731 K of a 1,644 K binary

The generated Message impl for each type contains three per-field code sequences, compute_size, write_to, and merge_field, plus merge_to_limit and merge_length_delimited wrappers. I measured the whatsapp waproto schema (334 top-level messages, 3,477 fields, owned types only) built at opt-level = "z". These five functions total 731 K of a 1,644 K binary: merge_field 211 K, merge_to_limit 174 K, write_to 147 K, compute_size 131 K, and merge_length_delimited 67 K. Across 3,477 fields, that is about 210 bytes per field. The changes in jlucaso1/buffa#1 reduce the duplication inside each of these sequences and in the wrappers around them. The table strategy removes the per-field sequences themselves.

The table strategy replaces per-field code with 12 bytes per field and one shared interpreter

Each message gets a static MessageTable holding a sorted array of 12-byte entries {tag, offset, kind, tag_len, aux}, a dense number-to-index lookup for small field numbers, and the offset of the unknown-fields slot. kind is one of 80 values (field type crossed with cardinality), so the interpreter dispatches through a single jump table per field. Field offsets come from core::mem::offset_of!. Message-typed, repeated, and enum fields have a small per-field-type accessor (place, push, get, set), referenced from aux. The generated Message impl forwards compute_size, write_to, merge_field, merge_to_limit, and merge_length_delimited to interpreters in buffa::table. Decode runs over a contiguous &[u8], gathering a non-contiguous Buf once first.

At z the five per-field functions shrink from 731 K to about 100 K: 42 K of tables (12 bytes per field), 13 K of shared interpreter, and 45 K of per-type accessors, which comes to about 29 bytes per field. Drop glue (37 K) and Default construction stay per message and set the size that remains after this change.

Table roughly halves the binary at z

Whatsapp waproto, x86-64, fat LTO, codegen-units = 1, panic = "abort", static relocation model, text + data in units of 1,000 bytes. "Main" is v0.9.2 (main was within 16 bytes of it). "PR minus outlined writers" is jlucaso1/buffa#1 with the outlined-writer change omitted, because outlining the writers cost 8–35% of encode throughput at O3 in my measurements.

opt-level main PR minus outlined writers Table
z 1,644 K 1,328 K 817 K
s 2,479 K 2,135 K 874 K
3 3,408 K 3,159 K 960 K

The size runs flatten the schema's roughly 150 oneof members into optional fields and its 3 map fields into repeated entry messages, because the proof-of-concept does not yet support either. Both transformations preserve the wire format, and they are applied to every column so the comparison is like for like.

Table costs up to 3.5x on dense small messages and nothing on bulk data

The interpreter runs about 10 instructions of dispatch per field where the unrolled code runs about 2, so the cost concentrates in messages made of many small scalar fields and disappears in messages dominated by bulk data. On google_message1 encode, perf stat shows an IPC of 4.1 and 72 K branch misses in 8.7 G branches, so the loss is instruction count and not misprediction.

These were measured on an AWS c7i.metal-24xl spot instance with a pinned core, each benchmark built under three code-alignment layouts. As a control, I also built the unrolled binary under the other two layouts. It differs from the first layout by at most 2% on 35 of 54 measurements and at most 7% on 48 of 54, and the extremes are 0.91 and 1.16 (log_record and analytics_event decode). Ratios of 1.1 or below are within layout noise. Ratios below are time with Table divided by time with Unrolled, both built at O3, averaged over the three layouts. The three shapes marked * were built with oneofs and maps flattened as described under the size table, so hash-map costs are not measured.

shape decode encode compute_size
google_message1 (42 fields) 1.40 3.49 3.34
api_response 1.07 1.50 1.76
packed_tile 1.06 1.10 1.11
mesh (tiny nested messages) 1.60 (median) 2.18 2.72
column_batch 0.98 0.95 1.01
analytics_event* 1.14 1.68 2.11
log_record* 1.10 1.29 0.85
media_frame* 1.01 1.13 1.16
whatsapp corpus (2,657 records, average 55 bytes) 1.00 1.78 2.27

Compiling the interpreter at O3 recovers most of the speed for 42 K

The interpreter's encode path is generic over the sink, and merge_to_limit is generic over the input Buf, so these functions are compiled in the crate that calls them, at that crate's opt-level. A size-constrained user builds at z, so the interpreter is also compiled at z, and that is where Table looked worst: on the whatsapp corpus it took 1.22x for decode, 2.52x for encode, and 2.52x for compute_size against unrolled at O3.

Setting [profile.release.package.buffa] opt-level = 3 while the generated crate stays at z recovers most of it, and the setting survives fat LTO. Encode needed one more change, a non-generic Vec<u8> entry point inside buffa, because the generic write path would otherwise be instantiated back in the user's crate at z. The proof-of-concept detects the sink with a type_name comparison; a real implementation would override encode_to_vec in Table messages to call a non-generic function in buffa. Times are relative to unrolled at O3:

whatsapp corpus size decode encode compute_size
Unrolled, everything at z 1,644 K 1.12 1.80 1.18
Table, everything at z 817 K 1.22 2.52 2.52
Table, buffa at O3, generated crate at z 859 K 1.03 1.82 2.28
Table, everything at O3 960 K 1.00 1.78 2.27

The same configuration on the other shapes, with unrolled at z beside it for reference (one layout, so ±5–10%):

shape decode: Unrolled z / Table + O3 runtime encode: Unrolled z / Table + O3 runtime compute_size: Unrolled z / Table + O3 runtime
google_message1 1.94 / 1.50 5.85 / 3.60 2.15 / 3.29
api_response 1.89 / 1.32 5.59 / 1.59 2.99 / 1.68
packed_tile 1.36 / 1.10 5.20 / 1.10 4.47 / 1.10
mesh 2.68 / 2.76 5.73 / 2.12 2.40 / 2.67
column_batch 1.81 / 0.99 5.65 / 0.96 3.86 / 1.02

Unrolled code at z is 1.1–5.9x slower than at O3, so a project that builds everything at z today already pays more than Table with an O3 runtime would cost. The table build is faster on encode for every shape and on decode for every shape except mesh, where the two are equal. It is slower only on compute_size for the small-message shapes (google_message1 1.5x, mesh 1.1x, and the whatsapp corpus 1.9x). The consequence for the configuration is that an Unrolled message that must stay fast needs to be compiled at O3, which under a z build means placing it in a separate crate, because Cargo sets opt-level per crate and not per module.

Decision: a configuration option, Unrolled by default

Unrolled remains the default, so existing users see no change. Table is opt-in, either for every message or for named messages.

Config::new()
    .codec_strategy(CodecStrategy::Table)
    .codec_strategy_in(CodecStrategy::Unrolled, &[".wa.Message", ".wa.Receipt"])
  • The plugin equivalents are codec_strategy=table and a repeatable codec_strategy_in=<path>=unrolled, spelled like override_feature_in, so buf.gen.yaml opts work.
  • Paths use the existing prefix grammar (matches_proto_prefix), nested messages included, and the last matching rule wins. A rule that matches no generated message produces a warning, as preserve_unknown_fields_in does.
  • The option selects only the binary Message implementation. It never changes the encoding, and JSON, text, and view code are generated per message either way. The name is codec_strategy and not wire_codec so that it cannot be read as changing the wire format.
  • A message that Table cannot handle follows the unbox_oneof_in precedent. The unsupported cases are oneofs, maps, and groups in the proof-of-concept, and custom pointer, string, or bytes representations. A prefix rule keeps such a message Unrolled and reports a CodeGenWarning, and a rule that specifies the message by its exact path is an error.
  • The interpreter requires a buffa feature because it uses unsafe. A crate with #![forbid(unsafe_code)] cannot opt in.
  • An Unrolled child inside a Table parent needs a small thunk set (size, merge over a slice, write). A static table cannot hold a function that is generic over the sink, so the write thunk would take &mut Vec<u8> and other sinks would encode into a scratch Vec and put_slice it, at the cost of one copy for Rope and BytesMut. A Table child inside an Unrolled parent needs nothing beyond the existing Message trait. The bridge is not built.

Keeping a hot Unrolled message costs about 700 bytes per field at O3 (the difference between the unrolled and table builds divided by the field count), so 20 hot messages of 10 fields cost about 140 K.

The tests so far cover five benchmark schemas and all 334 whatsapp top-level messages

Differential tests against the Unrolled codec on five of the benchmark schemas cover round trips, merging twice, non-contiguous input, unknown fields, truncation, and 3,000 corrupted inputs with identical Ok and Err results. On the whatsapp schema, 16,032 arbitrary-generated instances across all 334 top-level messages match on round trip, encoded_len, double merge, and error result. Re-encoding the flattened benchmark datasets produces identical bytes. Miri, a 32-bit target, and any platform other than x86-64 have not been tried.

What is not yet built or decided

  • The generated code uses offset_of!, stable since Rust 1.77, so the option needs rustversion checks or a higher MSRV than the current 1.75.
  • Table messages implement clear() by assigning Default, so clear() no longer keeps allocated capacity.
  • Decode of a non-contiguous Buf gathers it into one buffer first.
  • Oneof, map, and group support is not implemented. Oneofs need a per-variant accessor pair. Maps need generic key and value handling that avoids the sink-generic problem above.
  • A separate decode and encode setting is possible, since decode is the cheaper half (1.0–1.5x against up to 3.6x for encode) and merge is about 60% of the per-field code in main. I would start with one option.

@jlucaso1: what opt-level and target are you building for, and would setting buffa to O3 with the generated crate at z fit your build? Your workload's mix of oneof, map, and dense scalar messages also decides how much of the measured encode cost you would see.

Dominant language
Rust
Stars
900
Forks
95
Avg merge
5d 10h
Merged PRs (30d)
57

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from anthropics/buffa

All issues in anthropics/buffa

Similar issues

More Rust issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.