Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[RFC] ExecuTorch Persisting Device Specialized Delegate Artifacts

Aperta
#23,192 1 commento 1 reazione 1 assegnatario Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@JacobSzwejbka ci sta già lavorando.

Dal 28/9/2026.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

🚀 The feature, motivation and pitch

Problem

Some delegates receive portable weights and transform them during runtime initialization (XNNPack) others perform more extensive jit compilation. Repeating that work on every model load increases initialization latency. For XNNPack this is generally bearable though undesirable, but for other upcoming backends/features it will not be tolerable. Every backend could come up with its own individual caching solution, but 1. That's the sort of problem a framework should solve and 2. for LLMs or other large models this is still bad because we are wasting a ton of disk size.

For deployments where a model is tied to a particular device configuration, the runtime should be able to replace the portable representation with the processed representation.

These initialization operations only need to occur once per model per device.

Requirements

  • Support a delegate returning replacements for values identified by NamedDataMap key;
  • Reclaim space by rewriting and compacting the original serialized container
  • Have user defined writer api as a parallel to dataloader
  • Don’t break BC
  • Minimal regression for embedded users who don’t need this.

Proposal

The app calls a one-time per model per device API. Inside the API ExecuTorch finds each backend in the model and calls that backends initialize_backend_data once. Through a dedicated context, the backend walks its delegate blobs one at a time as well as any named data blobs it needs. If the backend produces replacement bytes, it submits them; otherwise ExecuTorch writes the original blob as it advances.

If the write fails before finishing the original file is untouched and we must start over. No checkpoints or state resuming are in scope.

Each blob is written exactly once at its final location.

After all blobs are processed, ExecuTorch copies the original FlatBuffer metadata and changes only existing segment offsets and sizes. Keys, segment count, and table structure cannot change, so the FlatBuffer should remain exactly the same size.

Finally, ExecuTorch publishes the completed temporary PTE. Future normal model loads read the already-prepared backend data and avoid preprocessing the weights again. Backends need to know on their own if they are reading device compiled or ahead of time generated blobs.

Alternative Solutions Rejected

Modify the existing delegate init to optionally allow it to write. This has quite a few problems. At the time of method load there are a lot of data structures involved that assume the underlying data is read only. Init is also run many many times over the life of a model on device when we would want to run the write back only once per model per device. We also have encouraged teams to reuse a Program object to load the same method on multiple threads if needed which introduces a reader writer and staleness problem into the mix. It's simpler if we just pursue an entirely separate api and tell the user to make sure of its exclusive access at the time.

I also elected to ignore in-place updates to the original file. I think it would require a more complicated solution and the cost of dying mid in-place update is pretty high. For most applications temporary disk pressure from the tmp file -> rename solution should be fine I think. That solution also lets a user /not/ override the original if you really don’t care about disk and want to preserve the original incase you need to rerun this again later (xnnpack runtime version update)

Layers

  • The backend only sees processed bytes, existing named keys, and compile specs
  • The user calls the coordinator layer which reconstructs the new .pte by iterating over unique backends and calls their relevant apis. The BackendDataWriter handles staging the updates to the pte through calls to Datawriter
  • DataWriter is ET agnostic and just writes provided bytes to the specified offsets

Backend API

BackendInterface gains one optional method. Existing backends inherit the default NotSupported implementation.

struct BackendData {
  Span<const uint8_t> bytes;
  std::optional<size_t> alignment;
  std::optional<TensorLayout> layout
};

struct BackendData {
  string_view key; // Must already exist in the supplied PTE.
  variant<BackendData, TensorData> data;
};

struct BackendDataInput {
  string_view method_name;
  FreeableBuffer processed;
  ArrayRef<CompileSpec> compile_specs;
};

class BackendDataInitContext {
 public:
  // Lazily loads and returns one input. nullopt means end of input.
  Result<std::optional<BackendDataInput>> next_backend_data();

  // Returns the model-wide named-data map used by normal backend init.
  const NamedDataMap* get_named_data_map() const;

  // Returns temporary storage that ET may reset on next_backend_data().
  MemoryAllocator* get_temp_allocator();
};

class BackendDataWriter {
 public:
  virtual ~BackendDataWriter() = default;

  // Replaces values for existing named-data pairs. Keys may not be added,
  // removed, or renamed.
  virtual Error write_named_data(
      Span<string_view>, Span<const NamedBackendData> values) = 0;
};

class BackendInterface {
 public:
  virtual Error initialize_backend_data(
      BackendDataInitContext& context,
      BackendDataWriter& output) const {
    return Error::NotSupported;
  }
};

Storage API

DataWriter handles writes to one PTE. When constructing one a user can decide if this will eventually delete the original pte on disk.

class DataWriter {
 public:
  virtual ~DataWriter() = default;

  // Synchronously copies size bytes to offset in the unpublished output.
  virtual Error write(
      const void* data,
      size_t size,
      size_t offset) = 0;

  // Completes the output and makes it visible. No writes are allowed after a
  // successful call.
  virtual Error publish() = 0;

A writer destroyed before successful publish() discards its unpublished output.

Coordinator Layer

Error initialize_and_save_backend_data(
    std::unique_ptr<DataLoader> pte_loader,
    std::unique_ptr<DataWriter> pte_writer,
    MemoryAllocator* allocator,
    EventTracer* event_tracer = nullptr);

Writing New PTE Flatbuffer

If we enforce the keyset to not change then this is pretty simple. The new flatbuffer segments are byte identical except the fields indicating a segments size or offset and the header indicating the total segments size. We probably can just update those with flatbuffers built in mutable apis and skip changing anything else.

If we do end up needing to modify the keyset some other things in the earlier api will need to change, but here we will choose a “guess” for the flatbuffer size and start serializing the segments from there, then if the real flatbuffer is smaller then the guess we can just have padding in the empty space. If the real flatbuffer is larger then we have to shift everything down and that sucks. I would imagine that if we have a new keyset its because weights are getting fused though, so the total size of the flatbuffer section should decrease. It would be uncommon for the guess to be too small.

Performance & Resource Comparison

Strategy Disk Size Needed Init Time
JIT Compile Every Time 1x Slow
JIT Compile Once (Lose/Replace Original) ~1x after processing 2x needed during JIT compilation process. Fast with one time cost
JIT Compile Once (Save Original) ~2x Fast with one time cost

Extensions

If other backends need we can extend this to allow rewriting the processed bytes or writing to the ptd.

Alternatives

No response

Additional context

No response

RFC (Optional)

No response

cc @larryliu0820 @lucylq

Lingua principale
Python
Stelle
5k
Fork
1.2k
Merge medio
2g 7h
PR unite (30g)
521

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di pytorch/executorch

Tutte le issue di pytorch/executorch

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.