Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

[Bug]: [RunInference] max_models_per_worker_hint is not enforced

Aberta
#40,468 0 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
4/5
Tempo estimado
3-5 dias
Facilidade para iniciantes
48/100
Tipo de issue
Bug
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
python
Domínio
backend

Direção de pesquisa

Start in sdks/python/apache_beam/ml/inference/base.py around the max-workers logic at lines 852–857, and trace how locks and deserialized RunInference handlers interact with the shared _ModelHandlerManager. Use the issue’s reproduction as a starting point and check existing Python SDK inference tests. Done when the model limit remains at the configured hint across multiple handler copies and the regression is covered by a test.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

awaiting triage bug P2 python
What happened?

The intent of max_models_per_worker hint is to limit the number of models that will be loaded per SDK process.

The logic to increment allowed max number of workers https://github.com/apache/beam/blob/dabcf50ffbf532bb30a1233927cd1edb6ae067bb/sdks/python/apache_beam/ml/inference/base.py#L852-L857 appears to be flawed since:

  1. The lock acquisition always succeeds (it's a new lock instance)
  2. If there is more than 1 process bundle descriptor over the life time of the SDK process, we might have more than 1 copy of the deserialized RunInference DoFn with unpickled ModelHandler instance, which won't persist self._max_models_per_worker_hint = None from a prior initialization:

AI repro:

from apache_beam.internal import pickler
from apache_beam.ml.inference import base
class Model:
  def predict(self, x):
    return x
class Handler(base.ModelHandler):
  def load_model(self):
    return Model()
  def run_inference(self, batch, model, inference_args=None):
    return [model.predict(x) for x in batch]
mhs = [base.KeyModelMapping([k], Handler()) for k in ('a', 'b', 'c')]
keyed_handler = base.KeyedModelHandler(mhs, max_models_per_worker_hint=1)
# RunInference shares one _ModelHandlerManager per transform across all
# DoFn instances (and processes) via MultiProcessShared.
manager = keyed_handler.load_model()
# 5 DoFn instances in one process (harness threads, re-created bundle
# processors); each deserializes its own copy of the model handler.
for _ in range(5):
  handler_copy = pickler.roundtrip(keyed_handler)
  handler_copy.override_metrics('ns')
  handler_copy.run_inference([('a', 1), ('b', 2), ('c', 3)], manager)
print('model limit:', manager._max_models)  # 5, expected 1
print('models in memory:', len(manager._tag_map))  # 3

Issue Priority

Priority: 2 (default / most bugs should be filed as P2)

Issue Components
  • Component: Python SDK
  • Component: Java SDK
  • Component: Go SDK
  • Component: Typescript SDK
  • Component: IO connector
  • Component: Beam YAML
  • Component: Beam examples
  • Component: Beam playground
  • Component: Beam katas
  • Component: Website
  • Component: Infrastructure
  • Component: Spark Runner
  • Component: Flink Runner
  • Component: Prism Runner
  • Component: Twister2 Runner
  • Component: Hazelcast Jet Runner
  • Component: Google Cloud Dataflow Runner
Linguagem predominante
Java
Estrelas
8.7k
Forks
4.7k
Merge médio
2d 8h
PRs com merge (30d)
246

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de apache/beam

Todas as issues de apache/beam

Issues semelhantes

Mais issues de Java

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.