[Bug]: [RunInference] max_models_per_worker_hint is not enforced
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 48/100
Direção de pesquisa
Start in sdks/python/apache_beam/ml/inference/base.py around the max-workers logic at lines 852–857, and trace how locks and deserialized RunInference handlers interact with the shared _ModelHandlerManager. Use the issue’s reproduction as a starting point and check existing Python SDK inference tests. Done when the model limit remains at the configured hint across multiple handler copies and the regression is covered by a test.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
What happened?
The intent of max_models_per_worker hint is to limit the number of models that will be loaded per SDK process.
The logic to increment allowed max number of workers https://github.com/apache/beam/blob/dabcf50ffbf532bb30a1233927cd1edb6ae067bb/sdks/python/apache_beam/ml/inference/base.py#L852-L857 appears to be flawed since:
- The lock acquisition always succeeds (it's a new lock instance)
- If there is more than 1 process bundle descriptor over the life time of the SDK process, we might have more than 1 copy of the deserialized RunInference DoFn with unpickled ModelHandler instance, which won't persist
self._max_models_per_worker_hint = Nonefrom a prior initialization:
AI repro:
from apache_beam.internal import pickler
from apache_beam.ml.inference import base
class Model:
def predict(self, x):
return x
class Handler(base.ModelHandler):
def load_model(self):
return Model()
def run_inference(self, batch, model, inference_args=None):
return [model.predict(x) for x in batch]
mhs = [base.KeyModelMapping([k], Handler()) for k in ('a', 'b', 'c')]
keyed_handler = base.KeyedModelHandler(mhs, max_models_per_worker_hint=1)
# RunInference shares one _ModelHandlerManager per transform across all
# DoFn instances (and processes) via MultiProcessShared.
manager = keyed_handler.load_model()
# 5 DoFn instances in one process (harness threads, re-created bundle
# processors); each deserializes its own copy of the model handler.
for _ in range(5):
handler_copy = pickler.roundtrip(keyed_handler)
handler_copy.override_metrics('ns')
handler_copy.run_inference([('a', 1), ('b', 2), ('c', 3)], manager)
print('model limit:', manager._max_models) # 5, expected 1
print('models in memory:', len(manager._tag_map)) # 3
Issue Priority
Priority: 2 (default / most bugs should be filed as P2)
Issue Components
- Component: Python SDK
- Component: Java SDK
- Component: Go SDK
- Component: Typescript SDK
- Component: IO connector
- Component: Beam YAML
- Component: Beam examples
- Component: Beam playground
- Component: Beam katas
- Component: Website
- Component: Infrastructure
- Component: Spark Runner
- Component: Flink Runner
- Component: Prism Runner
- Component: Twister2 Runner
- Component: Hazelcast Jet Runner
- Component: Google Cloud Dataflow Runner
- Linguagem predominante
- Java
- Estrelas
- 8.7k
- Forks
- 4.7k
- Merge médio
- 2d 8h
- PRs com merge (30d)
- 246
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de apache/beam
-
[Bug]: Row.toString throws for an ITERABLE field that is not backed by a ListTalvez já em andamento @PDGGK assumiu há 56 dias. Abertajava P3
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 76/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
apache/beam#39624 · 2 reações ·
Mantenedores costumam responder em até 1 dia
-
[Failing Test]: JmsIOTest. testCheckpointMark flakyTalvez já em andamento @mxtymoshyk assumiu há 14 dias. Abertabug failing test flake P2 pinned tests
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
apache/beam#30225 · 2 comentários ·
Mantenedores costumam responder em até 1 dia
-
[Bug]: PubsubIO used in batch incorrect batch cutoff sizeTalvez já em andamento @1fanwang assumiu há 45 dias. Abertabug io P3 pinned pubsub
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
apache/beam#28011 · 4 comentários ·
Mantenedores costumam responder em até 1 dia
-
build P3 sub-task
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
Todas as issues de apache/beam
Issues semelhantes
-
[BUG] S3 CORS responses omit Access-Control-Allow-Credentials for matched originsTalvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
floci-io/floci#5369 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
sqlcipher/sqlcipher-android#97 · 1 comentário ·
-
area-integrations
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
bug IIIF interoperability
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100