[Bug]: [RunInference] max_models_per_worker_hint is not enforced
维护者通常 1 天内回复
评估
调研方向
Start in sdks/python/apache_beam/ml/inference/base.py around the max-workers logic at lines 852–857, and trace how locks and deserialized RunInference handlers interact with the shared _ModelHandlerManager. Use the issue’s reproduction as a starting point and check existing Python SDK inference tests. Done when the model limit remains at the configured hint across multiple handler copies and the regression is covered by a test.
由索引模型根据 Issue 内容生成。
描述
What happened?
The intent of max_models_per_worker hint is to limit the number of models that will be loaded per SDK process.
The logic to increment allowed max number of workers https://github.com/apache/beam/blob/dabcf50ffbf532bb30a1233927cd1edb6ae067bb/sdks/python/apache_beam/ml/inference/base.py#L852-L857 appears to be flawed since:
- The lock acquisition always succeeds (it's a new lock instance)
- If there is more than 1 process bundle descriptor over the life time of the SDK process, we might have more than 1 copy of the deserialized RunInference DoFn with unpickled ModelHandler instance, which won't persist
self._max_models_per_worker_hint = Nonefrom a prior initialization:
AI repro:
from apache_beam.internal import pickler
from apache_beam.ml.inference import base
class Model:
def predict(self, x):
return x
class Handler(base.ModelHandler):
def load_model(self):
return Model()
def run_inference(self, batch, model, inference_args=None):
return [model.predict(x) for x in batch]
mhs = [base.KeyModelMapping([k], Handler()) for k in ('a', 'b', 'c')]
keyed_handler = base.KeyedModelHandler(mhs, max_models_per_worker_hint=1)
# RunInference shares one _ModelHandlerManager per transform across all
# DoFn instances (and processes) via MultiProcessShared.
manager = keyed_handler.load_model()
# 5 DoFn instances in one process (harness threads, re-created bundle
# processors); each deserializes its own copy of the model handler.
for _ in range(5):
handler_copy = pickler.roundtrip(keyed_handler)
handler_copy.override_metrics('ns')
handler_copy.run_inference([('a', 1), ('b', 2), ('c', 3)], manager)
print('model limit:', manager._max_models) # 5, expected 1
print('models in memory:', len(manager._tag_map)) # 3
Issue Priority
Priority: 2 (default / most bugs should be filed as P2)
Issue Components
- Component: Python SDK
- Component: Java SDK
- Component: Go SDK
- Component: Typescript SDK
- Component: IO connector
- Component: Beam YAML
- Component: Beam examples
- Component: Beam playground
- Component: Beam katas
- Component: Website
- Component: Infrastructure
- Component: Spark Runner
- Component: Flink Runner
- Component: Prism Runner
- Component: Twister2 Runner
- Component: Hazelcast Jet Runner
- Component: Google Cloud Dataflow Runner
- 主要语言
- Java
- 星标
- 8.7k
- 派生
- 4.7k
- 平均合并
- 2 天 7 小时
- 30 天内合并 PR
- 242
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
apache/beam 的其他 Issue
-
[Bug]: Row.toString throws for an ITERABLE field that is not backed by a List可能已有人在做 @PDGGK 于 58 天前认领。 未关闭java P3
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
apache/beam#39624 · 2 个 reaction ·
维护者通常 1 天内回复
-
[Bug]: PubsubIO used in batch incorrect batch cutoff size可能已有人在做 @1fanwang 于 47 天前认领。 未关闭bug io P3 pinned pubsub
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
build P3 sub-task
难度 1/5 1 小时以内 新手友好度 68/100
维护者通常 1 天内回复
-
bug gcp io java P3
难度 1/5 1 小时以内 新手友好度 65/100
维护者通常 1 天内回复
相似的 Issue
-
BoxAttachmentMulti parsing leaks IOException / ArrayIndexOutOfBoundsException on malformed content instead of IllegalArgumentException可能已有人在做 @Kshot3000 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 74/100
ergoplatform/ergo-appkit#272 ·
-
难度 2/5 1-3 小时 新手友好度 64/100
-
难度 2/5 1-3 小时 新手友好度 66/100
-
难度 2/5 1-3 小时 新手友好度 64/100
utopia-rise/godot-jvm#1004 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 82/100
spring-projects/spring-grpc#442 ·