workflowcheck OOM issues on large codebases
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 38/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- java
- 領域
- performance, tooling
調査の方向性
workflowcheck の findWorkflowClasses エントリポイントから始め、ClassPath の列挙、Loader.loadClass のキャッシュ、findWorkflowImplInfo、processMethodValidity を追跡します。scan-target/resolution-classpath とアノテーションフィルタリングに関する提案を、提案されている --scan-classpath の動作も含めて比較します。大規模なコードベースで eager スキャンと保持されるクラスを削減し、意図した workflow チェックを失わないことが完了の条件です。
索引モデルが issue の本文から書いたものです。
説明
workflowcheck scans the entirety of the runtimeClasspath and stores a cache of all visited classes in a hashmap. For large codebases, this is wasteful and can OOM CI jobs.
I suspect the current behavior was chosen because (1) it's simple, and (2) to potentially capture workflows that are brought in by dependencies. For the second point, it feels like the burden of determinism checks for workflows provided by a library should be on the library provider, rather than the consumer.
I suspect that suggestion 2 below may be the best fix and most-correct, but it might be behaviorally backwards-incompatible from what exists today
Current Behavior
temporal-workflowcheck uses a single classpath concept. When you call findWorkflowClasses(String... classPaths):
- ClassPath constructor enumerates every .class file across all JARs and directories (filtering only
java/,javax/,jdk/,com/sun/). For a larges app, this could be tens of thousands of classes. - The main loop calls
loader.loadClass()on every one of them, using ASM ClassReader to parse bytecode. Each class gets cached inLoader.classes(aHashMap<String, ClassInfo>). - For each class, it checks whether any public non-static non-abstract method is a workflow implementation (by walking the superclass/interface chain via
findWorkflowImplInfo). This triggers moreloadClasscalls. - For workflow methods found,
processMethodValiditytraces the call chain through dependencies to detect non-deterministic operations, triggering yet moreloadClasscalls.
The result for a project with a few workflow classes buried among ~20,000 dependency classes, the tool eagerly loads all 20,000 just to find those few classes. The Loader.classes cache retains every ClassInfo for the lifetime of the scan.
Suggestions
(1) Separate scan targets from resolution classpath
Add a new overload:
public List<ClassInfo> findWorkflowClasses(
String[] scanPaths, // enumerate classes from here
String[] resolutionPaths // available for call-chain resolution only
) throws IOException
ClassPath would enumerate class names only from scanPaths, but create a URLClassLoader covering both. The main loop iterates a fraction of the classes; Loader.loadClass() still works for call-chain tracing because the full classpath is on the classloader.
Instead of eagerly loading 20,000 classes in the discovery loop, only module classes are enumerated. Dependencies are loaded on-demand during processMethodValidity tracing, and only the onesactually in the call chain. The Loader.classes cache might end up with ~1k entries instead of ~20k.
We'd need to add a --scan-classpath <path> flag (repeatable). When present, the tool would only enumerate from those path(s). When absent, current behavior would be preserved (scan everything).
A much less durable alternative would be to support configurable classpath exclusions. While flexible, it'd put a toil-level burden on developers to whack-a-mole classpaths as a module's dependency graph changes.
(2) Lazy load with annotation filtering
Instead of loading every class through ASM's full ClassReader, do a two-pass approach:
- Pass 1: Quick scan of workflow-related annotations (
WorkflowInterface,WorkflowMethod, etc), which can be done by reading the class constant pools instead of the full class. Classes without these annotations (or that don't implement interfaces containing them) would be skipped entirely. - Pass 2: Full ASM analysis only on classes identified by pass 1.
- 主要言語
- Java
- スター
- 433
- フォーク
- 249
- 平均マージ
- 6日 5時間
- マージ済み PR(30日)
- 25
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
temporalio/sdk-java のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
temporalio/sdk-java#2676 · コメント 8 件 · リアクション 2 件 ·
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
temporalio/sdk-java#1825 ·
-
test server
難易度 4/5 3〜5日 初心者へのやさしさ 42/100
temporalio/sdk-java#3088 · コメント 2 件 ·
-
temporalio/sdk-java#3059 · 担当者 1 名 ·
-
enhancement
temporalio/sdk-java#3058 · 担当者 1 名 ·
temporalio/sdk-java の issue をすべて見る
似ている issue
-
area-deployment area-integrations triage:bot-seen
難易度 2/5 半日 初心者へのやさしさ 86/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
apache/flink-agents#1156 ·
-
[source-shopify] FAILED bulk operation without partialDataUrl is silently treated as successful オープンarea/connectors autoteam community connectors/source/shopify needs-triage team/use type/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100