Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

workflowcheck OOM issues on large codebases

未关闭
#2,846 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
38/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
冷清
技术栈
java

调研方向

从 workflowcheck 的 findWorkflowClasses 入口点开始,跟踪 ClassPath 枚举、Loader.loadClass 缓存、findWorkflowImplInfo 和 processMethodValidity。比较关于 scan-target/resolution-classpath 和注解过滤的建议,包括提议的 --scan-classpath 行为。完成的标准是在大型代码库中减少急切扫描和保留的类,同时不丢失预期的 workflow 检查。

由索引模型根据 Issue 内容生成。

描述

workflowcheck scans the entirety of the runtimeClasspath and stores a cache of all visited classes in a hashmap. For large codebases, this is wasteful and can OOM CI jobs.

I suspect the current behavior was chosen because (1) it's simple, and (2) to potentially capture workflows that are brought in by dependencies. For the second point, it feels like the burden of determinism checks for workflows provided by a library should be on the library provider, rather than the consumer.

I suspect that suggestion 2 below may be the best fix and most-correct, but it might be behaviorally backwards-incompatible from what exists today

Current Behavior

temporal-workflowcheck uses a single classpath concept. When you call findWorkflowClasses(String... classPaths):

  1. ClassPath constructor enumerates every .class file across all JARs and directories (filtering only java/, javax/, jdk/, com/sun/). For a larges app, this could be tens of thousands of classes.
  2. The main loop calls loader.loadClass() on every one of them, using ASM ClassReader to parse bytecode. Each class gets cached in Loader.classes (a HashMap<String, ClassInfo>).
  3. For each class, it checks whether any public non-static non-abstract method is a workflow implementation (by walking the superclass/interface chain via findWorkflowImplInfo). This triggers more loadClass calls.
  4. For workflow methods found, processMethodValidity traces the call chain through dependencies to detect non-deterministic operations, triggering yet more loadClass calls.

The result for a project with a few workflow classes buried among ~20,000 dependency classes, the tool eagerly loads all 20,000 just to find those few classes. The Loader.classes cache retains every ClassInfo for the lifetime of the scan.

Suggestions

(1) Separate scan targets from resolution classpath

Add a new overload:

public List<ClassInfo> findWorkflowClasses(
  String[] scanPaths,         // enumerate classes from here
  String[] resolutionPaths    // available for call-chain resolution only
) throws IOException

ClassPath would enumerate class names only from scanPaths, but create a URLClassLoader covering both. The main loop iterates a fraction of the classes; Loader.loadClass() still works for call-chain tracing because the full classpath is on the classloader.

Instead of eagerly loading 20,000 classes in the discovery loop, only module classes are enumerated. Dependencies are loaded on-demand during processMethodValidity tracing, and only the onesactually in the call chain. The Loader.classes cache might end up with ~1k entries instead of ~20k.

We'd need to add a --scan-classpath <path> flag (repeatable). When present, the tool would only enumerate from those path(s). When absent, current behavior would be preserved (scan everything).

A much less durable alternative would be to support configurable classpath exclusions. While flexible, it'd put a toil-level burden on developers to whack-a-mole classpaths as a module's dependency graph changes.

(2) Lazy load with annotation filtering

Instead of loading every class through ASM's full ClassReader, do a two-pass approach:

  • Pass 1: Quick scan of workflow-related annotations (WorkflowInterface, WorkflowMethod, etc), which can be done by reading the class constant pools instead of the full class. Classes without these annotations (or that don't implement interfaces containing them) would be skipped entirely.
  • Pass 2: Full ASM analysis only on classes identified by pass 1.
主要语言
Java
星标
434
派生
252
平均合并
6 天 5 小时
30 天内合并 PR
25

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

temporalio/sdk-java 的其他 Issue

查看 temporalio/sdk-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。