[Bug][tempo] Worklog collector only fetches the first 1000 worklogs of a rolling 90-day window; converted ids don't join jira
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 45/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
- 技術スタック
- go
調査の方向性
The bug is in backend/plugins/tempo/tasks/worklog_collector.go. Start by reading the collector's pagination logic and the Tempo v4 API spec. Check how the time range is set for team scopes and how raw params are shared. For the conversion issues, examine the ID generation in the worklog converter. Run the existing tests and create a test connection to verify fixes.
索引モデルが issue の本文から書いたものです。
説明
Search before asking
- I had searched in the issues and found no similar issues.
What happened
On a Tempo connection with team scopes, collect_worklogs never collects more than 1000 worklogs per team, and they are always the oldest ones of a rolling 90-day window — recent worklogs never arrive. Seen on v1.0.3-beta15; the collector is unchanged on main.
Evidence from _raw_tempo_api_worklogs after ~50 daily runs:
- every request URL has
offset=0, e.g.https://api.tempo.io/4/worklogs/team/4?from=2026-06-25&limit=1000&offset=0&to=2026-09-23, and each returns exactly 1000 results, i.e. there were more pages that were never requested; fromis always today − 90 days, regardless of the blueprint'stimeAfter(2026-01-28 here);- resulting
_tool_tempo_worklogsby start month: May 5040, Jun 5109, Jul 2208, Aug 461, Sep 46 — nothing before May, and a steady decline towards today, while the same teams log a roughly constant volume.
Root causes in backend/plugins/tempo/tasks/worklog_collector.go:
- Pagination —
GetTotalPagescomputes the page count frommetadata.total, but the Tempo v4 API does not return a total: per the official OpenAPI spec (https://apidocs.tempo.io/tempo-openapi.yaml),PageableWorklog.metadataisPageableMetadata=count, limit, next, offset, previous.totalunmarshals as 0 → 0 pages → only the first page is fetched. - Time range — for team scopes the query uses
from = now − 90d/to = nowunlessfromDate/toDatetask options are set (the blueprint never sets them), ignoring the sync policy (timeAfter, incremental state). - Shared state across teams — the collector's raw params only carry
ConnectionId(TeamIdis always 0, althoughTempoTeam.GetParams()declaresConnectionId+TeamId), so all team scopes of a connection share one raw-data/collector-state key. With (2) fixed, the second team collected in a pipeline would run incrementally from the first team's start time.
Separately, convert_worklogs emits ids that never join the jira domain layer:
issue_worklogs.issue_idisjira:JiraIssues:<conn>:<id>(plural), while the jira plugin generatesjira:JiraIssue:<conn>:<id>;issue_worklogs.author_idis the bare Atlassian account id instead ofjira:JiraAccount:<conn>:<id>, so worklogs don't joinaccounts/user_accounts.
What do you expect to happen
All worklogs from timeAfter onwards are collected (then incrementally via updatedFrom), and issue_worklogs rows join issues and accounts like the jira plugin's own worklogs do.
How to reproduce
- Create a Tempo connection and add a team scope whose members log more than 1000 worklogs in 90 days.
- Run a blueprint with
timeAfterolder than 90 days. SELECT url, COUNT(*) FROM _raw_tempo_api_worklogs GROUP BY url;→ one URL per run,offset=0, 1000 rows each;SELECT MIN(start_date), MAX(start_date) FROM _tool_tempo_worklogs;→ starts 90 days back, sparse towards today.SELECT COUNT(*) FROM issue_worklogs w JOIN issues i ON i.id = w.issue_id WHERE w.id LIKE 'tempo:%';→ 0.
Anything else
Happens on every run. Worklogs collected by the jira plugin for time logged through Tempo carry the "Timesheets by Tempo" app as author, so the tempo plugin is the only source of per-person time — which makes this quite visible for anyone building per-team/per-person dashboards.
Version
v1.0.3-beta15 (collector unchanged on main)
Are you willing to submit PR?
- Yes I am willing to submit a PR!
Code of Conduct
- I agree to follow this project's Code of Conduct
- 主要言語
- Go
- スター
- 3.1k
- フォーク
- 812
- 平均マージ
- 1日 23時間
- マージ済み PR(30日)
- 51
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
apache/devlake のほかの issue
-
type/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
-
type/bug
難易度 4/5 3〜5日 初心者へのやさしさ 52/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
-
Bob Shell support オープンenhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
santhosh-tekuri/jsonschema#276 ·