Docs and operating-model requests 504: content hydration crawls 2,179 GitHub blobs serially inside a 30s Lambda
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- aws, github, typescript
- 領域
- backend, cloud, observability, performance
調査の方向性
Start with backend/src/docs/githubStore.ts and trace how ContentsApiGithubStore.sync() is called by operatingModel/loader.ts. Reproduce the direct docs endpoint timeout, then choose and scope a hydration approach that completes within the Lambda budget and makes GitHub requests fail fast. Also inspect seedRuntimeUsers for the existing-user role backfill; done means docs and operating-model endpoints recover, hangs return promptly, and missing roles cannot recur.
索引モデルが issue の本文から書いたものです。
説明
Symptom
The portal Today page shows "Process documents are unavailable — HTTP 504", an empty action queue, and My Plan / operating-model data missing. Work API calls were also failing (fixed separately, see below).
Diagnosis (verified against production 2026-09-26)
GET /docs/process-qualityinvoked directly on the backend Lambda returnsSandbox.Timedoutafter 30.00s;/healthon the same container returns 200 in ~2s. CloudFront surfaces the timeout as 504.- Every docs/operating-model request calls
ContentsApiGithubStore.sync()(backend/src/docs/githubStore.ts), which on a cold container hydrates the wholecontent/tree one GitHub API call per blob, serially, with no fetch timeout. DataTalksClub/dataops-knowledgecontent/currently holds 2,179 blobs (443 md, 1,516 jpg, 220 png). At ~0.3–0.5s per blobs call that is ~10–15 minutes of work inside a 30s Lambda, so hydration never completes in one invocation. Each new container restarts the crawl (progress persists only in per-container /tmp), so no container ever becomes warm.- Blast radius:
/docs/*,/search,/api/operating-model,/api/my-plan, and the Today page's process-quality call (operatingModel/loader.tscalls the samestore.sync()). - Side damage: each doomed crawl burns up to ~2,200 GitHub requests; the knowledge token's rate limit already showed 1,301/5,000 used within one hour.
Work API 403s (fixed in data on 2026-09-26, needs a repo-level guard)
- The #164 role gate denies users without a supported
role. The three live user items predate the role attribute (created 2026-06-28, norole), andseedRuntimeUsersskips existing users as "unchanged", so the role was never backfilled. Every team read/work mutation returned 403 ("Team access requires an active admin or operator role"). - Applied the declared seed state (
role: 'admin'for grace/valeriia/alexey) directly todataops-v1-usersvia UpdateItem;/api/cardsand/api/tasksverified 200 after. - Follow-up needed: make
seedRuntimeUsersupdate missing attributes on existing users (or migrate in place), so this cannot recur. Data migration was done by hand this time only because the gate shipped without it.
Fix directions (for grooming)
- Stop eager per-blob hydration. Candidates: hydrate only
.mdeagerly and fetch images lazily viaensureFile; or download a single tarball and extractcontent/; or sync the knowledge repo to S3 from CI and hydrate from S3 in one/few calls. - Add an explicit timeout (
AbortSignal.timeout) torequest()/fetchImplingithubStore.tsso a GitHub hang fails fast with 502 instead of eating the whole Lambda budget. - Consider a startup/deploy-time hydration step (offline snapshot in the artifact) instead of on-demand crawling.
Related observability gap
Backend Lambda log streams stop on 2026-08-11 even though the function is actively invoked (CloudWatch metrics show hundreds of invocations/day; direct invokes today produced no streams). The execution role still has AWSLambdaBasicExecutionRole attached. Root cause unknown — needs its own investigation; until fixed, production debugging is blind (this diagnosis had to be reproduced by direct invokes).
- 主要言語
- TypeScript
- スター
- 2
- フォーク
- 0
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートなし
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
DataTalksClub/dataops のほかの issue
-
backend bug needs grooming
難易度 3/5 1〜2日 初心者へのやさしさ 75/100
DataTalksClub/dataops#227 · コメント 1 件 ·
-
bug needs grooming process-docs testing
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
DataTalksClub/dataops#225 · コメント 1 件 ·
-
design enhancement frontend P1 portal testing
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
DataTalksClub/dataops#219 · コメント 13 件 ·
-
design enhancement frontend P1 portal testing
難易度 5/5 1週間以上 初心者へのやさしさ 15/100
DataTalksClub/dataops#221 · コメント 17 件 ·
-
design docs enhancement frontend P1 portal testing
難易度 5/5 1週間以上 初心者へのやさしさ 18/100
DataTalksClub/dataops#220 · コメント 25 件 ·
DataTalksClub/dataops の issue をすべて見る
似ている issue
-
documentation
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
inu-appcenter/memorIN-frontend#106 ·
メンテナーはふだん 1 日以内に返信
-
kind/bug
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 7 日以内に返信
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
CSCfi/sd-search-ui#145 ·
メンテナーはふだん 1 日以内に返信
-
Add: Cbeebies pl SDオープンcheck:passed streams:add
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
メンテナーはふだん 1 日以内に返信