Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug] Brownfield discovery offers the wrong Terraform state bucket: it looks for gs://<prefix>-tn-<env>-<tenant>-0, but a tenant's own stage state lives in gs://<prefix>-<env>-<tenant>-iac-0

クローズ
#230 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
74/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
gcp, shell, terraform

調査の方向性

deploy.sh の discover_infrastructure() から始め、特に 584 行目付近のバケット検索、690-700 行目の CMEK チェック、718 行目付近のフォールバックを確認します。文書化された gcloud storage の一覧を使って brownfield discovery を再現し、期待される *.tfstate オブジェクトを含むテナント状態バケットが選択されることを検証します。既存のデプロイを再実行したときに、重複するリソースを提案するのではなく no-op プランが生成されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Bug Gemini - Government Level of Effort - Low Priority - High

Bug Description

Both names in this report are Cloud Storage buckets, not projects. They are easy to misread,
because a bucket name here differs from a project name by one word: the bucket
<prefix>-<env>-<tenant>-iac-0 lives inside the project <prefix>-<env>-<tenant>-iac-core-0.
No project carries -tn- in its name at all.

PR #217 changed the state-bucket lookup in discover_infrastructure():

# deploy.sh:584 (main)
POTENTIAL_BUCKET="${PREFIX}-tn-${ENVIRONMENT}-${TENANT}-0"

# previously (v2.13.0 and every release before it)
POTENTIAL_BUCKET="${PREFIX}-${ENVIRONMENT}-${TENANT}-iac-0"

The stated reason was that 1-resman creates the tenant state bucket as <prefix>-tn-<env>-<tenant>-0
"so the bucket was never found and a second one was created."

1-resman creates three buckets per tenant environment, with three different roles, and the stage
picked the wrong one.
This is not a guess about intent — 1-resman/outputs-tenants.tf names them:

# outputs-tenants.tf — local.tenant_tfvars
automation = {
  core_bucket    = module.tenant-core-gcs[k].name             # <prefix>-tn-<env>-<tenant>-0
  outputs_bucket = module.tenant-self-iac-gcs-outputs[k].name # <prefix>-<env>-<tenant>-iac-outputs-0
  state_bucket   = module.tenant-self-iac-gcs-states[k].name  # <prefix>-<env>-<tenant>-iac-0
}

The bucket upstream itself calls state_bucket is <prefix>-<env>-<tenant>-iac-0 — the one the
lookup used to find, and the one PR #217 moved away from.

The same file builds two distinct backend configurations from the same template, and they are not
interchangeable because they run as different identities:

Local Backend bucket Runs as Lives in
tenant_core_providers core_bucket = <prefix>-tn-<env>-<tenant>-0 tenant-core-sa the organization automation project
tenant_self_providers state_bucket = <prefix>-<env>-<tenant>-iac-0 tenant-self-iac-sa the tenant's own IaC core project

A Gemini Enterprise deployment is tenant-side work, so its state belongs in state_bucket — which is
where the pre-#217 lookup pointed and where every existing deployment has it. core_bucket is the
backend for the organization-side configuration that manages the tenant; on a landing zone where
nobody has applied that configuration it is simply empty, which is exactly what makes the wrong
choice silent.

The consequence is not a missing bucket. It is a silently different backend: on a landing zone
where both buckets exist, the wizard now binds stage 0 to the organization-level bucket, which holds
no Gemini state, so Terraform sees an empty state and plans to create a CmekConfig, a reserved VIP, a
Discovery Engine app and data stores that already exist in the project.

Nothing changed on the 1-resman side — this is a regression in deploy.sh alone. The three
buckets, their names, and the core_bucket / state_bucket labels in outputs-tenants.tf are
byte-identical at v2.9.0, v2.11.0, v2.12.0, v2.13.0, v3.0.0 and main; tenant-core-gcs
has existed since 2024-09-16 (a39c7ff4). So the premise behind the change — that 1-resman names
the tenant state bucket <prefix>-tn-<env>-<tenant>-0 — was never true at any release, and the
lookup it replaced had been correct for the life of the stage.

The hyphenated-tenant half of PR #217 (ten-1 no longer parsed as ten) is a genuine fix and is not
in question.

Environment and Deployment Context

  • Stellar Engine Version/Commit: main @ 6d7d08c0 (2026-09-09). Introduced by PR #217 (60a81387, merged 2026-09-05). v2.13.0 (8f5b67a6) and v3.0.0 (f64ce6cd) both use the previous name.
  • Deployment Type:
    • US Region Restricted (e.g., Access Policy constraint)
    • FedRAMP Moderate
    • FedRAMP High
    • DoD IL4
    • DoD IL5
    • Stand-alone / Custom
  • FAST Stage (if applicable): the gemini-enterprise blueprint's brownfield discovery, against a landing zone built by Stages 0-3
    • Stage 0 (Bootstrap)
    • Stage 1 (Resource Management)
    • Stage 2 (Networking)
    • Stage 3 (Security)

Steps to Reproduce

  1. Deploy Stages 0-1 with any prefix and a tenant (for example g4g), so that 1-resman creates both <prefix>-tn-<env>-<tenant>-0 and <prefix>-<env>-<tenant>-iac-0.
  2. Deploy the Gemini Enterprise blueprint stage 0 from a release at or before v3.0.0. Its state is written to gs://<prefix>-<env>-<tenant>-iac-0/.
  3. Update the checkout to main and re-run ./deploy.sh in brownfield mode against the same project.
  4. Read the bucket it reports at "Checking for Terraform State Bucket", and the backend it writes.

Expected Behavior

Discovery finds the bucket that holds the deployment's existing state, so a re-run is a no-op plan.

Actual Behavior

Discovery reports Found Terraform State Bucket: <prefix>-tn-<env>-<tenant>-0 and offers it as the
default of the state-bucket prompt. Nothing rejects it: the bucket exists, so no fallback fires,
and it passes the CMEK check at deploy.sh:690-700 because 1-resman encrypts both tenant buckets
with the same key. An operator who accepts the offered default — the documented path — binds the
stage to a bucket holding no state, and the plan proposes to create the whole stage a second time on
top of live resources.

Observed on a FedRAMP High landing zone, prefix alaska, tenants g4g and report, all read-only:

$ gcloud storage ls --recursive 'gs://alaska-prod-g4g-iac-0/**'
gs://alaska-prod-g4g-iac-0/terraform/state/stage-0/default.tfstate

$ gcloud storage ls --recursive 'gs://alaska-int-g4g-iac-0/**'
gs://alaska-int-g4g-iac-0/terraform/state/stage-0/default.tfstate
gs://alaska-int-g4g-iac-0/terraform/state/stage-1/default.tfstate

$ gcloud storage ls --recursive 'gs://alaska-tn-prod-g4g-0/**'
ERROR: (gcloud.storage.ls) One or more URLs matched no objects.

$ gcloud storage ls --recursive 'gs://alaska-tn-int-g4g-0/**'
ERROR: (gcloud.storage.ls) One or more URLs matched no objects.

The -iac-0 buckets hold the blueprint's own state, at a path that names the Gemini stages
explicitly. The -tn- buckets — the ones the new code selects — are completely empty.

Both are CMEK-encrypted with the same key and both are versioned, so nothing downstream distinguishes
them either:

$ gcloud storage buckets describe gs://alaska-prod-g4g-iac-0  --format='value(default_kms_key,versioning_enabled)'
projects/alaska-prod-g4g-iac-core-0/locations/us-west1/keyRings/Prod-g4g-keyring/cryptoKeys/gcs   True

$ gcloud storage buckets describe gs://alaska-tn-prod-g4g-0   --format='value(default_kms_key,versioning_enabled)'
projects/alaska-prod-g4g-iac-core-0/locations/us-west1/keyRings/Prod-g4g-keyring/cryptoKeys/gcs   True

And the two buckets live in different projects, which is the clearest statement of the difference:
<prefix>-<env>-<tenant>-iac-0 is in the tenant's own IaC project
(alaska-prod-g4g-iac-core-0), while <prefix>-tn-<env>-<tenant>-0 is in the organization
automation project (alaska-prod-iac-core-0) alongside the resman buckets.

Relevant Logs and Errors

No error. The run looks clean until the plan output is read.

Additional Context

  • Suggested fix: look for state_bucket (<prefix>-<env>-<tenant>-iac-0) rather than core_bucket; if it is absent, fall back to core_bucket, and if both exist and both hold state, prompt rather than choose. A bucket that exists but contains no *.tfstate should not be accepted silently.
  • Worth guarding generally: the wizard treats "bucket exists" as "state found". Checking for an object under the expected prefix would have caught this and the original problem #217 set out to fix — one gcloud storage ls on the bucket would distinguish them.
  • For completeness, there is a third name in play: if the CMEK check does reject the chosen bucket, deploy.sh:718 falls back to creating <prefix>-<env>-<tenant>-tfstate-0, which matches neither of the two 1-resman produces. That path is not what happens here — both tenant buckets carry the key — but it means a rejected bucket also does not lead back to the right one.
  • The remaining brownfield assumption the PR author set aside — the derived <Env>-<tenant>-keyring in the tenant IaC project, which 1-resman does not create — is reported separately as #132 and #106.
主要言語
HCL
スター
51
フォーク
21
平均マージ
1日 17時間
マージ済み PR(30日)
29

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

google/stellar-engine のほかの issue

google/stellar-engine の issue をすべて見る

似ている issue

Cloud の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。