Bug: Data loss if CSV created on node added to cluster, then node evicted later.
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
- 技術スタック
- azure, powershell
- 領域
- cloud, infrastructure
調査の方向性
Azure Local、WAC、CSV の拡張、および remove-mocphysicalnode と remove-clusternode -cleanupdisks コマンドを含む 5 段階の再現手順から始めます。観測された失敗状態とアクセス エラーを想定される動作と比較し、その後、issue で要求されている Azure Local のドキュメントを確認します。ノードのエビクションと CSV の拡張について、原因とサポートされている修正方法が確立されれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Bug description
If a CSV is created on a four node cluster and then the fourth node is evicted from the cluster, the CSV that was created while the cluster had four hosts will have the wrong number of columns. As a result, the CSV becomes a loaded time bomb and will eventually cause workloads to seize up if enough data is added and/or changed on the CSV and the CSV will fail if it is expanded, causing data loss.
Any attempt at trying to bring the CSV online produces either "Access Denied" errors or Error Code 0x8007054f "An internal error occurred".
Repro steps
- Join a node into a 3 node Azure Local cluster to make it into a 4 node cluster.
- Create a new CSV with the fourth node joined and operable. Make sure that the fourth node owns it for a little while.
- Move the ownership of the fourth newly created CSV to one of the other three nodes.
- Evict the fourth node from the cluster using the remove-mocphysicalnode and remove-clusternode -cleanupdisks commands. (As if the cluster was being permanently shrunk.)
- Take the fourth node offline.
- Expand the size of the newly created fourth CSV by some amount (Such as 5 TB) using the WAC
- Observe that newly created CSV goes into failed state and any running VM's on it will also fail.
(Note that all CSV's were encrypted, this may nor may not be reproducible with unencrypted CSV's.)
Expected behavior
- Expanding a CSV should not cause a CSV to go offline. Nor should added CSV's have mismatched columns to the other CSV's, causing workloads to eventually seize up and fail after a certain period of time.
- Mismatched columns between CSV's should be automatically corrected with the addition or subtraction of nodes in the cluster. There also should be code that automatically detects CSV's with mismatched columns and automatically corrects for it without user input, treating it no more differently than an automatic array repair after a disk is replaced.
- The node eviction may have been done improperly. Where is the documentation for Azure Local?
- Attempts at bringing the CSV online with Hyper-V tools shouldn't have produced errors, nor should attempts at trying to bring a failed CSV back online fail.
Environment (please complete the following information):
Build 12.2512.1002.16
4 node cluster
Production
East US
- 主要言語
- PowerShell
- スター
- 78
- フォーク
- 60
- 平均マージ
- 1日 6時間
- マージ済み PR(30日)
- 5
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Azure/AzureLocal-Supportability のほかの issue
-
難易度 1/5 1〜3時間 初心者へのやさしさ 88/100
-
難易度 1/5 1時間未満 初心者へのやさしさ 82/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
Azure/AzureLocal-Supportability の issue をすべて見る
似ている issue
-
Language: Terraform :globe_with_meridians: Needs: Triage :mag: Type: Bug :bug:
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
Azure/terraform-azurerm-avm-res-containerregistry-registry#230 · コメント 1 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
-
level/task module/gcp type/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100