Make Writes Idempotent
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 32/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 活発
- 技術スタック
- kafka, python
調査の方向性
Start by tracing the EventGate Lambda path that performs the DB, EventBridge, and Kafka writes, then inspect the existing DB tables and processing-start data. Reproduce or reason through message redelivery after one write fails. Done means redelivery is handled idempotently, arrival timing is documented or captured as needed, and tests or manual evidence cover the scenario.
索引モデルが issue の本文から書いたものです。
説明
Feature Description
When the same event arrives multiple times, there is no guarantee that it will not be re-inserted into the DB, or re-delivered to our consumers multiple times. There is no idempotency guarantee.
Problem / Opportunity
The EventGate Lambda handling events and performing writes, which is the central to the EventGate system, currently perform triple write: DB, EventBridge, and Kafka. If there is a failure, e.g. Kafka write fails, the Lambda finishes (the two writes are done) but it constructs response payload with HTTP status 500.
Then, the producer is notified about 500 and maybe it stores the msg into DLQ or redelivers the message again immediatelly, doesn't matter - the message is processed again and the triple write happens again - even if it previously succeeded for some of the writes. Therefore, we'll have duplicit items in our DB and our EventBridge/Kafka consumers might receive the message multiple times potentially.
Something like this happened with Unify recently (Unify as consumer), although no producer complained about this. But this certainly is something we should deal with and redesign.
Note: DB tables contain almost no constraints and very little auxiliary timestamp-based columns - like when a message arrives, a new column arrivedAt with default NOW() should be present - an idea - because otherwise we DO NOT KNOW when the message actually arrived (well, in this case, when it was stored in DB, the arrival happens seconds earlier). The processing start, column that is present, is set by producer, not by eventgate system. On producer's retries, this is the same value!
Acceptance Criteria
- Scenario with msg-redelivery is handled in better and more expected way, idempotency based ideally.
- Documentation
- Either programmatic tests exist, or the final solution was tested at least manually and evidence is attached somewhere, maybe into this ticket even.
- 主要言語
- Python
- スター
- 4
- フォーク
- 0
- 平均マージ
- 3日 12時間
- マージ済み PR(30日)
- 9
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドなし
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
AbsaOSS/EventGate のほかの issue
-
refactoring type:tech-debt
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 2 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 3/5 1〜2日 初心者へのやさしさ 64/100
メンテナーはふだん 2 日以内に返信
-
infrastructure type:tech-debt
難易度 3/5 1〜2日 初心者へのやさしさ 70/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 3/5 1〜2日 初心者へのやさしさ 70/100
メンテナーはふだん 2 日以内に返信
AbsaOSS/EventGate の issue をすべて見る
似ている issue
-
feedback simulation workshop
難易度 2/5 1〜3時間 初心者へのやさしさ 73/100
githubnext/gh-aw-workshop#4455 ·
メンテナーはふだん 1 日以内に返信
-
Triage 🩺
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitオープンneeds-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 77/100
krkn-chaos/krkn#1627 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
NousResearch/hermes-agent#136483 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信