OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry records
メンテナーはふだん 5 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- cpp, sqlite
調査の方向性
Start with lib/system/TelemetrySystem.cpp, lib/http/HttpResponseDecoder.cpp, and lib/offline/StorageObserver.cpp, then trace the accepted, network-failure, and aborted request paths. Reproduce an interrupted acknowledgement with persistent SQLite storage and compare the resulting records. Done means an accepted logical event produces one destination record even when the response outcome is ambiguous.
索引モデルが issue の本文から書いたものです。
説明
Describe your environment
- Client: Microsoft Teams for iOS
- Platform: iOS and iPadOS
- Device scope: Not isolated to a specific device model
- OS scope: Not isolated to a specific iOS or iPadOS version
- App scope: Multiple Microsoft Teams iOS versions
- OneDS SDK scope: Multiple versions from
3.4.295.1through3.10.173.1 - Current OneDS SDK:
3.10.173.1 - Current telemetry SDK field:
EVT-iOS-C++-No-3.10.173.1 - Event source:
OneCollector - Offline storage: OneDS persistent SQLite cache
- Transmission profile: Real-time
The issue is broadly distributed across the Teams iOS population. It is not isolated to one device model, OS version, app version, telemetry event, or OneDS SDK version.
One concrete affected environment is:
Device: iPhone18,2
OS: iOS 27.0.1
Teams version: 1417/8.17.77.2026172901
OneDS SDK: 3.10.173.1
Telemetry SDK field: EVT-iOS-C++-No-3.10.173.1
Event source: OneCollector
The same duplicate-delivery pattern was observed with multiple OneDS SDK versions:
3.10.173.1
3.10.40.1
3.9.324.1
3.9.309.1
3.9.267.1
3.9.212.1
3.8.249.1
3.6.187.1
3.5.200.1
3.5.189.1
3.5.148.1
3.5.127.1
3.5.25.1
3.4.295.1
We checked the current main branch of microsoft/cpp_client_telemetry. The relevant retry behavior is still present:
- Confirmed accepted requests delete records from offline storage.
- Network failures release records back to offline storage.
- Aborted requests release records back to offline storage.
- We could not find an idempotency mechanism that prevents an event already accepted by Collector from being ingested again if the client does not receive or process the acknowledgement.
Relevant current-main files:
- https://github.com/microsoft/cpp_client_telemetry/blob/main/lib/system/TelemetrySystem.cpp
- https://github.com/microsoft/cpp_client_telemetry/blob/main/lib/http/HttpResponseDecoder.cpp
- https://github.com/microsoft/cpp_client_telemetry/blob/main/lib/offline/StorageObserver.cpp
Steps to reproduce
The issue appears when an upload has an ambiguous result—for example, Collector accepts the request, but the client does not receive or process the HTTP success response.
A likely reproduction flow is:
- Initialize the OneDS C++ SDK with persistent SQLite offline storage.
- Log a telemetry event.
- Allow OneDS to begin uploading the event to Collector.
- After the request has reached Collector but before the HTTP success callback is processed, interrupt the response path by doing one of the following:
- terminate the application;
- call
pauseTransmission()or otherwise cancel active requests; - interrupt the network connection after the request is sent.
- Relaunch the application or restore network connectivity.
- Allow OneDS to process its persistent offline cache.
- Query the destination telemetry table using the event timestamp, sequence number, session ID, and pipeline record ID.
- Observe that the same serialized event was ingested multiple times with different pipeline ingestion timestamps.
The following production query demonstrates one occurrence:
scenario
| where Scenario_Step == "STOP"
| where Scenario_ID == "11FC326D-9FC3-47E0-B242-C7E2D592EBE4"
| project
Scenario_ID,
Scenario_Step,
Scenario_Status,
Scenario_ScenarioTimeTaken,
EventInfo_Time,
EventInfo_Sequence,
Session_Id,
PipelineInfo__550_55Id,
EventInfo_SdkVersion,
PipelineInfo_IngestionTime
| order by PipelineInfo_IngestionTime asc
The affected event has the following identity:
Scenario_ID: 11FC326D-9FC3-47E0-B242-C7E2D592EBE4
Scenario_Step: STOP
Scenario_Status: OK
Scenario_ScenarioTimeTaken: 6527.508974
EventInfo_Time: 2026-09-29 15:12:43.272 UTC
EventInfo_Sequence: 632
OneDS SDK: 3.10.173.1
Event source: OneCollector
The same event was ingested at the following times:
2026-09-29 16:10:29.882 UTC
2026-09-29 18:47:54.833 UTC
2026-09-29 20:02:04.765 UTC
2026-09-29 20:25:11.776 UTC
2026-09-29 21:08:53.200 UTC
All five rows have the same:
Scenario_IDEventInfo_TimeEventInfo_SequenceSession_IdPipelineInfo__550_55IdEventInfo_SdkVersion- status
- duration
- event payload
Only PipelineInfo_IngestionTime differs.
The identical sequence number, session ID, and pipeline record ID indicate that the application created the event once and that the same serialized OneDS record was delivered repeatedly.
What is the expected behavior?
One logical client event should produce one logical record in the destination telemetry table.
If OneDS retries an upload because the upload outcome is ambiguous, one of the following should prevent duplicate ingestion:
- Collector should deduplicate retries using a stable event or record identifier.
- The upload protocol should provide an idempotency key that remains stable across retries.
- The SDK and Collector should use another mechanism that prevents an already-accepted event from creating a second logical record.
Retries are expected and necessary for telemetry reliability. However, an event already accepted by Collector should not be counted multiple times in downstream telemetry.
What is the actual behavior?
The same serialized OneDS event is ingested multiple times.
Each copy retains the same:
- event timestamp;
- sequence number;
- session ID;
- client event ID;
- pipeline record ID;
- SDK version;
- event source;
- payload.
Each copy receives a different PipelineInfo_IngestionTime.
In the concrete example, the same event was ingested five times over approximately 298 minutes.
This inflates:
- telemetry event volume;
- success and failure counts;
- reliability metrics;
- latency metrics;
- usage metrics;
- telemetry ingestion cost.
Additional context
Production impact
A five-minute production sample from September 29, 2026 contained:
| Metric | Result |
|---|---|
| Logical telemetry events evaluated | 34,725,838 |
| Logical events affected by duplicate delivery | 346,712 |
| Extra ingested rows | 366,734 |
| Approximate affected logical-event rate | 1.0% |
| Maximum observed copies of one event | 7 |
This indicates a systemic issue rather than an isolated event-instrumentation problem.
Impact by OneDS SDK version
| OneDS SDK version | Affected logical events | Approximate affected rate |
|---|---|---|
3.10.40.1 |
208,469 | 0.96% |
3.10.173.1 |
129,372 | 1.07% |
3.8.249.1 |
2,193 | 0.91% |
3.9.324.1 |
1,965 | 0.72% |
3.6.187.1 |
1,251 | 1.57% |
3.9.309.1 |
1,223 | 0.83% |
3.9.267.1 |
988 | 0.84% |
3.9.212.1 |
825 | 0.73% |
The presence of duplicates across multiple SDK generations indicates that this is not a regression isolated to SDK 3.10.173.1.
Suspected failure mode
The observed behavior is consistent with at-least-once delivery:
- OneDS reserves a record from offline storage.
- OneDS sends the record to Collector.
- Collector accepts and persists the event.
- The response is lost, interrupted, or canceled before the SDK processes the success acknowledgement.
- The SDK classifies the request as a network failure or aborted request.
- The SDK releases the record back to offline storage.
- OneDS retries the same serialized record later.
- Collector ingests the record again without deduplicating it.
The relevant current-main routing is:
httpDecoder.eventsAccepted
>> storage.deleteRecords
>> stats.onUploadSuccessful
>> tpm.eventsUploadSuccessful;
httpDecoder.temporaryNetworkFailure
>> storage.releaseRecords
>> stats.onUploadFailed
>> tpm.eventsUploadFailed;
httpDecoder.requestAborted
>> storage.releaseRecords
>> stats.onUploadFailed
>> tpm.eventsUploadAborted;
This behavior avoids losing telemetry when the client cannot confirm success. However, without end-to-end idempotency, it produces duplicate ingestion when Collector accepted the original request but the acknowledgement was not received.
- 主要言語
- C
- スター
- 102
- フォーク
- 66
- 平均マージ
- 5日 14分
- マージ済み PR(30日)
- 9
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
microsoft/cpp_client_telemetry のほかの issue
-
C API enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
microsoft/cpp_client_telemetry#628 ·
メンテナーはふだん 5 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
microsoft/cpp_client_telemetry#1504 ·
メンテナーはふだん 5 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
microsoft/cpp_client_telemetry#1413 · コメント 2 件 ·
メンテナーはふだん 5 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 38/100
microsoft/cpp_client_telemetry#1387 ·
メンテナーはふだん 5 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
microsoft/cpp_client_telemetry#1353 ·
メンテナーはふだん 5 日以内に返信
microsoft/cpp_client_telemetry の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
darktable-org/darktable#22455 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
area:ci kind:gate-defect
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
InauguralSystems/EigenScript#1448 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
BasedHardware/omi#20084 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 64/100
raspberrypi/pico-sdk#3225 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100