Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry records

未关闭
#1,542 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 5 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
cpp, sqlite

调研方向

Start with lib/system/TelemetrySystem.cpp, lib/http/HttpResponseDecoder.cpp, and lib/offline/StorageObserver.cpp, then trace the accepted, network-failure, and aborted request paths. Reproduce an interrupted acknowledgement with persistent SQLite storage and compare the resulting records. Done means an accepted logical event produces one destination record even when the response outcome is ambiguous.

由索引模型根据 Issue 内容生成。

描述

bug

Describe your environment

  • Client: Microsoft Teams for iOS
  • Platform: iOS and iPadOS
  • Device scope: Not isolated to a specific device model
  • OS scope: Not isolated to a specific iOS or iPadOS version
  • App scope: Multiple Microsoft Teams iOS versions
  • OneDS SDK scope: Multiple versions from 3.4.295.1 through 3.10.173.1
  • Current OneDS SDK: 3.10.173.1
  • Current telemetry SDK field: EVT-iOS-C++-No-3.10.173.1
  • Event source: OneCollector
  • Offline storage: OneDS persistent SQLite cache
  • Transmission profile: Real-time

The issue is broadly distributed across the Teams iOS population. It is not isolated to one device model, OS version, app version, telemetry event, or OneDS SDK version.

One concrete affected environment is:

Device: iPhone18,2
OS: iOS 27.0.1
Teams version: 1417/8.17.77.2026172901
OneDS SDK: 3.10.173.1
Telemetry SDK field: EVT-iOS-C++-No-3.10.173.1
Event source: OneCollector

The same duplicate-delivery pattern was observed with multiple OneDS SDK versions:

3.10.173.1
3.10.40.1
3.9.324.1
3.9.309.1
3.9.267.1
3.9.212.1
3.8.249.1
3.6.187.1
3.5.200.1
3.5.189.1
3.5.148.1
3.5.127.1
3.5.25.1
3.4.295.1

We checked the current main branch of microsoft/cpp_client_telemetry. The relevant retry behavior is still present:

  • Confirmed accepted requests delete records from offline storage.
  • Network failures release records back to offline storage.
  • Aborted requests release records back to offline storage.
  • We could not find an idempotency mechanism that prevents an event already accepted by Collector from being ingested again if the client does not receive or process the acknowledgement.

Relevant current-main files:

Steps to reproduce

The issue appears when an upload has an ambiguous result—for example, Collector accepts the request, but the client does not receive or process the HTTP success response.

A likely reproduction flow is:

  1. Initialize the OneDS C++ SDK with persistent SQLite offline storage.
  2. Log a telemetry event.
  3. Allow OneDS to begin uploading the event to Collector.
  4. After the request has reached Collector but before the HTTP success callback is processed, interrupt the response path by doing one of the following:
    • terminate the application;
    • call pauseTransmission() or otherwise cancel active requests;
    • interrupt the network connection after the request is sent.
  5. Relaunch the application or restore network connectivity.
  6. Allow OneDS to process its persistent offline cache.
  7. Query the destination telemetry table using the event timestamp, sequence number, session ID, and pipeline record ID.
  8. Observe that the same serialized event was ingested multiple times with different pipeline ingestion timestamps.

The following production query demonstrates one occurrence:

scenario
| where Scenario_Step == "STOP"
| where Scenario_ID == "11FC326D-9FC3-47E0-B242-C7E2D592EBE4"
| project
    Scenario_ID,
    Scenario_Step,
    Scenario_Status,
    Scenario_ScenarioTimeTaken,
    EventInfo_Time,
    EventInfo_Sequence,
    Session_Id,
    PipelineInfo__550_55Id,
    EventInfo_SdkVersion,
    PipelineInfo_IngestionTime
| order by PipelineInfo_IngestionTime asc

The affected event has the following identity:

Scenario_ID: 11FC326D-9FC3-47E0-B242-C7E2D592EBE4
Scenario_Step: STOP
Scenario_Status: OK
Scenario_ScenarioTimeTaken: 6527.508974
EventInfo_Time: 2026-09-29 15:12:43.272 UTC
EventInfo_Sequence: 632
OneDS SDK: 3.10.173.1
Event source: OneCollector

The same event was ingested at the following times:

2026-09-29 16:10:29.882 UTC
2026-09-29 18:47:54.833 UTC
2026-09-29 20:02:04.765 UTC
2026-09-29 20:25:11.776 UTC
2026-09-29 21:08:53.200 UTC

All five rows have the same:

  • Scenario_ID
  • EventInfo_Time
  • EventInfo_Sequence
  • Session_Id
  • PipelineInfo__550_55Id
  • EventInfo_SdkVersion
  • status
  • duration
  • event payload

Only PipelineInfo_IngestionTime differs.

The identical sequence number, session ID, and pipeline record ID indicate that the application created the event once and that the same serialized OneDS record was delivered repeatedly.

What is the expected behavior?

One logical client event should produce one logical record in the destination telemetry table.

If OneDS retries an upload because the upload outcome is ambiguous, one of the following should prevent duplicate ingestion:

  1. Collector should deduplicate retries using a stable event or record identifier.
  2. The upload protocol should provide an idempotency key that remains stable across retries.
  3. The SDK and Collector should use another mechanism that prevents an already-accepted event from creating a second logical record.

Retries are expected and necessary for telemetry reliability. However, an event already accepted by Collector should not be counted multiple times in downstream telemetry.

What is the actual behavior?

The same serialized OneDS event is ingested multiple times.

Each copy retains the same:

  • event timestamp;
  • sequence number;
  • session ID;
  • client event ID;
  • pipeline record ID;
  • SDK version;
  • event source;
  • payload.

Each copy receives a different PipelineInfo_IngestionTime.

In the concrete example, the same event was ingested five times over approximately 298 minutes.

This inflates:

  • telemetry event volume;
  • success and failure counts;
  • reliability metrics;
  • latency metrics;
  • usage metrics;
  • telemetry ingestion cost.

Additional context

Production impact

A five-minute production sample from September 29, 2026 contained:

Metric Result
Logical telemetry events evaluated 34,725,838
Logical events affected by duplicate delivery 346,712
Extra ingested rows 366,734
Approximate affected logical-event rate 1.0%
Maximum observed copies of one event 7

This indicates a systemic issue rather than an isolated event-instrumentation problem.

Impact by OneDS SDK version

OneDS SDK version Affected logical events Approximate affected rate
3.10.40.1 208,469 0.96%
3.10.173.1 129,372 1.07%
3.8.249.1 2,193 0.91%
3.9.324.1 1,965 0.72%
3.6.187.1 1,251 1.57%
3.9.309.1 1,223 0.83%
3.9.267.1 988 0.84%
3.9.212.1 825 0.73%

The presence of duplicates across multiple SDK generations indicates that this is not a regression isolated to SDK 3.10.173.1.

Suspected failure mode

The observed behavior is consistent with at-least-once delivery:

  1. OneDS reserves a record from offline storage.
  2. OneDS sends the record to Collector.
  3. Collector accepts and persists the event.
  4. The response is lost, interrupted, or canceled before the SDK processes the success acknowledgement.
  5. The SDK classifies the request as a network failure or aborted request.
  6. The SDK releases the record back to offline storage.
  7. OneDS retries the same serialized record later.
  8. Collector ingests the record again without deduplicating it.

The relevant current-main routing is:

httpDecoder.eventsAccepted
    >> storage.deleteRecords
    >> stats.onUploadSuccessful
    >> tpm.eventsUploadSuccessful;

httpDecoder.temporaryNetworkFailure
    >> storage.releaseRecords
    >> stats.onUploadFailed
    >> tpm.eventsUploadFailed;

httpDecoder.requestAborted
    >> storage.releaseRecords
    >> stats.onUploadFailed
    >> tpm.eventsUploadAborted;

This behavior avoids losing telemetry when the client cannot confirm success. However, without end-to-end idempotency, it produces duplicate ingestion when Collector accepted the original request but the acknowledgement was not received.

主要语言
C
星标
102
派生
66
平均合并
5 天 14 分钟
30 天内合并 PR
9

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/cpp_client_telemetry 的其他 Issue

查看 microsoft/cpp_client_telemetry 的全部 Issue

相似的 Issue

更多 C Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。