Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

OneDS C++ SDK retries already-ingested iOS events, causing duplicate telemetry records

Abierto
#1,542 2 comentarios 0 reacciones 1 asignado Ver en GitHub

Los mantenedores suelen responder en 3 días

@bmehta001 ya está trabajando en esto.

Desde el 1/10/2026.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
cpp, sqlite

Línea de trabajo

Start with lib/system/TelemetrySystem.cpp, lib/http/HttpResponseDecoder.cpp, and lib/offline/StorageObserver.cpp, then trace the accepted, network-failure, and aborted request paths. Reproduce an interrupted acknowledgement with persistent SQLite storage and compare the resulting records. Done means an accepted logical event produces one destination record even when the response outcome is ambiguous.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug

Describe your environment

  • Client: Microsoft Teams for iOS
  • Platform: iOS and iPadOS
  • Device scope: Not isolated to a specific device model
  • OS scope: Not isolated to a specific iOS or iPadOS version
  • App scope: Multiple Microsoft Teams iOS versions
  • OneDS SDK scope: Multiple versions from 3.4.295.1 through 3.10.173.1
  • Current OneDS SDK: 3.10.173.1
  • Current telemetry SDK field: EVT-iOS-C++-No-3.10.173.1
  • Event source: OneCollector
  • Offline storage: OneDS persistent SQLite cache
  • Transmission profile: Real-time

The issue is broadly distributed across the Teams iOS population. It is not isolated to one device model, OS version, app version, telemetry event, or OneDS SDK version.

One concrete affected environment is:

Device: iPhone18,2
OS: iOS 27.0.1
Teams version: 1417/8.17.77.2026172901
OneDS SDK: 3.10.173.1
Telemetry SDK field: EVT-iOS-C++-No-3.10.173.1
Event source: OneCollector

The same duplicate-delivery pattern was observed with multiple OneDS SDK versions:

3.10.173.1
3.10.40.1
3.9.324.1
3.9.309.1
3.9.267.1
3.9.212.1
3.8.249.1
3.6.187.1
3.5.200.1
3.5.189.1
3.5.148.1
3.5.127.1
3.5.25.1
3.4.295.1

We checked the current main branch of microsoft/cpp_client_telemetry. The relevant retry behavior is still present:

  • Confirmed accepted requests delete records from offline storage.
  • Network failures release records back to offline storage.
  • Aborted requests release records back to offline storage.
  • We could not find an idempotency mechanism that prevents an event already accepted by Collector from being ingested again if the client does not receive or process the acknowledgement.

Relevant current-main files:

Steps to reproduce

The issue appears when an upload has an ambiguous result—for example, Collector accepts the request, but the client does not receive or process the HTTP success response.

A likely reproduction flow is:

  1. Initialize the OneDS C++ SDK with persistent SQLite offline storage.
  2. Log a telemetry event.
  3. Allow OneDS to begin uploading the event to Collector.
  4. After the request has reached Collector but before the HTTP success callback is processed, interrupt the response path by doing one of the following:
    • terminate the application;
    • call pauseTransmission() or otherwise cancel active requests;
    • interrupt the network connection after the request is sent.
  5. Relaunch the application or restore network connectivity.
  6. Allow OneDS to process its persistent offline cache.
  7. Query the destination telemetry table using the event timestamp, sequence number, session ID, and pipeline record ID.
  8. Observe that the same serialized event was ingested multiple times with different pipeline ingestion timestamps.

The following production query demonstrates one occurrence:

scenario
| where Scenario_Step == "STOP"
| where Scenario_ID == "11FC326D-9FC3-47E0-B242-C7E2D592EBE4"
| project
    Scenario_ID,
    Scenario_Step,
    Scenario_Status,
    Scenario_ScenarioTimeTaken,
    EventInfo_Time,
    EventInfo_Sequence,
    Session_Id,
    PipelineInfo__550_55Id,
    EventInfo_SdkVersion,
    PipelineInfo_IngestionTime
| order by PipelineInfo_IngestionTime asc

The affected event has the following identity:

Scenario_ID: 11FC326D-9FC3-47E0-B242-C7E2D592EBE4
Scenario_Step: STOP
Scenario_Status: OK
Scenario_ScenarioTimeTaken: 6527.508974
EventInfo_Time: 2026-09-29 15:12:43.272 UTC
EventInfo_Sequence: 632
OneDS SDK: 3.10.173.1
Event source: OneCollector

The same event was ingested at the following times:

2026-09-29 16:10:29.882 UTC
2026-09-29 18:47:54.833 UTC
2026-09-29 20:02:04.765 UTC
2026-09-29 20:25:11.776 UTC
2026-09-29 21:08:53.200 UTC

All five rows have the same:

  • Scenario_ID
  • EventInfo_Time
  • EventInfo_Sequence
  • Session_Id
  • PipelineInfo__550_55Id
  • EventInfo_SdkVersion
  • status
  • duration
  • event payload

Only PipelineInfo_IngestionTime differs.

The identical sequence number, session ID, and pipeline record ID indicate that the application created the event once and that the same serialized OneDS record was delivered repeatedly.

What is the expected behavior?

One logical client event should produce one logical record in the destination telemetry table.

If OneDS retries an upload because the upload outcome is ambiguous, one of the following should prevent duplicate ingestion:

  1. Collector should deduplicate retries using a stable event or record identifier.
  2. The upload protocol should provide an idempotency key that remains stable across retries.
  3. The SDK and Collector should use another mechanism that prevents an already-accepted event from creating a second logical record.

Retries are expected and necessary for telemetry reliability. However, an event already accepted by Collector should not be counted multiple times in downstream telemetry.

What is the actual behavior?

The same serialized OneDS event is ingested multiple times.

Each copy retains the same:

  • event timestamp;
  • sequence number;
  • session ID;
  • client event ID;
  • pipeline record ID;
  • SDK version;
  • event source;
  • payload.

Each copy receives a different PipelineInfo_IngestionTime.

In the concrete example, the same event was ingested five times over approximately 298 minutes.

This inflates:

  • telemetry event volume;
  • success and failure counts;
  • reliability metrics;
  • latency metrics;
  • usage metrics;
  • telemetry ingestion cost.

Additional context

Production impact

A five-minute production sample from September 29, 2026 contained:

Metric Result
Logical telemetry events evaluated 34,725,838
Logical events affected by duplicate delivery 346,712
Extra ingested rows 366,734
Approximate affected logical-event rate 1.0%
Maximum observed copies of one event 7

This indicates a systemic issue rather than an isolated event-instrumentation problem.

Impact by OneDS SDK version

OneDS SDK version Affected logical events Approximate affected rate
3.10.40.1 208,469 0.96%
3.10.173.1 129,372 1.07%
3.8.249.1 2,193 0.91%
3.9.324.1 1,965 0.72%
3.6.187.1 1,251 1.57%
3.9.309.1 1,223 0.83%
3.9.267.1 988 0.84%
3.9.212.1 825 0.73%

The presence of duplicates across multiple SDK generations indicates that this is not a regression isolated to SDK 3.10.173.1.

Suspected failure mode

The observed behavior is consistent with at-least-once delivery:

  1. OneDS reserves a record from offline storage.
  2. OneDS sends the record to Collector.
  3. Collector accepts and persists the event.
  4. The response is lost, interrupted, or canceled before the SDK processes the success acknowledgement.
  5. The SDK classifies the request as a network failure or aborted request.
  6. The SDK releases the record back to offline storage.
  7. OneDS retries the same serialized record later.
  8. Collector ingests the record again without deduplicating it.

The relevant current-main routing is:

httpDecoder.eventsAccepted
    >> storage.deleteRecords
    >> stats.onUploadSuccessful
    >> tpm.eventsUploadSuccessful;

httpDecoder.temporaryNetworkFailure
    >> storage.releaseRecords
    >> stats.onUploadFailed
    >> tpm.eventsUploadFailed;

httpDecoder.requestAborted
    >> storage.releaseRecords
    >> stats.onUploadFailed
    >> tpm.eventsUploadAborted;

This behavior avoids losing telemetry when the client cannot confirm success. However, without end-to-end idempotency, it produces duplicate ingestion when Collector accepted the original request but the acknowledgement was not received.

Lenguaje dominante
C
Estrellas
104
Forks
67
Merge medio
4 d 8 h
PR fusionados (30 d)
12

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de microsoft/cpp_client_telemetry

Todos los issues de microsoft/cpp_client_telemetry

Issues similares

Más issues de C

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.