Bulk event ingestion (POST /track/batch) for offline-first SDKs — IoT, mobile, edge
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- typescript
Hướng nghiên cứu
Bắt đầu từ entry point POST /track hiện có và pipeline theo từng loại của nó. Theo dõi cách hoạt động của việc validation, các giá trị __timestamp trong lịch sử, việc suy ra session và xử lý session_start trước khi xác định luồng batch. Được coi là hoàn tất khi các tiêu chí chấp nhận đều đạt, bao gồm từ chối một phần, giới hạn request, hành vi khi chạy đồng thời và hành vi không thay đổi của sự kiện đơn lẻ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
TL;DR
Add a POST /track/batch endpoint that accepts multiple events in a single HTTP request, back-dated to their original timestamps. This unblocks any SDK that captures events while offline and syncs them later — IoT firmware, mobile apps in airplane mode, edge devices on a wake schedule, etc.
The problem in one sentence
Every event today costs one HTTP request, and a buffered event's __timestamp is honoured for the row but not for session derivation — so a 6-hour-old event ends up in whatever session is currently active, not the session it belonged to when it was captured.
Why this matters for IoT
A real example. A solar-powered environmental sensor wakes every 10 minutes, takes a reading, and tries to upload. The cell modem has 70% uptime in its deployment area, so often the upload fails and the reading is buffered locally. When connectivity returns, the device has 50–500 buffered readings to flush.
With OpenPanel today, the device options are:
- Send one HTTP request per buffered reading. 500 round-trips. Every cell radio wake-up burns ~3 seconds of full-power TX. A 500-event flush is ~25 minutes of radio-on time. Battery dies faster than the panel can recharge it. This is why IoT teams build their own analytics pipelines instead of using product analytics tools.
- Stamp everything with arrival time. Loses the why — you have a count, but not a meaningful timeline. Any cohort analysis, retention curve, or session-based metric is wrong because all events get bucketed into the moment of network reconnection rather than when they actually happened.
- Build an aggregator. Have the firmware compute its own counters and send a daily summary. Works, but defeats the point of using a product analytics tool — you are back to dashboards built on pre-aggregated data, can't slice by user property, can't backfill a new dimension without redeploying firmware.
The same shape applies to:
- Mobile apps in airplane mode, subway, low-battery push deferral, OS-level network restrictions
- Wearables that sync over Bluetooth bursts when paired with a phone
- Kiosks that batch-sync on a wake schedule (e.g., overnight)
- Server-side replays importing from a different system
Mixpanel solved this years ago — they accept up to 2000 events per /import call, dedupe by insert_id, and respect a 5-day historical window. PostHog has similar (/batch endpoint, no hard cap but recommends 1000). The fact that we don't have it is one of the bigger gaps when evaluating OpenPanel as a Mixpanel replacement for product teams that have any flavour of offline-first behaviour.
What I'm proposing
POST /track/batch
Authorization: client-id + client-secret (same as /track)
Content-Type: application/json
Body:
{
"events": [
{ "type": "track", "payload": { "name": "...", "properties": { "__timestamp": "...", "__deviceId": "..." } } },
{ "type": "identify", "payload": { ... } },
{ "type": "group", "payload": { ... } },
{ "type": "increment", "payload": { ... } },
...
]
}
Per-request limits (Mixpanel parity):
- Up to 2000 events per request
- Up to 10 MB uncompressed body
- Beyond either, return 400/413
Acceptance window: events with a __timestamp up to 5 days in the past are accepted; older events are rejected per-row with reason: 'validation'. 1-minute future tolerance, beyond that the server clamps to wall-clock now (matches existing single-event behaviour).
Behaviour:
-
Each event in the batch is processed as if sent individually through
/track. Same validation, same per-type handlers (track,identify,increment,decrement,group,assign_group,replay). Thealiastype is rejected per-row with the same error single-event/trackreturns. -
Per-item validation failures don't fail the whole batch. Response is always 202 once auth + envelope pass:
{ "accepted": 1998, "rejected": [ { "index": 12, "reason": "validation", "error": "payload.name: Too small: expected string to have >=1 characters" }, { "index": 47, "reason": "validation", "error": "event timestamp older than 5 days" } ] }The caller can fix and retry only the bad indices instead of having to re-send 1998 good events.
-
Per-event timestamp respected for session derivation. A batch covering 5 days of buffered readings produces the right cluster of historical sessions, with
session_startrows back-dated to each event's actual timestamp. This is the part that makes the dashboard look correct after a backfill — without it, all 500 IoT readings collapse into one session at upload time, retention curves are meaningless, and any timeline-based analysis breaks.
Why now (vs. workarounds)
The argument against doing this is "users can hit /track 500 times in a loop." That's true on paper but it has three real costs that show up in production:
- Network: 500 connection setups instead of 1. With keep-alive on the server side that's ~50ms × 500 = 25s of just TLS handshake / HTTP framing overhead, before any code runs.
- Backpressure on the device: a 500-element queue with no batch endpoint means each event has to await its own HTTP response, or you fire-and-forget 500 requests and overwhelm the device's network stack.
- Session attribution wrong by default: even if you handle the network, you still get the timestamp problem unless the SDK and server collaborate on deriving session_id from
__timestamp. Right now the server treats wall-clock-now as the bucket key for__deviceIdoverrides, so a buffered event lands in whatever session is currently active for that device.
Solving all three at once with a batch endpoint that respects timestamps is much cleaner than asking SDK authors to work around them.
What this issue does NOT cover
- Idempotent retries (insert_id / messageId-based dedup). A separate concern: what should happen when a flaky network causes the same batch to be sent twice? Two reasonable designs (Mixpanel insert_id vs Segment messageId) and the trade-offs (storage overhead, lookup cost, replay semantics) deserve their own discussion. The batch endpoint as proposed here writes both copies if you send it twice — fine for reliable networks, not safe for at-least-once retry loops.
- Compression (gzip request bodies). Should be straightforward to add via Fastify; not blocking.
- SDK changes. API-only for now. Once the endpoint is in, SDK PRs can adopt batch on a per-platform schedule. Web/Node SDKs probably don't need it; React Native, mobile, and any custom IoT SDK would benefit immediately.
Open questions
- Is 5 days the right historical window, or should it be configurable per-project? Mixpanel uses 5 days. Going longer makes the deterministic session bucket more expensive to dedup (need a wider Redis lock TTL), going shorter excludes some legitimate offline-first use cases (e.g., devices that only sync weekly).
- Should
rejected[]items get a stable enum ('validation' | 'internal' | 'rate_limited') or freeform string? Current proposal:'validation' | 'internal'so callers can distinguish "I sent bad data" from "your server hiccupped." - Should the response include the queued
deviceId/sessionIdper accepted item? Currently it's just a count + rejected list. Including per-item identities would be useful for SDKs that want to update their local cache, but it doubles the response size for the common all-success case.
Acceptance criteria
-
POST /track/batchaccepts up to 2000 events / 10 MB and dispatches each via the same per-type pipeline as/track. - Per-item validation failures don't fail the batch; response is 202 with
{ accepted, rejected[] }. - Events with historical
__timestampget asession_idderived from that timestamp (deterministic 30-min bucket), not from wall-clock now. -
session_startis emitted exactly once per(projectId, sessionId)even when multiple workers / batches see the same bucket simultaneously. - Historical events do not extend the live
sessionEndjob or push current-session state forward. - Events with
__timestampolder than 5 days are rejected with a clear error. - Existing single-event
/trackbehaviour is unchanged — no regressions.
I have an implementation that's been running in production against a self-hosted instance with the changes verified across 24 scenarios (IoT 7-day backlog, cross-bucket boundary, multi-device household, kiosk, concurrent batches, etc.). Opening a PR alongside this issue.
- Ngôn ngữ chính
- TypeScript
- Star
- 7.1k
- Fork
- 510
- Merge trung bình
- 7 ngày 2 giờ
- Pull request đã merge (30 ngày)
- 7
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Không có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của Openpanel-dev/openpanel
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Openpanel-dev/openpanel#532 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Period comparison badge shows wrong percentage for decreases (100 → 50 shows ↓100%)Có thể đã có người làm @sarmah-rup đã nhận 9 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Openpanel-dev/openpanel#526 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Self-hosted missing op1-replay.jsCó thể đã có người làm @houstona đã nhận 9 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Openpanel-dev/openpanel#512 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
login page needs refinementCó thể đã có người làm @anandghegde đã nhận 23 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Openpanel-dev/openpanel#495 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
Openpanel-dev/openpanel#528 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của Openpanel-dev/openpanel
Issue tương tự
-
DB-plane provider_chat_options.* is accepted by config set but never merged into the loaded configĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Bump Firebase JS SDK (12.19.0 → 13.0.0)Có thể đã có người làm @SelaseKay đã nhận hôm nay. Đang mởNeeds Attention type: enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
invertase/react-native-firebase#9364 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
clouflaure de fernandoĐang mởenhancement
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
cloudflare/mcp#271 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 4 ngày
-
[fullsend] E2E: rhdh-version-override — run-e2e.sh overrides RHDH_VERSION to non-existent 2.1Đang mởe2e-failure ready-to-code
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
redhat-developer/rhdh-plugin-export-overlays#4261 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày