`usage.get` and `simlock stats`: the figures from the event history, on a worker and a gateway
I maintainer di solito rispondono entro 1 giorno
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Part of #329.
Scope
After this PR, simlock stats prints the usage figures for a window, as a table and as --json, on a worker and on a gateway, from the event history and nothing else. The programmatic client gets usage.get. ADR 0016 §1, §2, §5, §6, §7 and §8 fix how the figures are computed.
Technical spec
Modules touched
src/core/usage/(new, pure, platform-agnostic) —computeUsage(events, window, options)takes event envelopes in time order, including the carried steps below, and returns the figures.options.fleetselects the gateway rule (ADR 0016 §6);options.labelsis a map from requester id to label the caller fills.- Which window a fact belongs to: a request belongs to the window its
lease.requestedfalls in; its grant, rejection and end are joined from events up toto, so an end aftertois unseen and the lease is open attowith no held or turnaround sample. Provisioning, boot and incident counts go by the device event's timestamp. Utilisation and queue depth come from the steps; where no step precedes the window, the series points before the first step in it arenulland excluded from peak and mean.requestscountslease.requestedin the window;rejectedcounts the rejections of those requests, plus rejections refused before admission (nolease.requested), which belong to the window theirlease.rejectedfalls in; sogranted + rejectedmay exceedrequestsanddocs/CLI.mdsays so. A request from beforefromthat is granted or rejected inside the window is in no count. - Incidents:
quarantinedcountsdevice.quarantined;crashRecoveredcountsdevice.recovered;quarantineRecoveredcountsdevice.quarantine-recovered;lostcountsdevice.recovery-failed,device.quarantine-abandonedanddevice.quarantine-stranded. - Fleet rule (ADR 0016 §6, amended by this PR): a fleet request's outcome is the first relayed
lease.grantedorlease.rejectedfor its namespaced requester at or after the gateway's ownlease.requestedfor it and before the gateway's nextlease.requestedfor that requester;request.dispatchednames the worker when it exists, else the relayed event'sworkerId; a rejection carries the worker's reason; a request with no outcome in those bounds is open. Every other relayedlease.rejectedis ignored. The grant may arrive beforerequest.dispatched, since a warm device is granted before the dispatch answer reaches the gateway. This PR rewords ADR 0016 §6 to say so. - Series: the bucket width is the smallest of 1 minute, 5 minutes, 15 minutes, 1 hour, 6 hours and 1 day that keeps the series at or under 200 points.
src/bus/event-file.tsandEventHistory.replay— acarryoption: for each named event (capacity.changed,queue.changed), the latest envelope at or beforesinceTsis kept in the result, one perworkerIdwhere present. The reader scans every line already, so this adds no I/O.src/contract/errors.ts—HISTORY_NOT_KEPTjoins the error table as adomainerror withcliExitCode: 12,httpStatus: 422and details{ oldestTs }, so the CLI and the HTTP route take their behaviour from the table.src/contract/operations.ts—usage.get, roleadmin, effectread, input{ from, to }as epoch milliseconds withfrom < toandto - frombounded (at most 90 days).src/contract/schemas.tsfor the output shape.src/contract/operations.test.tsrole matrix.src/daemon/dispatcher.ts— worker handler: reads the history for[from, to]throughEventHistory(file plus ring, asevents.replaywithsinceTs, withcarryfor the two step events), fills labels from the token store, callscomputeUsagewithfleet: false, and answersHISTORY_NOT_KEPTwhen the oldest held timestamp is later thanto. It roundsfromandtodown tobucketMs, memoises the last answer by (rounded window, newest event id), and answers a call inside the same bucket while no event arrived from the memo without reading the history; the answer'swindowis the rounded one.src/gateway/dispatcher.ts— the same from the gateway's merged history withfleet: true, stripping its owngw:<instance>:prefix before the label lookup, with the same memo. The memo exists once, shared by both handlers.src/daemon/server.tssocket switch;src/simlock-client/client.tsandtypes.ts—usage(window).src/cli/index.ts—simlock stats [--since <duration> | --from <ISO> [--to <ISO>]] [--json]; default window the last 24 hours;--sinceand--fromtogether is a usage error. Human output: a header with the window andcoversFrom, the "partial" note when set, then the totals, the per-platform rows, the per-worker rows, the per-requester rows.docs/CLI.md,docs/CLIENT.md,docs/HTTP-API.mderror code list if it is shared.
Contract and event changes
- New operation
usage.get. Output, one object:window { from, to },coversFrom,partial: boolean,bucketMs.totals:requests,granted,bySource { warm, booted, provisioned },rejected { total, byReason },wait { p50, p95, max, count },held { p50, p95, max, count },turnaround { p50, p95, max, count },provisioning { p50, p95, max, count },boot { p50, p95, max, count },utilisation { slots { peak, mean, max }, ram? { peakBytes, meanBytes, limitBytes } },queue { peakDepth, meanDepth },incidents { quarantined, crashRecovered, quarantineRecovered, lost },failures { byEvent }: a count per failure event name (device.purge-failed,device.recovery-failed,component.install-failed, and any other*-failedevent) in the window by the event's timestamp.platforms: the same figures keyed byiosandandroid.workers: one entry per worker id with the same figures and alabel; on a worker, one entry for itself.requesters:{ id, label?, requests, granted, rejected, heldTotalMs }[], byrequestsdescending.series:{ at, slotsUsed, slotsMax, ramUsedBytes?, queueDepth, waiting }[], one point per bucket,waitingthe number of requests waiting at the bucket's end.- Durations and percentiles in milliseconds; a percentile over an empty set is
null.
- New error code
HISTORY_NOT_KEPTin the contract's error table (domain,cliExitCode: 12,httpStatus: 422) witholdestTsin its details. - Protocol range unchanged: the gateway does not call workers for this.
Rules in play
architecture.mdrules 1 and 10: the figures module has no platform or transport imports; the fleet dedupe rule and the plan-source map exist once.safety.mdrule 10:fromandtoare wire input, bounded before use.testing.mdrule 1: each percentile test states the input set and the expected number.
Tests
computeUsagecounts a request granted, held and released as one request, one grant, and one sample in wait, held and turnaround with the expected millisecondscomputeUsagereports p50, p95 and max of a set of twenty known waitscomputeUsagecounts a rejected request under its reason and not under grantscomputeUsageattributes grants towarm,bootedandprovisionedfromlease.granted.sourcecomputeUsagereads provisioning and boot durations fromdevice.provisionedanddevice.readycomputeUsagederives peak and time-weighted mean slot utilisation fromcapacity.changedsteps, including the step in force before the window startscomputeUsagereports RAM utilisation whencapacity.changedcarriesramBudgetand omits it otherwisecomputeUsagederives peak and mean queue depth fromqueue.changedcomputeUsagecountsquarantined,crashRecovered,quarantineRecoveredandlostfrom the named device events and nothing elsecomputeUsagecounts a request whoselease.requestedis beforefromand whose grant is inside the window under neither requests nor grantscomputeUsagegives no held or turnaround sample for a lease still open atto, and counts its request and grantcomputeUsagecounts a rejection with no precedinglease.requestedunderrejectedand not underrequestscomputeUsagecounts a request from beforefromthat is rejected inside the window under neither requests nor rejectionscomputeUsagereportsnullseries points and excludes them from peak and mean before the first step when no step precedes the windowcomputeUsagegroups by platform and byworkerIdcomputeUsagewithfleet: truecounts a fleet request once, from the gateway's own events, and attributes its grant to the worker named byrequest.dispatchedcomputeUsagewithfleet: trueignores relayedlease.requested,lease.queuedandqueue.changed, and a relayedlease.rejectedthat matches no dispatched requestcomputeUsagewithfleet: truesettles a dispatched request as rejected, with the worker's reason, from the relayedlease.rejectedof the worker it was dispatched tocomputeUsagewithfleet: truedoes not join a requester's second dispatch's grant to its first request, which stays opencomputeUsagesetspartialandcoversFromwhen the oldest event is inside the windowcomputeUsagepicks 1-minute buckets for one hour, 15-minute buckets for one day and 1-hour buckets for seven days, and never more than 200 points for a 90-day windowreadEventFilewithcarryreturns the latestcapacity.changedandqueue.changedat or beforesinceTs, one per worker id, and none when there is none- the worker handler answers a second call inside the same bucket from the memo without reading the history, and reads again after an event arrives or the bucket moves
computeUsagelists requesters by request count with the label the caller supplied- the worker handler answers
HISTORY_NOT_KEPTwith the oldest held timestamp when the window ends before it - the gateway handler strips its own
gw:prefix before the label lookup and leaves another prefix alone usage.getis anadminoperation anagenttoken cannot callsimlock statsrejects--sincewith--from, and--fromlater than--to, as usage errorssimlock stats --jsonprints the operation's output unchanged- (e2e) after a scripted run of leases against the fake driver,
simlock stats --since 1h --jsonreports the counts a test derives independently fromsimlock events --since 1h - (e2e)
simlock stats --since 1hreturns the same figures before and aftersimlock daemon stopandsimlock daemon start - (e2e) on a gateway with two workers,
simlock statshas fleet totals and one row per worker, and each worker's ownsimlock statscovers itself only
Done when
- After a scripted run against the fake driver with waits, a
--no-waitrejection and a timeout,simlock stats --since 1hprints request count, grants by source, held time and turnaround, wait p50, p95 and max, provisioning and boot durations, peak utilisation, queue depth, rejections by reason and incidents;--jsonprints the same figures; and the counts match what a reader counts by hand fromsimlock events --since 1h. - The same window prints the same figures after
simlock daemon stopandsimlock daemon start. - On a gateway with two workers, the output has fleet totals and one row per worker; each worker's own output has one row, itself.
- A request that waited shows in the wait figures; a request that timed out and one cancelled are counted under rejections, not grants.
- The per-requester rows show each requester's leases in the window, with the token label beside the id for a lease taken over HTTP.
- After a
--no-waitrefusal,simlock stats --since 1hshowsrequestsandrejected.byReason.no-waiteach one higher than before it andgrantedunchanged. simlock stats --from <a day ago> --to <two days ago>is a usage error;--since 30don a history whose oldest event is a minute old prints the figures with "Figures cover from " above them; a window that ends before the oldest held event prints theHISTORY_NOT_KEPTmessage and exits non-zero.docs/CLI.mddescribessimlock stats, its flags, the partial note and the error.
Out of scope
- The HTTP route and the CSV export (next task).
- The console view.
Depends on
- #345
- #346
Approval
- Approved for delivery
Written by an agent.
- Lingua principale
- TypeScript
- Stelle
- 15
- Fork
- 1
- Merge medio
- 8h 23m
- PR unite (30g)
- 136
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di callstackincubator/simlock
-
bug:new
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
callstackincubator/simlock#424 ·
I maintainer di solito rispondono entro 1 giorno
-
bug:new
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
callstackincubator/simlock#422 ·
I maintainer di solito rispondono entro 1 giorno
-
bug:new
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
callstackincubator/simlock#420 ·
I maintainer di solito rispondono entro 1 giorno
-
bug:new
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
callstackincubator/simlock#350 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Flaky e2e: gateway-fleet reads a worker's capacity before the worker has left startingForse già presa @V3RON l’ha presa 2 giorni fa. Apertabug:ready flaky-test
Difficoltà 4/5 1-2 giorni Idoneità per principianti 48/100
callstackincubator/simlock#414 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di callstackincubator/simlock
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 84/100
answerLoops/answerLoops#345 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
siyuan-note/siyuan#20313 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
LanternOps/breeze#8254 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 82/100
gofish-graphics/gofish-graphics#1084 ·
I maintainer di solito rispondono entro 1 giorno