test(importer): the paged hermes session enumeration is unverified against a real store
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- testing-qa
Research direction
Start with list_exportable_sessions in raven/importer/scanners/hermes.py, then inspect _collect_window and _uncovered. Run raven import scan and raven import run --tier full against a Hermes home with more than 100 ended sessions on macOS or Linux and Windows, recording probes, timing, uncovered counts, and cron/ACP ids. Pin observed row shapes in tests/test_importer_hermes_scanner.py and address any failures found.
Written by the indexing model from the issue text.
Description
Problem
list_exportable_sessions (raven/importer/scanners/hermes.py) has two paths. When
hermes counts at most the 100 rows its dry-run listing prints, one unwindowed probe answers
the question. Above that, the started_at axis is partitioned into half-open,
minute-aligned windows and probed recursively, and _uncovered reconciles the collected
ids against the header count.
#264 ships that second path, but it has only ever run against synthetic stores: the
install it was developed against had fewer than 100 exportable sessions. The PR
description flags it as one of two unverified paths. What is unproven on real data:
- Listing rows. The parser takes the first token of any indented line. Real stores
carry ids the timestamped default does not look like -- cron ids (cron_<job>_<stamp>)
and bare uuid4 from ACP -- and a row that fails to parse is caught only by the
len(ids) > expectedarithmetic, which cannot see a row that parses into something
wrong. - The reconciliation under movement.
reopen_session()clearsended_atand only
ended sessions are candidates, so a session can leave the candidate set mid-scan. The
code counts that rather than raising, which is right, but it also means a genuine
coverage bug and a benign race are indistinguishable from the outside. Nobody has seen
what the count actually looks like on a store in use. - Cost. Probe counts and wall clock were measured against synthetic stores in the #264
discussion (17 / 27 / 79 / 159 probes for 150 / 500 / 2000 / 5000 sessions), not against
a real one. - Windows.
_PARTITION_FLOORmoved to 1970-01-02 inbf27da2, which closes the
pre-epochOSErrorclass, but the paged path itself has never run on Windows.
Proposal
Run the paged path against a real Hermes home with more than 100 ended sessions:
raven import scanandraven import run --tier full, on macOS or Linux and once on
Windows.- Record the probe count, the wall clock, and whatever
_uncoveredreports; confirm the
ids that come back cover the cron and ACP shapes. - Fix whatever that turns up, and pin the real row shapes as fixtures in
tests/test_importer_hermes_scanner.pyso the parser is anchored to observed output
rather than to the format as read.
Alternatives considered
Bounded concurrency was raised and measured in the #264 discussion: asyncio.gather on
the two halves saves 10-22s at 2000-5000 sessions but, because _collect_window is
recursive, expands the tree exponentially -- peak concurrent hermes processes 28 and 72 for
those two sizes. A Semaphore(4) version keeps most of the gain. It was correctly left out
of #264: enumeration is under 1% of a full-tier import, and the correctness of the bounded
version would rest on the same synthetic data this issue exists to replace. Worth revisiting
only after the path has run for real.
Area
Memory / skills
Additional context
Introduced by #264. The two-path split is at list_exportable_sessions; the recursion is
_collect_window; the reconciliation is _uncovered. The 100-row cap is upstream
(hermes_cli/sessions_cmd.py, candidates[:100]), so the paged path is reachable by any
user with a long Hermes history, not an edge case.
- Dominant language
- Python
- Stars
- 4.1k
- Forks
- 94
- Avg merge
- 9h 54m
- Merged PRs (30d)
- 370
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EverMind-AI/Raven
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#798 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
EverMind-AI/Raven#797 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
EverMind-AI/Raven#640 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EverMind-AI/Raven#479 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 75/100
EverMind-AI/Raven#474 · 2 comments ·
Maintainers usually reply within 1 day
All issues in EverMind-AI/Raven
Similar issues
-
customer-reported
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Azure/azure-cli#34150 · 1 comment ·
Maintainers usually reply within 1 day
-
community-request
Difficulty 1/5 Under an hour Newbie friendliness 95/100
NVIDIA-NeMo/Curator#2464 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
WeblateOrg/translation-finder#1099 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
trezor/trezor-firmware#7997 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day