Add cursor-based continuation to `/query` so paging depth is unbounded and cost is flat
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 45/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- javascript, mongodb
- Domain
- api, backend, databases, documentation
Research direction
Start by tracing the /query handler and its paging and Link-header generation paths, then inspect controllers/utils.js:70 for ID handling. Review the cursor-related requirements against the dev collection, and update public/API.html and the OpenAPI contract; done means all listed acceptance criteria pass, including mixed _id types and malformed-token handling.
Written by the indexing model from the issue text.
Description
Status: deferred. This was prototyped on the
302-query-optimizationbranch and removed before merge, so it can be considered later as its own feature. #302 uses askip-basedrel="next"instead, which is still sent past theskipmaximum and answered there with the #301400. The design and measurements below are what to start from.
Summary
skip cannot page past the configured maximum. Since #301, a request past it is a 400 instead of a repeating page. That is honest, but it also tells a client it may not read past 100000 records, which is a correct contract and a worse one than clients think they have today.
Keyset (cursor) continuation removes the ceiling. /query already sorts on a unique, indexed key since #300, so a cursor is "resume after this _id":
db.find({ "$and": [props, { "_id": { "$gt": lastId } }] }).sort({ _id: 1 }).limit(limit + 1)
Cost is flat with depth, because Mongo seeks the index rather than counting past skip documents. Depth is unbounded, because there is no offset to cap.
skip stays. Decision recorded on the parent thread: keep both modes and document them — skip as the bounded, random-access mode, cursors as the unbounded, sequential one.
Why this matters
It is what makes the skip rejection complete. Rejecting skip > 100000 with no alternative narrows what the API can do. #302's next link now ends a longer walk in that 400, which is honest. A next link that carries a cursor would let the same walk continue, with no change for a client that already follows the link.
Deep skip is expensive even inside the cap. Mongo walks and discards every skipped document. The parent draft measured /query at 340-445 ms across skip=0 through skip=99000, so this is not currently a crisis on /query — but it is linear work that a cursor makes constant, and it is the shape that matters once collections grow.
Clients cannot build a cursor themselves, and should not. idNegotiation() deletes _id from every response body (controllers/utils.js:180) and only reattaches it as id when the @context matches a known mapping. So the server must mint the token and hand it back. That is the right design anyway: an opaque, server-minted token keeps the encoding an implementation detail, and a client that only follows rel="next" needs no code change when the encoding changes.
Proposed design
The token
/query would accept a cursor query parameter. It is opaque to clients and appears only in the server's rel="next" link:
Link: <https://store.rerum.io/v1/api/query?limit=100&cursor=eyJfaWQiOiI2OWU5MTBkZjQ0YWNhMDQxNGMwZWYwMTYifQ>; rel="next"
The token is base64url of a small JSON envelope holding the _id of the last record served: {"_id":"69e910df44aca0414c0ef016"} above.
- Every
_idis a string.newID()returnsnew ObjectId().toHexString(), and every_idon production now and in the future is a string. The dev collection also holds legacy v0 documents whose_idis anObjectIdor an embedded object. That is old bad dev data, and cursor paging would not support it. - Anything the server did not issue would be a 400, not a guess:
- a value that is not base64url or not JSON
- an envelope other than exactly
{"_id": …} - an
_idthat is not a string - a repeated
cursorparameter
The filter
Because every supported _id is a string, "resume after this _id" is a plain $gt in the same order the sort serves:
{ "$and": [props, { "_id": { "$gt": lastId } }] }
The client's filter is combined under $and rather than spread with _id added, so an _id clause in the request body would still apply on cursor pages.
MongoDB comparison operators only match values of the same BSON type as the operand, so a cursor walk would never reach the legacy non-string _id documents on dev. A skip walk still pages into them, so on dev the two modes would differ for a query that matches them. That is expected. If a skip page on dev ends on one, its next link should advance skip instead of carrying a cursor.
Interaction with skip
cursorandskiptogether would be a 400, includingskip=0. They are two different positioning schemes and combining them has no coherent meaning.limitwould apply to both modes and be clamped identically.- Cursor pages would report
Pagination-Limit,Pagination-Limit-Max, andPagination-Skip-Max, but notPagination-Skip, because no offset was applied and reporting0would mislead a client that computes offsets. rel="next"on/querywould always carry acursor, including on a page requested byskip, replacing theskipvalue #302's link carries today. A client that follows the link moves onto the unbounded mode with no change on its side. That transparency is why #302 could ship first.skipwould keep working, keep its maximum, and be documented as the random-access mode for jumping to a known offset within the cap.
Scope
/query only. HEAD /query is removed by #304 rather than extended. See the note below on /search.
Prototype measurements
Measured on the prototype before it was removed, read-only, against the local pm2 app reading the dev collection (rerum-test):
- Past the
skipmaximum. Following onlyrel="next"atlimit=500over{"__rerum.history":{"$exists":true}}walked 102001 records in 205 pages, with no duplicates, and stopped whennextwas absent. - Same records as
skip. Pages atskip=0,50000, and99500were identical to the same slices of the cursor walk. - Flat latency. Cursor pages averaged 158 ms at depth 0-10k, 125 ms at 45-55k, 121 ms at 90-100k, and 97 ms past 100k.
skippages on the same query took 149 ms, 224 ms, and 365 ms atskip=0,50000, and99500, and cannot go further. - A token the server did not issue. An
ObjectId-shaped cursor,{"_id":{"$oid":"…"}}, was a 400.
Notes
/searchis deliberately excluded. Under the current in-memory paging (controllers/search.js:280), a cursor could only encode an offset into the merged array. Same cost, same ceiling, dressed up as a cursor — worse than not shipping one, because it would advertise a guarantee the implementation does not provide./searchpaging is tracked under #306; revisit a cursor with #309, where Atlas Search's ownsearchAfter/searchSequenceTokenpaging is the genuine equivalent. Availability on our cluster tier and server version needs confirming before that goes in any plan — do not assume it.- Depends on #300 (a sort key, closed) and #302 (somewhere to put the token, in place).
- The prototype was small. It added a
cursoroption togetPagination()that decoded the token, rejectedcursorwithskip, and skipped reportingPagination-Skip. It also added a token encode/decode pair incontrollers/utils.js, the$and/$gtfilter and cursor minting in the/querycontroller, and anextCursorinput to therel="next"helper (setNextPageLink()in #302). On the documentation side it added acursorrow and askip-versus-cursorparagraph inpublic/API.html, and aPageCursorparameter on/api/queryin the OpenAPI contract. - A keyset cursor is stable under insert in a way
skipis not: a document inserted before the cursor position does not shift the remaining pages. Given RERUM's mark-deleted-never-remove policy, this makes cursor walks meaningfully more correct than offset walks over a live collection, which is a second reason to prefer them beyond cost. - The cursor is not bound to the query it came from. Sending it with a different body means "records matching this body after that
_id", which is coherent and harmless. - A final page still scans to the end of the matching range, because the one-record over-fetch has to prove nothing follows. That is the same work a short final
skippage does today. - This is #252's third recommendation. Worth linking there.
Acceptance criteria
-
/queryaccepts an opaquecursorparameter and returns the page following it - The
rel="next"link on/queryalways carries acursor, including on a page requested byskip, and a client following only that link pages past theskipmaximum and terminates correctly - Cursor paging returns every matching document exactly once, over string
_idvalues; non-string_idlegacy dev data is not supported - A cursor walk and a
skipwalk of the same query return the same records in the same order, within the range whereskipis legal -
/querylatency is flat with respect to depth under cursor paging, measured to the depthskipcannot reach - A malformed, unparseable, or repeated
cursorreturns 400, as does a token whose_idis not a string -
cursorcombined withskip, includingskip=0, returns 400 - An
_idclause in the request body still applies on cursor pages - Cursor pages report
Pagination-Limitand both maximums, and noPagination-Skip -
skipcontinues to work within its maximum, and both modes are documented inpublic/API.htmland in the OpenAPI contract as acursorparameter on/api/query
- Dominant language
- JavaScript
- Stars
- 3
- Forks
- 6
- Avg merge
- 4d 9h
- Merged PRs (30d)
- 5
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from CenterForDigitalHumanities/rerum_server_nodejs
-
bug documentation
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
backend dependencies easy
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
CenterForDigitalHumanities/rerum_server_nodejs#290 · 1 comment ·
-
Difficulty 4/5 3-5 days Newbie friendliness 50/100
-
Difficulty 3/5 1-2 days Newbie friendliness 74/100
All issues in CenterForDigitalHumanities/rerum_server_nodejs
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
apache/rocketmq-dashboard#6067 ·
Maintainers usually reply within 4 days
-
severity: low
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
luainkernel/lunatik#1853 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
factory-active factory-automatic harness/claude-code task-identify-harness-labels-done task-identify-issue-type-done
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day