Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Data layer: current-state map of data sources and paging candidates

Đang mở
#403 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
18/100
Loại issue
Tài liệu
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
typescript
Lĩnh vực
backend, databases

Hướng nghiên cứu

Start with the data source table and check each row against the pinned commit 9a4249b, beginning with projectStorage.ts for the serialization queues and the draft shard logic (getDraft, saveDraft), then updateAnalysis and updateProjectMetadata for the project blob. Also read useInterlinearizerBookData, fw-lite-lexicon.ts and the validators to confirm the remaining rows. Done means each claim in the inventory and diagrams is verified against the code at that commit, and the team has agreed the paging candidates.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Purpose

We want to move the extension's data to a database backend that answers focused queries: no whole-book saves and no full-book or full-project queries. The main reason is performance with a secondary goal of supporting conflict free collaboration through a CRDT backend. This issue is the starting point for a plan the team will develop together. It maps where every piece of data comes from today and every place that would need paging.

It describes the current state only and proposes no design. All code links are pinned to commit 9a4249b.

Two facts drive most of what follows:

  1. The analysis is held in memory for the whole project. A single TextAnalysis covering every book is held in the draft hook and mirrored into Redux. Reducers scan its flat arrays linearly. Selectors build memoized indexes over the whole project, and those indexes are rebuilt whenever an array changes.
  2. The text is held in memory one whole book at a time. Each book arrives as a single USJ document. Rendering is windowed, but the data is not.

Data source inventory

Source Where it comes from Fetched / stored as What drives its size Paging candidate?
Scripture text platformScripture.USJ_Book PDP, via useInterlinearizerBookData The whole book, keyed by scrRef.book only, then extracted and tokenized synchronously Book length Yes
Draft analysis papi.storage keys draft:{src}, draft:{src}:analysis:{BOOK} and draft:{src}:shards (keys) Stored as an envelope plus one shard per book, plus a no book shard for payloads no link references. getDraft merges every shard into one object. saveDraft receives the whole draft JSON, rewrites only the shards whose JSON differs from what the backend last wrote or loaded, and always rewrites the envelope Whole project (all books, all analyses) Yes
Saved project papi.storage key project:{id} One monolithic blob: metadata + analysis + links + segmentation. Every write is a read-modify-write of the whole blob (updateAnalysis, updateProjectMetadata) Whole project Yes
Project index papi.storage keys projectIds and pendingCleanup Whole arrays Number of projects (typically under 10, at worst the low hundreds) No; but listing reads every project in full (see below)
Book list platformScripture.booksPresent project setting (useProjectBookIds) One string Small No
Project settings continuousScroll, showMorphology, showFreeTranslation (validators); platform.languageTag Single values Small No
WebView state useWebViewState keys: activeProject, analysisLanguageChoices, sidePanel, sidePanelLayout, dismissedStaleAnalyses, dismissedLostBoundaries Small per-tab values Small No
Lexicon registry The FieldWorks Lite lexicon link, read from a project setting (fw-lite-lexicon.ts) A link id, plus lookups made through the provider Not ours to page No

Diagram: data flow

Thick arrows (==>) carry a whole book, a whole project or a whole blob.

flowchart TB
  subgraph HOST["Platform.Bible host"]
    USJ["USJ_Book PDP"]
    STORE[("papi.storage<br/>one file per key")]
    SET["Project settings"]
  end

  subgraph EXT["Extension host"]
    CMD["Commands<br/>getDraft, saveDraft, getProject,<br/>saveAnalysis, getProjectsForSource, ..."]
    PS["projectStorage.ts<br/>serialization queues per project,<br/>per draft, and for the index"]
  end

  subgraph WV["WebView"]
    BD["useInterlinearizerBookData<br/>extract + tokenize"]
    DR["useDraftProject<br/>draft ref, undo snapshots,<br/>debounced autosave"]
    AS["AnalysisStore (Redux)<br/>analysis for every book"]
    VIEW["Views: segment list, continuous,<br/>catalog, stale notice"]
    CON["Concordance index"]
  end

  USJ ==>|"whole book"| BD
  USJ ==>|"every book present"| CON
  PS ==>|"project blob, draft shards"| STORE
  STORE ==>|"project blob, all draft shards"| PS
  PS --- CMD
  CMD ==>|"getDraft: merged draft"| DR
  DR ==>|"saveDraft: whole draft JSON<br/>Save: whole analysis JSON"| CMD
  DR ==>|"seed; replaceAnalysis on<br/>undo, redo, reanchor"| AS
  AS ==>|"full analysis on every edit"| DR
  BD ==>|"whole Book: reanchor pass"| DR
  BD ==>|"whole Book"| VIEW
  AS --> VIEW
  AS ==>|"whole analysis"| CON
  SET --> BD

Diagram: entities

Book, Segment and Token are not stored; they are rebuilt from USJ on every load. Links point at them by tokenRef (e.g. GEN 1:1:0) or segmentId. The book prefix of the ref is the key the draft shards use today.

Payload, as used throughout this issue, means one record in a TextAnalysis's tokenAnalyses, phraseAnalyses or segmentAnalyses array: a TokenAnalysis, PhraseAnalysis or SegmentAnalysis. A payload holds what a span of text means (gloss, morphemes, sense, or free and literal translation), its own id, and the surface form it analyzes. It records no location: it doesn't say which token, phrase or segment it applies to. It also has no status: a payload itself is never approved, suggested or stale. Both of those live on links, the per-occurrence records (TokenAnalysisLink, PhraseAnalysisLink, SegmentAnalysisLink) that attach a payload to one location by analysisId and carry that occurrence's status.

Token and phrase payloads are shared by every occurrence analyzed identically, including across books. Segment payloads are one-to-one with their links. Elsewhere in this issue, "the analysis" or "the whole analysis" means the entire TextAnalysis object, which holds both payloads and links.

Every token, phrase and segment link lies within one book; a phrase may span segments, but not books. The type puts no book constraint on alignment links, but nothing creates them yet.

erDiagram
  INTERLINEAR_PROJECT ||--|| TEXT_ANALYSIS : "analysis"
  DRAFT_PROJECT ||--|| TEXT_ANALYSIS : "analysis"
  INTERLINEAR_PROJECT ||--o{ ALIGNMENT_LINK : "links"
  TEXT_ANALYSIS ||--o{ TOKEN_ANALYSIS : "tokenAnalyses"
  TEXT_ANALYSIS ||--o{ TOKEN_ANALYSIS_LINK : "tokenAnalysisLinks"
  TEXT_ANALYSIS ||--o{ PHRASE_ANALYSIS : "phraseAnalyses"
  TEXT_ANALYSIS ||--o{ PHRASE_ANALYSIS_LINK : "phraseAnalysisLinks"
  TEXT_ANALYSIS ||--o{ SEGMENT_ANALYSIS : "segmentAnalyses"
  TEXT_ANALYSIS ||--o{ SEGMENT_ANALYSIS_LINK : "segmentAnalysisLinks"
  TOKEN_ANALYSIS ||--o{ TOKEN_ANALYSIS_LINK : "analysisId"
  TOKEN_ANALYSIS ||--o{ MORPHEME_ANALYSIS : "embedded"
  PHRASE_ANALYSIS ||--o{ PHRASE_ANALYSIS_LINK : "analysisId"
  SEGMENT_ANALYSIS ||--o{ SEGMENT_ANALYSIS_LINK : "analysisId"
  TOKEN_ANALYSIS_LINK }o--|| TOKEN : "token.tokenRef"
  PHRASE_ANALYSIS_LINK }o--|{ TOKEN : "tokens[].tokenRef"
  SEGMENT_ANALYSIS_LINK }o--|| SEGMENT : "segmentId"
  SEGMENT ||--o{ TOKEN : "rebuilt from USJ"

  INTERLINEAR_PROJECT {
    string id PK
    string sourceProjectId
    number modelVersion
    json segmentation
  }
  DRAFT_PROJECT {
    string sourceProjectId PK
    number modelVersion
    boolean dirty
  }
  TOKEN_ANALYSIS {
    string id PK "shared payload"
    json gloss "MultiString"
  }
  TOKEN_ANALYSIS_LINK {
    string analysisId FK
    string tokenRef "book prefix"
    string surfaceText "drift check"
    string status
  }
  PHRASE_ANALYSIS_LINK {
    string id PK "one per occurrence"
    string analysisId FK
  }
  SEGMENT_ANALYSIS_LINK {
    string analysisId FK
    string segmentId
    string status
  }
  TOKEN {
    string ref "not stored"
  }
  SEGMENT {
    string id "not stored"
  }

How to read this diagram

Each line joins two entities. The symbol at each end of a line says how many of the entity at that end go with one of the entity at the other end.

Line-end symbol Meaning
Two short bars Exactly one
Circle and crow's foot (three-pronged fork) Zero or more
Bar and crow's foot One or more

For example, the line between TOKEN_ANALYSIS and TOKEN_ANALYSIS_LINK has two bars at the TOKEN_ANALYSIS end and a circle with a crow's foot at the link end. So one token analysis has zero or more links, and each link points at exactly one token analysis.

The text on a line names the field or array that holds the relationship (for example, analysisId or tokenAnalyses). Inside a box, each row is a type and a field name. PK marks the field that identifies the record, FK marks a field that refers to another record, and quoted text is a note. Each box lists only the fields that matter here, not every field of the type.

How the counts grow: there is about one link per analyzed occurrence (plus competing links distinguished by status), so links grow with the amount of text analyzed. Token and phrase payloads grow with the number of distinct analyses; segment payloads grow with their links.

Read paths today

  • Book load.
    1. BookUSJ fetches the whole book.
    2. extractBookFromUsj and tokenizeBook run in a useMemo.
    3. Re-segmentation and re-heading run over the whole book.
    4. The reanchor pass calls reanchorAnalysisToBook, which scans every link in the project to pick out this book's.
  • Draft load. getDraft reads every shard listed in the manifest (plus any orphans in the journal), merges them with mergeAnalyses, and returns one object. That object seeds the Redux store, which then holds every book.
  • Project lists. listProjects and getProjectsForSource read the index and then every project in full, analysis included. The picker shows only metadata plus two facts worked out from each analysis (summarizeAnalysis): the books its token links reference and its token-analysis count.
  • Concordance. While the index is enabled, useConcordanceIndex reads and tokenizes the USJ of every book present in the project, four at a time. It indexes text only. The panel subscribes to the whole analysis and joins it against the rows on screen.

Write paths today

  • Edit and autosave.
    1. A reducer runs over the whole arrays.
    2. save() reads the full state.analysis.
    3. autosaveDraft pushes the previous full content onto the undo stack.
    4. After a 300 ms debounce, persist sends JSON.stringify(draft) as one message.
    5. The backend splits the draft by book and writes the shards whose JSON changed.
  • Save. handleSave sends the whole analysis to saveAnalysis, which rewrites the whole project:{id} blob.
  • Save As. createProject then saveAnalysis, or, when overwriting, saveAnalysis then updateProjectMetadata. Both are whole-blob.
  • Undo and redo. undo-history.ts keeps up to 100 whole-draft {analysis, segmentation} snapshots. These are references to immutable content, not deep copies. Undo restores a snapshot, re-anchored to the current text, and dispatches replaceAnalysis with the whole object.
  • Reanchor pass. The reanchor pass reruns whenever the loaded book's text, the segment boundaries or the whole draft change; gloss edits don't trigger it. If anything moved, it writes the whole healed draft and records the pass in the undo history.
  • Wipe. wipeBook and wipeAll replace the whole draft and persist it immediately, with no debounce.

Paging candidates by access pattern

1. Visible-range reads

What the screen shows is a window of segments. useSegmentWindow first mounts 12 segments either side of the anchor, then adds 24 at a time and culls by geometry; 400 mounted segments is only a runaway cap. But the data behind that window is the whole book's text and the whole project's analysis, and several derivations walk the whole book on every load.

Call sites
2. By id, across books

A token or phrase payload is shared by every occurrence analyzed identically, in any book. Looking one up, or writing to it, today means scanning or rewriting the whole project's arrays.

Call sites
3. Aggregates across books

These answers depend on every link in the project.

Call sites
4. Whole-book and whole-project derivations
Call sites
5. Whole-blob writes
Call sites
6. Project listing reads whole records

Projects are few, typically under 10 and at worst in the low hundreds, so the list itself doesn't need paging. The cost is that every listing reads each project's whole record, analysis included, only to compute summarizeAnalysis (the books referenced and the token-analysis count) for display. Opening the project picker therefore costs the total size of every project's analysis.

Call sites

Constraints any design must honor

These are facts about today's model, not proposals.

  • Payloads are shared across books. One TokenAnalysis or PhraseAnalysis can back occurrences in many books. The draft shards today copy a shared payload into each book's shard and dedupe on load.
  • At most one linked analysis per token, segment or phrase may be approved. The store's reducers maintain this when they write: they repoint or remove the earlier approved link, or refuse the approval. Nothing rejects a breach. validateTextAnalysis only logs it at the storage boundary.
  • A phrase occurrence is identified by PhraseAnalysisLink.id, never by its analysisId.
  • Book, Segment and Token are rebuilt from USJ on every load. For token and phrase links, the reanchor pass detects drift from TokenSnapshot.surfaceText when the baseline text changes. Segment analyses compare their own surfaceText with the segment's baselineText.
  • Records are gated by modelVersion. Every project and draft-envelope write stamps CURRENT_MODEL_VERSION, and a record stamped higher is refused. Draft shards, the shard journal and the index carry no stamp. There are no migrations yet.
  • Undo restores whole-draft snapshots, held in memory only. Undo history is never persisted.
  • A draft and a saved project are separate records. There is one draft per source project, and Save copies the draft's analysis into a project.
  • The backend serializes writes, in memory. Per-project and per-draft queues guard read-modify-write, alongside single queues for the project index and pendingCleanup. A shard journal recovers from interrupted saves, because papi.storage cannot list keys.
  • Future note (not in scope here): each user will eventually have their own send/receive changes, and those will be reconciled into a local snapshot.
PT9 import: out of scope, but it shares the storage

PT9 import is not a paging candidate and is left out of the maps above. It writes to the same storage, though, so the plan must account for it:

  • An import is stored as an ordinary project:{id} record carrying a pt9Import field (file hashes and import time). Each sync replaces the whole record (savePt9Import).
  • Finding a source's import scans every project (getPt9ImportForSource).
  • Imports are read-only. The functions that save an analysis or update metadata refuse them, and only a sync replaces one; an import can still be deleted.

Proposed baseline before implementation

Before any change, we should capture a baseline so the team can measure the effect of each step. Proposed metrics:

  • Book-load latency. From the book change to the first interactive render, split into USJ fetch, tokenize, reanchor and store mount.
  • Edit latency. From dispatch to re-render, including the selector index rebuild, plus the autosave round-trip.
  • Memory. The WebView heap with a large project loaded, including the undo stack.

The dataset and tooling are for the team to choose.

Notes from paranext-core

Facts from paranext-core at fcaa6d0, recorded here for planning and not as recommendations:

  • Text has chapter and verse getters. platformScripture offers USJ_Chapter and USJ_Verse alongside USJ_Book (types). Each chapter call converts USX to USJ in the extension host (getChapterUSJ). There is no range selector.
  • Data provider updates are not filtered by selector. An update names data types only. Each subscriber re-runs its get for its own selector and deep-compares the result (service). Any Scripture edit notifies every USFM, USX and plain-text data type for the project (AllScriptureDataTypes). The platform-scripture extender re-emits each USX update as the matching USJ update, so every USJ subscriber re-fetches, whatever its selector.
  • papi.storage writes one file per key (buildUserDataUri), with a plain writeFile (writeFile): no temp file and rename, and no size limit.
  • The bus closes the connection on any message over 100 MiB (MAX_WEBSOCKET_PAYLOAD_BYTES).
  • We found no performance harness or benchmark suite in core. The nearest are isolated timing assertions in e2e tests.

Open questions for the team

  • Which dataset and tooling should the baseline use?
  • Does the draft vs. saved-project split survive a database, or does it become something else?
  • Should the database reject a breach of the "at most one approved link" rule, rather than relying on reducers to maintain it and the storage boundary to log it?
  • Which aggregates that span books (suggestion pool, usage counts, stale summary, concordance) must be live, and which can be computed lazily or in the background?
  • Should a payload shared across books stay a single record, or be split per book the way the draft shards copy it today?
Ngôn ngữ chính
TypeScript
Star
2
Fork
0
Merge trung bình
2 ngày 3 giờ
Pull request đã merge (30 ngày)
46

Chuẩn bị môi trường

Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của sillsdev/interlinearizer-extension

Tất cả issue của sillsdev/interlinearizer-extension

Issue tương tự

Thêm issue về TypeScript

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.