Data layer: current-state map of data sources and paging candidates
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 18/100
- Loại issue
- Tài liệu
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- typescript
Hướng nghiên cứu
Start with the data source table and check each row against the pinned commit 9a4249b, beginning with projectStorage.ts for the serialization queues and the draft shard logic (getDraft, saveDraft), then updateAnalysis and updateProjectMetadata for the project blob. Also read useInterlinearizerBookData, fw-lite-lexicon.ts and the validators to confirm the remaining rows. Done means each claim in the inventory and diagrams is verified against the code at that commit, and the team has agreed the paging candidates.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Purpose
We want to move the extension's data to a database backend that answers focused queries: no whole-book saves and no full-book or full-project queries. The main reason is performance with a secondary goal of supporting conflict free collaboration through a CRDT backend. This issue is the starting point for a plan the team will develop together. It maps where every piece of data comes from today and every place that would need paging.
It describes the current state only and proposes no design. All code links are pinned to commit 9a4249b.
Two facts drive most of what follows:
- The analysis is held in memory for the whole project. A single
TextAnalysiscovering every book is held in the draft hook and mirrored into Redux. Reducers scan its flat arrays linearly. Selectors build memoized indexes over the whole project, and those indexes are rebuilt whenever an array changes. - The text is held in memory one whole book at a time. Each book arrives as a single USJ document. Rendering is windowed, but the data is not.
Data source inventory
| Source | Where it comes from | Fetched / stored as | What drives its size | Paging candidate? |
|---|---|---|---|---|
| Scripture text | platformScripture.USJ_Book PDP, via useInterlinearizerBookData |
The whole book, keyed by scrRef.book only, then extracted and tokenized synchronously |
Book length | Yes |
| Draft analysis | papi.storage keys draft:{src}, draft:{src}:analysis:{BOOK} and draft:{src}:shards (keys) |
Stored as an envelope plus one shard per book, plus a no book shard for payloads no link references. getDraft merges every shard into one object. saveDraft receives the whole draft JSON, rewrites only the shards whose JSON differs from what the backend last wrote or loaded, and always rewrites the envelope |
Whole project (all books, all analyses) | Yes |
| Saved project | papi.storage key project:{id} |
One monolithic blob: metadata + analysis + links + segmentation. Every write is a read-modify-write of the whole blob (updateAnalysis, updateProjectMetadata) |
Whole project | Yes |
| Project index | papi.storage keys projectIds and pendingCleanup |
Whole arrays | Number of projects (typically under 10, at worst the low hundreds) | No; but listing reads every project in full (see below) |
| Book list | platformScripture.booksPresent project setting (useProjectBookIds) |
One string | Small | No |
| Project settings | continuousScroll, showMorphology, showFreeTranslation (validators); platform.languageTag |
Single values | Small | No |
| WebView state | useWebViewState keys: activeProject, analysisLanguageChoices, sidePanel, sidePanelLayout, dismissedStaleAnalyses, dismissedLostBoundaries |
Small per-tab values | Small | No |
| Lexicon registry | The FieldWorks Lite lexicon link, read from a project setting (fw-lite-lexicon.ts) |
A link id, plus lookups made through the provider | Not ours to page | No |
Diagram: data flow
Thick arrows (==>) carry a whole book, a whole project or a whole blob.
flowchart TB
subgraph HOST["Platform.Bible host"]
USJ["USJ_Book PDP"]
STORE[("papi.storage<br/>one file per key")]
SET["Project settings"]
end
subgraph EXT["Extension host"]
CMD["Commands<br/>getDraft, saveDraft, getProject,<br/>saveAnalysis, getProjectsForSource, ..."]
PS["projectStorage.ts<br/>serialization queues per project,<br/>per draft, and for the index"]
end
subgraph WV["WebView"]
BD["useInterlinearizerBookData<br/>extract + tokenize"]
DR["useDraftProject<br/>draft ref, undo snapshots,<br/>debounced autosave"]
AS["AnalysisStore (Redux)<br/>analysis for every book"]
VIEW["Views: segment list, continuous,<br/>catalog, stale notice"]
CON["Concordance index"]
end
USJ ==>|"whole book"| BD
USJ ==>|"every book present"| CON
PS ==>|"project blob, draft shards"| STORE
STORE ==>|"project blob, all draft shards"| PS
PS --- CMD
CMD ==>|"getDraft: merged draft"| DR
DR ==>|"saveDraft: whole draft JSON<br/>Save: whole analysis JSON"| CMD
DR ==>|"seed; replaceAnalysis on<br/>undo, redo, reanchor"| AS
AS ==>|"full analysis on every edit"| DR
BD ==>|"whole Book: reanchor pass"| DR
BD ==>|"whole Book"| VIEW
AS --> VIEW
AS ==>|"whole analysis"| CON
SET --> BD
Diagram: entities
Book, Segment and Token are not stored; they are rebuilt from USJ on every load. Links point at them by tokenRef (e.g. GEN 1:1:0) or segmentId. The book prefix of the ref is the key the draft shards use today.
Payload, as used throughout this issue, means one record in a TextAnalysis's tokenAnalyses, phraseAnalyses or segmentAnalyses array: a TokenAnalysis, PhraseAnalysis or SegmentAnalysis. A payload holds what a span of text means (gloss, morphemes, sense, or free and literal translation), its own id, and the surface form it analyzes. It records no location: it doesn't say which token, phrase or segment it applies to. It also has no status: a payload itself is never approved, suggested or stale. Both of those live on links, the per-occurrence records (TokenAnalysisLink, PhraseAnalysisLink, SegmentAnalysisLink) that attach a payload to one location by analysisId and carry that occurrence's status.
Token and phrase payloads are shared by every occurrence analyzed identically, including across books. Segment payloads are one-to-one with their links. Elsewhere in this issue, "the analysis" or "the whole analysis" means the entire TextAnalysis object, which holds both payloads and links.
Every token, phrase and segment link lies within one book; a phrase may span segments, but not books. The type puts no book constraint on alignment links, but nothing creates them yet.
erDiagram
INTERLINEAR_PROJECT ||--|| TEXT_ANALYSIS : "analysis"
DRAFT_PROJECT ||--|| TEXT_ANALYSIS : "analysis"
INTERLINEAR_PROJECT ||--o{ ALIGNMENT_LINK : "links"
TEXT_ANALYSIS ||--o{ TOKEN_ANALYSIS : "tokenAnalyses"
TEXT_ANALYSIS ||--o{ TOKEN_ANALYSIS_LINK : "tokenAnalysisLinks"
TEXT_ANALYSIS ||--o{ PHRASE_ANALYSIS : "phraseAnalyses"
TEXT_ANALYSIS ||--o{ PHRASE_ANALYSIS_LINK : "phraseAnalysisLinks"
TEXT_ANALYSIS ||--o{ SEGMENT_ANALYSIS : "segmentAnalyses"
TEXT_ANALYSIS ||--o{ SEGMENT_ANALYSIS_LINK : "segmentAnalysisLinks"
TOKEN_ANALYSIS ||--o{ TOKEN_ANALYSIS_LINK : "analysisId"
TOKEN_ANALYSIS ||--o{ MORPHEME_ANALYSIS : "embedded"
PHRASE_ANALYSIS ||--o{ PHRASE_ANALYSIS_LINK : "analysisId"
SEGMENT_ANALYSIS ||--o{ SEGMENT_ANALYSIS_LINK : "analysisId"
TOKEN_ANALYSIS_LINK }o--|| TOKEN : "token.tokenRef"
PHRASE_ANALYSIS_LINK }o--|{ TOKEN : "tokens[].tokenRef"
SEGMENT_ANALYSIS_LINK }o--|| SEGMENT : "segmentId"
SEGMENT ||--o{ TOKEN : "rebuilt from USJ"
INTERLINEAR_PROJECT {
string id PK
string sourceProjectId
number modelVersion
json segmentation
}
DRAFT_PROJECT {
string sourceProjectId PK
number modelVersion
boolean dirty
}
TOKEN_ANALYSIS {
string id PK "shared payload"
json gloss "MultiString"
}
TOKEN_ANALYSIS_LINK {
string analysisId FK
string tokenRef "book prefix"
string surfaceText "drift check"
string status
}
PHRASE_ANALYSIS_LINK {
string id PK "one per occurrence"
string analysisId FK
}
SEGMENT_ANALYSIS_LINK {
string analysisId FK
string segmentId
string status
}
TOKEN {
string ref "not stored"
}
SEGMENT {
string id "not stored"
}
How to read this diagram
Each line joins two entities. The symbol at each end of a line says how many of the entity at that end go with one of the entity at the other end.
| Line-end symbol | Meaning |
|---|---|
| Two short bars | Exactly one |
| Circle and crow's foot (three-pronged fork) | Zero or more |
| Bar and crow's foot | One or more |
For example, the line between TOKEN_ANALYSIS and TOKEN_ANALYSIS_LINK has two bars at the TOKEN_ANALYSIS end and a circle with a crow's foot at the link end. So one token analysis has zero or more links, and each link points at exactly one token analysis.
The text on a line names the field or array that holds the relationship (for example, analysisId or tokenAnalyses). Inside a box, each row is a type and a field name. PK marks the field that identifies the record, FK marks a field that refers to another record, and quoted text is a note. Each box lists only the fields that matter here, not every field of the type.
How the counts grow: there is about one link per analyzed occurrence (plus competing links distinguished by status), so links grow with the amount of text analyzed. Token and phrase payloads grow with the number of distinct analyses; segment payloads grow with their links.
Read paths today
- Book load.
BookUSJfetches the whole book.extractBookFromUsjandtokenizeBookrun in auseMemo.- Re-segmentation and re-heading run over the whole book.
- The reanchor pass calls
reanchorAnalysisToBook, which scans every link in the project to pick out this book's.
- Draft load.
getDraftreads every shard listed in the manifest (plus any orphans in the journal), merges them withmergeAnalyses, and returns one object. That object seeds the Redux store, which then holds every book. - Project lists.
listProjectsandgetProjectsForSourceread the index and then every project in full, analysis included. The picker shows only metadata plus two facts worked out from each analysis (summarizeAnalysis): the books its token links reference and its token-analysis count. - Concordance. While the index is enabled,
useConcordanceIndexreads and tokenizes the USJ of every book present in the project, four at a time. It indexes text only. The panel subscribes to the whole analysis and joins it against the rows on screen.
Write paths today
- Edit and autosave.
- A reducer runs over the whole arrays.
save()reads the fullstate.analysis.autosaveDraftpushes the previous full content onto the undo stack.- After a 300 ms debounce,
persistsendsJSON.stringify(draft)as one message. - The backend splits the draft by book and writes the shards whose JSON changed.
- Save.
handleSavesends the whole analysis tosaveAnalysis, which rewrites the wholeproject:{id}blob. - Save As.
createProjectthensaveAnalysis, or, when overwriting,saveAnalysisthenupdateProjectMetadata. Both are whole-blob. - Undo and redo.
undo-history.tskeeps up to 100 whole-draft{analysis, segmentation}snapshots. These are references to immutable content, not deep copies. Undo restores a snapshot, re-anchored to the current text, and dispatchesreplaceAnalysiswith the whole object. - Reanchor pass. The reanchor pass reruns whenever the loaded book's text, the segment boundaries or the whole draft change; gloss edits don't trigger it. If anything moved, it writes the whole healed draft and records the pass in the undo history.
- Wipe.
wipeBookandwipeAllreplace the whole draft and persist it immediately, with no debounce.
Paging candidates by access pattern
1. Visible-range reads
What the screen shows is a window of segments. useSegmentWindow first mounts 12 segments either side of the anchor, then adds 24 at a time and culls by geometry; 400 mounted segments is only a runaway cap. But the data behind that window is the whole book's text and the whole project's analysis, and several derivations walk the whole book on every load.
Call sites
useInterlinearizerBookData: whole-book USJ fetchuseBookIndexes: seven whole-book maps and one array, built in one pass, called fromInterlinearizer.tsxuseSegmentHeights: predicts the height of every segmentSegmentListView: labels, chapter list and joins overbook.segmentsSegmentListView: per-index free-translation and stale lookupsContinuousView: flattens every token in the book- Project-wide selector indexes read per token (
selectAnalysisById,selectApprovedIdByTokenRef,selectPendingAnalysesByTokenRef,selectRejectedAnalysisIdsByTokenRef) liveTokensByRef: map over every token in the book
2. By id, across books
A token or phrase payload is shared by every occurrence analyzed identically, in any book. Looking one up, or writing to it, today means scanning or rewriting the whole project's arrays.
Call sites
mergeIntoIdenticalPayload: searches everytokenAnalysesrecord for a duplicate and scans every link, repointing the ones that held the collapsed payload- Catalog-wide reducers keyed by analysis id:
writeAnalysisGloss,deleteAnalysis,mergeAnalysesInto selectPhraseGloss/selectSegmentFreeTranslation: linear.findper callreadEditOutcome: scans the payloads after each catalog edit, and the links too when the edit merged the record into another
3. Aggregates across books
These answers depend on every link in the project.
Call sites
selectPoolIndex: suggestion pool rebuilt from every token payload with at least one approved link, on each analysis change. Every mounted chip that shows suggestions reads it throughselectResolvedTokenAnalysis;selectSuggestionAfterClearingreads it only for the token being cleared (selectors)selectCatalogRowsandbuildCatalogRows: usage counts and locations for every analysisselectAnalysisDeletionOutcome: scans every link. It marks the outcomeuncertainwhen an affected token in an unloaded book has no analysis of its own to fall back to, or when the affected tokens would fall back to different analysesselectStaleTokenLinks/selectStaleFreeTranslations: every stale link in the project, filtered to the book inStaleAnalysesReporterhasFreeTranslations: scans every segment analysis- Concordance: tokenized USJ for every book present plus whole-analysis maps in the panel
4. Whole-book and whole-project derivations
Call sites
- Reanchor pass whenever the book text, boundaries or whole draft change, which runs
reanchorAnalysisToBookover every link in the project wipeBook/wipeAll:wipeBookrebuilds the whole analysis without one book (removeBookFromAnalysis);wipeAllswaps in an empty analysisuseStaleLocationReclaims: dry-runs the reducer over the full stateuseAnalysis(): subscribes to the whole object
5. Whole-blob writes
Call sites
- Autosave:
persistsends the whole draft, triggered fromautosaveDraft - Save:
saveAnalysiswith the whole analysis - Save As
updateAnalysisandupdateProjectMetadata: whole-blob read-modify-write- Undo snapshots
6. Project listing reads whole records
Projects are few, typically under 10 and at worst in the low hundreds, so the list itself doesn't need paging. The cost is that every listing reads each project's whole record, analysis included, only to compute summarizeAnalysis (the books referenced and the token-analysis count) for display. Opening the project picker therefore costs the total size of every project's analysis.
Call sites
listProjectsgetProjectsForSource, which feeds the project picker throughuseProjectsForSourceand summarizes each project inmain.ts
Constraints any design must honor
These are facts about today's model, not proposals.
- Payloads are shared across books. One
TokenAnalysisorPhraseAnalysiscan back occurrences in many books. The draft shards today copy a shared payload into each book's shard and dedupe on load. - At most one linked analysis per token, segment or phrase may be
approved. The store's reducers maintain this when they write: they repoint or remove the earlier approved link, or refuse the approval. Nothing rejects a breach.validateTextAnalysisonly logs it at the storage boundary. - A phrase occurrence is identified by
PhraseAnalysisLink.id, never by itsanalysisId. Book,SegmentandTokenare rebuilt from USJ on every load. For token and phrase links, the reanchor pass detects drift fromTokenSnapshot.surfaceTextwhen the baseline text changes. Segment analyses compare their ownsurfaceTextwith the segment'sbaselineText.- Records are gated by
modelVersion. Every project and draft-envelope write stampsCURRENT_MODEL_VERSION, and a record stamped higher is refused. Draft shards, the shard journal and the index carry no stamp. There are no migrations yet. - Undo restores whole-draft snapshots, held in memory only. Undo history is never persisted.
- A draft and a saved project are separate records. There is one draft per source project, and Save copies the draft's analysis into a project.
- The backend serializes writes, in memory. Per-project and per-draft queues guard read-modify-write, alongside single queues for the project index and
pendingCleanup. A shard journal recovers from interrupted saves, becausepapi.storagecannot list keys. - Future note (not in scope here): each user will eventually have their own send/receive changes, and those will be reconciled into a local snapshot.
PT9 import: out of scope, but it shares the storage
PT9 import is not a paging candidate and is left out of the maps above. It writes to the same storage, though, so the plan must account for it:
- An import is stored as an ordinary
project:{id}record carrying apt9Importfield (file hashes and import time). Each sync replaces the whole record (savePt9Import). - Finding a source's import scans every project (
getPt9ImportForSource). - Imports are read-only. The functions that save an analysis or update metadata refuse them, and only a sync replaces one; an import can still be deleted.
Proposed baseline before implementation
Before any change, we should capture a baseline so the team can measure the effect of each step. Proposed metrics:
- Book-load latency. From the book change to the first interactive render, split into USJ fetch, tokenize, reanchor and store mount.
- Edit latency. From dispatch to re-render, including the selector index rebuild, plus the autosave round-trip.
- Memory. The WebView heap with a large project loaded, including the undo stack.
The dataset and tooling are for the team to choose.
Notes from paranext-core
Facts from paranext-core at fcaa6d0, recorded here for planning and not as recommendations:
- Text has chapter and verse getters.
platformScriptureoffersUSJ_ChapterandUSJ_VersealongsideUSJ_Book(types). Each chapter call converts USX to USJ in the extension host (getChapterUSJ). There is no range selector. - Data provider updates are not filtered by selector. An update names data types only. Each subscriber re-runs its
getfor its own selector and deep-compares the result (service). Any Scripture edit notifies every USFM, USX and plain-text data type for the project (AllScriptureDataTypes). The platform-scripture extender re-emits each USX update as the matching USJ update, so every USJ subscriber re-fetches, whatever its selector. papi.storagewrites one file per key (buildUserDataUri), with a plainwriteFile(writeFile): no temp file and rename, and no size limit.- The bus closes the connection on any message over 100 MiB (
MAX_WEBSOCKET_PAYLOAD_BYTES). - We found no performance harness or benchmark suite in core. The nearest are isolated timing assertions in e2e tests.
Open questions for the team
- Which dataset and tooling should the baseline use?
- Does the draft vs. saved-project split survive a database, or does it become something else?
- Should the database reject a breach of the "at most one approved link" rule, rather than relying on reducers to maintain it and the storage boundary to log it?
- Which aggregates that span books (suggestion pool, usage counts, stale summary, concordance) must be live, and which can be computed lazily or in the background?
- Should a payload shared across books stay a single record, or be split per book the way the draft shards copy it today?
- Ngôn ngữ chính
- TypeScript
- Star
- 2
- Fork
- 0
- Merge trung bình
- 2 ngày 3 giờ
- Pull request đã merge (30 ngày)
- 46
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của sillsdev/interlinearizer-extension
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
sillsdev/interlinearizer-extension#407 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
sillsdev/interlinearizer-extension#402 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 8/100
sillsdev/interlinearizer-extension#401 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
sillsdev/interlinearizer-extension#388 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
sillsdev/interlinearizer-extension#383 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của sillsdev/interlinearizer-extension
Issue tương tự
-
[Feature]: [P3] engine-rs: the package source hash should ignore line endings and untracked filesĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
maniator/verticopolis#880 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
siyuan-note/siyuan#20353 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
black-forest-labs/skills#17 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Albert-Weasker/niubigeo#168 ·
Maintainer thường phản hồi trong vòng 1 ngày