Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Proposal: enrich existing Pure publications with multi-source metadata

未关闭
#7 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
30/100
Issue 类型
功能
描述清晰度
需要澄清
活跃度
活跃
技术栈
python

调研方向

先阅读 backend/app/services/pure_enrichment.py,尤其是 _apply_publication_updates 和 _set_volume,然后检查现有的 job 类型和 review-table 流程。第一个里程碑是执行一次不写入数据的 dry run,按 publication 报告缺失字段、源值以及与 Pure 的不一致;实现范围取决于对写入策略、源可信度、记录选择以及 connector layer 放置位置的决策。

由索引模型根据 Issue 内容生成。

描述

Idea

doi_resolver compares publication metadata across several scholarly sources (Crossref, OpenAlex, DataCite, Europe PMC, Unpaywall, CORE), keeps provenance for every value, and can write selected values back to Pure. Today that runs one DOI at a time through a UI.

Proposal: a BackToPure job that does the same for existing Pure publications in bulk — filling gaps in records that are already there.

Sibling of #6 (which covers full-text PDFs). Depends on #2, #3 and #5.

What doi_resolver already enriches

_apply_publication_updates (backend/app/services/pure_enrichment.py:755) maps these fields onto Pure:

title, journal, abstract, volume, issue (→ journalNumber), pages, open_access_status, license, keywords, subjects, language, published_year, published_date

It also handles identifiers, and matches persons and organizations against existing Pure entities, reporting what matched and what was skipped and why.

The important risk: this overwrites

The setters replace the value in Pure whenever it differs. _set_volume (:955) is representative — it returns early only if the incoming value is identical, otherwise it writes.

That is fine in doi_resolver, where a person looks at one DOI and picks each value deliberately. It is not fine in a bulk job: a run across a faculty could silently overwrite curated Pure data with a worse value from an external source.

This needs a deliberate decision before any implementation:

  • Fill-only — write only where Pure has no value. Safe, and probably the right default.
  • Overwrite with review — propose differences, require a human to accept each one via the NEEDS_REVIEW / review-table flow.
  • Overwrite on trust rules — e.g. trust Crossref over OpenAlex for pages. Most useful, most work, needs agreement on the ranking.

Recommend starting fill-only, with a dry-run report showing where sources disagree with Pure. That disagreement report is valuable on its own as a data-quality signal, separately from whether anything is written.

Scope: avoid duplicating existing jobs

BackToPure already has EXTERNAL_PERSONS and EXTERNAL_ORGS job types that enrich persons and organisations. doi_resolver's person/organisation matching overlaps with those.

This issue should therefore cover publication-level metadata only — the scalar fields and identifiers listed above. Person and organisation enrichment stays with the existing jobs unless there is a clear reason to merge them, which would be its own issue.

Fit with the job architecture

As with #6: a new JobType, identity_columns=("doi",), NEEDS_REVIEW plus the review-table for the human check, and apply / rollback for the write phase. Rollback matters more here than usual, since this touches existing curated records rather than adding new ones.

Where should this functionality live? (decide later)

#6 argues for porting code out of doi_resolver rather than depending on it, because the PDF code is small and the dependency is heavy. That reasoning does not automatically carry over here: the connector layer is six sources plus provenance and merge logic, which is a lot to maintain in two places.

There is a third option worth keeping open — BackToPure may simply be the better long-term home for this functionality.

Arguments for consolidating it here:

  • BackToPure is what colleagues actually run. doi_resolver is not in general use.
  • The job architecture already provides what bulk enrichment needs: a job registry, NEEDS_REVIEW with a review table, apply/rollback, logs, artifacts and cancellation. doi_resolver would have to grow all of that to do this at scale.
  • Multi-source comparison with provenance is useful to every job type here, not only to this one. As a shared capability inside BackToPure it could also improve the research-outputs import, which currently relies on OpenAlex alone.

Arguments against, or at least for care:

  • doi_resolver's single-DOI UI is genuinely useful for interactive quality control, and that use case should not be lost.
  • It is a working, tested codebase; moving it is not free.

Possible shapes, if consolidation wins: move the connector layer into BackToPure and leave doi_resolver as a thin UI over it; or move it and retire doi_resolver; or keep both and accept the duplication.

No decision needed now. Flagging it so the choice is made deliberately when this issue is picked up, rather than defaulted into by whoever writes the first line of code.

Open questions

  1. Fill-only or overwrite? (see above — needs a decision before implementation)
  2. Which sources are trusted for which fields?
  3. Which Pure records are in scope — a faculty, a date range, or records missing specific fields?
  4. Port, depend, or consolidate? See the section above — this is the significant architectural decision in this issue, and it is deliberately left open.

Suggested first step

A dry-run job that reports, per publication: which fields Pure is missing, what each source offers, and where sources contradict Pure. No writes. That gives a concrete picture of how much this would actually improve the data before anyone commits to the write path.

主要语言
Python
星标
0
派生
0
平均合并
8 小时 3 分钟
30 天内合并 PR
16

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

UtrechtUniversity/BackToPure 的其他 Issue

查看 UtrechtUniversity/BackToPure 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。