acl-org/acl-anthology

Extract abstracts from PDF

开放

#395 创建于 2019年6月6日

 (28 条评论) (2 个反应) (1 位负责人)Python (386 个派生)github user discovery
enhancementhelp wanted

仓库指标

星标
 (733 个星标)
PR 合并指标
 (平均合并 7天 20小时) (30 天内合并 52 个 PR)

描述

The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.

贡献者指南