acl-org/acl-anthology

Extract abstracts from PDF

オープン

#395 opened on 2019/06/06

 (28 件のコメント) (2 件のリアクション) (1 人の担当者)Python (386 件のフォーク)github user discovery
enhancementhelp wanted

Repository metrics

Stars
 (733 個のスター)
PR merge metrics
 (PR metrics pending)

説明

The anthology currently only shows the abstracts if there is an authoritative version in the XML. It would be nice if we could scrape the PDF using some off-the-shelf software to extract the abstracts and dump them into a different file (to not tamper with handcrafted information). Having an abstract on the web pages makes quickly searching through literature much faster.

コントリビューターガイド