scrapinghub/dateparser

Wrong prioritization of languages

已关闭

#770 创建于 2020年8月21日

 (6 条评论) (0 个反应) (2 位负责人)Python (443 个派生)batch import
good first issue

仓库指标

星标
 (2,318 个星标)
PR 合并指标
 (平均合并 397天 12小时) (30 天内合并 13 个 PR)

描述

I think there is something wrong in dateparser prioritization of languages, as introducing 'en' even in the last position hurts extraction of dates that were extracted properly when English was not there.

import dateparser
dateparser.parse("11/12", languages=['en'])
Out[3]: datetime.datetime(2020, 11, 12, 0, 0)

This is right

dateparser.parse("11/12", languages=['es'])
Out[4]: datetime.datetime(2020, 12, 11, 0, 0)

This is also right, because the standard in Spain is DD/MM But now if we add English to the languages list in the last position...

dateparser.parse("11/12", languages=['es', 'en'])
Out[5]: datetime.datetime(2020, 11, 12, 0, 0)

We got it parsed like in English, even if Spanish is first in the list of languages. This is unexpected to me, I would have expected prioritizing Spanish instead.

贡献者指南