py-pdf/pypdf

can not decode afii characters (ISO 10036)

Ouverte

#1 381 ouverte le 5 oct. 2022

 (1 commentaire) (0 réaction) (0 personne assignée)Python (1 258 forks)batch import
help wantedworkflow-arabic-text-extraction

Métriques du dépôt

Stars
 (6 413 étoiles)
Métriques de merge PR
 (Merge moyen 8j) (49 PRs mergées en 30 j)

Description

extracted from #1379 PS : in the extraction result, the arabic characters are replaced with /afiinnnn. this is because the data uses the iso 10036 standard that I've not been able to find any free information on how to do transcoding file 02voc.pdf test code:

import PyPDF2;
PyPDF2.PdfReader("e:/02voc.pdf").pages[2].extract_text()

Originally posted by @pubpub-zz in https://github.com/py-pdf/PyPDF2/issues/1379#issuecomment-1268897035

Guide contributeur