help wantedworkflow-arabic-text-extraction
Repository-Metriken
- Stars
- (6.413 Sterne)
- PR-Merge-Metriken
- (Durchschn. Merge 8T) (49 gemergte PRs in 30 T)
Beschreibung
extracted from #1379 PS : in the extraction result, the arabic characters are replaced with /afiinnnn. this is because the data uses the iso 10036 standard that I've not been able to find any free information on how to do transcoding file 02voc.pdf test code:
import PyPDF2;
PyPDF2.PdfReader("e:/02voc.pdf").pages[2].extract_text()
Originally posted by @pubpub-zz in https://github.com/py-pdf/PyPDF2/issues/1379#issuecomment-1268897035