good first issuescraper
仓库指标
- Star
- (56 star)
- PR 合并指标
- (PR 指标待抓取)
描述
Summary
Now that we're scraping the text of longer bills from the PDFs, we're seeing occasional errors where a Bill would exceed the maximum acceptable size of a Firebase document (~1 MiB). This prevents scraping of the bill altogether - which is
Instead of failing altogether, we should handle this case more gracefully:
- Catch failed Firestore writes around
maximum allowed sizeand re-try saving the Bill document without the DocumentText (basically, the current behavior where we only have the PDF)
Success Criteria
- Bills with excessively long DocumentText can be mostly successfully scraped (by excluding the long DocumentText field)
Additional Info
-
Error: In
fetchBillBatch:Error: 3 INVALID_ARGUMENT: Document 'projects/digital-testimony-dev/databases/(default)/documents/generalCourts/194/bills/H5500' cannot be written because its size (1,373,153 bytes) exceeds the maximum allowed size of 1,048,576 bytes. -
Example Bill: Bill H5500 in Court 194