Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Query Tool: first 4 MB of a large file is loaded twice when a later chunk isn't valid UTF-8

Open Beginner friendly
#10,462 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@dev-hari-prasad is already working on this.

Since Oct 7, 2026.

  • #10508 by @dev-hari-prasad — open

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Start in the pgadmin.misc.file_manager module at read_file_generator() and its use by load_file(), then run the supplied reproduction with a file larger than 4 MB containing a later invalid UTF-8 byte. Done means the content is yielded exactly once using the Latin-1 fallback, while preserving line endings and addressing the reported codecs.open() deprecation.

Written by the indexing model from the issue text.

Description

Describe the bug

When a file larger than 4 MB is opened in the Query Tool, and a byte that isn't valid UTF-8 appears after the first 4 MB (for example, a Latin-1 é in an otherwise UTF-8 SQL dump), the file's first 4 MB are sent to the editor twice.

load_file streams read_file_generator(file_path, enc). enc comes from check_file_for_bom_and_binary(), which only looks at the first 1024 bytes, so it is utf-8 for such a file. read_file_generator() then reads the file in 4 MB chunks and yields each chunk as it goes. When a later chunk raises UnicodeDecodeError, the except branch reopens the file with latin-1 and yields it again from the start, after the earlier chunks have already been sent. If the user then saves the file, the duplicated content is written back.

To Reproduce

  1. Create a file of just over 5 MB that is UTF-8 except for one Latin-1 byte near the end:
    data = b"SELECT 1;\r\n" * (5 * 1024 * 1024 // 11) + b"-- caf\xe9\r\n"
    open("big.sql", "wb").write(data)
    
  2. Read it the way load_file does:
    from pgadmin.misc.file_manager import read_file_generator
    out = "".join(read_file_generator("big.sql", "utf-8"))
    print(len(data), len(out))  # 5242884 9437188
    
    The output is 9,437,188 characters for a 5,242,884-byte file: the first 4 MB appear twice.

I checked this by calling the function on current master; I haven't gone through the UI with a file this size.

Expected behavior

The file's content appears exactly once, decoded as Latin-1 when it isn't valid in the detected encoding, which is what the fallback is meant to do.

Desktop

  • pgAdmin version: current master
  • Mode: any (the file is read on the server)

Additional context

The same function uses codecs.open(), which is deprecated as of Python 3.14.

Would you like me to open a PR that fixes both? It would check which encoding applies before yielding anything, then read the file with the built-in open() (with newline='', so line endings stay exactly as they are).

Dominant language
Python
Stars
3.9k
Forks
904
Avg merge
23h 17m
Merged PRs (30d)
30

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from pgadmin-org/pgadmin4

All issues in pgadmin-org/pgadmin4

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.