Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Query Tool: first 4 MB of a large file is loaded twice when a later chunk isn't valid UTF-8

Đang mở Phù hợp với người mới
#10,462 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@dev-hari-prasad đang làm issue này rồi.

Từ ngày 7/10/2026.

  • #10508 của @dev-hari-prasad — đang mở

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
82/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
backend

Hướng nghiên cứu

Bắt đầu trong module pgadmin.misc.file_manager tại read_file_generator() và nơi nó được load_file() sử dụng, sau đó chạy reproduction được cung cấp với một tệp lớn hơn 4 MB chứa một byte UTF-8 không hợp lệ xuất hiện về sau. Hoàn tất khi nội dung được yield chính xác một lần bằng Latin-1 fallback, đồng thời giữ nguyên các ký tự kết thúc dòng và xử lý việc codecs.open() đã được báo cáo là deprecated.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Describe the bug

When a file larger than 4 MB is opened in the Query Tool, and a byte that isn't valid UTF-8 appears after the first 4 MB (for example, a Latin-1 é in an otherwise UTF-8 SQL dump), the file's first 4 MB are sent to the editor twice.

load_file streams read_file_generator(file_path, enc). enc comes from check_file_for_bom_and_binary(), which only looks at the first 1024 bytes, so it is utf-8 for such a file. read_file_generator() then reads the file in 4 MB chunks and yields each chunk as it goes. When a later chunk raises UnicodeDecodeError, the except branch reopens the file with latin-1 and yields it again from the start, after the earlier chunks have already been sent. If the user then saves the file, the duplicated content is written back.

To Reproduce

  1. Create a file of just over 5 MB that is UTF-8 except for one Latin-1 byte near the end:
    data = b"SELECT 1;\r\n" * (5 * 1024 * 1024 // 11) + b"-- caf\xe9\r\n"
    open("big.sql", "wb").write(data)
    
  2. Read it the way load_file does:
    from pgadmin.misc.file_manager import read_file_generator
    out = "".join(read_file_generator("big.sql", "utf-8"))
    print(len(data), len(out))  # 5242884 9437188
    
    The output is 9,437,188 characters for a 5,242,884-byte file: the first 4 MB appear twice.

I checked this by calling the function on current master; I haven't gone through the UI with a file this size.

Expected behavior

The file's content appears exactly once, decoded as Latin-1 when it isn't valid in the detected encoding, which is what the fallback is meant to do.

Desktop

  • pgAdmin version: current master
  • Mode: any (the file is read on the server)

Additional context

The same function uses codecs.open(), which is deprecated as of Python 3.14.

Would you like me to open a PR that fixes both? It would check which encoding applies before yielding anything, then read the file with the built-in open() (with newline='', so line endings stay exactly as they are).

Ngôn ngữ chính
Python
Star
3.9k
Fork
904
Merge trung bình
1 ngày 3 giờ
Pull request đã merge (30 ngày)
30

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của pgadmin-org/pgadmin4

Tất cả issue của pgadmin-org/pgadmin4

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.