Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug] Expand fails to recognize user-defined tablespaces

Đang mở Phù hợp với người mới
#1,885 0 bình luận 2 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

@nix-oss đang làm issue này rồi.

Từ ngày 18/8/2026.

  • #1901 của @nix-oss — đang mở

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
76/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Ít trao đổi
Công nghệ
postgresql, python
Lĩnh vực
databases

Hướng nghiên cứu

Bắt đầu trong gpMgmt/bin/gpexpand bằng cách đọc read_tablespace_file() và generate_tablespace_inputfile(), tập trung vào cách các giá trị của os.listdir() được so sánh với kết quả của get_tablespace_oid_names(). Chạy bản tái hiện được cung cấp với tablespace do người dùng định nghĩa và gpexpand không tương tác, sau đó xác minh rằng newTableSpaceInfo.json được tạo và các liên kết tượng trưng của tablespace được sửa trên các segment mới.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

type: Bug
Apache Cloudberry version

2.1.0-incubating

What happened

When user-defined tablespaces exist in the cluster, gpexpand does not create the newTableSpaceInfo.json file in the new segment template when running with a config file (non-interactive mode).

This leads to broken symlinks for tablespaces and causes failures during the prepare_schema() stage.

The problem is present in two methods within the same file:

  1. read_tablespace_file() in gpMgmt/bin/gpexpand (lines 1219–1285)
  2. generate_tablespace_inputfile() in gpMgmt/bin/gpexpand (lines 1114–1148)
  • read_tablespace_file() (line 1244):
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
    if oid in tblspc_oid_names:  <---
        flag = True
if not flag:
    return None  
  • generate_tablespace_inputfile() (line 1124)
tblspc_oid_names = self.get_tablespace_oid_names()
tblspc_info = {}

for oid in tblspc_oids:
    if oid not in tblspc_oid_names:  <---
        continue
    location = os.path.dirname(os.readlink(os.path.join(coordinator_tblspc_dir,
                                                        oid)))
    tblspc_info[oid] = {"location": location,
                        "name": tblspc_oid_names[int(oid)]}

Root cause: type mismatch between string values returned by os.listdir() and integer keys returned by the SQL query SELECT oid, spcname FROM pg_tablespace.

When attaching to the process and debugging, the types differ, causing this issue:

  • type mismatch
Image
  • newTableSpaceInfo is None
Image

Impact:

  • newTableSpaceInfo is always None.
  • _handle_tablespace_template() is never invoked.
  • newTableSpaceInfo.json is not included in the template tar.
  • gpconfigurenewsegment on new hosts does not fix tablespace symlinks.
  • new segments start with broken symlinks.
What you think should happen instead

To fix the issue, unify the data types in both methods:

  • In generate_tablespace_inputfile():
for oid in tblspc_oids:
  if int(oid) not in tblspc_oid_names:
  • In read_tablespace_file():
tblspc_oids = os.listdir(coordinator_tblspc_dir)
tblspc_oid_names = self.get_tablespace_oid_names()
flag = False
for oid in tblspc_oids:
  if int(oid) in tblspc_oid_names:

After this change, newTableSpaceInfo is no longer None:

Image
How to reproduce
  1. Connect to the database:
gpadmin@cbdb-mdw:~$ psql warehouse
  1. Create a user-defined tablespace:
warehouse=# CREATE TABLESPACE ts_stage_logs LOCATION '/tblspc_stage_logs';
warehouse=# SELECT * FROM pg_tablespace;
  oid  |    spcname    | spcowner | spcacl | spcoptions | spcfilehandlersrc | spcfilehandlerbin 
-------+---------------+----------+--------+------------+-------------------+-------------------
  1663 | pg_default    |       10 |        |            |                   | 
  1664 | pg_global     |       10 |        |            |                   | 
 17019 | ts_stage_logs |       10 |        |            |                   | 
(3 rows)
  1. Create an append-optimized columnar table and populate it with test data:
 warehouse=# CREATE TABLE logs_aot (
     id              BIGSERIAL,
     log_timestamp   TIMESTAMP WITH TIME ZONE DEFAULT NOW(),
     log_level       VARCHAR(10)
 )
 WITH (
     APPENDONLY = TRUE,
     ORIENTATION = COLUMN
 )
 TABLESPACE ts_stage_logs
 DISTRIBUTED BY (id);

 warehouse=# INSERT INTO logs_aot (
     log_timestamp,
     log_level
 )
 SELECT
     NOW() - (random() * INTERVAL '90 days'),
     (ARRAY['INFO', 'WARN', 'ERROR', 'DEBUG'])[floor(random() * 4 + 1)]
 FROM generate_series(1, 8000);
  1. Prepare the expansion configuration file for adding 2 new hosts to the cluster:
sdw3-new|sdw3-new|6000|/primary_data/gpseg4|10|4|p
sdw4-new|sdw4-new|7000|/mirror_data/gpseg4|11|4|m
sdw3-new|sdw3-new|6001|/primary_data/gpseg5|12|5|p
sdw4-new|sdw4-new|7001|/mirror_data/gpseg5|13|5|m
sdw4-new|sdw4-new|6000|/primary_data/gpseg6|14|6|p
sdw3-new|sdw3-new|7000|/mirror_data/gpseg6|15|6|m
sdw4-new|sdw4-new|6001|/primary_data/gpseg7|16|7|p
sdw3-new|sdw3-new|7001|/mirror_data/gpseg7|17|7|m
  1. Start the expansion process (non-interactive mode):
gpadmin@cbdb-mdw:~$ gpexpand -i expand.cfg
  1. The following error appears in the logs (log truncated for readability.)
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Heap checksum setting consistent across cluster
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Syncing Apache Cloudberry extensions
20260805:18:33:04:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locking catalog
20260805:18:33:05:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Locked catalog
20260805:18:33:06:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating segment template
20260805:18:33:07:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying postgresql.conf from existing segment into template
20260805:18:33:08:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Copying pg_hba.conf from existing segment into template
20260805:18:33:09:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Creating schema tar file
...
20260805:18:33:26:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating gpexpand.status_detail with data from database postgres
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Populating gpexpand.status_detail with data from database warehouse
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand failed: ERROR:  could not open file "pg_tblspc/17019/GPDB_3_302606111/17018/16386": No such file or directory  (seg6 203.0.113.5:6000 pid=21449)

Exiting...
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[ERROR]:-gpexpand is past the point of rollback. Any remaining issues must be addressed outside of gpexpand.
20260805:18:33:27:052024 gpexpand:cbdb-mdw:gpadmin-[INFO]:-Shutting down gpexpand...

Hope this analysis is helpful.

Operating System

Ubuntu 22.04

Anything else

No response

Are you willing to submit PR?
  • Yes, I am willing to submit a PR!
Code of Conduct
Ngôn ngữ chính
C
Star
1.4k
Fork
260
Merge trung bình
4 ngày 9 giờ
Pull request đã merge (30 ngày)
51

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/cloudberry

Tất cả issue của apache/cloudberry

Issue tương tự

Thêm issue về C

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.