Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Label name resolved from adjacent memory under concurrent edge writes: relation "<graph>.<garbage>" does not exist

Đang mở
#2,588 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
30/100
Loại issue
Lỗi
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
c, postgresql
Lĩnh vực
databases

Hướng nghiên cứu

The payload names no source file, test, or entry point. Start with the supplied concurrent edge-upsert workload and rerun it with jit = off, then capture a debug-symbol stack trace at failure. Done means the workload no longer raises relation errors containing property data and the underlying label-resolution failure is identified.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Summary

Under sustained concurrent Cypher writes, AGE intermittently fails to resolve a label name and raises an undefined-table error for a relation whose name is not a label at all. The name is made of fragments of the data being written: agtype property keys, property values, and occasionally a single control byte equal to a label id.

The graph's label catalog is correct and unchanged throughout, so this does not look like catalog corruption. It looks like the label name is being read through a pointer that no longer refers to the label.

Version
  • Apache AGE 1.7.0 (PG17/v1.7.0-rc0)
  • PostgreSQL 17.10 (Debian 17.10-1.pgdg12+1)
  • Linux x86_64, AGE loaded via shared_preload_libraries
What happens

The error surfaces from a CREATE of an edge:

ERROR:  relation "mygraph.\x01" does not exist
ERROR:  relation "mygraph.<entity name fragment>" does not exist
ERROR:  relation "mygraph.dfile_pathsource_idcreated_atdescriptionentity_type<entity name><filename><chunk id>" does not exist
ERROR:  relation "mygraph.<text fragment><SEP><text fragment><binary bytes>" does not exist

The third example is the clearest. file_path, source_id, created_at, description and entity_type are the property keys of the edge being created, concatenated, followed by a property value and an identifier. \x01 is chr(1), which is the id of _ag_label_vertex in ag_catalog.ag_label for this graph.

So the resolver appears to be producing a name either from the raw label id or from memory adjacent to the agtype payload, instead of from the label catalog.

The catalog is fine

Queried during and after the failures, unchanged throughout:

SELECT name, kind, id FROM ag_catalog.ag_label l
  JOIN ag_catalog.ag_graph g ON g.graphid = l.graph
 WHERE g.name = 'mygraph';

      name        | kind | id
------------------+------+----
 _ag_label_vertex | v    |  1
 _ag_label_edge   | e    |  2
 base             | v    |  3
 DIRECTED         | e    |  4

Four labels, exactly as expected. information_schema.tables for the graph schema likewise shows four tables. No label named anything like the strings in the errors has ever existed.

Workload shape

The write is an idempotent edge upsert. Edge properties are inlined in the CREATE clause because SET r += {...} does not persist edge properties on AGE (the endpoint ids are parameterised):

SELECT r FROM cypher('mygraph', $$
  MATCH (source:base {entity_id: $src_id})
  WITH source
  MATCH (target:base {entity_id: $tgt_id})
  WITH source, target
  OPTIONAL MATCH (source)-[old:DIRECTED]-(target)
  DELETE old
  WITH source, target
  CREATE (source)-[r:DIRECTED {`file_path`: "...", `source_id`: "...", `created_at`: 1790828551, `description`: "...", `entity_type`: "..."}]->(target)
  RETURN r
$$, $1) AS (r agtype);

Conditions when it occurs:

  • 3 concurrent writers on separate pooled connections, each running the statement above in its own transaction
  • Graph at roughly 500,000 vertices and 1,400,000 edges, and growing
  • search_path includes ag_catalog, set per connection
  • Transaction-scoped advisory locks serialise writers on the same endpoint pair, so two sessions do not write the same edge concurrently. Different edges do proceed concurrently.
  • Always on high-degree endpoints: the vertices involved are the most connected in the graph, appearing in a large share of all edges. Low-degree endpoints have not produced it.
  • Some property values are large: source_id can hold a few hundred separator-joined identifiers.
Frequency

Roughly 3 to 5 failures per 100 documents ingested, where each document performs many edge upserts. It is not deterministic and not tied to any particular endpoint pair.

It appeared only late in a large ingest, once the graph had grown to the size above. Early in the same ingest, with the same code and concurrency, it did not occur.

Diagnostics

Gathered on the live instance that reproduces this.

Build and configuration
PostgreSQL 17.10 (Debian 17.10-1.pgdg12+1) on x86_64-pc-linux-gnu,
  compiled by gcc (Debian 12.2.0-14+deb12u1) 12.2.0, 64-bit
age 1.7.0
shared_preload_libraries = age

Settings that seemed most relevant:

jit                             on
jit_above_cost                  100000
jit_inline_above_cost           500000
jit_optimize_above_cost         500000
shared_buffers                  128MB
work_mem                        4MB
maintenance_work_mem            64MB
max_connections                 100
max_locks_per_transaction       64
max_parallel_workers_per_gather 2
dynamic_shared_memory_type      posix
huge_pages                      try

JIT is enabled. We have not yet tested with jit = off, but given the symptom is a pointer that stops referring to the label, it seemed worth surfacing early in case it narrows things for you. We can run that experiment and report back.

Graph scale at time of capture
vertices  531,452
edges   1,492,885
labels          4   (_ag_label_vertex, _ag_label_edge, base, DIRECTED)
Edge property sizes

The faulting operation inlines properties into CREATE, so payload size is a plausible factor. Distribution over all edges, as length(properties::text):

avg      350
p50      289
p99    1,332
max   73,715

So most payloads are small but the tail is three orders of magnitude larger.

The bad names, categorised

24 occurrences over 48 hours, classified by what the invalid relation name appears to contain. Values are described rather than quoted, since the real strings contain customer data, which is itself the point: the memory being read is the payload.

 8  text fragment with non-ASCII or raw bytes
 5  chunk / task identifier fragment
 4  property VALUE including our field separator
 3  agtype PROPERTY KEYS concatenated, in declaration order
 3  plain text fragment (entity name or filename)
 1  single control byte equal to a label id

Length of the invalid name:

n=24  min=0  p50=54  max=327

The min=0 case is an empty name. The single-control-byte case was \x01, which equals the id of _ag_label_vertex in this graph.

Not one bad backend

Occurrences are spread across backends rather than concentrated in a poisoned session:

distinct backends that hit it      ~10
distinct backends in the window    397
most occurrences in one backend      2

So a session does not appear to "go bad" and stay bad. Each occurrence looks independent.

No STATEMENT is logged with the error

log_min_error_statement = error, yet these errors appear in the log with no accompanying STATEMENT: line, unlike other errors from the same instance in the same window. We do not know whether that is meaningful to you, but it was unexpected and may indicate the error is raised in a context where the statement is not attached.

The faulting operation is known from the client side regardless: it is always the edge upsert shown in the original report, never a read.

What we ruled out
  • Catalog corruption. ag_label is correct before, during and after.
  • A DDL storm. Our client library was calling create_graph() on every pool checkout, around 31,000 failed calls an hour against an existing graph. We suspected the repeated failed DDL was invalidating cached label information and removed it entirely. The failure rate did not change. Measured per unit of work it did not improve at all:
3.3 failures per 100 ingested documents   (day the DDL was removed)
5.0 failures per 100 ingested documents   (the following day)

So the DDL storm is not the trigger.

  • A missing or renamed label. The labels in the errors never existed.
  • Our own string building. The endpoint ids are bound parameters, and the property literal is JSON-escaped. The strings in the error are not a quoting artefact of ours; they are the contents of the properties we are legitimately writing, appearing where a label name should be.
Impact

Each occurrence aborts the transaction and fails the unit of work. In a long-running ingest this means permanent, silent data loss unless the caller retries, and the error is not distinguishable from a genuine schema problem by its type.

Reproduction and what we can offer

We do not have a minimal standalone reproduction, and we want to be straightforward about why: the fault needs a graph of substantial size and sustained concurrency before it appears. Early in the same ingest, with identical code, concurrency and data shape, it does not occur. It began only once the graph passed roughly half a million vertices.

What we do have is a workload that reproduces it reliably and repeatedly at about 3 to 5 failures per 100 units of work, on an instance we control and can instrument. The diagnostics above were gathered proactively on that instance rather than waiting to be asked.

We are willing to:

  • run with jit = off and report whether the rate changes, which seems the most informative single experiment available to us
  • capture a stack trace with debug symbols, if you can suggest where to break
  • capture pg_stat_activity, lock state, or memory-context dumps at the point of failure
  • try a specific patch or build against the live workload

Tell us which of these is most useful and we will run it. If none of it helps without a minimal reproduction, say so and we will put effort into building one instead.

Ngôn ngữ chính
C
Star
4.9k
Fork
529
Merge trung bình
8 ngày 15 giờ
Pull request đã merge (30 ngày)
3

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/age

Tất cả issue của apache/age

Issue tương tự

Thêm issue về C

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.