Label name resolved from adjacent memory under concurrent edge writes: relation "<graph>.<garbage>" does not exist
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 30/100
Hướng nghiên cứu
The payload names no source file, test, or entry point. Start with the supplied concurrent edge-upsert workload and rerun it with jit = off, then capture a debug-symbol stack trace at failure. Done means the workload no longer raises relation errors containing property data and the underlying label-resolution failure is identified.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
Under sustained concurrent Cypher writes, AGE intermittently fails to resolve a label name and raises an undefined-table error for a relation whose name is not a label at all. The name is made of fragments of the data being written: agtype property keys, property values, and occasionally a single control byte equal to a label id.
The graph's label catalog is correct and unchanged throughout, so this does not look like catalog corruption. It looks like the label name is being read through a pointer that no longer refers to the label.
Version
- Apache AGE 1.7.0 (
PG17/v1.7.0-rc0) - PostgreSQL 17.10 (Debian 17.10-1.pgdg12+1)
- Linux x86_64, AGE loaded via
shared_preload_libraries
What happens
The error surfaces from a CREATE of an edge:
ERROR: relation "mygraph.\x01" does not exist
ERROR: relation "mygraph.<entity name fragment>" does not exist
ERROR: relation "mygraph.dfile_pathsource_idcreated_atdescriptionentity_type<entity name><filename><chunk id>" does not exist
ERROR: relation "mygraph.<text fragment><SEP><text fragment><binary bytes>" does not exist
The third example is the clearest. file_path, source_id, created_at, description and entity_type are the property keys of the edge being created, concatenated, followed by a property value and an identifier. \x01 is chr(1), which is the id of _ag_label_vertex in ag_catalog.ag_label for this graph.
So the resolver appears to be producing a name either from the raw label id or from memory adjacent to the agtype payload, instead of from the label catalog.
The catalog is fine
Queried during and after the failures, unchanged throughout:
SELECT name, kind, id FROM ag_catalog.ag_label l
JOIN ag_catalog.ag_graph g ON g.graphid = l.graph
WHERE g.name = 'mygraph';
name | kind | id
------------------+------+----
_ag_label_vertex | v | 1
_ag_label_edge | e | 2
base | v | 3
DIRECTED | e | 4
Four labels, exactly as expected. information_schema.tables for the graph schema likewise shows four tables. No label named anything like the strings in the errors has ever existed.
Workload shape
The write is an idempotent edge upsert. Edge properties are inlined in the CREATE clause because SET r += {...} does not persist edge properties on AGE (the endpoint ids are parameterised):
SELECT r FROM cypher('mygraph', $$
MATCH (source:base {entity_id: $src_id})
WITH source
MATCH (target:base {entity_id: $tgt_id})
WITH source, target
OPTIONAL MATCH (source)-[old:DIRECTED]-(target)
DELETE old
WITH source, target
CREATE (source)-[r:DIRECTED {`file_path`: "...", `source_id`: "...", `created_at`: 1790828551, `description`: "...", `entity_type`: "..."}]->(target)
RETURN r
$$, $1) AS (r agtype);
Conditions when it occurs:
- 3 concurrent writers on separate pooled connections, each running the statement above in its own transaction
- Graph at roughly 500,000 vertices and 1,400,000 edges, and growing
search_pathincludesag_catalog, set per connection- Transaction-scoped advisory locks serialise writers on the same endpoint pair, so two sessions do not write the same edge concurrently. Different edges do proceed concurrently.
- Always on high-degree endpoints: the vertices involved are the most connected in the graph, appearing in a large share of all edges. Low-degree endpoints have not produced it.
- Some property values are large:
source_idcan hold a few hundred separator-joined identifiers.
Frequency
Roughly 3 to 5 failures per 100 documents ingested, where each document performs many edge upserts. It is not deterministic and not tied to any particular endpoint pair.
It appeared only late in a large ingest, once the graph had grown to the size above. Early in the same ingest, with the same code and concurrency, it did not occur.
Diagnostics
Gathered on the live instance that reproduces this.
Build and configuration
PostgreSQL 17.10 (Debian 17.10-1.pgdg12+1) on x86_64-pc-linux-gnu,
compiled by gcc (Debian 12.2.0-14+deb12u1) 12.2.0, 64-bit
age 1.7.0
shared_preload_libraries = age
Settings that seemed most relevant:
jit on
jit_above_cost 100000
jit_inline_above_cost 500000
jit_optimize_above_cost 500000
shared_buffers 128MB
work_mem 4MB
maintenance_work_mem 64MB
max_connections 100
max_locks_per_transaction 64
max_parallel_workers_per_gather 2
dynamic_shared_memory_type posix
huge_pages try
JIT is enabled. We have not yet tested with jit = off, but given the symptom is a pointer that stops referring to the label, it seemed worth surfacing early in case it narrows things for you. We can run that experiment and report back.
Graph scale at time of capture
vertices 531,452
edges 1,492,885
labels 4 (_ag_label_vertex, _ag_label_edge, base, DIRECTED)
Edge property sizes
The faulting operation inlines properties into CREATE, so payload size is a plausible factor. Distribution over all edges, as length(properties::text):
avg 350
p50 289
p99 1,332
max 73,715
So most payloads are small but the tail is three orders of magnitude larger.
The bad names, categorised
24 occurrences over 48 hours, classified by what the invalid relation name appears to contain. Values are described rather than quoted, since the real strings contain customer data, which is itself the point: the memory being read is the payload.
8 text fragment with non-ASCII or raw bytes
5 chunk / task identifier fragment
4 property VALUE including our field separator
3 agtype PROPERTY KEYS concatenated, in declaration order
3 plain text fragment (entity name or filename)
1 single control byte equal to a label id
Length of the invalid name:
n=24 min=0 p50=54 max=327
The min=0 case is an empty name. The single-control-byte case was \x01, which equals the id of _ag_label_vertex in this graph.
Not one bad backend
Occurrences are spread across backends rather than concentrated in a poisoned session:
distinct backends that hit it ~10
distinct backends in the window 397
most occurrences in one backend 2
So a session does not appear to "go bad" and stay bad. Each occurrence looks independent.
No STATEMENT is logged with the error
log_min_error_statement = error, yet these errors appear in the log with no accompanying STATEMENT: line, unlike other errors from the same instance in the same window. We do not know whether that is meaningful to you, but it was unexpected and may indicate the error is raised in a context where the statement is not attached.
The faulting operation is known from the client side regardless: it is always the edge upsert shown in the original report, never a read.
What we ruled out
- Catalog corruption.
ag_labelis correct before, during and after. - A DDL storm. Our client library was calling
create_graph()on every pool checkout, around 31,000 failed calls an hour against an existing graph. We suspected the repeated failed DDL was invalidating cached label information and removed it entirely. The failure rate did not change. Measured per unit of work it did not improve at all:
3.3 failures per 100 ingested documents (day the DDL was removed)
5.0 failures per 100 ingested documents (the following day)
So the DDL storm is not the trigger.
- A missing or renamed label. The labels in the errors never existed.
- Our own string building. The endpoint ids are bound parameters, and the property literal is JSON-escaped. The strings in the error are not a quoting artefact of ours; they are the contents of the properties we are legitimately writing, appearing where a label name should be.
Impact
Each occurrence aborts the transaction and fails the unit of work. In a long-running ingest this means permanent, silent data loss unless the caller retries, and the error is not distinguishable from a genuine schema problem by its type.
Reproduction and what we can offer
We do not have a minimal standalone reproduction, and we want to be straightforward about why: the fault needs a graph of substantial size and sustained concurrency before it appears. Early in the same ingest, with identical code, concurrency and data shape, it does not occur. It began only once the graph passed roughly half a million vertices.
What we do have is a workload that reproduces it reliably and repeatedly at about 3 to 5 failures per 100 units of work, on an instance we control and can instrument. The diagnostics above were gathered proactively on that instance rather than waiting to be asked.
We are willing to:
- run with
jit = offand report whether the rate changes, which seems the most informative single experiment available to us - capture a stack trace with debug symbols, if you can suggest where to break
- capture
pg_stat_activity, lock state, or memory-context dumps at the point of failure - try a specific patch or build against the live workload
Tell us which of these is most useful and we will run it. If none of it helps without a minimal reproduction, say so and we will put effort into building one instead.
- Ngôn ngữ chính
- C
- Star
- 4.9k
- Fork
- 529
- Merge trung bình
- 8 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 3
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/age
-
Mark agtype_string_match_starts_with / _ends_with IMMUTABLE (contains already is)Có thể đã có người làm @mmustafasenoglu đã nhận 12 ngày trước. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
apache/age#2576 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
TRUNCATE fails with `schema "ag_catalog" does not exist` in databases without AGE, when AGE is in shared_preload_librariesCó thể đã có người làm @crdv7 đã nhận 13 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/age#2520 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Integer modulo (%) by a zero divisor errors with `floating-point exception` (22P01) instead of `division by zero`Có thể đã có người làm @cocofabio đã nhận 19 ngày trước. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
apache/age#2514 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
Maintainer thường phản hồi trong vòng 1 ngày
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
EchoTools/nevr-runtime#171 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[P2] Workspace updates silently ignore forbidden assignments while staging the rowCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 4 ngày
-
sdl3-image update to 3.4.8Đang mởcategory:port-update
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
microsoft/vcpkg#54338 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
area:http-gateway good first issue priority:low type:docs
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
crazy-goat/php-fpm-ng#828 ·
Maintainer thường phản hồi trong vòng 1 ngày