Run the warehouse implementations against real Snowflake and Databricks
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
Hướng nghiên cứu
Bắt đầu bằng cách đọc pkg-r/tests/testthat/test-live-warehouses.R, helper-live-warehouses.R và pkg-r/tests/testthat/README.md để hiểu các tùy chọn và skip hiện có cho các live test. Xem lại các case dùng chung catalog-access-errors.json và definition-warehouse-sql.json trước khi xác định warehouses và thông tin xác thực nên được lấy từ đâu. Được xem là hoàn tất khi đã quyết định mô hình thực thi và có kế hoạch triển khai bao quát việc provisioning, các harness Python và R, xử lý thông tin xác thực, các skip phù hợp và các live assertion.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Everything commons knows about Snowflake and Databricks has been verified against a fake. Nothing in pkg-py has ever run a statement against a real warehouse, and the R live tests have no way to run anywhere except one person's laptop. The catalog readers, the access and session checks, and the warehouse SQL emitters are all specified against what we believe those systems return.
What exists today
pkg-r/tests/testthat/test-live-warehouses.R (744 lines) plus helper-live-warehouses.R connect through ODBC (odbc::snowflake(), a DSN named Databricks) and skip unless R options name real objects: a test table, a denied table, an alternate Snowflake role, semantic views, parameterized models. pkg-r/tests/testthat/README.md documents the options. Nothing provisions the objects those options point at, no CI job runs any of it, and pkg-py has no equivalent at all: every warehouse test there drives a fake backend that replays canned rows.
The open question
Where does the warehouse come from? This is the decision the rest depends on, so it should be made first. Options worth pricing:
- a shared Posit-owned Snowflake account and Databricks workspace, with credentials in GitHub secrets;
- per-developer trial accounts plus a provisioning script;
- Databricks free edition or a Snowflake trial, created and torn down in CI;
- accepting that this stays manual, and lives in a documented pre-release checklist rather than in CI.
Then the mechanics
- Provisioning as code. The fixtures are not just a table. The access tests need an object the test principal genuinely cannot read, and the session test needs a second usable role. A SQL script (or Terraform) that creates the schema, tables, views, roles, and grants would bring any account to a known state and stop the test options being hand-typed identifiers.
- A Python harness of the same shape. Python connects through SQLAlchemy rather than ODBC/DBI, so this needs
snowflake-sqlalchemyanddatabricks-sqlalchemyas optional dev dependencies, an environment-variable or config equivalent of the R options, and a clean skip when they are unset. - Credential handling. What reaches CI, what stays local, and what a contributor without warehouse access sees. Today they get a clean skip, which should stay true.
- Scope of what gets asserted. Prefer what a fake cannot tell us over restating unit tests: identifier case folding per backend, the real row shapes of
SHOW OBJECTS,DESC TABLE,system.information_schemaandDESCRIBE TABLE, the SQLSTATE and message a genuine permission refusal carries, the session identity queries, and whether the Snowflake and Databricks SQL the definition emitters produce actually executes.
Why it matters
tests/shared/catalog-access-errors.json decides whether a driver failure is an authorization refusal, a transient fault, or neither, and commons caches the refusal or retries based on that answer. Its cases were written from documentation rather than from a captured failure. tests/shared/definition-warehouse-sql.json pins SQL that no warehouse has ever parsed, and it is hand-maintained precisely because there is no upstream authority to check it against. A live run is what turns both from plausible into verified. The same applies to the native semantic-model probes for Snowflake semantic views and Databricks metric views, which cannot be written honestly without one.
Not urgent for the conference demo, which runs on DuckDB, and not a blocker for the data layer milestone. It is what stands between "the warehouse support is implemented" and "the warehouse support is known to work".
- Ngôn ngữ chính
- Python
- Star
- 58
- Fork
- 2
- Merge trung bình
- 2 ngày 5 giờ
- Pull request đã merge (30 ngày)
- 50
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của posit-dev/commons
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 45/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Subagent workflowsĐang mở
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
posit-dev/commons#398 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement py r
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
posit-dev/commons#391 · 1 bình luận · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 76/100
posit-dev/commons#376 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của posit-dev/commons
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement good first issue Stellar Wave trivial
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
StellarCanary/ProtocolCanary-Fixtures#258 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
IBM/ai-atlas-nexus#295 ·
Maintainer thường phản hồi trong vòng 6 ngày
-
github_actions
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
Hochfrequenz/aibap.mcp#578 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
mishraprafful/multihull#150 ·
Maintainer thường phản hồi trong vòng 1 ngày