Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[FEATURE] Improve PIT usage for large queries over wildcard index patterns

オープン
#5,698 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
活発
技術スタック
java
領域
backend, data

調査の方向性

CreatePitRequest と #3879 で参照されているクエリの pushdown 作業から始め、その後、提案されている 5 つのアプローチをワイルドカード展開および PIT の制限と比較してください。完了した変更では、定義された緩和策を選択して実装し、明示的な制限の例と無制限スキャンの例の両方を対象にし、広範なワイルドカードクエリが失敗しなくなったことを検証する必要があります。

索引モデルが issue の本文から書いたものです。

説明

enhancement

Is your feature request related to a problem?

The plugin creates a PIT whenever a non-aggregate request needs more rows than index.max_result_window (default 10,000). When the query uses a wildcard index pattern, OpenSearch expands the wildcard and opens one reader context per matching shard, so a broad pattern over hundreds of daily indices can exhaust the per-node search.max_open_pit_context limit (default 300) and the query fails:

Trying to create too many point in time contexts. Must be less than or equal to: [300].

Examples

Two common query shapes trigger it:

  1. Explicit large limit: source=logs-* | head 100000 - the limit is pushed into the scan and exceeds the window, so a PIT is created and ten pages are fetched.
  2. Unbounded scan: source=logs-* | streamstats count() as seen - no explicit limit, so the plugins.query.size_limit cap sits above an operator it cannot be pushed through. The scan is left unbounded and a PIT is created, but only 10,000 rows are returned and no second page is ever requested, so the snapshot serves no purpose.

What solution would you like?

Option Approach Pros Cons Notes
1. Predicate-based index pruning Narrow the wildcard to only the indices that can match, then open the PIT over that set. Snapshot semantics unchanged. No consistency change. Smallest change. Works even for shapes that cannot avoid a PIT. Reimplements pruning core already does, and misses core's ongoing work in this area. Core could instead expose can_match, or new index/field-level stats, as an internal API. See opensearch-project/OpenSearch#21865, #22483, #22451.
2. Incremental execution Split the resolved index set into batches and scan batch by batch, opening and closing one PIT per batch. Bounds context count regardless of how broad the pattern is. Complements option 1 when many indices still match after pruning. Snapshot is per batch rather than global. More PIT create/delete calls and longer wall-clock time. Could later extend to progressive result delivery, returning rows as each batch completes.
3. Stateless pagination Drop the PIT and page with a value-based search_after cursor. Each page is a plain search holding no server state. Nothing accumulates against the cap. Every page gets can_match and coordinator pruning automatically. No cross-page snapshot, so concurrent writes may cause missed or duplicated rows. Needs a stable, unique sort key, and none is both cheap and globally unique. Known as keyset pagination. _shard_doc requires a PIT (opensearch-project/OpenSearch#18924); _seq_no is the closest alternative but is unique only per shard. Paginate docs
4. Full pipeline pushdown Compile the whole pipeline into one search so the cluster returns a finished result — no pagination, like aggregation queries today. Removes the failure mode entirely. Fixes a far broader translation gap than this issue. Largest effort. Feasibility unverified, and coverage can never be complete, so a fallback is still needed. #3879 and #5646 are prior art for widening pushdown coverage. Scripted metrics are a possible escape hatch for pipelines with no aggregation equivalent.
5. New search primitive Add the missing primitive in core: a PIT scoped by predicate, or a search that owns its own pagination state. Clean for every client, not just SQL/PPL. A core contribution, so timeline and effort are unknown. CreatePitRequest has no query body or can-match phase today, and no upstream proposal exists. opensearch-project/OpenSearch#22530 is a precedent for adding a can-match phase to an engine.

What alternatives have you considered?

Mitigations and adjacent directions considered, none of which we treat as a fix:

Alternative Effect Drawback
Raise search.max_open_pit_context Query untouched; more contexts permitted. Each context pins segment readers and blocks merged-segment deletion.
Raise index.max_result_window and use from + size Avoids the PIT entirely. Every matching shard then returns up to size top hits for coordinator-side reduction, a far larger memory spike than paged reads.
Request fewer rows than max_result_window No PIT; the query succeeds immediately. Truncates the input to downstream row-consuming operators, so results change.
Fewer primary shards for new indices Cuts fan-out as indices roll over. Only affects indices created afterwards, so it does not resolve an active failure.
Precompute with rollups or transforms Removes the query shape entirely for recurring dashboards. Only suits known, repeated queries, and adds a pipeline to maintain.
Offload to the async query path Sidesteps coordinator limits for very large fetches. Changes the interaction model to submit-and-poll; not a drop-in for dashboard traffic.

Do you have any additional context?

Related work in this repo:

  • #3879 avoided PIT for queries under max_result_window — prior art for the pushdown direction.
  • #5634, #5220 are symptom reports of the same failure.
  • #5631 surfaces the exhaustion with an actionable error message.
主要言語
Java
スター
176
フォーク
230
平均マージ
2日 11時間
マージ済み PR(30日)
32

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

opensearch-project/sql のほかの issue

opensearch-project/sql の issue をすべて見る

似ている issue

Java の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。