[Optimize]: Resolution budgets, crop/bbox coordinate tools, and image-workflow guidance for vision tools
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 38/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- rust
- 領域
- ai, computer-vision, tooling
調査の方向性
まず、ImageLimits と coordinate_note の処理を含め、既存の analyze_image、view_image、optimize_image_with_size_limit、ImageAnalyzer の各パスをたどります。次に、組み込みの SKILL.md ファイルがどのように登録されているかを確認します。budget プリセット、crop_image と draw_bbox、budget-path のアップスケーリング、image-workflow skill が連携し、従来の動作が変更されないことをもって完了とします。
索引モデルが issue の本文から書いたものです。
説明
Summary
Upgrade BitFun's built-in vision tools with four capabilities borrowed from Qwen-MM-Plugins' design:
- Resolution budget presets for
analyze_image/view_image(small/normal/large) - Coordinate closed loop: new
crop_imageanddraw_bboxtools using 0-1000 normalized coordinates - Small-image upscaling to a minimum pixel floor (capped at 2x linear) so tiny crops stay legible for VLMs
- Built-in decision skill (
image-workflow) teaching the agent when to view vs analyze vs crop, and how to pick a budget
Background
The current vision path is two fixed-limit tools plus a batch pre-analysis path:
analyze_imagesends the image + prompt to the configured image-understanding model. Every call pays full image tokens, and the only size control is the hard provider cap (1MB tool cap / provider dimension limits). There is no way for the agent to trade detail against cost.view_imageattaches an image to the primary model context, also with no resolution control.optimize_image_with_size_limitnever upscales, so a tiny crop (e.g. a 64x64 region) is sent as-is and OCR/detail suffers.- The only coordinate handling is a prose
coordinate_notewarning the model that positions in the analysis are estimates. There is no tool to actually crop or annotate a static image, so the estimate cannot be turned into an action loop. - Tool descriptions are one-liners; the model has no guidance on which tool to use when.
Qwen-MM-Plugins solves these with token-budget-driven resolution presets (256/1024/2048 visual tokens), a crop/draw_bbox pair that consumes grounding output in the same 0-1000 normalized coordinate system, an upscale floor, and declarative SKILL.md decision documents.
Proposed changes
1. Budget presets (optional parameter, legacy behavior unchanged)
- Add
ImageBudget(small/normal/large) mapping to pixel targets via token budgets (256/1024/2048 x 32^2 = ~512^2 / ~1024^2 / ~1448^2), clamped by the providerImageLimitsceiling. - Add an optional
budgetparameter toanalyze_imageandview_image. When omitted, the existing path runs unchanged (no behavior change for existing callers/configs). - Fix the
resize_notewording which currently hard-codes "downscaled" and would misreport upscaling.
2. Coordinate closed loop: crop_image + draw_bbox
crop_image(path, box[x1,y1,x2,y2 in 0-1000], output_path?): validates and clamps the box, saves the cropped region next to the source by default, and returns pixel-mapped coordinates. Attaches a preview when the primary model supports multimodal tool output; degrades to a text summary otherwise.draw_bbox(path, boxes[{bbox,label?}], output_path?): draws rectangles with an adaptive line width and a fixed palette. v1 draws boxes only (no text labels, no new dependencies).- Both support remote workspace paths (read + write through workspace filesystem services).
- Closed loop: the model reports a region in normalized coordinates,
crop_imagecuts it, thenanalyze_image(crop_path, budget="large")re-inspects the region at high resolution.
3. Small-image upscaling
- In the budget path only: images below the minimum pixel floor (~512^2) are upscaled, capped at 2x linear so tiny images do not burn tokens pointlessly (larger regions should be re-cropped instead).
4. Built-in decision skill
- New builtin skill
image-workflow(SKILL.md): tool-selection decision table, budget guidance (small preview / normal default / large detail; large on crops), coordinate discipline (normalized 0-1000, model coordinates are estimates, use reported width/height/was_resized to convert), and cost tips (image tokens dominate; crop before repeated full-image calls).
Non-goals (follow-ups)
- 32px patch-grid snapping for Qwen-style providers
- Text label rendering in
draw_bbox(needs font rasterization) - Budget support for the
ImageAnalyzerbatch pre-analysis path
- 主要言語
- Rust
- スター
- 2.3k
- フォーク
- 236
- 平均マージ
- 2時間 42分
- マージ済み PR(30日)
- 466
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
GCWing/OpenBitFun のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
GCWing/OpenBitFun#3213 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
GCWing/OpenBitFun#2363 ·
メンテナーはふだん 1 日以内に返信
-
question
難易度 1/5 1時間未満 初心者へのやさしさ 78/100
GCWing/OpenBitFun#2340 ·
メンテナーはふだん 1 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 1/100
GCWing/OpenBitFun#3273 ·
メンテナーはふだん 1 日以内に返信
-
[Feature]: 生态兼容逻辑优化オープン
難易度 3/5 1〜2日 初心者へのやさしさ 58/100
GCWing/OpenBitFun#3271 ·
メンテナーはふだん 1 日以内に返信
GCWing/OpenBitFun の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
dani-garcia/vaultwarden#7801 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
boxlite-ai/boxlite#1814 ·
メンテナーはふだん 1 日以内に返信
-
tech-debt
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
xingkongliang/skills-manager#516 ·
メンテナーはふだん 5 日以内に返信