Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Optimize]: Resolution budgets, crop/bbox coordinate tools, and image-workflow guidance for vision tools

オープン
#2,248 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
38/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
静か
技術スタック
rust

調査の方向性

まず、ImageLimits と coordinate_note の処理を含め、既存の analyze_image、view_image、optimize_image_with_size_limit、ImageAnalyzer の各パスをたどります。次に、組み込みの SKILL.md ファイルがどのように登録されているかを確認します。budget プリセット、crop_image と draw_bbox、budget-path のアップスケーリング、image-workflow skill が連携し、従来の動作が変更されないことをもって完了とします。

索引モデルが issue の本文から書いたものです。

説明

enhancement

Summary

Upgrade BitFun's built-in vision tools with four capabilities borrowed from Qwen-MM-Plugins' design:

  1. Resolution budget presets for analyze_image / view_image (small / normal / large)
  2. Coordinate closed loop: new crop_image and draw_bbox tools using 0-1000 normalized coordinates
  3. Small-image upscaling to a minimum pixel floor (capped at 2x linear) so tiny crops stay legible for VLMs
  4. Built-in decision skill (image-workflow) teaching the agent when to view vs analyze vs crop, and how to pick a budget

Background

The current vision path is two fixed-limit tools plus a batch pre-analysis path:

  • analyze_image sends the image + prompt to the configured image-understanding model. Every call pays full image tokens, and the only size control is the hard provider cap (1MB tool cap / provider dimension limits). There is no way for the agent to trade detail against cost.
  • view_image attaches an image to the primary model context, also with no resolution control.
  • optimize_image_with_size_limit never upscales, so a tiny crop (e.g. a 64x64 region) is sent as-is and OCR/detail suffers.
  • The only coordinate handling is a prose coordinate_note warning the model that positions in the analysis are estimates. There is no tool to actually crop or annotate a static image, so the estimate cannot be turned into an action loop.
  • Tool descriptions are one-liners; the model has no guidance on which tool to use when.

Qwen-MM-Plugins solves these with token-budget-driven resolution presets (256/1024/2048 visual tokens), a crop/draw_bbox pair that consumes grounding output in the same 0-1000 normalized coordinate system, an upscale floor, and declarative SKILL.md decision documents.

Proposed changes

1. Budget presets (optional parameter, legacy behavior unchanged)
  • Add ImageBudget (small/normal/large) mapping to pixel targets via token budgets (256/1024/2048 x 32^2 = ~512^2 / ~1024^2 / ~1448^2), clamped by the provider ImageLimits ceiling.
  • Add an optional budget parameter to analyze_image and view_image. When omitted, the existing path runs unchanged (no behavior change for existing callers/configs).
  • Fix the resize_note wording which currently hard-codes "downscaled" and would misreport upscaling.
2. Coordinate closed loop: crop_image + draw_bbox
  • crop_image(path, box[x1,y1,x2,y2 in 0-1000], output_path?): validates and clamps the box, saves the cropped region next to the source by default, and returns pixel-mapped coordinates. Attaches a preview when the primary model supports multimodal tool output; degrades to a text summary otherwise.
  • draw_bbox(path, boxes[{bbox,label?}], output_path?): draws rectangles with an adaptive line width and a fixed palette. v1 draws boxes only (no text labels, no new dependencies).
  • Both support remote workspace paths (read + write through workspace filesystem services).
  • Closed loop: the model reports a region in normalized coordinates, crop_image cuts it, then analyze_image(crop_path, budget="large") re-inspects the region at high resolution.
3. Small-image upscaling
  • In the budget path only: images below the minimum pixel floor (~512^2) are upscaled, capped at 2x linear so tiny images do not burn tokens pointlessly (larger regions should be re-cropped instead).
4. Built-in decision skill
  • New builtin skill image-workflow (SKILL.md): tool-selection decision table, budget guidance (small preview / normal default / large detail; large on crops), coordinate discipline (normalized 0-1000, model coordinates are estimates, use reported width/height/was_resized to convert), and cost tips (image tokens dominate; crop before repeated full-image calls).

Non-goals (follow-ups)

  • 32px patch-grid snapping for Qwen-style providers
  • Text label rendering in draw_bbox (needs font rasterization)
  • Budget support for the ImageAnalyzer batch pre-analysis path
主要言語
Rust
スター
2.3k
フォーク
236
平均マージ
2時間 42分
マージ済み PR(30日)
466

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

GCWing/OpenBitFun のほかの issue

GCWing/OpenBitFun の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。