Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Rubric design may be too strict for open-ended workspace tasks

オープン
#11 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
静か
技術スタック
huggingface, python
領域
ai, testing-qa

調査の方向性

リンクされた Hugging Face データセットから task_lite_clean_cn/152/metadata.json と task_lite_clean_cn/158/metadata.json を読み始め、各タスクの説明をそのルーブリックと比較します。現在の評価基準が代替ファイル名と、根拠のある計画上の判断をどのように扱っているかを確認します。オープンエンドのタスクで、あらかじめ正確に定義された回答だけに依存せず、必要な出力、網羅性、意味的正確性、品質を評価できれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Hi, thanks for building and releasing Workspace-Bench Lite. I found the benchmark very useful for evaluating agent behavior in realistic workspace environments.

I would like to raise one concern about some open-ended tasks: the task descriptions are relatively broad and allow multiple reasonable solutions, but the rubrics sometimes enforce very specific output details as if there were only one correct answer. This may cause valid task completions to be marked as failures.

Example 1: Task 152
https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite/blob/main/task_lite_clean_cn/152/metadata.json

The task asks the agent to organize several scientific drawing icons into a new folder and rename them so that their contents can be clearly identified.

This is an open-ended file organization task. A user would likely accept multiple reasonable names, such as:

  • 对话气泡-问号.png
  • 问号-气泡.png
  • 灯泡创意-问号.png
  • 绿色数据表格.png

However, the rubric requires exact filenames such as 问号-气泡.png, 问号-人.png, 表格-绿色.png, etc. This makes the evaluation closer to exact-answer matching rather than checking whether the agent correctly understood and renamed the icons.
Also, the task explicitly asks to create a new folder, but the rubric focuses heavily on filenames and image details, while the folder-organization requirement seems less directly evaluated.

Example 2: Task 158
https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite/blob/main/task_lite_clean_cn/158/metadata.json

The task asks the agent to generate an operational work plan based on several user communication records and team responsibilities.

This is also an open-ended planning task. The agent needs to extract user needs, prioritize them, assign responsibilities, and produce an actionable plan. However, the rubric requires very specific conclusions, such as fixed priority levels, fixed department assignments, fixed follow-up mechanisms, and exact demand counts.

Some of these requirements may be reasonable, but they are not always uniquely implied by the task description. For example, different priority assignments may be valid if the agent provides a clear rationale. A good answer should be judged by coverage, reasoning quality, structure, and actionability, not only by whether it matches a predefined plan.

Suggestion

For open-ended workspace tasks, it may be better to split rubrics into different levels:

Hard constraints: required output file exists, correct format, no missing input files, images remain valid, etc.
Content coverage: all key input information is extracted or processed.
Semantic correctness: filenames, summaries, or plans are reasonable and clearly reflect the input.
Quality criteria: consistency, actionability, prioritization rationale, and absence of obvious mismatches.

For tasks like 152, the rubric could allow semantically equivalent filenames instead of requiring exact names.

For tasks like 158, the rubric could evaluate whether the priority assignment is justified and complete, rather than requiring one fixed priority mapping.

This adjustment may make the benchmark better reflect real-world agent performance in open-ended environments, while still keeping the evaluation reliable and diagnosable.

主要言語
Python
スター
72
フォーク
7
平均マージ
7分
マージ済み PR(30日)
6

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

OpenDataBox/Workspace-Bench のほかの issue

OpenDataBox/Workspace-Bench の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。