Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Four CN Lite tasks have missing or mismatched source data (tasks 33, 207, 380, and 381)

オープン
#23 コメント 1 件 リアクション 1 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発

調査の方向性

まず、tasks 33、207、380、381 の manifests と file_dep_graphs を、Issue に記載された実際の Workspace ファイルと比較します。影響を受ける Prompts と Rubric-Grounding audits を確認し、そのうえで各 task に修正済みの Assets が必要なのか、改訂された Contract が必要なのかを判断します。各必須結果が task から見える Inputs から導出可能になり、影響を受ける Rubrics が再生成されれば完了です。

索引モデルが issue の本文から書いたものです。

説明

Summary

Four tasks in the Chinese Workspace-Bench Lite dataset appear to have missing or mismatched source data:

  • opendatabox-workspace-bench-33
  • opendatabox-workspace-bench-207
  • opendatabox-workspace-bench-380
  • opendatabox-workspace-bench-381

I checked:

  1. The task metadata, data_manifest, and file_dep_graph
  2. The actual files distributed with each Lite task
  3. The complete filename index of filesys_cn.zip from Workspace-Bench/Workspace-Bench-Workspaces (28,023 entries)
  4. The provided rubric-grounding audit

The issues are described below.


1. opendatabox-workspace-bench-33: hospital-grade data is unavailable

The task asks the agent to calculate the numbers of level-1, level-2, and level-3 hospitals in eastern, central, and western China.

The supplied inputs are:

  • 1-12 2023年各地区按床位数分组的社区卫生服务中心(站)数.xlsx
  • 1-2 2023年各地区医疗卫生机构数.xlsx
  • 1-3 2023年各类医疗卫生机构数.xlsx

The files contain:

  • Regional and province-level hospital totals
  • Hospital counts by institution type
  • Community health center/station bed-count groups
  • National institution classifications

They do not contain regional or province-level hospital-grade fields.

A search of the complete Operations Manager workspace found no hospital-grade distribution workbook. The closest file is:

  • 医疗分析/医院收支/4-12 2023年三级公立医院收入与支出.xlsx

That workbook contains national financial indicators by hospital grade, not regional or province-level hospital-grade counts.

As a result:

  • The eastern, central, and western hospital totals are grounded.
  • The community health center/station statistics are grounded.
  • The required level-1, level-2, and level-3 counts for the three regions are not independently derivable.
  • The Beijing and Shanghai hospital totals are grounded, but their grade-specific counts are not.

The current audit marks 10 of 19 rubrics as grounded and 9 as ungrounded.

Suggested maintainer action

Either:

  1. Add the intended regional hospital-grade source workbook to the Operations Manager workspace and the task manifest; or
  2. Remove the hospital-grade calculations and regenerate the affected rubrics using the available institution-type and bed-group data.

2. opendatabox-workspace-bench-207: required scoring-model workbook is missing

The task explicitly says:

基于5-通用人才画像(模型).xlsx评价四份简历,并且生成人才评价.xlsx到桌面

However, the task manifest contains only four resume files:

  • 张浩然简历.docx
  • 王佳宁简历.docx
  • 赵思远简历.docx
  • 李雨辰简历.docx

The required 5-通用人才画像(模型).xlsx is absent from both the task manifest and file_dep_graph.

A full search of the Logistics Manager workspace found no file named:

  • 5-通用人才画像(模型).xlsx
  • 5-通用人才画像(模型).xlsx

It also found no filename containing 通用人才画像.

The workspace contains only these related templates:

  • 人才画像/模板/2-人才画像矩阵图.xlsx
  • 人才画像/模板/3-人才画像_履历模型_.xlsx
  • 人才画像/模板/4-人才画像_冰山模型_.xlsx
  • 人才画像/模板/6-核心岗位人才画像_模版_.docx

None can safely be assumed to be the missing model.

Without the model workbook, an agent cannot independently recover:

  • The five scoring dimensions and their exact weights
  • The grade thresholds
  • The intended scoring formulas
  • Exact candidate scores such as 82, 85, 88, and 92

Resume-derived facts remain grounded, but the scoring-model results do not. The current audit marks 10 of 20 rubrics as grounded and 10 as
ungrounded.

Suggested maintainer action

Recover and add 5-通用人才画像(模型).xlsx to:

  • The Logistics Manager workspace
  • data_manifest
  • file_dep_graph
  • Any generated setup/input manifest

If the original workbook contains expected candidate answers, please publish a sanitized template containing only task-visible criteria,
weights, formulas, and grade thresholds, and regenerate rubrics that depend on hidden scores.


3. opendatabox-workspace-bench-380: prompt and source assets describe different business scenarios

The prompt claims the following assets are provided:

  • activity_photos/
  • impact_chart.csv
  • feedback_wordcloud.png
  • future_plan.md
  • recruitment_poster.png

None of these paths exists in the complete Chinese workspace image.

The actual task manifest contains:

  • post_1.json
  • market_analysis_1.json
  • user_feedback_category_202601.md
  • strategic_plan_1.md
  • training_module_1.md

These files describe corporate marketing operations rather than a volunteer association:

  • post_1.json is an Instagram marketing post and contains only a remote example.com image URL.
  • market_analysis_1.json describes market size, growth, and competitors.
  • user_feedback_category_202601.md contains product, UX, and customer-service feedback.
  • strategic_plan_1.md contains customer-growth, product, and revenue targets.
  • training_module_1.md is a Marketing Operations training course.

Importantly, the files are not unrelated to the rubrics. Most exact rubric values come directly from these actual marketing files, including:

  • Reach of 87,170,000
  • 2,070,100 new followers
  • A 3.51% engagement rate
  • Market size 318990B, growth 6%, and CAGR 11%
  • Feedback counts and percentages
  • The 2024–2026 revenue targets
  • The four-hour online Marketing Operations course

The current audit therefore marks 21 of 25 rubrics as grounded.

The defect is a semantic mismatch among:

  • A volunteer-association prompt
  • Corporate marketing inputs
  • Rubrics primarily generated from the marketing inputs

Renaming the current files would not resolve the mismatch: a marketing post is not an activity-photo collection, and a marketing training
module is not a recruitment poster.

Suggested maintainer action

Either:

  1. Add the five intended volunteer-association assets and regenerate all content rubrics from them; or
  2. Rewrite the prompt as a corporate marketing annual-review task, retaining the current inputs and grounded rubric values while removing the
    unsupported volunteer, activity-photo, word-cloud, and recruitment-poster requirements.

Given the existing rubric content, the second option may require fewer changes.


4. opendatabox-workspace-bench-381: filename, year, format, data-grain, and scoring-model mismatch

The prompt names five 2025 CSV files:

  • hospital_finance_2025.csv
  • liabilities_2025.csv
  • drug_production_2025.csv
  • medical_expenses_2025.csv
  • chronic_disease_2025.csv

None exists in the complete Chinese workspace image.

The actual task inputs are five Chinese-named 2023 XLSX workbooks:

  • 4-12 2023年三级公立医院收入与支出.xlsx
  • 4-6 2023年各类医疗卫生机构资产与负债.xlsx
  • 4-4-2 2023年全国药品生产流通.xlsx
  • 4-20 2023年各地区公立医院门诊和住院病人次均医药费用.xlsx
  • 9-10 2023年调查地区15岁及以上居民慢性病患病率.xlsx

This is not only a filename mismatch. The data grains are incompatible with the requested province-level correlation analysis:

  • Medical expenses: 31 province-level regions
  • Drug production/distribution: 31 province-level regions
  • Chronic disease: national, urban/rural, and broad eastern/central/western aggregates only
  • Hospital income and expenditure: hospital-grade aggregates, not provinces
  • Assets and liabilities: institution-type aggregates, not provinces

Therefore, the inputs cannot be merged into a 31-region dataset containing both financial-health and chronic-disease variables.

The task also does not define:

  • The financial-health score
  • The balance/surplus ratio formula
  • Risk-score weights
  • Low/medium/high-risk thresholds
  • How the exact 5/23/3 risk split should be produced

There are also apparent rubric inconsistencies:

  • A value of 86.06% appears more consistent with an expense-to-income ratio than a conventional surplus ratio. A conventional surplus ratio
    would instead be approximately 13.94%.
  • If drug-production licenses are used as the industry-scale variable, the correlation with outpatient cost is approximately r = +0.108,
    not negative.
  • Using drug-distribution-business licenses instead gives approximately r = -0.137, showing that the required negative conclusion depends
    on an unspecified metric choice.

The current audit marks only 3 of 20 rubrics as grounded and 17 as ungrounded.

Suggested maintainer action

This task likely needs to be regenerated rather than fixed through filename aliases alone.

Either:

  1. Add the intended 2025 province-level CSV datasets, explicitly define all formulas and risk thresholds, and regenerate the rubrics; or
  2. Rewrite the task around the available 2023 XLSX workbooks, restrict analysis to supported aggregation levels, define all derived metrics,
    and regenerate the correlation and risk rubrics.

Expected outcome

Please either correct the affected source assets and manifests or revise the task contracts and regenerate rubrics so that every required
result can be independently derived from task-visible inputs.

主要言語
Python
スター
72
フォーク
7
平均マージ
7分
マージ済み PR(30日)
6

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

OpenDataBox/Workspace-Bench のほかの issue

OpenDataBox/Workspace-Bench の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。