[rig-tasks] Daily rig evaluation — 2026-10-10 — 10/10 passed
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 1-2 dias
- Facilidade para iniciantes
- 15/100
- Tipo de issue
- Documentação
- Clareza
- Precisa de esclarecimento
- Status de atividade
- Ativa
- Stack de tecnologia
- nodejs, typescript
- Domínio
- documentation, tooling
Direção de pesquisa
This is a daily evaluation report, not a defined task, so there is no single change to make. The payload names skills/rig/samples.md, SKILL.md, .github/agents/task-generator.agent.md and .github/agents/rig-expander.agent.md, plus package-lock.json. A maintainer would need to decide which documentation gaps to act on before a first-timer can start. Done is not stated in the payload.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Summary
Theme seed: 7da8a996fd41eaa6ff6bddbfb9f68fa3.
All 10 seeded domain/artifact/constraint combinations were accepted and turned into novel samples; none were rejected or left unused:
| Theme | Status |
|---|---|
| weather and climate / free-form reports / bounded resource budgets | accepted → 541-station-report-budget |
| manufacturing / resource schedules / multilingual labels | accepted → 542-shift-schedule-localizer |
| public transit / linked records / multilingual labels | accepted → 543-transit-stop-translation-joiner |
| ecology / free-form reports / multilingual labels | accepted → 544-species-observation-digest |
| food and recipes / tabular measurements / incremental updates | accepted → 545-recipe-conversion-delta |
| travel planning / free-form reports / bounded resource budgets | accepted → 546-itinerary-budget-pipeline |
| scientific experiments / free-form reports / multilingual labels | accepted → 547-lab-result-steered-translation |
| inventory and logistics / resource schedules / bounded resource budgets | accepted → 548-warehouse-slot-allocator |
| media production / linked records / partial observations | accepted → 549-shot-log-gap-reconciler |
| community events / resource schedules / incremental updates | accepted → 550-venue-booking-diff-writer |
The 18-entry cached backlog pool was checked first. Every entry mapped onto a lesson already published under samples 511–540 (commit-type classification, nested tsconfig validation, repair-addon demos, healthcheck timeouts, lint/test gate workflows, glob-tagger fan-out, port-scanner tool stub, dependency side-effect report, dotenv completeness, commit churn classification, review-readiness report, basename similarity, parallel word-frequency aggregation, bug-template stub writer, line-count report writer, semver-satisfies tool, TODO/FIXME parallel workflow, and steering+nullable commit). None were novel, so all capacity went to the 10 freshly seeded themes.
| Task | Description | Typecheck | Novelty / closest sample | Key finding |
|---|---|---|---|---|
| 1 (new) | timeout()-budgeted multi-station weather summarizer with truncation flag | ✅ pass | new lesson vs. 524-healthcheck-timeout.md, 49-timeout-signal-helper.md | Clean first-try generation; correctly combined timeout() with a multi-item array output and an explicit truncatedDueToBudget boolean instead of the usual single-probe status enum. |
| 2 (new) | shift roster localized into an s.record keyed by 4 fixed locale codes | ✅ pass | new lesson vs. 41-parse-coverage.md, 32-command-planner.md | Model used s.record(s.array(s.object(...))) correctly on the first attempt; instructions explicitly forbid partitioning workers across locales, which is the key novelty-preserving constraint. |
| 3 (new) | defineTool-mediated transit-stop/translation join with nullable missing lookups | ✅ pass | new lesson vs. 538-semver-satisfies-tool.md, 517-custom-tool-port-scanner-stub.md | Tool handler correctly returns { label: undefined } for missing table entries and the agent is told to map that to null in the typed array output — exercises s.optional at the tool boundary and s.nullable at the agent boundary simultaneously. |
| 4 (new) | multilingual species-survey digest with frequency-based local-name disambiguation + p.write | ✅ pass | new lesson vs. 518-write-sideeffect-report-generator.md, 537-line-count-report-writer.md | Embedded CSV-like log inside a p.bash printf heredoc; output mostCommonLocalName record requires picking the single most frequent variant per species, a disambiguation step absent from prior write-report samples. |
| 5 (new) | baseline recipe + sparse update list merged with s.optional previousAmount | ✅ pass | new lesson vs. 508-dotenv-template-validator.md, 42-json-repair.md | Generated program correctly instructs the model to omit previousAmount (not null/0) for unchanged rows, matching the s.optional semantics precisely. |
| 6 (new) | workflow() with phase-filtered variable-size fan-out to a free-form report stage | ✅ pass | new lesson vs. 525-lint-test-gate-workflow.md, 515-sequential-release-gate-workflow.md | The workflow body does the array filter (.filter(d => d.overBudget)) in plain TypeScript between phases rather than inside a prompt — the cleanest demonstration yet of deterministic post-processing between call() invocations. |
| 7 (new) | steering()+repair() over an array of enum-locale-tagged lab measurements | ✅ pass | new lesson vs. 540-steering-nullable-commit.md, 126-tsconfig-options-auditor.md | Correctly orders addons: [steering(), repair()] and applies the s.enum("en","ja") constraint per-array-item rather than at the top level, which is the intended structural novelty. |
| 8 (new) | model-driven delegation of two subagents sharing one fixed capacity budget | ✅ pass | new lesson vs. 58-genaiscript-travel-plan-port.md, 50-end-to-end-release-agent.md | Uses the agents: {...} field (true model-driven delegation) rather than workflow(); both subagents are given the identical embedded SLOTS string so they reason against the same shared constraint. |
| 9 (new) | defineTool partial-observation reconciliation with derived completeness percentage | ✅ pass | new lesson vs. 77-env-key-checker.md, 419-dotenv-drift-detector.md | Tool handler models "not found" as { found: false } (no durationSec), and the agent is asked to compute completenessPct as a ratio over the aggregated found/missing tool results — a derived metric not present in prior tool-lookup samples. |
| 10 (new) | agent self-computes a diff between two embedded full snapshots before writing a changelog | ✅ pass | new lesson vs. 2-review-git-diff.md, 518-write-sideeffect-report-generator.md | Embeds both "yesterday" and "today" schedules via two separate p.bash heredocs and explicitly instructs the model that "no external diff tool has been run" — forcing snapshot-to-snapshot reasoning rather than consuming a pre-computed diff. |
Problems encountered
Custom subagent tool incompatibility (not a rig API issue). The task-generator custom agent failed on both invocation attempts with 400 Cannot translate Copilot request feature 'frequency_penalty' between Responses and Chat Completions. before producing any output. This appears to be a harness/model-routing configuration issue for that specific custom agent definition (likely a default model or sampling-parameter mismatch between the Responses and Chat Completions APIs), not a problem with the rig codebase. Workaround: task proposal and novelty verification were performed directly by reading /tmp/gh-aw/agent/sample-catalog.json, skills/rig/samples.md, and the nearest sample files, which worked fine and produced the same structured task objects the subagent would have returned.
The rig-expander custom agent hit the identical 400 ... frequency_penalty ... error on its default model for the first invocation, but succeeded on every subsequent call once an explicit model: "claude-sonnet-5" override was passed to the task tool. This strongly suggests the custom agent's configured default model/routing (not the rig harness) is the root cause — worth checking whether .github/agents/task-generator.agent.md and .github/agents/rig-expander.agent.md pin an incompatible model identifier or sampling parameter set.
npm install infrastructure (environment-level, not rig). npm ci repeatedly failed with "Exit handler never called!" because package-lock.json's resolved URLs point to (msfeed25.pkgs.visualstudio.com/redacted), a corporate package-feed mirror presenting a self-signed TLS certificate in this sandbox, while registry.npmjs.orgitself was directly reachable and fully functional. Root cause isolated vianpm config get replace-registry-host(default"npmjs", which is supposed to replace npmjs URLs with the configured registry — but the *lockfile's* resolvedfields already point at the broken mirror host, and no combination of--replace-registry-host=never/always, --omit-lockfile-registry-resolved, or a scoped --userconfigfixednpm cidirectly). Workaround: locally rewrotepackage-lock.json's resolvedURLs from the mirror host toregistry.npmjs.org(same versions, only host changed), rannpm cisuccessfully, thengit checkout package-lock.jsonto restore the original committed lockfile afterward —node_modules` remained installed and correct, and the original lockfile was never modified in the final commit. This is purely a local sandbox/network artifact and was not included in the PR.
No rig program typecheck failures occurred this run — all 10 generated samples passed --typecheck on the first attempt.
Improvement opportunities
Missing or undiscoverable schema helpers (s.*)
No gaps surfaced this run — s.record, s.nullable, s.optional, s.enum, s.path, and s.int/s.number all composed cleanly for every proposed data contract (locale-keyed records, nullable tool-join results, sparse-update optional fields, enum-constrained array items). No new helper was needed.
Missing or undiscoverable prompt helpers (p.*)
p.write with a non-empty placeholder second argument (e.g. p.write("species-digest-report.md", "# Species Observation Digest\n")) is a slightly awkward convention — several existing samples (69, 530) use a near-empty placeholder like "" while newly generated programs sometimes supply a more "real-looking" starter string. It would help SKILL.md to state explicitly whether the second argument is purely a type-inference placeholder (ignored at runtime, replaced by the model's actual write) or an actual seed/fallback value, since model-generated programs are inconsistent about how much content to put there.
Error message quality
No unclear --typecheck messages were encountered this run (none of the 10 generated programs failed typecheck), so there is no direct evidence to report here today.
API ergonomics
The workflow() body pattern of const x = await call(...) followed by plain-TypeScript array filtering (as in sample 546) worked well and is a good idiom, but it is not yet documented as a first-class pattern anywhere in skills/rig/samples.md's workflow bucket description or SKILL.md — every existing deterministic-workflow example combines/aggregates fixed-shape results rather than filtering an array to a variable-size subset before the next call(). Making this "filter-then-call" idiom an explicit, named pattern (with a short canonical snippet) would likely reduce future model confusion between workflow fan-out (fixed parallel calls) and workflow fan-out with dynamic subsetting.
Candidate lint rules
No single repeated model mistake was observed across the 10 tasks (all typechecked on the first attempt), so no new lint rule is proposed this run. One soft observation: nothing currently flags a p.write(path, placeholder) call where placeholder is a long, fully-formed-looking document rather than a short stub — this is a style nit, not a correctness bug, and not clearly autofixable without knowing authorial intent, so it is not proposed as a rule.
Documentation gaps
- SKILL.md would benefit from one sentence clarifying the intended semantics of
p.write's second (placeholder/seed) argument, as noted above. - The workflow-bucket section of
skills/rig/samples.mdcould add a one-line "read when" pointer for "an earlier phase's output determines the size/shape of a later phase's workload" now that sample 546 exists, so future duplicate-detection passes can find it by description rather than only by family name. - Consider documenting the
.github/agents/*.agent.mdcustom-agent model/routing requirements (or at least noting known model-compatibility caveats) given thefrequency_penaltytranslation failures encountered against their default-configured model in this run.
Tasks run today
- (backlog, rejected duplicate)
59924b5fconventional-commit classifier — duplicate of 511-commit-type-enum-classifier.md - (backlog, rejected duplicate)
849f4c19nested tsconfig validator — duplicate of 512-nested-tsconfig-validator.md - (backlog, rejected duplicate)
29b3efd2repair() addon strict-schema demo — duplicate of 513-json-repair-addon-demo.md - (backlog, rejected duplicate)
00975351timeout() health-check — duplicate of 524-healthcheck-timeout.md - (backlog, rejected duplicate)
cb7220easequential lint/test/gate workflow — duplicate of 525-lint-test-gate-workflow.md - (backlog, rejected duplicate)
fb40f271glob fan-out tagger aggregator — duplicate of 536-glob-tagger-fanout.md - (backlog, rejected duplicate)
f28162bdcustom port-reachability tool — duplicate of 517-custom-tool-port-scanner-stub.md - (backlog, rejected duplicate)
fcc95898package.json dependency report + p.write — duplicate of 518-write-sideeffect-report-generator.md - (backlog, rejected duplicate)
e8b4be8c.env-like config parser with source enum — duplicate of 508-dotenv-template-validator.md - (backlog, rejected duplicate)
51b73010git churn risk classifier (3-step) — duplicate of 157-commit-churn-classifier.md - (backlog, rejected duplicate)
d60fbc4freview-readiness report (diff stats + CONTRIBUTING.md) — duplicate of 527-review-readiness-report.md - (backlog, rejected duplicate)
c7604627basename similarity scorer tool — duplicate of 528-basename-similarity-finder.md - (backlog, rejected duplicate)
c45587b3parallel word-frequency aggregator workflow — duplicate of 529-parallel-wordfreq-aggregator.md - (backlog, rejected duplicate)
9f72e395conditional bug-report template stub writer — duplicate of 530-bug-template-stub-writer.md - (backlog, rejected duplicate)
9904c748line-count report writer — duplicate of 537-line-count-report-writer.md - (backlog, rejected duplicate)
0fee113asemver-satisfies custom tool — duplicate of 538-semver-satisfies-tool.md - (backlog, rejected duplicate)
32be38c1parallel TODO/FIXME counter workflow — duplicate of 539-todo-fixme-parallel-workflow.md - (backlog, rejected duplicate)
ff840d87steering()+repair() nullable commit field — duplicate of 540-steering-nullable-commit.md - (new) 541-station-report-budget — runtime bucket, timeout() + partial-completion truncation flag, lesson: self-reported budget truncation over a multi-item array
- (new) 542-shift-schedule-localizer — schemas bucket, locale-keyed s.record fan-out, lesson: fixed-key record re-stating a full dataset per key
- (new) 543-transit-stop-translation-joiner — tools bucket, defineTool join with nullable lookup miss, lesson: two-dataset join surfacing "not found" as s.nullable
- (new) 544-species-observation-digest — context bucket, p.write + frequency disambiguation, lesson: picking the most common of several competing labels before writing
- (new) 545-recipe-conversion-delta — schemas bucket, sparse update merge, lesson: s.optional conditioned on a changed boolean across a full+sparse dataset merge
- (new) 546-itinerary-budget-pipeline — workflows bucket, phase-filtered variable fan-out, lesson: an earlier phase's boolean array dynamically sizes a later phase's workload
- (new) 547-lab-result-steered-translation — runtime bucket, steering()+repair() over array/enum, lesson: addon validation surface across array items, not a single top-level field
- (new) 548-warehouse-slot-allocator — delegation bucket, shared-budget subagent coordination, lesson: two subagents reasoning against one shared numeric constraint, root produces a remediation plan
- (new) 549-shot-log-gap-reconciler — tools bucket, partial-observation tool reconciliation, lesson: per-item tool calls whose absent-result drives a derived completeness percentage
- (new) 550-venue-booking-diff-writer — context bucket, self-computed snapshot diff, lesson: agent computes its own diff between two full snapshots rather than consuming a pre-diffed artifact
Companion PR: adds all 10 passing, novel samples and the refreshed skills/rig/samples.md catalog.
Generated by Daily Rig Task Generator · copilot · auto · 455.9 AIC · ⌖ 17.6 AIC · ⊞ 11.4K · ◷
- Linguagem predominante
- TypeScript
- Estrelas
- 13
- Forks
- 0
- Merge médio
- 1h 40min
- PRs com merge (30d)
- 26
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de githubnext/rig
-
agentic-workflows
Dificuldade 4/5 1-2 dias Facilidade para iniciantes 15/100
githubnext/rig#619 ·
Mantenedores costumam responder em até 1 dia
-
ai-agent automation
Dificuldade 3/5 Meio dia Facilidade para iniciantes 15/100
githubnext/rig#618 ·
Mantenedores costumam responder em até 1 dia
-
agentic-workflows
Dificuldade 4/5 1-2 dias Facilidade para iniciantes 18/100
githubnext/rig#617 ·
Mantenedores costumam responder em até 1 dia
-
agentic-workflows
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 18/100
githubnext/rig#614 ·
Mantenedores costumam responder em até 1 dia
-
ai-agent automation
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 12/100
githubnext/rig#611 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de githubnext/rig
Issues semelhantes
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 77/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 72/100
cline/mcp-marketplace#2932 ·
-
community first-timers-only good first issue hacktoberfest help wanted low hanging fruit up-for-grabs
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 82/100
lingdojo/kana-dojo#32188 · 1 comentário · 5 reações ·
Mantenedores costumam responder em até 1 dia
-
triage
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 76/100
Portkey-AI/gateway#1844 ·
Mantenedores costumam responder em até 1 dia
-
needs triage
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 61/100
rjsf-team/react-jsonschema-form#5495 ·
Mantenedores costumam responder em até 2 dias