feat(aidd-dev/test): enforce test effectiveness and test-double discipline
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 82/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- markdown
- Domain
- documentation
Research direction
Read plugins/aidd-dev/skills/06-test/actions/01-test.md first, then compare its contract with plugins/aidd-dev/skills/06-test/SKILL.md and the cited testing guidance in cli/aidd_docs/memory/testing.md. Update the action and parent skill only as needed so generated tests require observable behavior, justified doubles, and an explicit regression-effectiveness check while preserving project conventions; verify every acceptance criterion without adding dependencies.
Written by the indexing model from the issue text.
Description
Problem
aidd-dev:06-test asks the agent to identify untested behavior, generate tests, iterate until they pass, and review them against quality criteria. The current contract says to test functional behavior rather than implementation details, but it does not explicitly define a minimum discipline for test doubles or how to check that a generated test protects the intended behavior.
An agent can therefore produce a technically green but low-value test: configure a mock to return a value, call the SUT, assert that same value, and optionally assert that the mock was called. The test passes while exercising little or none of the behavior it is meant to protect. The relevant question is not only whether the test passes, but whether it would fail if the intended behavior or regression were broken.
This is not a request to ban mocks. It is a request to make observable behavior and test effectiveness explicit in the generic skill.
Scope
- Update
plugins/aidd-dev/skills/06-test/actions/01-test.mdand, only if needed for consistency, the parentSKILL.md. - Define framework-independent quality criteria for generated tests: assert observable behavior, never replace the behavior under test with a double, use doubles at justified seams or boundaries, prefer real domain/pure objects when practical, and justify interaction mocks when the interaction itself is the behavior under test.
- Require an explicit effectiveness check before considering a generated test complete: would the test fail if the intended behavior/regression were broken?
- Reject tautological mock-return/assert arrangements that provide no meaningful regression protection.
- Preserve existing project testing conventions as authoritative.
Acceptance criteria
-
06-testrequires assertions against observable behavior rather than implementation details or only a collaboration graph. - The behavior under test is not itself replaced by a mock or other test double.
- Test doubles are introduced only at justified boundaries/seams, or when the interaction itself is the behavior being verified.
- Generated tests avoid tautological mock-return/assert patterns that provide no meaningful regression protection.
- Before considering a generated test complete, the skill explicitly evaluates whether breaking the intended behavior would cause the test to fail.
- Existing project testing commands, frameworks, and conventions remain authoritative.
- No mocking, coverage, or mutation-testing dependency is installed merely to satisfy this contract.
- Focused structural checks or documentation tests, if required by the repository, verify the new contract without turning it into a language-specific testing guide.
Prior art in this repo
plugins/aidd-dev/skills/06-test/SKILL.md#L22-L26already requires functional behavior and rejects coupling to implementation details.plugins/aidd-dev/skills/06-test/actions/01-test.md#L13-L19already asks the agent to review passing tests against quality criteria, but those criteria are not yet explicit about test doubles or regression protection.cli/aidd_docs/memory/testing.md#L7-L10distinguishes unit, integration, and E2E layers.cli/aidd_docs/memory/testing.md#L17-L21already states:Doubles from tests/helpers/ports/. Substitute at the seam; never mock functional behaviour.- #907 reserves
06-testfor project-owned automated tests, classifies test layers, and reports critical behavior left uncovered. This issue is complementary: it defines what makes an individual generated test meaningful enough to be considered complete. - #45 and #53 provide historical prior art around testing practices,
no-mocks-without-reason, and avoiding fragile or implementation-coupled tests.
Out of scope
- Banning mocks, stubs, fakes, spies, or other test doubles in general.
- Choosing one testing framework, mocking library, coverage threshold, or mutation-testing tool for all projects.
- Making mutation testing mandatory; an existing project configuration may be used according to its conventions, but no new dependency is required here.
- Reworking
test-journeyor browser acceptance QA, which are covered by the #907/#919 workstream. - Replacing the scope of #907: test-layer classification, coverage discovery, and reporting critical uncovered behavior remain there.
- Dominant language
- TypeScript
- Stars
- 481
- Forks
- 45
- Avg merge
- 11h 4m
- Merged PRs (30d)
- 93
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from ai-driven-dev/framework
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
ai-driven-dev/framework#940 ·
Maintainers usually reply within 1 day
-
refactor(aidd-orchestrator): the check zone says when to stop, and reviews its axes in one roundOpen
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
ai-driven-dev/framework#887 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
ai-driven-dev/framework#873 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
ai-driven-dev/framework#625 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
ai-driven-dev/framework#467 · 1 comment ·
Maintainers usually reply within 1 day
All issues in ai-driven-dev/framework
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
prime-radiant-inc/evener#3726 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
FuRongJun-1999/dsh-memory#56 ·
Maintainers usually reply within 1 day
-
bug via-triage
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
pingdotgg/t3code#15682 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
openwatersio/slackwater.xyz#152 ·
Maintainers usually reply within 1 day
-
[BUG] 请修改标题为您遇到的问题Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
OpenListTeam/OpenList-Worker#103 ·
Maintainers usually reply within 1 day