Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

feat(aidd-dev/test): enforce test effectiveness and test-double discipline

Open Beginner friendly
#952 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
markdown
Domain
documentation

Research direction

Read plugins/aidd-dev/skills/06-test/actions/01-test.md first, then compare its contract with plugins/aidd-dev/skills/06-test/SKILL.md and the cited testing guidance in cli/aidd_docs/memory/testing.md. Update the action and parent skill only as needed so generated tests require observable behavior, justified doubles, and an explicit regression-effectiveness check while preserving project conventions; verify every acceptance criterion without adding dependencies.

Written by the indexing model from the issue text.

Description

Problem

aidd-dev:06-test asks the agent to identify untested behavior, generate tests, iterate until they pass, and review them against quality criteria. The current contract says to test functional behavior rather than implementation details, but it does not explicitly define a minimum discipline for test doubles or how to check that a generated test protects the intended behavior.

An agent can therefore produce a technically green but low-value test: configure a mock to return a value, call the SUT, assert that same value, and optionally assert that the mock was called. The test passes while exercising little or none of the behavior it is meant to protect. The relevant question is not only whether the test passes, but whether it would fail if the intended behavior or regression were broken.

This is not a request to ban mocks. It is a request to make observable behavior and test effectiveness explicit in the generic skill.

Scope

  • Update plugins/aidd-dev/skills/06-test/actions/01-test.md and, only if needed for consistency, the parent SKILL.md.
  • Define framework-independent quality criteria for generated tests: assert observable behavior, never replace the behavior under test with a double, use doubles at justified seams or boundaries, prefer real domain/pure objects when practical, and justify interaction mocks when the interaction itself is the behavior under test.
  • Require an explicit effectiveness check before considering a generated test complete: would the test fail if the intended behavior/regression were broken?
  • Reject tautological mock-return/assert arrangements that provide no meaningful regression protection.
  • Preserve existing project testing conventions as authoritative.

Acceptance criteria

  • 06-test requires assertions against observable behavior rather than implementation details or only a collaboration graph.
  • The behavior under test is not itself replaced by a mock or other test double.
  • Test doubles are introduced only at justified boundaries/seams, or when the interaction itself is the behavior being verified.
  • Generated tests avoid tautological mock-return/assert patterns that provide no meaningful regression protection.
  • Before considering a generated test complete, the skill explicitly evaluates whether breaking the intended behavior would cause the test to fail.
  • Existing project testing commands, frameworks, and conventions remain authoritative.
  • No mocking, coverage, or mutation-testing dependency is installed merely to satisfy this contract.
  • Focused structural checks or documentation tests, if required by the repository, verify the new contract without turning it into a language-specific testing guide.

Prior art in this repo

  • plugins/aidd-dev/skills/06-test/SKILL.md#L22-L26 already requires functional behavior and rejects coupling to implementation details.
  • plugins/aidd-dev/skills/06-test/actions/01-test.md#L13-L19 already asks the agent to review passing tests against quality criteria, but those criteria are not yet explicit about test doubles or regression protection.
  • cli/aidd_docs/memory/testing.md#L7-L10 distinguishes unit, integration, and E2E layers.
  • cli/aidd_docs/memory/testing.md#L17-L21 already states: Doubles from tests/helpers/ports/. Substitute at the seam; never mock functional behaviour.
  • #907 reserves 06-test for project-owned automated tests, classifies test layers, and reports critical behavior left uncovered. This issue is complementary: it defines what makes an individual generated test meaningful enough to be considered complete.
  • #45 and #53 provide historical prior art around testing practices, no-mocks-without-reason, and avoiding fragile or implementation-coupled tests.

Out of scope

  • Banning mocks, stubs, fakes, spies, or other test doubles in general.
  • Choosing one testing framework, mocking library, coverage threshold, or mutation-testing tool for all projects.
  • Making mutation testing mandatory; an existing project configuration may be used according to its conventions, but no new dependency is required here.
  • Reworking test-journey or browser acceptance QA, which are covered by the #907/#919 workstream.
  • Replacing the scope of #907: test-layer classification, coverage discovery, and reporting critical uncovered behavior remain there.
Dominant language
TypeScript
Stars
481
Forks
45
Avg merge
11h 4m
Merged PRs (30d)
93

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from ai-driven-dev/framework

All issues in ai-driven-dev/framework

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.