proposal(ai-sdk): goal-driven step loop with a model-agnostic decide step
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 42/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- typescript
- Domain
- ai, mobile-dev
Research direction
Start with dependency #2656 and the existing agent-device/ai-sdk entry point around createAgentDeviceTools, then review ADR 0014 and ADR 0017. Compare the proposed fake-device and fake-decide tests with examples/sdk/goal-loop.ts and website/docs/docs/ai-sdk.md; done means the iOS and Android sign-in flows, step table, pinned-ref debug log, unit fixtures, and documentation are covered.
Written by the indexing model from the issue text.
Description
Purpose
PR #2654 showed a real gain and the wrong home for it. The gain: a loop that asks a model one narrow question per step, "which on-screen element advances this goal", over the interactive snapshot, turns N agent turns into one call and cuts decide time per action from seconds to well under one second. The wrong home: a vendor HTTP client, a vendor API key, a pricing constant, and two always-listed CLI commands and MCP tools in core that refuse without the key.
agent-device core is the device side of an agent. Model handling belongs in agent-device/ai-sdk, where the host already brings its own model. That decision is independent of which vendor is behind the head.
The gain is not vendor-specific. On the PR's three sign-in screens plus a real iOS artifact, a small general model (gpt-5.4-nano, reasoning off) with a schema-forced choice among candidate refs chose correctly 16/16 at 620-730 ms median and about 200 input tokens per decision. A decision-only vendor can plug into the same seam from outside this repo.
Proposed shape
- Lives under
agent-device/ai-sdknext tocreateAgentDeviceTools, using the optionalaipeer. No new CLI command, MCP tool, registry entry, flag or env contract in core. runGoal({ client, goal, model | decide, inputs, maxSteps, minConfidence, onStep }):- candidates come from the structured interactive snapshot:
kind, name, value, enabled, ref (needs #2656); decidedefaults to AI SDKgenerateObjectagainst the host's model with a schema of{ target: enum(refs | none), done, blocked }; a host may pass its owndecide(state) => decision;- acts through
client.interactions.press/fillwithsettle: true, every mutation pinned to the snapshot'srefsGeneration(ADR 0014); - "did anything change" is read from the settle
diffandtailonSettleObservation, not from a second snapshot; - text is supplied, never generated; sensitive values go through the ADR 0017 channel (
recordAs,AD_VAR_*), not a new--inputor env family; - returns the step table (screen, decision, outcome, snapshot/decide/action ms) and a typed status:
done,blocked,escalated,max-steps.
- candidates come from the structured interactive snapshot:
- Out of scope: scroll, back, alert and gesture planning; multi-screen plans; text generation.
Completion conditions
examples/sdk/goal-loop.tsor theai-sdkexport drives a sign-in flow on an iOS simulator and on an Android emulator with a host-configured model, with the step table and the--debugrequest log showing~sNpinned refs attached.- Unit coverage over a fake device port and a fake
decide, with fixtures that carry the errors production emits. - Docs: one section in
website/docs/docs/ai-sdk.md.
Open decisions
- In-tree
ai-sdkexport, orexamples/sdk/goal-loop.tsfirst. - Confidence floor and unproductive-step limit defaults; whether
blockedis ever terminal on its own. - Whether
kind-based candidate filtering is enough or the loop needs ahittable/interactionBlockedgate.
Dependencies
Blocked by: #2656. Related: #2634 (fill on fields that normalize their input), ADR 0014, ADR 0017.
- Dominant language
- TypeScript
- Stars
- 4.8k
- Forks
- 315
- Avg merge
- 11h 17m
- Merged PRs (30d)
- 536
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from callstack/agent-device
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
callstack/agent-device#1869 ·
Maintainers usually reply within 1 day
-
needs-triage refactor
Difficulty 5/5 Over a week Newbie friendliness 25/100
callstack/agent-device#3116 · 4 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
callstack/agent-device#3106 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
callstack/agent-device#3105 ·
Maintainers usually reply within 1 day
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
callstack/agent-device#3104 ·
Maintainers usually reply within 1 day
All issues in callstack/agent-device
Similar issues
-
bug HemiStake
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
hemilabs/ui-monorepo#2413 ·
Maintainers usually reply within 1 day
-
component/ui framework/react kind/bug language/javascript
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
meshery/meshery#22216 · 3 comments ·
Maintainers usually reply within 1 day
-
type/bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
paperclipai/paperclip#14982 ·
Maintainers usually reply within 1 day
-
community first-timers-only good first issue hacktoberfest help wanted low hanging fruit up-for-grabs
Difficulty 1/5 Under an hour Newbie friendliness 95/100
lingdojo/kana-dojo#31515 · 1 comment · 5 reactions ·
Maintainers usually reply within 1 day