[Feedback]: Model during a plan is routinely partially completing tasks as a subagent
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Active
- Tech stack
- vscode
- Domain
- ai
Research direction
The report concerns MAI-Code-1-Flash used through the VS Code Copilot surface, with no repository files, tests, or entry points identified. First reproduce the model's behavior using a plan and backend-writing task, then compare its completion claims with the implemented feature; done requires determining whether the model consistently reports partially completed criteria as complete.
Written by the indexing model from the issue text.
Description
Model
MAI-Code-1-Flash
Feedback
I figured the best way to use this model would have been to write a plan with a frontier and use MAI-Code for writing backend.
This seems to have partially succeeded consistently. I have observed the MAI model out right lying saying criteria has been completed but when checked for review the feature is partially implemented.
Copilot surface / environment
VS Code
- Dominant language
- No language data
- Stars
- 44
- Forks
- 2
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/MAI-Code
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
-
[Feedback]: When told to open a pull request in Azure DevOps, the model does not link the work item Open
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
All issues in microsoft/MAI-Code
Similar issues
-
triage/confirmed
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
agentscope-ai/agentscope#2775 ·
-
area/sessions comp/cron comp/gateway P2 sweeper:risk-message-delivery sweeper:risk-session-state type/bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NousResearch/hermes-agent#118863 ·
-
automation missing-model model-sync provider:pioneer
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
anomalyco/models.dev#7701 ·
-
external-plugin ready-for-review
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
github/awesome-copilot#3586 · 2 comments ·
-
bug-unconfirmed
Difficulty 2/5 1-3 hours Newbie friendliness 76/100