Feature Request: Make GitHub Copilot Coding Agent as Reliable, Autonomous, and Effective as Its Full Model Potential Allows #209973
Replies: 3 comments
|
💬 Your Product Feedback Has Been Submitted 🎉 Thank you for taking the time to share your insights with us! Your feedback is invaluable as we build a better GitHub experience for all our users. Here's what you can expect moving forward ⏩
Where to look to see what's shipping 👀
What you can do in the meantime 💻
As a member of the GitHub community, your participation is essential. While we can't promise that every suggestion will be implemented, we want to emphasize that your feedback is instrumental in guiding our decisions and priorities. Thank you once again for your contribution to making GitHub even better! We're grateful for your ongoing support and collaboration in shaping the future of our platform. ⭐ |
|
Strongly agree with the loop detection and verification-before-completion points especially, those are the two that actually cost me the most time in practice. The repetitive tool call thing is real and it's not just annoying, it burns through the task budget before the agent gets to the part that actually matters. I've had sessions where it re-reads the same three files four or five times across a single task, each time as if starting fresh, then runs out of room to actually finish the edit. A simple "have I already looked at this exact file in this session" check would go a long way even without anything fancier. The verification point (G in your list) is the one I'd push hardest on though. The failure mode I hit most isn't the agent getting stuck, it's the agent confidently saying a task is done when it only edited the file and never actually ran anything to confirm the change works. If there's a test suite or a build step available in the repo, completion should mean "I ran this and here's the result," not "I wrote code that should do this." Even a lightweight version, where the agent states explicitly whether it verified its own change or is just asserting it should work, would change how much I trust the "done" signal without requiring the full task-integrity architecture you're describing. One thing I'd add to your list under context loss: when an agent resumes after compaction, it's often not just that it forgets what it already tried, it sometimes reopens an approach that was already explicitly rejected earlier in the conversation. That's worse than just losing progress, since it can quietly undo a decision the user already corrected it on. Durable state needs to track not just "completed steps" but "things the user told me not to do," and that list should survive compaction with higher priority than almost anything else, since silently re-doing a reverted mistake is more damaging than just being slow. Well written proposal overall, hope it gets real traction instead of just the auto-reply. |
|
The Great Wall Of Instructions (GWOI) that copilot prefers to assemble is partially responsible. Keep in mind that Copilot is just a MCP. ChatYou probably understand that "chat" is a stateless LLM (model) using a context window to pretend to be stateful. "Chat with Copilot", code-review, and "coding agent" are all cloud-agent in different containers/scopes:
That makes it look intelligent/alive/capable but:
It tries every possible combination:
The one that remains is, at best, "probably right."
Currently, LLMs "brute force" their way through programming/everything, and appear semi-intelligent, some of the time. "Chat with Copilot" has psuedo-sudo scope:
When you tell it "read this", it is compelled to summarize it for you (by default):
Anything involving an LLM is a "chatbot with tools," shuffling context around.
This dictates how you must instruct/prompt them. Which is what MCP and native skills are for!
Instead of searching through conditional prose (GWOI) to "find" instructions, native skills inject only relevant instructions directly into the context window. BreadcrumbsEach PR is a new chat:
Turns:
You have to instruct agents to communicate with each other across the turn boundary, or they won't. Instead of trying to instruct away LLM mistakes (not possible)You must create minimal governance, in case [github docs wrong, copilot api error, etc.].
Remember that agents know nothing on turn 1 except their
The agent learns to build one JIT to satisfy the
A closure matrix is not needed for every session, every time:
Every domain has a manual to read
Native skillCreate a native skill ( |
Uh oh!
There was an error while loading. Please reload this page.
🏷️ Discussion Type
Bug
💬 Feature/Topic Area
Copilot in GitHub
Body
1. Summary
I would like GitHub to improve the overall reliability, autonomy, speed, and task-completion ability of GitHub Copilot coding agent.
Powerful AI models are becoming increasingly capable at software development. However, the practical results users receive can differ significantly between coding-agent products. Even when a product offers access to a highly capable model, the surrounding agent workflow can prevent that model from delivering its full potential.
GitHub Copilot coding agent should reliably understand a task, plan the work, modify the correct files, execute tools efficiently, recover from errors, preserve context, run appropriate checks, and complete the requested task with minimal unnecessary human intervention.
The goal is not simply to provide access to a more powerful model. The goal is to make the entire coding-agent experience powerful, dependable, and consistent.
2. Problems users experience
A. Repetitive tool calls and unproductive loops
The agent sometimes searches the same files repeatedly, runs similar commands, revisits previously explored approaches, or continues a sequence of actions without meaningful progress. It should recognize when an action has already been attempted and determine whether repeating it is justified. If an approach is not working, the agent should change strategy instead of continuing indefinitely.
B. Slow execution, even for simple tasks
Some straightforward changes can take unnecessarily long because of excessive planning, repeated exploration, redundant tool calls, or unnecessary review cycles. The agent should adapt its effort to the complexity of the task. A small code change should not require the same execution process as a large architectural change.
C. Incomplete task execution
The agent may modify some files but leave other explicit requirements unfinished. It can focus on one part of a task while overlooking related changes, edge cases, documentation, or tests. It should maintain a checklist of requirements and reconcile that checklist against the actual repository changes before declaring the task complete.
D. Weak debugging and self-correction
When the agent introduces an error, encounters a failed command, or breaks a test, it should use the available evidence to identify the root cause and attempt a targeted correction. It should not repeatedly apply superficial fixes, ignore error output, or continue as though the task succeeded when the problem remains unresolved.
E. Context loss during long tasks
As a session grows, important requirements, previous decisions, investigated approaches, and remaining work can be lost or inconsistently recalled. Context compaction should preserve the information needed to continue the task correctly. The agent should maintain durable task state instead of relying entirely on its latest conversational context.
F. Excessive dependence on human intervention
Users should not have to repeatedly remind the agent of the original request, explain what it already attempted, redirect it away from loops, or prompt it to finish work that was part of the original task. The agent should be able to carry out delegated coding tasks autonomously within the permissions, safeguards, and available tools provided by the product.
G. Insufficient verification and misleading completion signals
The agent should not equate generating code or editing files with successfully completing a task. Before claiming completion, it should review the requested requirements, inspect the changes, run relevant tests or checks when available, and clearly report anything that remains incomplete or unverified.
H. Inconsistent performance across coding-agent experiences
Users may experience different levels of speed, autonomy, context retention, and correctness across coding-agent products and environments. GitHub should evaluate and improve the full Copilot coding-agent workflow so that users receive consistent, dependable results across supported Copilot experiences.
3. The larger issue: Powerful models do not automatically create powerful coding agents
One of the most important concerns is the difference users can observe between products such as Codex, Claude Code, and GitHub Copilot.
A capable model can still produce a disappointing coding experience if the surrounding agent system does not use that capability effectively.
The overall result depends on more than the underlying model. It can also depend on:
These are potential contributing factors, not a claim that the products use identical model configurations or that one particular component is definitively responsible.
GitHub should evaluate the complete coding-agent system, rather than treating access to a powerful model as sufficient evidence of a powerful coding experience.
4. Proposed solution: A progress-aware task integrity controller
I propose introducing a task-integrity layer into GitHub Copilot coding agent.
This component would track task progress, preserve requirements, detect ineffective behavior, support recovery, and verify outcomes.
A. Durable task state
Maintain a structured, continuously updated record containing original requirements and constraints; decisions and assumptions; planned and completed steps; files inspected and modified; commands and checks executed; errors and recovery attempts; remaining requirements and blockers. This state should survive context compaction and session continuation.
B. Progress and loop detection
Detect repeated or equivalent tool calls, repeated searches, recurring failures, and sequences that produce no meaningful progress. When a stall is detected, assess what has already been learned, avoid unnecessary repetition, and select a more useful next action. Repeated calls should remain possible when justified by new evidence.
C. Adaptive planning and execution
Use a task-appropriate execution strategy. For simple tasks, minimize unnecessary exploration and workflow overhead. For complex tasks, use deeper planning, appropriate decomposition, and stronger validation. Reconsider the plan when new information contradicts assumptions or progress stalls.
D. Evidence-based debugging and recovery
Use actual command output, test failures, diagnostics, and repository state to identify problems. Determine what failed, identify the likely root cause, choose a targeted correction, rerun the relevant check, and preserve lessons learned. If recovery is impossible because of permissions, unavailable dependencies, or another genuine blocker, explain it clearly.
E. Requirement-to-result reconciliation
Maintain a traceable relationship between each user requirement and the resulting implementation or verification. Before finishing, check which requirements are complete, tested, unverified, or blocked.
F. Verification before completion
When appropriate and available, run relevant tests, builds, lint checks, type checks, and other validation. Inspect final changes for omissions and unintended modifications. Report the checks performed and actual results. Distinguish code that was written, code that was checked, and behavior that has genuinely been verified.
G. Reliable context compaction
Before compressing context, preserve the original objective, constraints, decisions, current state, failed approaches, remaining tasks, and next steps. After compaction or resumption, continue without unnecessarily repeating completed work or forgetting requirements.
H. Reduced unnecessary interruptions
Continue autonomously when there is sufficient information and permission. Request user input when a meaningful decision, missing information, security boundary, or genuine blocker requires it—not because the execution process lost track of the task.
5. Expected end-to-end workflow
The process should be iterative and adaptive rather than rigid. The agent should return to earlier steps when evidence shows further work is necessary.
6. Acceptance criteria
7. How GitHub should measure improvement
Evaluate representative real-world repository tasks using end-to-end task completion rate; correctness and quality; time to verified solution; redundant tool calls; recovery rate after failed commands or tests; requirement retention across long sessions; regressions; unnecessary user interventions; and verification accuracy. Where possible, compare versions under controlled conditions.
8. Why this matters
A coding agent is more than a model connected to a repository. Its planning, memory, tools, recovery behavior, and verification determine whether it can reliably deliver useful software changes.
Developers want to delegate a task and trust that the agent will work through it responsibly, recover from ordinary failures, and report the truth about the result.
The objective is not unlimited autonomy or eliminating human oversight. It is reducing avoidable failures and wasted effort while preserving user control, repository safety, and transparency.
9. Final request
Please consider prioritizing improvements to the underlying GitHub Copilot coding-agent architecture and workflow, including progress awareness, durable task state, adaptive execution, evidence-based recovery, and completion verification.
The full potential of capable AI models should translate into a dependable coding experience—not just impressive individual responses.
The goal: an agent that understands the task, makes meaningful progress, corrects its mistakes, finishes the requested work, verifies what it can, and clearly explains anything it could not complete.
All reactions