Where agentic AI ships work. Choose the workflow first.
Concrete workflows where Codex, Claude Code, MCP and related tools earn their place: shape, gate and outcome.
Use-case map
Choose the workflow before the tool.
Large refactor with full test coverage
A multi-file refactor where tests are the contract and the agent does the mechanical work. Shape: Read-only review, agreed scope, working branch, test gate, diff reviewed in passes. Outcome: Large refactors done as agent work plus review, with tests as the safety net.
Repository question-and-answer for new joiners
New joiners ask repo questions with grounded references before taking senior time. Shape: Codex or Claude Code gets read-only repo access plus a short guide for good questions and escalation. Outcome: New joiners reach a first useful pull request sooner, and senior time goes on the questions the agent cannot answer well.
Deterministic migration across the codebase
A version bump, API rename or dependency swap where each file follows the same rule. Shape: Rule and example first. Agent applies it file by file with tests and a small checklist. Outcome: Migrations that would stall in a backlog get done, and the written rule is kept for the next one.
Codebase-specific eval suite for agent use
A small set of repeatable tasks with known-good outputs for every model or workflow change. Shape: Three to five tasks, written prompts, known-good outputs and an automated comparator in the repo. Outcome: Model and prompt changes stop being a vibes call. The team has a short list of tasks it knows the workflow should pass.
MCP-mediated internal tool
An internal search, deploy command or private dataset exposed through scoped MCP tools. Shape: API design, server skeleton, identity scopes and call-pattern logs. The agent uses it on real tasks. Outcome: Internal knowledge becomes reachable in agent flows without sprawl. The server is reviewed monthly against the call log.
Code review triage for the senior reviewer
A senior reviewer uses an agent to triage long PRs without delegating approval. Shape: Agent reads the diff and returns a triage note. The reviewer reads the note, then the diff. Outcome: Long PRs get reviewed in less time without losing the senior judgement. The triage notes themselves become a corpus the team learns from.
Why each use case has a check
Almost-right output is the most common problem developers report. Every use case here ends in a test, diff or review.
- Almost right, but not quite66%
- Debugging AI code takes longer45.2%
- Less confident in own problem-solving20%
- Hard to understand how the code works16.3%
Agents are early
About three in ten developers use AI agents at work, and more than eight in ten worry about accuracy and data security.
- Use agents daily14.1%
- Use agents weekly9%
- Use agents monthly or less7.8%
- Plan to17.4%
- Autocomplete only13.8%
- No plans37.9%
- Accuracy of what agents produce86.9%
- Security and privacy of data81.4%
Need a broader route?
Review the department workflows and the AI readiness score.
Keep exploring
