Choosing a Partner

How to Vet an AI-Assisted Development Agency: 9 Questions to Ask in 2026

7 min read | Business

Every development agency now says it “builds with AI.” In 2026 that sentence carries almost no information. It could mean a disciplined engineering process with review gates and test coverage, or it could mean a junior developer pasting prompts into a terminal and shipping whatever compiles.

The gap between those two is where your money goes. The industry data is not subtle about it: research summarised by Veracode found that 45% of AI-generated code samples failed security benchmarks across OWASP Top-10 categories, and a CodeRabbit analysis reported AI-written code carrying 1.7× more issues than human-written code, including 75% more logic errors. Meanwhile Stack Overflow’s survey shows developer trust in AI output falling from around 40% to 29% in a single year.

None of that means avoid AI-assisted agencies. We are one, and we would not go back. It means the vetting questions have changed. Here are the nine we would ask if we were on your side of the table.

45%
of AI-generated code samples failed OWASP security benchmarks (Veracode)
1.7×
more issues in AI-generated code vs. human-written (CodeRabbit)
29%
of developers now trust AI code accuracy, down from ~40% (Stack Overflow)
90%
of developers use at least one AI tool at work (JetBrains, Jan 2026)

1. The Nine Questions

1
“Which tools, and at which stage of the work?”
You want a specific answer: Claude Code for refactors and test generation, a human for architecture, review before merge. A vague “we use AI throughout” means nobody has drawn the line.
2
“Who reviews AI-written code before it merges?”
The only good answer names a human and a gate. SonarSource found 96% of developers do not fully trust AI output — yet only 48% always review it. Ask which half they are in.
3
“What is your test coverage policy, and is it enforced in CI?”
AI makes writing tests nearly free, which removes the last excuse. If coverage is aspirational rather than gated in the pipeline, it is not a policy.
4
“Does our code or data enter a model provider’s training set?”
Ask for the specific plan tier and the data-retention terms in writing. Enterprise tiers of the major coding tools exclude customer content from training — consumer tiers historically have not.
5
“Who owns the output, and what is your IP indemnity?”
You want unambiguous assignment of all deliverables to you, plus a statement on how they handle third-party licence contamination in generated code.
6
“Show me a security scan of something you shipped.”
SAST output, dependency audit, or a pen-test summary from a recent project (redacted is fine). An agency that runs these routinely can produce one in a day.
7
“How do you handle a large legacy codebase?”
A METR study found experienced developers were 19% slower with AI assistance on mature codebases over a million lines. An agency that claims uniform speedups everywhere has not worked on one.
8
“What does handover include?”
Repo access, architecture docs, runbook, environment configs, and the project’s AI context files (the CLAUDE.md or equivalent). That last one is new, and it is the difference between a maintainable codebase and a haunted one.
9
“Which engineer will actually be on my project?”
Unchanged from the pre-AI era, and more important than ever. AI raises the floor of what a junior can produce and the ceiling of what a senior can. You are paying for the judgment, not the typing.

2. What Good and Bad Answers Sound Like

Green flags
On toolingNames tools per task
On reviewHuman gate, named
On limits“Here’s where it fails”
On dataPlan tier in writing
On speedFaster on some phases
Red flags
On tooling“AI-powered” brochure
On review“The AI checks itself”
On limitsNo downsides given
On data“It’s fine, don’t worry”
On speed“10× on everything”

3. The Pricing Question Nobody Asks Out Loud

If AI genuinely makes an agency faster, should you be paying less? Sometimes. But be careful what you are optimising for.

AI compresses the parts of the work that were always mechanical: boilerplate, migrations, test scaffolding, documentation, multi-file refactors. It does not compress discovery, architecture, security review, or the judgment calls about what to build. If an agency’s price dropped proportionally to its typing speed, it would be cutting the wrong side of the ledger.

The honest version is that AI shifts the shape of an estimate more than its total: less time on build, proportionally more on discovery and verification. Our breakdown of what a web app costs in 2026 walks through where the money actually sits.

The single best test

Ask the agency to describe a time AI-assisted work went wrong on a real project, and what they changed afterwards. Anyone shipping production software with these tools has that story. Anyone who does not have it has not shipped.

4. Put It in the Brief, Not the Contract Negotiation

The cheapest place to settle all of this is your initial brief. Add three lines: your data-sensitivity level, whether AI-assisted development is acceptable to your stakeholders, and who on your side reviews code. Our guide to briefing an app development agency covers the rest of the document, and if you want the practitioner’s view of where these tools genuinely earn their keep, we wrote up the lessons from building production apps with Claude Code — including the parts that still need a human.


Ask us all nine.

We will answer every question on this list in writing, including the uncomfortable ones about where AI-assisted development has failed us and what we changed as a result.

Start the Conversation →

Engineering Insights

Latest from Syntaxa Studio.

Loading latest posts