← Back to blog

Buying guide ยท AI Agency

How to Choose an AI Agency: 12 Questions That Expose Demo-Ware

In short

  • Judge an agency on systems that have been live for months, not on the demo. Ask what broke after launch and how it was fixed.
  • Insist on two numbers: build cost and cost to run per month. Any vendor who only quotes one is hiding the other.
  • IP ownership, handover documentation and an exit path should be written into the contract before work starts.
  • In-house makes sense when AI is your product. For workflow and voice automation, a senior studio plus your own operators is usually faster and cheaper.

The AI services market in 2026 is crowded with firms that added "AI" to a web-development menu eighteen months ago. Most can build something impressive in a two-week sprint. Far fewer have taken an agent into production, watched it fail on real callers, and kept it running. The questions below are designed to tell those two groups apart before you have signed anything.

Questions about track record

1. Show me one system that has been live for more than six months. What broke after launch? Every production AI system breaks after launch. Data changes, an API changes, callers say things nobody scripted. A team that has been through this will answer with specifics and a story about what they changed. A team that has only shipped demos will say nothing broke.

2. How do you test something that does not behave the same way twice? Language models are probabilistic. Ask how the agency evaluates a conversation flow before release: scripted scenario suites, transcript review, automated scoring against criteria, regression runs when a prompt changes. If the answer is "we test it manually," you will be the test suite.

3. Can I speak to two references, and can I ask them what went wrong? One reference is a friend. Two references who will tell you about the hard parts are evidence.

Questions about scope and delivery

4. What is the smallest useful version, and how fast can I see it working on my data? The right first project is narrow: one call type, one workflow, one integration. Be wary of any proposal that opens with a multi-agent platform. Our own model is a 3-day discovery sprint that ends with a proof of concept on your real inputs, because a working artifact on day three settles more arguments than a 40-page proposal.

5. What does the timeline look like, and what has to be true on your side for it to hold? A credible answer names dependencies: API access, sample data, a decision-maker who can approve conversation copy, a test group of real users. A 6-week build is realistic when those are in place and fictional when they are not. We have written elsewhere about why we ship in six weeks and what the client side of that commitment involves.

6. Who exactly will work on this, and will they still be here in month four? Named senior engineers, not a rotating bench. Ask whether the people in the sales call are the people who will build.

Questions about money

7. What will it cost to build, and what will it cost to run every month? These are different numbers and you need both. Run cost includes model inference, voice, telephony, hosting, monitoring and whoever maintains the integrations. A vendor who quotes only build cost is deferring the conversation you should be having now. Our guide to AI voice agent pricing walks through how to model both.

8. What happens to the price if the underlying model or voice provider changes theirs? Model pricing moves quarterly. You want to know whether the agency passes costs through transparently or has margin built into a bundled rate.

Questions about ownership and exit

9. Who owns the code, the prompts, the workflows, the fine-tuned models and the call recordings? The answer should be you, without qualification. Vague answers on IP are the single most reliable red flag in this market. Full IP transfer is written into every Claudeter engagement for exactly this reason.

10. What does handover look like if we part ways? A serious answer includes architecture documentation, prompt version history, credentials handover, and an overlap period where the outgoing team answers questions. A vendor who makes leaving expensive is counting on you staying reluctantly.

Questions about safety and compliance

11. How does the agent behave when it is unsure, and how does a human take over? Ask to see the escalation path in the demo: what the agent says, what context the human receives, how long the handoff takes. An agent that bluffs instead of escalating is a liability in any regulated conversation.

12. Which regulations apply to my use case, and how are they built into the system rather than bolted on? HIPAA and a signed BAA for US healthcare. PDPL for UAE customer data. DPDP and TRAI rules for outbound calling in India. TCPA consent and disclosure for AI voices in the US. A capable agency will name the relevant rules unprompted and show where in the architecture they are enforced. Our piece on what developers get wrong about HIPAA-compliant AI is a useful preview of the depth you should expect.

Build in-house, buy a platform, or hire a studio?

The honest framework:

  • Build in-house when AI is your product and the model behaviour is your competitive edge. You need the talent anyway, and outsourcing the core is a strategic error.
  • Buy a platform when the use case is a single, generic call type with no integration into your systems. Template agents from self-serve platforms are fast and cheap for that.
  • Hire a studio when the use case is a business workflow (voice, back office, revenue cycle) that must connect to your systems of record, meet regulatory requirements, and go live in weeks rather than quarters. The math usually favours a small senior team that has done it twenty times over a first-time internal build, provided you own everything at the end.

Whichever path you take, insist on the same discipline: a narrow first agent, shadow-mode rollout before autonomy, a maintenance plan agreed before launch, and a written exit. Those four habits predict success better than any portfolio.

Red flags, in one list

  • No named case studies and no references who will discuss failures
  • Promises of full autonomy on day one, with no shadow-mode phase
  • A single price with no separation between build and run
  • Hedged answers on IP ownership
  • Timelines that do not name your dependencies
  • No answer to "what breaks after launch"

If a vendor clears all twelve questions, the demo quality matters a lot less. If they fail three, the demo quality is the reason you are talking to them, and that is the problem.

Frequently asked questions

What should I look for when hiring an AI development agency?

Evidence of systems running in production for months, a clear testing methodology for non-deterministic behaviour, separate build and run pricing, unambiguous IP ownership, a documented handover process, and unprompted awareness of the regulations that apply to your use case.

Is it better to build AI agents in-house or hire an agency?

Build in-house when AI is your core product. For workflow, voice and back-office automation that must integrate with existing systems and go live quickly, a small senior studio is usually faster and less expensive, provided the contract gives you full ownership of code, prompts and data.

How long should an AI agent project take?

A narrow first agent with one or two integrations is realistic in about six weeks when the client provides API access, sample data and a decision-maker for conversation copy. Multi-agent platforms promised in the same window should be treated with suspicion.

What are the biggest red flags with AI agencies?

Vague IP terms, no references willing to discuss problems, promises of full autonomy without a shadow-mode phase, a single bundled price that hides run cost, and no concrete answer about what has broken after launch on previous projects.

Want AI That Ships?

Talk to our team about building AI that ships in 6 weeks. Full IP transfer, no lock-in, senior team.

Book a Free Discovery Call