Planning to Build AI Agents? Here’s How to Get Started 

The team, timeline, and budget it actually takes, and the data on where most of these builds quietly stall

Summarize with AI:

Stay Updated:

TL;DR

  • A complete AI support agent build runs through seven phases. Knowledge base readiness is usually the longest phase and the most underestimated.
  • Projected 6–10 months for a production-grade, multi-category deployment, longer where compliance gates each expansion.
  • The build needs five to nine distinct roles. This guide breaks down what each one owns and what it typically costs.
  • Plan around resolution, not deflection, from day one.

Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, and it isn’t blaming the models. The reasons it cites are escalating costs, unclear business value, and inadequate risk controls, according to Anushree Verma, the Gartner analyst behind the prediction. MIT’s Project NANDA reached a similar place from a different angle: in its “State of AI in Business 2025” report, 95 percent of enterprise generative AI pilots showed no measurable effect on the P&L, based on 300 public deployments and more than 150 leadership interviews.

Neither number is really a verdict on whether AI agents work for support. Plenty of them do. The gap between the failures and the successes isn’t the technology; it is the intent behind building AI agents. And that intent starts with a robust plan. 

This is that plan: the phases, the team, the realistic timeline, the budget ranges, and the specific places where these projects tend to quietly stall. If you’ve been handed the job of building this for your organization, treat it as a working document to create your project plan to build AI agents for customer support. 

Table Of Contents 

  1. How do you set the objectives?
  2. How to define the scope of customer support automation?
  3. How to structure the team to build AI agents for support?
  4. How to create a phased project plan?
  5. What a realistic timeline and budget actually look like
  6. Where are these projects actually stalled?
  7. The Bigger Picture

How do you set the objectives? 

Before anyone touches a model or a vector database, the project needs two decisions on paper: what the agent is actually responsible for, and how you’ll know if it worked.

Scope it by ticket category, not by ambition. “Build an AI agent for support” is not a scope. “Automate password resets, order status, and return initiation, and hand off everything else” is a scope. Analysts studying why agentic AI projects get canceled point to all-in rollouts as a recurring technical failure mode, alongside building production systems on top of proof-of-concept architecture. 

Bounded scope with measurable outcomes is one of the few things that shows up consistently on the success side of that research. Pick three to five ticket categories where a correct answer is unambiguous, and expand from there once the first batch is holding up.

Decide now whether you’re measuring deflection or resolution, because they are not the same number. Deflection counts any ticket that never reached a human, including the ones where the customer got a wrong answer and gave up. Resolution counts only the tickets where the problem actually went away. 

Gartner research found that AI deflects more than 45 percent of customer queries but only around 14 percent of those interactions reach genuine self-service resolution. That roughly 31-point gap is where a project can look successful on a dashboard while customers quietly re-contact through another channel a day later. 

Build your success metric around true resolution from week one: deflected tickets, minus wrong or incomplete answers, minus any same-issue re-contact within 48 hours, divided by total AI-handled tickets. 

Baseline before you build. You cannot show ROI on a number you never measured. Pull your current average handle time, CSAT, and ticket volume by category for the segments you’re scoping in. This becomes the comparison point at every stakeholder review for the next year, and it’s far cheaper to pull now than to reconstruct later.

How to define the scope of customer support automation? 

Before assigning roles or hours, get specific about the target capability set. A support Agentic AI system that’s actually production-ready, not a demo, typically needs to do all of the following:

  • Retrieve accurate answers from a unified knowledge base spanning help docs, past tickets, community threads, and product documentation, grounded through retrieval-augmented generation rather than static scripts. This requires relevance — at scale. 
  • Resolve or deflect common Tier-1 and regular Tier-2 troubleshooting and telemetry requests without human involvement
  • Read and write to CRM and ITSM records, not just read them, so account context and case history stay current
  • Complete the ticket workflow with proprietary escalation rules that focus on resolution, not deflection numbers. 
  • Hand off to a human agent with full conversation context when it hits its limits. Additionally give critical control levers to the human in the loop for sustained support operations. 
  • Operate across the channels your customers actually use: web widget, in-app, email, and often voice
  • Improve over time through transcript review, evaluation runs, and knowledge base updates, rather than staying static after launch
  • In built audit mechanisms and key insight engine to give reviewers the complete picture. 

The idea behind building an AI agent for support is to take the entire workflow and build Agentic systems to complement it. A working Agentic system is much more than a simple FAQ bot; rather, it can be a CX engine for growth. 

Deloitte estimates a 60–80% reduction in cost per resolution once that threshold is hit. That’s the bar a project plan should be built against, not a chatbot that answers three FAQ categories.

How to structure the team to build AI agents for support? 

The roles below show up in every serious enterprise AI agent build, whether the people filling them are full-time hires, contractors, or existing staff wearing an extra hat. What tends to catch teams off guard is not any single role. It’s how many of these nine functions need to exist simultaneously, and how few organizations have all of them sitting idle.

RoleWhat they actually ownUS median annual wageClosest BLS occupation
Executive sponsorThe business case, cross-team conflict resolution, protecting the budget past the first hard quarterTime, not a line item, but the project stalls without itn/a
Support ops or product ownerScope, success metrics, the roadmap after launchUsually an internal, part-time role at firstn/a
Knowledge manager or KCS leadTaxonomy, article templates, metadata, the freshness review cadenceNo clean BLS match; the role most AI project budgets forget to includen/a
AI or LLM engineerRAG pipeline, model selection, orchestration$140,910Computer and information research scientists
Data engineerIngestion, chunking, embeddings, vector database upkeep$123,100Database administrators and architects
Integration engineerConnecting the agent to ticketing, CRM, and backend systems so it can take real actions, not just answer$131,450Software developers, QA analysts, and testers
Conversation or prompt designerEscalation scripting, tone, edge case handlingOften folded into the AI engineer role on smaller teamsn/a
Security and compliance reviewerPII handling, SOC 2, GDPR, audit trail design$124,910Information security analysts
QA and evaluation leadEval pipelines, red-teaming, accuracy benchmarks$131,450Software developers, QA analysts, and testers

Wage data: US Bureau of Labor Statistics, Occupational Outlook Handbook, May 2024 median annual wages. For reference, the median across all US occupations is $49,500.

What does this team cost?

Sum the four technical roles in the table above, plus a second software developer or QA hire for AI-specific evaluation work, and base salaries for a core four-to-five-person build team land around $520,000 to $650,000 a year before benefits. The Bureau’s Employer Costs for Employee Compensation survey puts benefits at roughly 30 percent of total private-sector compensation, which, using national median pay, would gross out to a fully loaded cost of roughly $740,000 to $930,000 a year. Tech-sector pay typically runs higher still: the Bureau’s own industry breakout puts median pay for computer and information research scientists at $237,990 at software publishers specifically, well above the all-industry median, which is pulled down by lower-paying sectors like government and academia. None of this includes the knowledge manager, the executive sponsor, and support ops time layered on top.

That cost pressure isn’t easing, either. The Bureau projects 20 percent employment growth for computer and information research scientists through 2034, more than six times the average across all occupations, and Deloitte’s 2026 State of AI in the Enterprise survey of more than 3,200 global leaders found the AI skills gap is now the single biggest barrier to scaling AI, ahead of budget or technology maturity. Whatever this team costs on paper, expect to compete for it.

What does the total cost of ownership of AI Agents look like on the ledger?

Learn More

How to create a phased project plan? 

Phases Of The Project Plan

Seven phases get you from a scoping document to a support agent that’s still trustworthy a year after launch. Several run in parallel rather than strictly in sequence, and the last one never really ends.

Phase 1: Discovery and scoping 

Lock the ticket categories in scope, the success metrics, the baseline numbers, and who has authority to expand or pull back scope once the project is underway. Get explicit sign-off from your executive sponsor on what “working” looks like before phase two starts. This is also where you decide the scope of the Agent’s capabilities, escalation rules, your human + AI workflows, and more. Moreover, the agent’s risk tolerance: which actions it can take autonomously, and which ones always route to a human regardless of confidence score, is also planned here. 

Phase 2: Knowledge foundation and data readiness 

This is where most in-house builds quietly lose months, and it’s worth spending real time on. Separate 2026 research on self-service found that 84 percent of customers try to solve issues on their own before contacting support, and 91 percent say they’d use a knowledge base if it were accurate and relevant. Only about one in five companies rate their own knowledge base as very accurate. That gap between demand and content quality is exactly where an AI agent’s resolution rate quietly caps out.

How to get the Knowledge Foundation AI Agent Ready?

Read Our Whitepaper To Learn About A Five-Pillar Fix

Download Here

The practical work here is unglamorous: audit and consolidate every source the agent might need to draw from (help center, community forums, internal wikis, past ticket resolutions), retire duplicate or conflicting articles, restructure content into single-topic, templated articles that chunk cleanly for retrieval (400 to 600 tokens per chunk is a reasonable target), and build a real metadata schema, meaning topic, audience, effective date, owner, and review cadence on every article, not just a title and a body. Assign an actual owner to each content domain, and tie review cycles to product releases instead of a calendar reminder, because that’s what keeps content from getting stale between reviews.

If this phase isn’t planned and budgeted up front, it tends to appear later as a surprise. Data engineering prerequisites commonly add 20 to 35 percent to total program cost when teams discover mid-build that their content and data environment wasn’t actually ready. 

Phase 3: Architecture, models, and integration

This is where the RAG pipeline, model selection, and orchestration logic get built: retrieval, embeddings, a vector database, and, a multi-agent orchestration layer for handling multi-step requests. Building the architecture is simpler on the surface, but getting AI agents to work successfully, requires better planning. 

For instance, building a RAG layer is a defined scope, but building a RAG relevance engine to drive sustained precision at scale requires skilled engineering. Similarly, equally hard is the integration, i.e., connecting the agent to your ticketing system, CRM, and backend systems so it can actually take actions, requiring engineering prudence. Significantly, every additional system connection adds a new surface for auth token expiry, rate limits, and schema changes — to quietly break something months after launch.

Phase 4: Governance, guardrails, and human-in-the-loop design 

Escalation logic, audit trails, and compliance review belong in the project plan as their own phase, not as a checklist item inside engineering. Define explicit criteria for when the agent hands off: low confidence, high emotional sentiment, an irreversible action, or a category you’ve deliberately excluded, like billing disputes or anything touching a pending legal claim. Handoffs matter more than most teams expect. 

With Agentic systems capable of executing actions on their own, governance is table stakes. Ideally, plan a layered governance framework, deeply ingrained in the agentic system.

If the agent touches PII, financial data, or anything in a regulated category, budget for it explicitly: security and compliance review typically adds $3,000 to $10,000 and two to four weeks per cycle, and it recurs every time you expand scope, not just once before launch.

Phase 5: Build, tune, and evaluate (overlapping phases 3 and 4)

Getting an agent from roughly 80 percent accuracy to 95 percent typically takes three to five times longer than the initial build, according to 2026 cost benchmarking from bmdpat. That gap is prompt iteration, eval pipeline construction (test datasets, automated scoring, regression detection), and red-teaming, and it’s the phase most likely to get compressed under launch-date pressure. Set a retrieval quality target before you start tuning; a reasonable production floor is Precision@5 above 0.8, meaning at least four of the top five retrieved chunks for a given query are genuinely relevant.

Run shadow mode against real ticket volume before anything goes live: let the agent draft answers a human agent reviews and sends, without the customer ever seeing an unreviewed response. This is where you catch the failure modes that a clean demo never surfaces.

Phase 6: Pilot and phased rollout 

Launch in agent-assist mode first, where the AI drafts and a human approves, before granting full autonomy on any category. Expand category by category rather than switching everything on at once; the “all-in rollout” pattern shows up repeatedly in postmortems of canceled agentic AI projects. Train your support team explicitly on the new workflow, including how to read the context the agent hands off and when to override it. Set CSAT guardrails that auto-escalate below a sentiment or confidence threshold, and update your SOPs so the new process is documented, not tribal knowledge held by whoever built it.

Phase 7: Post-launch operations (ongoing)

This is the phase in which most project plans are under budget, because most project plans are built like the work ends at launch. It doesn’t. Ongoing prompt tuning typically runs 10 to 20 hours a month. Knowledge base and RAG pipeline maintenance commonly consumes 20 to 30 percent of an engineer’s sprint capacity on an ongoing basis, which works out to roughly $40,000 to $75,000 a year in pure upkeep per senior engineer at typical fully loaded salaries, per 2026 research from Digital Applied. Add model updates as providers ship new versions, connector maintenance as backend systems change their schemas, and content review cycles tied to every product release. None of this is optional if you want the resolution rate you launched with to still hold a year later. Budget for it as a permanent line item, not a contingency.

What a realistic timeline and budget actually look like

PhaseWhat drives the cost or the delay
1. Discovery and scopingStakeholder alignment, defining resolution as the metric that matters
2. Knowledge foundation and data readinessKB audit, consolidation, chunking, metadata, ownership; usually the longest phase
3. Architecture, models, and integrationRAG pipelines, Connector engineering, vector database setup, model selection
4. Governance and human-in-the-loop designCompliance review, escalation logic, audit trails
5. Build, tune, and evaluatePrompt iteration, eval pipeline, accuracy testing
6. Pilot and phased rolloutAgent-assist mode first, category-by-category expansion, team training
7. Post-launch operationsPrompt tuning, KB governance, model updates, connector maintenance

Add it up and a genuinely production-grade, action-taking support agent across several ticket categories typically takes 6 to 10 months to reach broad go-live, longer in regulated industries where compliance review gates each expansion. That lines up with 2026 estimates across industry and with MIT NANDA’s finding that mid-market organizations average roughly 90 days from pilot to full implementation, while large enterprises average nine months or longer.

On cost, once you add the team from the table above, the data engineering work in phase two, integration engineering in phase three, and compliance review in phase four, first-year all-in cost for an in-house build commonly lands in the mid six to seven figures for a mid-sized enterprise deployment, before counting the ongoing maintenance load in phase seven. That range will move a lot based on your industry, ticket volume, and how many systems the agent needs to touch, but it’s rarely close to the cost of the model access itself, which is usually the smallest line on the budget.

Where are these projects actually stalled?

Gartner’s three cited causes for agentic AI project cancellations (escalating costs, unclear business value, and inadequate risk controls) map almost exactly onto the phases above, and a few specific patterns show up repeatedly in the postmortems:

Scope creep from a bounded pilot into “automate everything” before the first version has proven itself. Treating the knowledge base as a one-time cleanup project instead of infrastructure with a permanent owner and a review cadence. Measuring deflection instead of resolution, so the dashboard stays green while re-contacts and CSAT quietly erode underneath it. Running compliance and governance as a launch-day gate rather than a continuous discipline that has to keep pace with every scope expansion. And budgeting for the build but not for the 20 to 30 percent of ongoing engineering capacity that phase seven actually consumes.

There’s a broader pattern underneath all of that. MIT’s NANDA research found that organizations partnering with specialized vendors succeed roughly 67 percent of the time, while internal builds succeed at about a third of that rate. The report’s own framing is direct: crossing what it calls the GenAI Divide requires a shift from building toward buying and from static tools toward systems that already know how to learn and adapt. That’s not an argument against ever building. It’s a data point worth weighing honestly against your own team’s capacity before you commit a year and seven figures to it.

Want To Learn More Before You Take The Call?

Get an In-depth Decision Framework Whitepaper

Download Now

The Bigger Picture 

Agentic AI is here and is changing the entire enterprise landscape. However, beyond the promise, it is still a new technology, very different than what organizations are accustomed to witnessing. While business transformations were earlier — adjusting to rising expectations, SaaS tools or newer cybersecurity norms. Agentic AI necessitates a complete overhaul — requiring enterprises to rewire and adjust — on multiple dimensions.  

There lie the challenges — on one end of the spectrum, the change is taking place on multiple fronts; on the other, teams are dealing with a new technology — one that throws in multiple variables — adding to the complexities. 

The project plan lays out a roadmap to build AI Agents. It lays out the variables in play when a team sets out to build an AI agent. That said, it’s about which layers are worth your team owning, and which ones are infrastructure that has already been laid out. 

Before you lock headcount and budget for a from-scratch build, it’s worth running your own ticket volume, team capacity, and compliance requirements against the ranges above and comparing that honestly to what a support-specific platform already handles out of the box. That comparison is really the whole decision.

FAQ

1. How long does it take to build an AI agent for customer support? 

A production-grade, action-taking support agent typically takes 6 to 10 months to reach broad go-live across seven phases: scoping, knowledge readiness, architecture, governance, build and tuning, pilot, and ongoing operations. Regulated industries usually run longer due to compliance review cycles.

2. What’s the difference between deflection and resolution in AI support metrics? 

Deflection counts any ticket that never reached a human agent, including ones where the customer got a wrong answer and gave up. Resolution counts only tickets where the problem was actually solved. Research cited by Gartner found AI deflects over 45% of queries but only about 14% reach genuine resolution, a roughly 31-point gap.

3. What team do you need to build an AI agent for support? 

Nine functions typically need to be covered: executive sponsor, support ops owner, knowledge manager, AI or LLM engineer, data engineer, integration engineer, conversation designer, security and compliance reviewer, and QA or evaluation lead. 

Begin your AI Transformation

ai-discover

Discover More Resources

Browse Library
ai-time

Experience SearchUnify Solutions

Schedule a Demo
ai-connect

Have any questions?