





























Key Insights
- Scope customer service automation by request type, not by channel, starting with a tagged sample of 200 to 500 recent tickets.
- A request type must clear four tests before an agent resolves it alone: volume, rule clarity, consequence, and data sensitivity.
- Status checks, document requests, address changes, reschedules, work orders, and policy questions typically clear all four.
- Disputes, complaints, clinical or legal questions, out-of-policy requests, and any ask for a person are handed off with full context by design.
- Expand scope one skill at a time, each with its own rules, actions, data path, and handoff triggers, after a paid pilot proves the resolution rate.
The first question most teams ask about customer service automation is "what can the AI handle?" It is the wrong question. The right one is "what should it handle alone, what should it handle with a confirmation step, and what should it never touch?"
Those are scoping decisions, and they are yours to make. The agent does not get to write its own job description any more than a new hire does. This post gives you a four-part framework for making those decisions, concrete examples on both sides of the line, and a method for widening the line over time.
Start with the shape of your queue
Most support queues are far more repetitive than they feel from inside them. McKinsey's 2025 contact center research found that 50 to 60 percent of interactions remain transactional despite years of effort to reduce them. One European bank in the study saw roughly half its calls fall into that category. A North American telecom saw about 40 percent.
Transactional does not mean trivial to the customer. It means the request has a known answer and a known action. Where is my order. Send me my policy document. Move my appointment. Update my address. These arrive by the hundreds, phrased a thousand ways.
Pull 200 to 500 recent tickets and tag each one by request type. You will end up with a short list that accounts for most of the volume and a long tail that accounts for most of the judgment calls. The framework below is for sorting the short list.
Four tests for every request type
Run each high-volume request type through four questions. Each one is a filter, and a request type needs to clear all four before an agent resolves it without a person in the loop.
- Volume: Does it arrive often enough to matter?
- Rule clarity: Can you write down how to resolve it so that two people would reach the same answer?
- Consequence: If the agent gets it wrong, what breaks, and can it be reversed?
- Data sensitivity: What must the agent read or write to finish, and under what rules?
A request type that clears all four is eligible for autonomous resolution. One that clears three is often eligible with a confirmation step or a cap. One that fails volume is not worth building yet. One that fails consequence or sensitivity goes to a person regardless of how often it shows up.
Test one: volume
Volume decides whether the request earns a place in the first scope. An agent that resolves a request type arriving 40 times a day changes your queue. One that resolves a request arriving twice a month changes nothing and still costs time to configure, test, and review.
Volume also predicts how fast you learn. High-frequency requests generate enough cases in the first two weeks of a paid pilot to show you the resolution rate, the handoff reasons, and the edge cases. Low-frequency ones leave you guessing for months.
Gartner's 2024 survey of 5,728 customers found that only 14 percent of issues were fully resolved in self-service, and only 36 percent of the issues customers called very simple. Those very simple, very frequent requests are the ones currently leaking from self-service onto your team's queue. They are where volume is highest and the gap is widest.
Test two: rule clarity
The clearest signal that a request is ready for automation is that you can write the procedure down and a new hire could follow it on their first day. If resolving the request depends on "you develop a feel for it," it is not ready.
Test this directly. Take ten recent tickets of the same type and ask two experienced reps to document how each one should be resolved. Where they agree, you have a rule. Where they differ, you have either an undocumented policy or a true judgment call. Document the first and route the second to a person.
Rule clarity also covers the information needed to apply the rule. A reschedule is clear if the calendar shows available slots and the policy says how much notice is required. It is unclear if availability lives in a technician's head.
Write the rules you find into the agent's instructions, not into a hope that the model will infer them.
Test three: consequence
Ask what happens if the agent resolves the request incorrectly, and whether the mistake can be undone. A wrong tracking link is embarrassing and reversible. A wrong refund amount is a financial event. A wrong answer about coverage can become a claim dispute.
Sort consequences into three tiers. Low-consequence, reversible actions can be completed without confirmation. Medium-consequence actions, such as changing a payment method or canceling a service visit, warrant a confirmation step with the customer before the agent commits. High-consequence actions, anything involving money above a threshold, contractual terms, or health, belong with a person.
This tier structure lets you automate more than a single yes-or-no rule would. The agent can prepare an action it is not permitted to submit, so a person spends thirty seconds approving rather than ten minutes working.
Test four: data sensitivity
Every request requires the agent to read something and usually write something. The question is what, and under which rules. Checking an order status requires reading one record. Changing an address requires writing to the system of record. Answering a question about a medical plan or a legal matter may involve data governed by regulation.
Map the data path for each request type. Note what the agent must verify before it reads account data, which fields it may write, and which systems hold regulated information. Those constraints become access controls and verification rules, not suggestions.
The NIST AI Risk Management Framework treats mapping, measuring, and managing risk as continuous functions rather than a one-time review before launch. For a customer service agent, that means the data map is something you revisit each time the scope grows.
Requests that pass all four
Across SaaS, insurance, property management, and subscription businesses, the same request types tend to clear every test:
- Status requests: order, shipment, claim, application, or ticket status.
- Document requests: proof of coverage, invoices, receipts, policy documents, statements.
- Contact and address changes: with verification, written directly to the account record.
- Reschedules: appointments, service visits, and deliveries within policy.
- Work orders: a tenant's maintenance request, created with the right priority and an entry window confirmed.
- Plan and policy questions: answered from approved content.
Each of these has high volume, a writable rule, a reversible or low-stakes outcome, and a bounded data path. Each ends in a system state your team can check: a document sent, a record updated, a work order created, a ticket closed.
Requests that always go to a person
The other list is just as important, and it is not a list of what the agent cannot do. It is a list of what your business has decided a person should do.
- Disputes: billing disagreements, chargebacks, contested charges above your threshold.
- Complaints: anything where the customer is unhappy with the company rather than asking for a task.
- Clinical or legal questions: anything touching diagnosis, treatment, liability, or contractual interpretation.
- Out-of-policy requests: exceptions, waivers, anything that requires someone with authority to say yes.
- Any request to speak to a person.
Gartner's 2026 survey of 3,566 customers found that 87 percent expect companies using generative AI in service to provide access to a human agent. The handoff is part of the design. The agent verifies the customer, gathers the history, and transfers the case with context so the person starts from the middle rather than the beginning.
Expanding scope one skill at a time
The first scope should be narrow on purpose: two or three request types that clear all four tests. Run them in a paid pilot, review every handoff and a sample of successes, and measure resolution rate against the manual baseline. When the numbers hold, give the agent its next skill.
A skill is a defined request type with its own rules, actions, data path, and handoff triggers. Adding one is a scoping decision, not a rebuild. Gartner's February 2026 survey of 321 service leaders found 91 percent under executive pressure to implement AI. The way to meet that pressure without breaking service quality is to expand deliberately, one proven skill at a time.
Vida's customer service agent is designed around this model. It verifies the customer against your rules, answers from approved content, completes eligible actions inside your web-based ticketing, CRM, and service systems, and hands off everything outside scope with full context. From the Vida AI Agent Operating System, you define what the agent can say, read, write, and complete, review its conversations, and track resolution and handoff rates per agent as you widen the scope.
Bring us your tagged queue and your always-handoff list. We will help you scope the first paid pilot. Design Your AI Agent.
Citations
- McKinsey & Company. "The contact center crossroads: Finding the right mix of humans and AI." 2025. Referenced for 50 to 60 percent of interactions remaining transactional and the bank and telecom examples. https://www.mckinsey.com/capabilities/operations/our-insights/the-contact-center-crossroads-finding-the-right-mix-of-humans-and-ai
- Gartner. "Gartner Survey Finds Only 14% of Customer Service Issues Are Fully Resolved in Self-Service." 2024. Referenced for the 14 percent full-resolution rate and 36 percent for very simple issues. https://www.gartner.com/en/newsroom/press-releases/2024-08-19-gartner-survey-finds-only-14-percent-of-customer-service-issues-are-fully-resolved-in-self-service
- National Institute of Standards and Technology. "Artificial Intelligence Risk Management Framework (AI RMF 1.0)." 2023. Referenced for mapping, measuring, and managing risk as continuous functions. https://doi.org/10.6028/NIST.AI.100-1
- Gartner. "Gartner Survey Finds 87% of Customers Say Companies Using GenAI for Customer Service Must Provide Access to a Human Agent." 2026. Referenced for the 87 percent human-access figure. https://www.gartner.com/en/newsroom/press-releases/2026-08-04-gartner-survey-finds-87-percent-of-customers-say-companies-using-genai-for-customer-service-must-provide-access-to-a-human-agent0
- Gartner. "Gartner Survey Finds 91% of Customer Service Leaders Under Pressure to Implement AI in 2026." 2026. Referenced for 91 percent of 321 service leaders reporting executive pressure to implement AI. https://www.gartner.com/en/newsroom/press-releases/2026-02-18-gartner-survey-finds-ninety-one-percent-of-customer-service-leaders-under-pressure-to-implement-ai-in-2026
