Do not switch support to AI-first yet. Run a four-week assisted-drafting pilot for billing and account questions, where evidence is strongest and escalation risk is bounded.1
- Draft assistance can reduce handling time without hiding a human from customers.3
- Cancellation, security, and outage cases should remain human-owned.
- Approve a broader rollout only if quality holds and repeat-contact rate does not rise.
Evidence reviewed
| Evidence | What it says | Limitation |
|---|---|---|
| 2,400 recent tickets | 46% are repetitive billing or account questions1 | One quarter of seasonality |
| Agent shadow study | Drafting consumes 38% of handling time2 | Six agents, two weeks |
| Vendor benchmark | 22-31% handling-time reduction3 | Vendor-selected customers |
| Escalation audit | Security and cancellation errors have high downside4 | Low frequency, high severity |
Recommendation
Use AI to draft replies for two queues: billing explanations and account administration. A human reviews every draft before sending. Retrieval must be limited to approved help-center and account-policy sources; the model should not infer refunds, credits, or security status.
Keep these queues out of the pilot:
- Account compromise and identity verification.
- Cancellation exceptions, refunds, and credits.
- Active incidents or degraded service.
- Any message carrying legal, medical, or regulated content.
Decision criteria
| Metric | Baseline | Pilot gate | Why |
|---|---|---|---|
| Median handling time | 11.4 min | at least 20% lower | Measures operating value |
| QA score | 92% | no more than 1 point lower | Prevents speed from masking quality loss |
| Repeat contact in 7 days | 14% | no increase | Detects superficially correct replies |
| Unsafe suggestion rate | not tracked | below 0.5% | Hard safety boundary |
Four-week pilot
| Week | Work | Exit condition |
|---|---|---|
| 1 | Build retrieval set and 100-ticket offline evaluation | No critical policy errors |
| 2 | Shadow mode with five agents | Reviewers agree on quality rubric |
| 3 | Human-reviewed live drafts at 25% of eligible volume | Unsafe suggestions below 0.5% |
| 4 | Expand to 50% and measure repeat contact | All decision gates reported |
Open questions
The largest uncertainty is not draft quality in common cases; it is whether agents become less attentive when most drafts are acceptable. Track edit distance, skipped review time, and reviewer disagreement to detect automation complacency.2