# Sizing pilots before committing to full-scale rollout
A regional call center in Ohio deploys an AI customer-service assistant to handle billing questions for 40,000 monthly calls. Six weeks later, leadership must decide: expand to all 22 call centers nationwide, or pull the plug. The decision hinges on numbers nobody agreed on before the pilot started. This is the most common, and most expensive, mistake in telecom AI adoption.
A pilot is not "AI on a small scale." It is a controlled experiment designed to answer one question: does this solution perform well enough, cheaply enough, and safely enough to justify the cost of scaling it?
Telecom operators like Verizon, Deutsche Telekom, and Vodafone run dozens of AI pilots per year, covering network operations, customer service, fraud detection, and churn prediction. Most never scale. Not because the technology fails, but because the pilot was never designed to produce a clear go/no-go answer. Teams pick a call center, run the tool for a few weeks, notice "customers seem happier," and greenlight a national rollout. Then costs balloon and results don't replicate.
Sizing a pilot properly means fixing four things before day one: scope, duration, success thresholds, and comparison baseline.
For an AI customer-service assistant (a system using natural language processing, or NLP, the branch of AI that lets software understand and generate human language, to handle chat or voice queries), scope means:
A pilot at one mid-performing regional center handling 40,000 monthly billing calls is a reasonable unit. Large enough to generate statistically meaningful data, small enough to contain the blast radius if something breaks.
Six weeks is the default pilot length in many corporate playbooks. That default is often wrong. Duration should be driven by how many interactions you need to detect a real difference in performance, not how many fit into a quarter.
Simple worked example:
Suppose your current human-agent-only resolution rate for billing calls is 78% (resolved on first contact, no callback within 7 days). You want to know if the AI assistant improves this to at least 83%, a threshold your team judges to be commercially meaningful.
To detect a 5-percentage-point difference with reasonable statistical confidence (using a standard two-proportion test at 80% power, 95% confidence), you need roughly 1,200 to 1,500 calls per group (AI-assisted vs. human-only), depending on variance assumptions. At 40,000 monthly calls, splitting even 10% of volume into an AI-assisted test group gives you 4,000 calls a month, enough to reachreachThe number of unique people exposed to your message in a given period. Unlike impressions, reach counts each person once, no matter how often they see it.View full definition → that threshold within 3 to 4 weeks, not six.
Running shorter than needed gives you noise dressed up as a decision. Running longer than needed just delays the rollout and burns budget on a question you already answered.
This is the step most pilots skip, and it's the one that matters most. Define numeric go/no-go thresholds in advance, across at least three dimensions:
1. Performance thresholds
2. Cost thresholds
3. Risk thresholds
A realistic threshold set for the Ohio pilot might read: "Expand only if containment rate exceeds 35%, CSATCSATCustomer Satisfaction Score, a direct measure of satisfaction captured right after a specific interaction or experience, usually on a short rating scale.View full definition → stays within 0.2 points of the human-agent baseline, and cost per resolved contact drops by at least 15%." Anything short of all three is a no-go or a redesign, not a partial win dressed as success.
Compare the AI pilot against current performance, measured in the same period and same site, not against a company average or last year's numbers. Call center performance drifts with staffing, seasonality, and call mix. A December baseline (holiday billing spikes) is not comparable to a March pilot.
Best practice: run a concurrent control group. Route a similar volume of calls to human agents only, in the same weeks, same site. This isolates the AI's effect from unrelated changes in call volume or staffing.
Public benchmarks are scarce and vendor-reported figures should be treated skeptically, but a few grounded reference points help calibrate expectations (estimates, as of 2024-2025):
The distinction matters: an AI assistant that helps agents resolve calls faster (agent-assist) has a different risk and ROIROIReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.View full definition → profile than one that resolves calls autonomously (containment). Pilots should specify which model they are testing, since thresholds and risks differ substantially.
For deeper grounding on evaluating conversational AI systems, the Stanford HAI AI Index Report publishes annual benchmarks on AI capability and deployment trends worth cross-checking against vendor claims.
Knowledge check
1. What is the core purpose of a properly designed AI pilot in a telecom setting?
2. Why does the lesson recommend starting a customer-service AI pilot with billing disputes rather than open-ended technical support?
3. According to the lesson, why is it risky to pilot an AI assistant across both voice and chat channels simultaneously?
4. Select ALL correct answers about why many telecom AI pilots fail to scale, according to the lesson.
Select all the correct answers.
5. Select ALL correct answers about the four elements the lesson says must be fixed before a pilot begins.
Select all the correct answers.
🎬 [VIDEO: "How Companies Are Using AI in Customer Service" - youtube.com/@BloombergTelevision - a grounded look at real enterprise deployments and the operational tradeoffs involved]