# Measuring outcomes when services take years to pay off
A workforce-training program graduates 500 people in June. By August, the agency director stands at a budget hearing and gets asked a brutal question: "Did it work?" The honest answer is that nobody knows yet. Wages, job retention, and reduced reliance on public benefits show up over two to four years, not two months. But the budget cycle does not wait.
This gap, between when you spend money and when the payoff appears, is the central measurement problem in the public sector. This lesson shows you how to measure honestly inside that gap.
In business, you can often see results fast: a sale closes, a customer churns, revenue lands this quarter. Public programs are different for three reasons.
Long lag. Job training, early childhood education, and reentry programs pay off over years. The 2011 Perry Preschool follow-up studies famously tracked participants into their 40s.
Attribution is murky. A trainee got a job. Was it your program, the strong local labor market, or their own hustle? You need to separate your effect from everything else.
Political pressure to look good now. Hearings reward visible activity. That pressure pushes agencies toward vanity metrics, numbers that look impressive but do not prove impact.
A vanity metric is a number that goes up when you do more work but does not tell you whether anyone is better off.
| Vanity metric (activity) | Outcome metric (impact) |
|---|---|
| People enrolled | People employed 12 months later |
| Training hours delivered | Wage gain versus before the program |
| Certificates issued | Retention in the job at 2 years |
| Workshops held | Reduced reliance on public assistance |
The rule of thumb: if the metric measures what *you* did, it is an output. If it measures what *changed for the participant*, it is an outcome.
Outputs are not useless. They tell you the program is running. But at a budget hearing, "we delivered 40,000 training hours" invites the reply, "so what?" You need to connect that activity to a result.
A logic model is a simple diagram that links what you put in to what you expect to change. It forces you to state your assumptions out loud, which is exactly what auditors and legislators want to see.
The standard chain:
Inputs → Activities → Outputs → Outcomes → Impact
For our workforce-training program:
The logic model does two things. It tells you *what* to measure at each stage. And it lets you report early signals (outputs, short-term outcomes) while honestly labeling the long-term impact as "to be confirmed."
The W.K. Kellogg Foundation Logic Model Development Guide is a free, widely used template if you want to build one.
🎬 [VIDEO: "Logic Models: The Foundation for Program Evaluation" — youtube.com — a clear walkthrough of building a logic model from inputs to impact]
Here is the question that separates real evaluation from wishful thinking: compared to what?
Suppose 70 percent of your graduates are employed a year later. Sounds great. But what if 65 percent of similar people who never entered your program also got jobs, because the local economy was booming? Then your true effect is only 5 percentage points, not 70.
The counterfactual is what would have happened without your program. You can never observe it directly for the same people, so you estimate it using a comparison group: people similar to your participants who did not receive the service.
Randomized controlled trial (RCT). Randomly assign eligible applicants to the program or a waitlist. This is the gold standard because randomization makes the two groups statistically similar. It is often feasible when demand exceeds slots, which is common in public programs.
Matched comparison. When you cannot randomize, find non-participants who look like your participants on age, prior earnings, education, and location. Statistical matching methods pair them up.
Regression discontinuity. If eligibility depends on a cutoff (for example, an income threshold), compare people just above and just below the line. They are nearly identical except for program access.
The point is not to master the statistics. It is to be able to say, in a hearing, "we compared our trainees to a similar group who did not get trained, and our people earned more." That sentence survives scrutiny. "Our trainees did great" does not.
Say you have participant outcomes and a matched comparison group. A basic difference-in-differences calculation compares the *change* for each group, not just the ending level. This controls for the general economic trend that affected everyone.
Before After Change
Program group $18,000 $30,000 +$12,000
Comparison group $18,000 $25,000 +$7,000
Program effect = $12,000 - $7,000 = $5,000 per participantThe comparison group's $7,000 gain reflects the improving job market. Only the extra $5,000 is plausibly attributable to your program. Reporting the full $12,000 would be dishonest, and a sharp auditor would catch it.
Multiply that $5,000 by 350 employed participants and you have a defensible impact figure. Attach it to the program cost and you can discuss cost per dollar of wage gain, which is what a budget committee actually wants.
Vérification des acquis
1. What is the central measurement problem the lesson identifies for public-sector programs?
2. According to the lesson's rule of thumb, how do you distinguish an output from an outcome?
3. Why does 'attribution' make outcome measurement difficult in the public sector?
4. Select ALL correct answers. Which of the following are examples of vanity metrics (outputs) rather than outcome metrics?
Sélectionnez toutes les réponses correctes.
5. Select ALL correct answers. Why do public programs face particular pressure toward vanity metrics?
Sélectionnez toutes les réponses correctes.
You still have that August hearing, years before the wage data matures. Here is how to report honestly without stalling.
Use leading indicators. These are early, measurable signals that reliably predict later outcomes. Research on workforce programs consistently links early job placement and first-quarter retention to longer-term earnings. So reporting "80 percent placed within 90 days, 90 percent still employed at quarter one" is meaningful, because those metrics forecast the outcome you care about.
Set staged milestones. Tell stakeholders up front: outputs reported at 3 months, short-term outcomes at 12 months, impact evaluation at 36 months. This turns "we don't know yet" into "we are on schedule."
Be explicit about what is confirmed versus projected. Label projected long-term impact as an estimate based on prior cohorts or published research. Never present a projection as a measured result.
Link administrative data early. Many agencies can match participant records to state unemployment insurance wage records, which update quarterly. This is far faster and cheaper than surveys and is a common backbone for workforce evaluation.
Auditors are not trying to trap you. They check whether your claims match your evidence. Three habits keep you safe:
Document your methods before you collect data. Write down your logic model, your outcome definitions, and your comparison-group approach in advance. Choosing your metrics after seeing the results is called "cherry-picking" and destroys credibility.
Define terms precisely. "Employed" needs a definition: any job, or 30-plus hours, or above a wage floor? Auditors will ask.
Keep a clean data trail. Record who was served, when, and how you matched them to outcomes. If you cannot reproduce a number, treat it as if it does not exist.