Skip to content
Question Vault?
Free to readNo accountNo email wallNo invented statisticsNo partial listsCopy or print any set and take it with you

Questions to Ask When Evaluating a Program

These are questions to ask when evaluating a program you run, fund or oversee, whether it is a nonprofit service, a public health project, a school initiative or a workplace training program. They follow the order a review usually takes: what the program was for, who it reached, how it was delivered, what changed and how much of that it can claim, what each result cost, and what to do next. Some are for staff, some for participants, and a good many can only be answered by the records.

53 questions

Want questions from the whole vault instead? Try the random question generator.

The questions

Each question, and why to ask it

Purpose

What was this program meant to change, and for whom?

Why ask it

Ask each person for one sentence before anyone opens a report. If three staff members give three different sentences, the first finding is that the program has no shared aim, and its outcomes will be hard to judge until it has one.

What problem was the program set up to solve, and is that problem still the same size and shape?

Why ask it

Programs outlive the conditions that produced them. Ask for the figures or the event that justified the launch, then for the same figures today. A need that has shrunk or moved is a reason to redesign even a program that runs smoothly.

How is the program supposed to work, step by step, from its activities to the change it is after?

Why ask it

This is the logic model or theory of change, whether or not anyone ever drew it. Have a staff member talk it through at a whiteboard and mark every arrow where the link is assumed and has never been checked. Those arrows are where the review should look hardest.

What would success have looked like by now, and who set that target?

Why ask it

A target written at the start can be held up against the results. One written after the numbers came in cannot. If there was none, agree on a reasonable one before you look at any outcomes, and record that you did it in that order.

Who asked for this evaluation, and what decision is waiting on it?

Why ask it

A funder weighing a renewal needs different answers from a team trying to improve its own work. Name the decision and its date first. Any question that could not affect that decision can be left for another time.

How does this program fit the organization's mission and its current priorities?

Why ask it

A board member's question, and a fair one even of a program that works. Ask what the organization could no longer say about itself if the program closed. If the answer is 'nothing much', the program may sit better with a partner whose main work it is.

Who has a stake in the findings, and which of them have we not heard from yet?

Why ask it

List participants, frontline staff, managers, funders, partner agencies and the people who were eligible but stayed away. The groups nobody has spoken to tend to hold the least flattering account, and often the most useful one.

What is the program being judged on that it was never designed to do?

Why ask it

Hopes pile up over the years: a job skills course ends up blamed for the local unemployment rate. Separating the original promise from the later expectations keeps the review fair, and it shows whether the stated aims need rewriting.

What has already been evaluated or reported about this program, and what came of it?

Why ask it

Earlier reviews save work and reveal habits. Ask which of the last report's recommendations were acted on. If the answer is none, plan how this one will be used before you write a word of it.

Reach

Who is the program meant for, and roughly how many of those people are there?

Why ask it

This is the denominator. Without an estimate of the eligible population, 'we served 400 people' cannot be read as a lot or a little. A rough figure with its source stated beats no figure.

Who actually enrolled, and how do they compare with the people the program was designed for?

Why ask it

Compare on whatever defined the intended group: age, income, site, job grade, level of need. Programs drift toward the people who are easiest to recruit, and those are often the ones who needed the help least.

Who is eligible but not taking part, and what do we know about why?

Why ask it

Staff can guess, but only the people themselves know. A handful of short conversations with non-participants about cost, timing, transport, trust or never having heard of the program can change the picture more than another survey of those who came.

Who else offers something similar to the same people, and how many use both?

Why ask it

A program can be well run and still be the third of its kind in one neighborhood, or the second course on one subject in the same company. Ask participants what else they have used and where they would turn if this closed. Overlap is not always waste, since two services with waiting lists are both needed, but it changes how much of the outcome this one can be credited with.

How does someone find out about the program and get a place on it?

Why ask it

Trace the route one recent participant took: the referral, the form, the wait, the first session. Every step that needs a daytime phone call, a document or a manager's sign-off is a point where a particular kind of person gives up.

How many people started, how many finished, and at what point do the others leave?

Why ask it

Ask for a count at each stage, not a single completion rate. Losses after the first session point at the welcome or the fit, and losses in the middle point at workload or at life getting in the way. Then ask whether anyone contacts the people who go.

When more people apply than there are places, how is the choice made?

Why ask it

The selection rule shapes the results. A program that picks its most motivated applicants can post better outcomes than one that takes people as they come, even when the sessions are identical. Write the rule down, because you will need it when you read the outcome figures.

How much of the program does a typical participant actually receive?

Why ask it

Being on the roll is not the same as being in the room. Get sessions attended, hours or modules completed per person, and look at the spread as well as the average. Better results among heavier users are a clue, though keener people also turn up more.

Delivery

What was the program designed to deliver, and what is delivered in practice?

Why ask it

Put the manual, the grant proposal or the original plan beside a description of last month. Go item by item: delivered as written, delivered differently, or dropped. A program that was never run as designed has not been tested yet, so poor outcomes say little about the design.

Where has the program been adapted, and who agreed to each change?

Why ask it

Some adaptation is good practice, such as new examples for a different audience. Cutting the part the model says does the work is another matter. Sort the changes into those two piles with the people who made them, and ask what prompted each one.

Does the program run the same way at every site and with every facilitator?

Why ask it

Ask for the same few figures by site or by person: attendance, completion, session length. Large differences mean you are evaluating several programs under one name. The strongest site is worth a visit, since it may be doing something the others could copy.

Do staff have the training, time and materials the design assumed they would?

Why ask it

A design that expects a caseload of twenty and gets one of fifty is unlikely to show the results it promised. Put this to frontline staff directly and not only to their managers. Turnover belongs here too, as each new hire delivers a slightly different program.

Which parts of the program work well enough that nobody should touch them?

Why ask it

Reviews hunt for faults, and a redesign can remove the one thing people came for. Ask staff and participants separately what they would fight to keep, and look for the item both name. Put it in the report by name so it survives the next budget round.

What do participants say about how the program is run, and would they send a friend to it?

Why ask it

Satisfaction is not an outcome, but it helps explain attendance and dropout. Ask about the times, the venue or platform and how they were treated, and take the reason behind a yes or no on recommending it as the useful part. If feedback forms are handed in to the person being rated, read the scores as generous.

What do frontline staff think is not working?

Why ask it

They often know, and may never have been asked in a setting where it felt safe to say. Speak to them without their manager in the room and say how their comments will be reported. Listen for workarounds, which show where the design and the real job disagree.

What gets in the way of delivery that the program does not control?

Why ask it

Late referrals from a partner, a venue lost halfway through, a hiring freeze, a policy change. Note each obstacle and whether it is a one-off or a standing condition. A recommendation that ignores a standing condition will not survive its first month.

What records are kept of what was delivered, and how far can they be trusted?

Why ask it

Look at the raw attendance sheets or system entries, not the summary. Check who fills them in, when, and whether anything rides on the numbers looking good. If the records are thin, say so in the report and start better ones now.

Outcomes

Which outcomes are measured, and are they the ones the program set out to change?

Why ask it

Lay the measures beside the stated aims. Programs often count what is easy, such as sessions held or leaflets handed out, and call it results. Mark each measure as an activity, an output or a change in someone's life, and see how many are the third kind.

Where did participants start, and how do we know?

Why ask it

A change needs a starting point taken before the program began, or very early in it. If no baseline was collected, ask whether intake forms, school records, HR data or a referral letter can stand in, and whether you are allowed to use them for this. Memories of how things used to be are the weakest substitute.

What changed for participants, and how large is the change?

Why ask it

Ask for the size in plain units: points on the test, days absent, people in work. A difference can be real and still too small to matter to anyone. Then ask what size of change the program's designers were hoping for.

What would have happened to these people without the program?

Why ask it

This is the hardest question in the review and the one a funder most needs answered. Children grow, many people out of work find jobs, and people in a crisis often steady themselves with time. Ask what a similar group did over the same period and where that information could come from.

What else was going on at the same time that could explain the change?

Why ask it

A new manager, a pay rise, another service opening nearby, a change in the local economy, a different intake policy. List the rivals with people who know the setting, and for each ask what you would expect to see if it, and not the program, were the cause.

Who are participants being compared with, and how alike are the two groups?

Why ask it

People who sign up tend to differ from people who do not, before anything is done for them. Check how the comparison group was chosen and whether the two looked similar at the start on the things that matter. With no comparison group at all, the honest wording is that outcomes improved, not that the program improved them.

Whose results are missing from the data?

Why ask it

Outcome figures often cover only the people who stayed to the end and filled in the last form. Those who left early may well have done worse. Ask how many started for every one counted, and how the result looks if the missing are assumed not to have improved.

Does the program work better for some participants than for others?

Why ask it

An average can hide a strong effect in one group and none in another. Break results down along the same lines used for enrollment, but be wary when a subgroup holds only a handful of people. A pattern that staff also recognize from daily work deserves more weight.

Do the gains last after the program ends?

Why ask it

Results taken on the final day tend to be the best they will ever look. Find out whether anyone has checked three, six or twelve months on. If nobody has, a short follow-up with a sample of past participants is often the cheapest new evidence you can add.

What has the program changed that nobody planned, for better or worse?

Why ask it

Ask participants and staff this openly, without a list to tick. Side effects run both ways: new friendships and confidence, or stigma, lost work hours and strain on people outside the program. Any report of harm goes to whoever handles safeguarding or complaints in your organization before the review goes further.

In their own words, what do participants say is different for them now?

Why ask it

Numbers show how much changed, and stories show how. Hold a few open conversations and listen for whether people credit the program or something else. Choose who to speak to yourself, since staff tend to put forward their success stories.

How was the outcome data collected, and who collected it?

Why ask it

Answers given to the person who ran the sessions tend to lean positive, and a survey returned by a fifth of participants speaks for that fifth. Note the method, the response rate and the timing beside every figure. If you plan to gather new data from people, ask what consent and data protection rules apply where you work.

Cost

What does the program cost to run for a year, including what never appears on its budget line?

Why ask it

The budget often leaves out a manager's time, donated space, volunteers' hours and shared services such as IT or finance. Have the finance lead build the fuller figure with you. A program that looks cheap because another department carries half of it is not cheap to the organization.

What does it cost for each participant, and for each participant who finishes?

Why ask it

Work out both, because they can sit far apart. A low cost per enrollee with a high cost per completer means the money is going on people who leave. Use the fuller cost figure, and state which year and which headcount you divided by.

What does the program spend for each result it can fairly claim?

Why ask it

Divide the full cost by the outcomes left after the comparison questions have been asked: jobs kept for six months, students who reached the reading standard, injuries avoided. The figure is only as sound as the outcome count, so give it as a range and show the working.

Which parts of the program take most of the money, and are those the parts that do the work?

Why ask it

Split spending by component and set it against the program's own account of how it works. Sometimes the element participants value most is the cheapest, and an expensive one survives only out of habit. This tells you what to protect if the budget is cut.

What else could the same money buy toward the same aim?

Why ask it

A cost per result means little alone. Look for one or two alternatives, such as a lighter version, another provider's approach or direct support, and compare on the same outcome. Be careful with figures from elsewhere, since costs are counted differently from place to place.

What does taking part cost the participants themselves?

Why ask it

Travel, childcare, unpaid time off, time away from a desk that a workplace then has to cover. These costs often fall hardest on the people a program most wants to reach, and they can explain much of the dropout counted under Reach.

How long is the current funding in place, and what conditions come with it?

Why ask it

A funder may require particular measures, a set number of people served or a fixed model, and those terms limit which changes are open to you. Read the agreement before recommending anything. If a renewal date is near, it sets the deadline for the whole review.

Decision

Taking everything together, should the program continue, change, grow or end?

Why ask it

Have each person in the room state a position and the single piece of evidence behind it before discussion begins. That stops the loudest voice from setting the frame. If the evidence is too thin to choose, say that plainly and name what would settle it.

If the program continues, which two or three changes would make the most difference?

Why ask it

A list of twenty recommendations rarely gets any of them done. Rank by expected effect and by how hard each is to carry out, and test the top few with the staff who would have to make them happen. Their objections are cheaper to hear now than after the report is signed.

What would have to be true before we expanded it?

Why ask it

Results at one site with one committed team often shrink at ten. Ask what the current version depends on, such as a particular person, a partner or a local condition, and whether that can be reproduced. A second site treated as a test is a safer step than a rollout.

If the program ends, what happens to the people who rely on it now?

Why ask it

Find out who is midway through, where they could be referred and what notice they and the staff are owed. Contracts, grant terms and employment rules on closing a service vary, so ask how it works in your organization and location. A badly handled ending can undo trust that other programs depend on.

What do we still not know, and is it worth finding out before we decide?

Why ask it

Every review ends with gaps. For each, ask whether a different answer would change the decision. If it would and can be had in weeks, wait; if it would not, decide now and note the gap for next time.

Who needs to hear the findings, and in what form?

Why ask it

A board wants a page and a recommendation, staff want detail they can act on, and participants are owed a plain account of what their answers led to. Plan each version before writing the long report. Share the difficult findings with the program team first so nobody meets them in a meeting.

Who owns each recommendation, and when will we look at it again?

Why ask it

A recommendation with no name and no date is a suggestion. Put both beside every item and set one meeting, three or six months out, to go down the list. That meeting is the difference between an evaluation that was done and one that was used.

What should we start recording now so the next review is easier?

Why ask it

Most of the gaps you met, such as no baseline, no follow-up and patchy attendance records, can be closed with a few fields added to routine paperwork. Pick the smallest set that would answer the questions under Outcomes. Check with staff that collecting it will not take time from the people they serve.

Using these questions in a program review

Practical guidance for the conversation itself

Before you ask anyone anything

Fix the decision and its date

An evaluation with no decision attached tends to grow until it runs out of time. Write down who will decide, what the options are and when the choice has to be made, then work backward. A renewal due in six weeks allows a careful read of existing records and a dozen conversations. It does not allow a new survey with a follow-up.

Match the questions to the program's age

A program in its first year has rarely had time to change much, so judging it on outcomes alone tells you little. Spend that review on Purpose, Reach and Delivery: is it reaching the right people and running as planned. An established program with settled delivery is the one to press on Outcomes and Cost.

Get the logic onto one page

Draw a simple chain with the program team: what goes in, what is done, who receives it, what should change first, and what should change in the end. It need not be a formal logic model. The page gives you something to test, and disagreements over it among staff are findings in themselves.

Agree in advance what would count

Before looking at results, settle with the people who commissioned the review what would lead them to continue, to change or to stop. Doing this first keeps the bar from moving once the numbers are known. It also shows early if someone has already made up their mind, which is better to learn at the start.

Collecting the answers

Put each question to three sources

Nearly every question here can go to staff, to participants and to the records. Staff tell you what was intended, participants tell you what it was like to receive, and the records show what was counted. Where all three agree you can be fairly confident. Where they differ, the difference is the thing to investigate.

Go to the people who are hard to find

Current participants who turn up to a feedback session tend to be the ones the program suits best. Make a deliberate effort to reach people who dropped out, people who were referred and never started, and people who were eligible and never applied. Even five such conversations will test the conclusions you were about to draw.

Watch the program happen

Sit in on a session, a home visit or a training day if that is appropriate and participants agree. An hour of observation can show what an interview leaves out: how long the useful part lasts, who speaks, and what the manual says that nobody does. Use the same short checklist at each site so the visits can be compared.

Think about who is doing the asking

People answer differently when the questioner controls their funding, their job or their place on the program. Where you can, have someone outside the delivery team hold the conversations, and tell people how their words will be reported. Rules on consent, confidentiality and ethical review differ by sector and place, so check which apply to you before collecting anything new.

Making sense of the evidence

Read reach and delivery before outcomes

Outcome figures mean little until you know who the program served and what they received. Weak results from a program that reached the wrong group, or delivered half its sessions, point to a recruitment or delivery fix. Weak results from one that reached the right people and ran as designed point at the idea itself.

Claim only what the design supports

There is a ladder of claims. 'Participants improved' needs a before and an after. 'Participants improved more than similar people' needs a comparison group. 'The program caused the improvement' needs a comparison that rules out the other explanations. Use the wording your evidence earns, and consider 'contributed to' where several things were at work.

Be careful with small numbers

With thirty participants, three people changing their answers moves the result by ten points. Report the counts beside the percentages, avoid slicing small groups into smaller ones, and lean more on interviews and observation where the numbers cannot carry the weight.

Check both halves of cost per result

The figure is a division, and either half can be off. The top is often too low because shared and donated costs were left out. The bottom is often too high because it counts everyone who improved, including those who would have improved anyway. Give a range, show how you built it, and compare only with figures built the same way.

Reporting and deciding

Lead with the answer

Open the report with the recommendation and the three or four findings that support it. Method, tables and caveats can follow for those who want them. A decision-maker who has to reach page thirty to learn what you think will often not get there.

Report what did not work

A review that finds nothing wrong is rarely believed and seldom useful. State the weak points in neutral terms, tied to evidence and not to individuals. Programs are run by people who care about them, and they are far more likely to act on a finding that reads as a problem to solve than as a verdict on them.

Keep continue, change and end all open

'Continue with modifications' is the easy landing place for a review, because it offends nobody. Before accepting that, ask the room to argue each of the other options in earnest for ten minutes. If ending or replacing the program cannot be argued at all, the middle course has been chosen on evidence and not on comfort.

Close the loop with the people who answered

Participants and frontline staff gave time and candor. Tell them, in a page or a short meeting, what was found and what will change as a result. It is fair to them, and it makes the next evaluation easier, because people answer more freely when they have seen their answers lead somewhere.

More on this topic