Skip to content
Question Vault?
Free to readNo accountNo email wallNo invented statisticsNo partial listsCopy or print any set and take it with you

Questions to Ask About Machine Learning

Questions to ask about machine learning when a model, a project proposal or a vendor demo is in front of you and you have to decide whether to rely on it or pay for it. They are written for managers, product owners, buyers and students with no technical background, and they run in six groups that follow the conversation: the problem, the data, accuracy, bias and failures, running it, and accountability. Bring a handful to the meeting, and when an answer comes back in jargon, ask for it again in plain words.

52 questions

Want questions from the whole vault instead? Try the random question generator.

The questions

Each question, and why to ask it

The problem

What decision or task will this model take over, and who handles it today?

Why ask it

A clear answer names one job: flagging invoices for review, ranking sales leads, routing support tickets. A reply that describes a capability, such as 'it finds patterns in our data', means the project has a technique and is still looking for a problem. The person who does that job now is the next one to talk to.

Why does this need machine learning, and how far would a simple rule get us?

Why ask it

The simple rule is the baseline, and it should have been measured before anything was built. 'Call every customer who has not ordered in 60 days' may get most of the way for almost nothing. What you are funding is the gap between that rule and the model, so get the gap as a number.

Did you build this model yourselves, or does it run on another company's model?

Why ask it

A product can be a thin layer over a model rented from a large provider. That is no bad thing, but it moves some of the risk outside the room: the provider's name, what gets sent to it, and what happens when it changes or withdraws the version in use. Terms differ from one provider and plan to the next, so the answer should point to a document.

What does the model actually put out: a score, a category, a ranking or written text?

Why ask it

People in the same meeting often picture different things. A score needs someone to pick a cutoff, a ranking needs someone to decide how far down the list to go, and generated text needs someone to read it. One real output, shown exactly as its user would see it, settles the matter faster than a description.

What happens to a prediction once it is made, and who or what acts on it?

Why ask it

Follow a single prediction through: it lands on a screen, in a queue or in another system, and then a person or a piece of software does something different because of it. However accurate it is, a prediction nobody acts on is worth nothing. If that last step is still to be designed, the model is the easy half of the project.

How will we measure success in the terms this business already uses?

Why ask it

Model accuracy is the team's measure. Yours is fewer refunds, faster handling or more renewals, and somebody has to have agreed how the first turns into the second. Without that conversion nobody will be able to say later whether the project paid for itself.

Has this been done before on a problem like ours, and may we talk to the people who did it?

Why ask it

Mostly for a vendor pitch. A reference customer in your industry, with volumes like yours, counts for more than a famous logo on a slide. On the call, ask what took longer than they were told it would.

The data

What data was the model trained on, and where did it come from?

Why ask it

Good answers come with names attached: which systems, which years, bought or collected. For a vendor's model the honest reply is often 'other customers' records, not yours', which leads straight to how it will do on yours. Vagueness at this point should slow down everything that follows.

How were the correct answers in the training data decided, and by whom?

Why ask it

A model that learns from examples copies whoever labeled them. If 'fraud' in the records means the cases an investigator happened to catch, the model learns what gets caught, which is not the same as what is fraudulent. A useful follow-up is how often two people labeling the same case agreed.

What period does the training data cover, and what has changed since then?

Why ask it

Prices, products, customer habits and the forms people fill in all move. Come with two or three changes you know about, such as a new pricing plan or a merger, and check whether the data includes them. Training data from before a big shift describes a business you no longer run.

Which customers, cases or situations are thin or missing in the training data?

Why ask it

The model will be weakest where it saw least: new products, small regions, rare but expensive events. People who know their data can list these without looking anything up. Set their list beside the cases you most need it to get right.

How closely does the training data match what the model will see once it is live?

Why ask it

Models are often built on tidy historical records and then fed live inputs that are messier, later or half filled in. One live record placed next to one training record shows the difference at a glance. Which fields are present at the moment of prediction matters more than how many rows there were.

Could the model be leaning on information that will not exist yet at the moment it has to predict?

Why ask it

Builders call this leakage, and it produces test results that look too good. A model predicting which customers will cancel should not be shown the field that records the cancellation call. Surprisingly high accuracy on a first attempt is the usual sign, and the team should be able to say what they did to rule it out.

How much data did you have, and how did you decide that was enough?

Why ask it

There is no general number, so listen for the reasoning: results stopped improving as more was added, or the rare cases were counted and judged sufficient. The count that matters is of the outcome you care about. Ten thousand records containing forty frauds is forty examples.

Are we permitted to use this data for this purpose, and who has confirmed that?

Why ask it

Consent, contracts and privacy rules differ by country, state, industry and what people were told when the data was collected, so the engineering team's view is not enough. Look for the name of someone in legal, privacy or compliance who reviewed it. A vendor owes you the same for the data on their side.

Where does our data go when we use the model, and will it train models that serve other customers?

Why ask it

A vendor question, and the answer belongs in the contract, not in the sales call. It should cover what is kept, for how long, whether you can opt out and what happens to it when you leave. Terms vary by provider and by plan, so read the ones that would apply to you.

Accuracy

How accurate is it, and what exactly does that figure count?

Why ask it

Have it put as counts: out of a thousand cases, how many it got right, and how many of those were the easy majority. A model that answers 'no fraud' every time is 99 percent accurate when one case in a hundred is fraud. Then ask which measure they would choose if the money were their own.

Was it tested on cases it never saw during training, and how were those set aside?

Why ask it

A model scored on its own training examples is being marked on an exam it has memorized. The usual discipline is to hold back a portion before any work begins and, for anything involving time, to test on a later period than the one it learned from. How many times that held-back portion was looked at while the model was being adjusted shows how clean the test really is.

Is it better than the people or the process doing this today, on the same cases?

Why ask it

Today's process may be a person, a rule or nothing at all, and it should have been scored the same way the model was. That step gets skipped when the people doing the job were never measured. A model slightly worse than your staff but far cheaper can still be worth having, and one slightly better may not be.

Which is the costlier mistake for us, a false alarm or a miss, and how often does each one happen?

Why ask it

The two errors trade against each other, and only the business can price them: a good customer's payment blocked, against a fraud let through. You need both rates separately, with a rough cost on each. One blended figure usually means nobody has had this conversation yet.

Where is the cutoff between yes and no set, and who chose it?

Why ask it

Many models produce a score, and someone picks the line above which a case gets flagged. Moving that line changes how many alarms your staff see each day. Results at two or three other settings make the trade visible, and the choice should sit with someone who owns the consequences.

Can you show me a batch of cases it got wrong?

Why ask it

Twenty real errors teach more than any chart. Look for a pattern you recognize: all from one branch, all new customers, all with the same field left blank. If examples cannot be produced quickly, the team has been studying the score and not the mistakes.

When the model says it is 90 percent sure, is it right about nine times in ten?

Why ask it

The property is called calibration, and an accurate model can still lack it. It matters whenever a person will read the confidence figure and decide how hard to check. If nobody has tested it, treat the scores as a ranking and ignore the percentages.

Has it run on live cases yet, or only on historical records?

Why ask it

Results on old data are a forecast of how it will do, and results from live use are the thing itself. For a live run, the details that count are how long it lasted, how many cases it covered and whether anyone was acting on the output at the time. A model that was watched but never acted on has not yet met the people who will use it.

Has anyone outside the team that built it checked these results?

Why ask it

Builders are the people least likely to find their own mistakes, through no fault of character. An internal analyst, an audit function or an independent evaluation all count. With a vendor, propose running the model on a sample of your own cases while you keep the true outcomes to yourself.

Bias and failures

Does it perform equally well for different groups of people, and how was that checked?

Why ask it

An overall score can hide a group for whom the model is much worse. Results should be broken out by the groups that matter in your setting, such as age band, region, language or account size. Which groups you are required to examine, and which attributes you may even record, depends on where you operate, so take that part to your own compliance people.

Could it be relying on a stand-in for something it should not use, such as a zip code in place of income or ethnicity?

Why ask it

Removing a sensitive field does not remove its influence when other fields carry the same information. The check is to see which inputs the model leans on most and whether any of them track a characteristic you would not want to decide by. Surprise at the question suggests nobody has looked.

Which kinds of input make it fail badly?

Why ask it

Every model has them: blurred scans, unusual names, very short messages, a product launched last month. A specific list means the team has gone looking. Take the two that are most common in your operation and work out what share of your volume they make up.

What does it do with a case unlike anything it was trained on?

Why ask it

The dangerous behavior is a confident answer where there should be none. Better systems are built to say 'not sure' and pass the case to a person. Bring an odd example of your own and watch what happens to it.

If it writes text or answers questions, how often does it state something false, and how was that counted?

Why ask it

Language models can produce fluent, confident statements with nothing behind them, and a demo rarely shows one. A serious answer describes test questions with known answers and a person who graded the replies, and says who reads the output before a customer does. Skip this for a model that only scores or sorts.

Can you explain why it gave this particular result, in words the affected person would accept?

Why ask it

Pick one real case and have the reason given for it. Some models can give a faithful account and others only an approximation produced by a second tool, and the team should say which you are getting. Whether an explanation is owed to a customer or an applicant depends on the use and the place, so find out what applies to yours.

How could someone trick it once they know it is there?

Why ask it

This matters most for fraud, spam, content moderation and anything else where the people being scored have a reason to adapt. What happened the last time behavior shifted in response to a rule is a good guide. Expect a description of how they would notice, and distrust a claim that it cannot be done.

Does the model end up learning from the results of its own decisions, and how do you keep that from skewing it?

Why ask it

A model that decides who gets a loan never sees how the people it turned down would have done. Over time it can grow more certain on thinner evidence. Listen for a concrete countermeasure, such as letting a small random share of cases through for comparison.

What is the worst outcome if it is wrong, and who bears it?

Why ask it

Push for the single worst realistic case, not the average one: a claim wrongly denied, a safety fault missed, an offensive reply sent to a customer. What stands between the model and that outcome is the second half of the answer. 'That would be very rare' invites one more question: how would they find out it had happened?

Running it

What will it cost to run each month once it is live, and what pushes that figure up?

Why ask it

Building is paid for once and running is paid for as long as the model is on: computing, licenses, storage, the people who watch it. Usage-based pricing can climb steeply with volume, so get the bill at twice and at ten times today's load. Prices and plans vary by provider, which is a reason to have the quote in writing.

How will we find out that it has started getting worse?

Why ask it

Models decay quietly as the world moves away from their training data, and nothing crashes when it happens. A real answer names what is tracked after launch, how often someone looks, and what level triggers action. 'Users will tell us' means the first sign will be a complaint.

Does it keep learning from new cases once it is live, or stay fixed until someone retrains it?

Why ask it

People often assume a model improves by itself with use. Many are frozen at the version that was tested, which makes them predictable, and one that updates itself needs someone checking that each update is no worse than the last. Either design can be right, but you should know which one you are buying.

How often will it need retraining, and who does that work?

Why ask it

Some models hold for years and others slip within weeks, depending on how fast the behavior underneath them changes. The team should have an expectation for this one and a reason for it. Ask whether retraining is a button or a project, because a project needs a budget line of its own.

Who looks after it when the people who built it have moved on?

Why ask it

The answer is a name or a team, and they should have agreed to it. Just as important is what exists in writing: how it was trained, on what, and how to rebuild it. A model only one person understands is on loan from that person.

What does it have to connect to, and who owns those systems?

Why ask it

Integration is where schedules slip, because the model depends on data arriving from systems run by other teams with their own priorities. Those owners should have been consulted already, and what they said is worth hearing. A quiet change to a field upstream is a common way for a working model to break.

What changes for the people who do this work now, and have they been told?

Why ask it

Often a model moves the work instead of removing it: fewer routine cases, more exceptions, and someone checking the output. The plan should say who does the checking and what training they get. Whether staff or their representatives must be consulted first depends on your country, contracts and employer policy, so put that to HR before the rollout is announced.

If we stop using it or change supplier, what do we walk away with?

Why ask it

The possibilities run from the trained model, the labeled data and the code down to screenshots of the dashboard. With a vendor this is a contract question, so raise it before signing. Exit terms are far easier to negotiate before you depend on the system than after.

Can we switch it off and go back to the old way within a day?

Why ask it

You want a fallback that still exists: the rule, the manual queue, the staff who remember how. Someone needs the authority to pull the switch, and it helps if doing so has been rehearsed. Projects that retire the old process on launch day have removed their own safety net.

What would a small trial look like before we commit to the full rollout?

Why ask it

A good proposal names one team or region, a fixed period, a comparison group and the result that would count as success. Agree on that result before the trial starts so nobody can redraw it afterward. Resistance to a trial from the people selling the project deserves a direct 'why'.

Accountability

Who is accountable when the model gets one wrong?

Why ask it

You are after a named role inside the organization, not 'the algorithm' and not the vendor's help desk. Where legal responsibility finally sits depends on the contract and on the law where you operate, so put that part to a lawyer. Inside the building, someone should own the outcome before launch, and it helps if they are in the room.

Does a person review its decisions, and do they have the time and the authority to overrule it?

Why ask it

'A human is in the loop' means little if that human approves two hundred cases an hour. The telling figure is how often reviewers disagree with the model today. A rate near zero suggests they have stopped checking, and a very high rate suggests the model is not saving anyone time.

How does someone affected by one of its decisions question it or ask for a second look?

Why ask it

Have it walked through from the customer's or employee's side: where they go, who reads the complaint, how long it takes. Some uses carry a required appeal or notice process and others do not, depending on the sector and the place. If no route exists, mistakes will surface as lost customers and never as feedback.

Will people be told that a model was involved in the decision?

Why ask it

Decide this on purpose. Disclosure rules differ by jurisdiction and by use, so check what applies to you, then try the plainer test of how it would read if customers learned it from a news story. Staff who work alongside the model need telling as well.

What record is kept of what the model decided, with which version, and on what inputs?

Why ask it

When a decision is challenged six months on, you will need to reconstruct it. That takes the inputs, the output, the model version and whether a person changed the result. How long the logs are kept has to fit both your record-keeping duties and your privacy commitments, which can pull in opposite directions.

Who signs off on a new version before it replaces the one in use?

Why ask it

An updated model is a different model, and it can be better on average while worse on the cases you care about. Look for the same tests run again, a comparison with the version being replaced, and an approver who is not the person who built it. With a vendor, the question becomes whether updates reach you automatically and how much notice you get.

Which laws, regulations or company policies apply to using a model for this kind of decision?

Why ask it

Hiring, lending, insurance, health and anything involving children tend to attract specific rules, and those rules differ from one country or state to the next and keep changing. A general assurance from the technical team or a vendor is not confirmation. The use should have been reviewed by your own legal or compliance people, recently enough that the review still describes the system.

What would make you advise us not to go ahead?

Why ask it

Hold this for the end and leave a pause after it. People who know their work can name the condition: too few examples, an error that costs too much, no one to maintain it. 'Nothing' means you are talking to someone selling, and the earlier answers should be weighed with that in mind.

How to question a machine learning project before you back it

Practical guidance for the conversation itself

Before the meeting

Write down the decision you have to make

Approving a budget, signing a contract, letting a model touch real customers and making sense of a lecture each need a different depth. Put the decision in one line at the top of your notes and choose the questions that bear on it. Ten asked properly beat forty read aloud.

Ask for the short written version first

Request a page beforehand covering what the model predicts, what it was trained on, how it was tested and what it will cost to run. Gaps on that page show where to spend the meeting, and the team gets a fair chance to bring real numbers instead of improvising them.

Bring cases of your own

Pick five or six real examples from your side of the business, including a couple of awkward ones, and ask to see what the model does with them. A demo is rehearsed on cases its builders chose. Yours were not chosen by them, which is the reason to bring them.

Invite the person who does the job today

Whoever currently makes the decision the model would make knows the exceptions, the workarounds and the cases that never reach the system. They will spot an unrealistic claim faster than anyone holding a budget, and the rollout will depend on their cooperation.

Matching the questions to the situation

A vendor pitch

Spend most of the time on The data and Running it. Settle early whether the product runs on a model the vendor rents from someone else, because that brings a third party's terms into the deal. The vendor's model learned from other people's records, so the central issue is how it performs on yours, and the surest way to find out is a trial on a sample you supply. Take data use, pricing at higher volume and exit terms into the contract discussion, where the answers get written down.

An internal project asking for money

Start with The problem and do not move on until the simple-rule question has a number attached. Then go to Running it, because internal proposals tend to budget the build and leave the upkeep to whoever is around, and the people whose work changes are often the last to hear. Have the team say in advance what result from a trial would make them stop.

A model that is already live

Accuracy, Bias and failures and Accountability are the groups to use when you have inherited something. Find out when it was last retrained, who approved the version now running, and what has been watched since launch, then read the log of overrides and complaints. A model nobody has re-examined since it went in is a finding in itself.

A student or someone new to the subject

Use the list as a reading frame for a paper, a course project or a news story about a model. Take one question from each group and see whether the source answers it. The ones left unanswered are good to put to a lecturer or a guest speaker, and they are the same ones a manager would raise.

Reading the answers

A score needs three things beside it

What it was measured on, what it is being compared with, and which kind of error it hides. A percentage offered without those is decoration. Ask for each in turn and write the answers next to the number.

Examples outrank summaries

Charts average away the cases you care about. Ask to see individual predictions, right and wrong, and read them the way the eventual user would. A quarter of an hour with real outputs can change a room's opinion more than the slides did.

Plain language is a fair request

If an answer arrives in terms you do not know, say so and ask for it again as it would be told to a customer. People who understand their model can do this. Hearing the same jargon repeated more slowly tells you they cannot, or would prefer not to.

An honest gap is workable

No project has checked everything. 'We have not tested that yet, and here is how we would' is a good answer, and it can become a condition of the funding. A team with no gaps at all to report is the one to worry about.

Where these conversations go wrong

Being won over by the demo

A demonstration shows the model on its best cases, at low volume, with nothing else connected. It answers whether the idea is possible. It does not answer whether the model holds up on your data, at your scale, eight months in.

Arguing about the algorithm

Which technique was chosen is rarely what decides whether a project succeeds, and a non-specialist cannot judge it anyway. The data, the test and the plan for running the model are where the outcome is settled, and all three can be examined without any mathematics.

Funding the build and forgetting the upkeep

Monitoring, retraining, mending broken data feeds and answering complaints go on for as long as the model runs. If the proposal ends at launch, ask for the second year's costs and its owner before approving the first.

Taking a rule from elsewhere as your own

What must be disclosed, recorded, explained or reviewed differs by country, state, sector and sometimes by employer policy. An assurance that satisfied another customer or another office settles nothing for you. Take the specific use to your own legal or compliance people and ask how it works where you are.

Asking everything as an accusation

The aim is to find out, not to catch someone. Teams that feel ambushed answer defensively and bury the gaps you most need to hear about. Say at the start that every proposal gets the same questions.

More on this topic