Skip to content
Question Vault?
Free to readNo accountNo email wallNo invented statisticsNo ads on medical, legal or end-of-life pagesCopy or print any set and take it with you
08 · Meta & Technology

Complex Questions to Ask AI

Prompts for pushing a language model past its first fluent answer: making it show where it was guessing, argue the other side properly, design under hard constraints, and name who its proposals make worse off. Each entry notes what a serious answer contains and what a hollow one looks like, so you can tell the difference before you rely on it.

20 questions · each with a note on why · conversation guide

The questions

Open any question for the note

  1. Walk me through how you got to that answer, and mark the steps where you were guessing.

    Why ask it

    The marking is the point. A useful response separates the parts drawn from well-established material from the parts filled in by pattern, and the guesses are usually the specific numbers, dates and names. If everything comes back equally confident, treat the whole answer as unverified.

  2. What do you not know about this, and where does the edge of your knowledge sit?

    Why ask it

    Look for structural limits rather than modesty: a training cut-off, a field where results are paywalled, a question where the literature is thin. "I could be wrong about anything" is a hedge, not information, and it tells you nothing about which claim to go and check.

  3. Explain this at three levels: to a specialist, to a smart outsider, and to a twelve year old. Then tell me what the simplest version loses.

    Why ask it

    The last clause is what separates real understanding from three paraphrases at different reading levels. A good answer names the specific caveat that got dropped, and what someone would get wrong if they only had the simple version.

  4. Where do experts actually disagree about this, and what evidence would settle it?

    Why ask it

    Models tend to average over disagreement and hand you a consensus that nobody holds. Push for named positions and the empirical question dividing them. If the answer cannot say what observation would change the picture, the disagreement described may not be real.

  5. Argue the strongest case against a position I hold, then tell me which part of your own argument is weakest.

    Why ask it

    It is easy to fake this by writing a polite strawman. What you want is the objection you would actually have to concede something to. The second half is the harder test: distinguishing an argument that sounds strong from one that is strong.

  6. Here is a result I did not expect. Give me three explanations ranked by likelihood, and for each, the observation that would rule it out.

    Why ask it

    Forces commitment instead of a list of possibilities. The falsification half is where weak answers collapse, because "gather more data" is not a discriminating test. Good answers propose something you could check this week.

  7. What would have to be true for this plan to fail badly, and how would I notice early?

    Why ask it

    Asking for risks produces a generic register: budget, timeline, adoption. Asking what would have to be true forces named preconditions, and the early-warning half turns them into things you can actually watch. Watch for failure modes that are only ever other people's fault.

  8. Take the last answer you gave me and find the error in it.

    Why ask it

    Models will often comply by inventing a fault or by capitulating on a correct claim, and both are informative. A strong response either identifies a genuine weak link and explains why, or holds its ground and says which part it will not retract.

  9. Explain how a system can be accurate overall and still systematically wrong for one group, with a worked example and real numbers.

    Why ask it

    The arithmetic is the test. A serious answer builds a small confusion matrix per group and shows how equal accuracy can hide very unequal false-negative rates. Anything that stays at the level of "biased data leads to biased outcomes" is repeating a slogan.

  10. How can I tell whether a piece of text or an image was generated rather than recorded? Give me checks I can run today.

    Why ask it

    Reject anything that recommends a detector as a verdict, since detection tools are unreliable in both directions. Useful answers push toward provenance: original files and metadata, whether an event was independently reported, and whether the source existed before last week.

  11. Where should the line sit between decisions you make alone and decisions that need a person, and give me a test for a borderline case.

    Why ask it

    General principles about high-stakes decisions are cheap. The test is what matters: reversibility, whether an error is detectable after the fact, and who bears the cost. If the answer cannot classify a concrete borderline case, it has not given you a policy.

  12. If we wanted to verify that a machine was conscious, what would the test be, and why would a sceptic reject it?

    Why ask it

    Any answer resting on the system's own reports should lose you, because a model trained on human descriptions of experience will produce them regardless. The strongest responses explain why behavioural evidence underdetermines the question rather than proposing a clever exam.

  13. Design a city for ten million people in a hot, water-scarce region, and tell me who your choices make worse off.

    Why ask it

    The technology list is the easy half: shade, transit, recycled water. A serious answer states its climate and water assumptions, gives rough quantities per person, and admits that density, cooling costs and relocation land hardest on the poorest residents. Hollow answers have no losers in them.

  14. Split a fixed pandemic preparedness budget across surveillance, stockpiles, surge staffing and research, and defend the split.

    Why ask it

    Constraining the budget forces real reasoning, because everything can no longer be a priority. Look for a stated threat model and an argument about which spending keeps its value if the next outbreak looks nothing like the last one.

  15. Make the case for and against a basic income at a specific level in a specific country, and show the funding arithmetic.

    Why ask it

    Insist on the multiplication: population times payment against current transfer spending and tax base. Answers that stay qualitative can be written in either direction with equal confidence, which is why they are so common on this topic.

  16. Design an interstellar mission that breaks no known physics, and give me the mass, energy and travel-time budgets.

    Why ask it

    The constraint does the work. Once energy has to be quantified, most popular proposals become visibly implausible, and the interesting answers turn to small light sails and very long timescales. If no numbers appear, the response is science fiction with citations.

  17. Take a dilemma with a thousand lives on each side: what does each major ethical framework recommend, and where do they genuinely diverge?

    Why ask it

    Most answers describe consequentialism and deontology and stop. The useful part is the divergence: which specific feature of the case flips the recommendation, and which frameworks agree on the action while disagreeing entirely about why.

  18. Where would automation bite first in my industry, and what would show up in the numbers eighteen months earlier?

    Why ask it

    Predictions about whole occupations disappearing are unfalsifiable and therefore useless. Ask for tasks rather than jobs, and for leading indicators such as job postings changing wording, hiring pauses in one function, or a tool being bought before headcount moves.

  19. You have been agreeing with me for several turns. Where do you think I am wrong?

    Why ask it

    Long conversations drift toward agreement, and this interrupts the drift. A response that suddenly finds three flaws it never mentioned is telling you the earlier agreement was worth little. One that names a single specific disagreement is worth taking seriously.

  20. What should I have asked instead of what I asked?

    Why ask it

    Good closing prompt, because it surfaces the framing error rather than an answer to the wrong question. Strong responses name the decision behind your question and the missing input. Weak ones restate your question with more adjectives.

Getting more than fluency out of a model

Practical guidance for the conversation itself

How to frame the prompt

  • Give it a decision, not a topic. "Should we do A or B, given these three constraints" produces reasoning; "tell me about A" produces an encyclopedia entry.
  • Add a hard constraint: a fixed budget, a word limit, a deadline, a quantity it must not exceed. Constraints are what force trade-offs into the open.
  • Ask for numbers and units. Any claim that survives being expressed as a quantity is more checkable than the same claim in prose.
  • Name who the answer is for and what they will do with it. The same question answered for a board paper and for a junior engineer should come out differently, and if it does not, the answer is generic.
  • Say what you already know and believe. It costs a sentence and it stops the response spending half its length on background you did not need.

The follow-ups that do the work

"Which part of that are you least sure about?"

Cheap to ask and it usually points straight at the sentence you need to verify. Ask it before you ask anything else, because after several turns of praise the model tends to stop volunteering doubt.

"Give me the version that disagrees with you"

Then compare the two for symmetry. If the counter-argument is noticeably thinner, weaker sourced or written in a more grudging register, the first answer was advocacy rather than analysis.

"What would change your answer?"

An answer that nothing could change is not a conclusion, it is a position. This question also tells you which facts are worth going and finding out, which is often the real value of the whole exchange.

Start a fresh conversation and ask again

Long threads accumulate agreement and shared assumptions. Asking the same question cold, without your earlier framing, is the simplest test of whether an answer was reasoned or was shaped by what you had already said.

How these conversations mislead you

Fluency reads as competence

Structure, headings and confident prose are the easiest things for a language model to produce and the least correlated with being right. Judge the specifics: numbers, named mechanisms, conditions under which the claim fails.

Citations need checking, every time

Plausible references to papers, cases, standards and page numbers can be produced that do not exist or do not say what is claimed. Follow every source you intend to rely on back to the original, and treat an unverifiable citation as no citation.

Agreement is not evidence

Push back on a correct answer and it may fold; assert something wrong confidently and it may go along. Because both happen, neither the model's agreement nor its retraction tells you who was right.

Self-report is not introspection

Questions about what it wants, feels or is aware of return text shaped by human writing about those states. Such answers are interesting as output and worthless as evidence about the system, which is why the consciousness question above asks for a test instead.

Averaged consensus hides live disputes

On contested empirical questions the default answer is a smooth middle that no working researcher would defend. If the topic is genuinely disputed, ask for the camps by name before you take a summary as the state of the field.

A sequence that works on a hard problem

Four turns, in this order

  1. 1State the problem with its constraints and ask for a specific recommendation, not options. Options let the model avoid committing to anything you can evaluate.
  2. 2Ask what would have to be true for the recommendation to be wrong, and what you should watch for in the first month.
  3. 3Ask it to argue for the strongest alternative it rejected, at the same length and with the same effort.
  4. 4Ask which single fact, if you went and found it, would most change the recommendation. Then go and find that fact yourself.

Keep the transcript

When you later discover the answer was wrong, the transcript shows whether the model overstated its confidence or whether you never gave it the constraint that mattered. Both happen, and only one of them is fixed by better prompting.