Why the default output is definitions
A definition question can be generated from a single sentence of source text. An application question requires inventing a scenario, deciding which principle it tests, and constructing plausible distractors — considerably more work, and nothing in "give me ten practice questions" asks for it. Models optimise for the request as stated.
The fix is to state the level explicitly. Once you name it, output quality changes immediately and dramatically, which is why this is the highest-value single change to your prompt.
The four things to specify
| Specify | Weak version | Strong version |
|---|---|---|
| Cognitive level | "Practice questions on enzyme kinetics" | "Application and analysis questions — give a scenario and ask what happens, not what a term means." |
| Format and marks | "Some questions" | "Four short-answer worth 5 marks each and one 15-mark essay question, in the style of a UK finals paper." |
| Source | (nothing — general knowledge) | "From the attached lecture slides only. If the material doesn't cover something, don't ask about it." |
| Marking | "With answers" | "With a mark scheme showing where each mark is awarded, and the most common wrong answer with why it's wrong." |
The fourth is the one people skip and it's worth as much as the other three. A mark scheme tells you what earns credit, which is different from a model answer — and "the most common wrong answer, and why" produces the distractor analysis that makes a question diagnostic rather than just a test.
Prompts that work
- 1
Anchor to your material, not to the topic
"Using only the attached lecture notes, generate…" A question generated from general knowledge tests the internet's version of the subject; your exam tests your lecturer's. Where those diverge is precisely where the marks are.
- 2
Name the exam and the level
"Second-year undergraduate pharmacology, in the style of a single-best-answer paper with five options and one clearly best answer." Specificity here is worth more than any amount of prompt cleverness.
- 3
Ask for a difficulty spread explicitly
"Three that most students would get right, four mid-difficulty, three that would separate a first from a 2:1." Otherwise you get a uniform middle, and the hard ones are the useful ones.
- 4
Demand a scenario stem for application questions
"Each question must begin with a specific case or dataset, not with 'which of the following'." This single instruction is most of the difference between exam-shaped and quiz-shaped questions.
- 5
Ask for the mark scheme separately, after you've attempted them
Generate the questions, close the answers, attempt them cold, then ask for marking. Seeing the answers alongside the questions destroys the retrieval that makes the whole exercise worth doing.
- 6
Then ask it to critique its own questions
"Which of these could be answered without knowing the material, by elimination or grammar?" Models are surprisingly good at catching this, and it removes the free questions that inflate your score.
Verifying the question bank
A wrong practice question is worse than no question, because you'll learn the wrong answer confidently and it will feel like knowledge. The verification burden is real and it's the reason ungrounded generation is risky for anything you'll be examined on.
- Check the answer against your source, not against the model. Asking "are you sure?" produces agreeableness, not verification.
- Be suspicious of numbers. Specific values, thresholds and dosages are where generation goes wrong most often and where the error is least visible.
- Watch for questions your syllabus doesn't cover. They're not harmful, but they consume revision time on material you won't be asked about.
- Check the distractors are actually wrong. A multiple-choice question with two defensible answers teaches you to doubt correct reasoning.
- Prefer generation with citations. If each question links to the page it came from, verification is a glance — see AI study tools that cite sources.
Question types worth asking for by name
| Ask for | What it trains |
|---|---|
| "Given this scenario, what happens and why?" | Application — the bulk of most exams' marks. |
| "Compare X and Y, and state when you'd choose each" | Discrimination between confusable concepts, which is what examiners target. |
| "What would happen if this step failed?" | Mechanism understanding, and it doubles as pathology in medical subjects. |
| "Here is a wrong answer — identify the error" | Error detection, which transfers directly to checking your own work. |
| "Which piece of evidence would change this conclusion?" | Evaluation, the top band in essay subjects. |
Use the questions properly once you have them
Generating questions is not studying. Attempt them cold, write full answers rather than thinking "I know this", mark them honestly against the scheme, and categorise every lost mark — knowledge, technique, or misreading. The categorisation is what tells you how to spend tomorrow.
Then re-attempt the ones you got wrong days later, not immediately. Immediate re-attempts test short-term memory of the correction; spaced re-attempts test whether you learned anything. That's spaced repetition applied to questions rather than to facts, and it's how a generated bank turns into retention.