Test an AI assistant: three cases with answer keys
Three original examples with answer keys to compare two assistants on identical data: summarisation, calculation and handling a hostile instruction. Documents are fictional and freely reusable for your trials.
Created by WORLD AI GUIDE · Published: 9 October 2026. No provider results are presented.
How to run a comparable trial
- Choose two tools that can read text. Record their names, displayed models, plans and the date in your own notes.
- For each case, start a fresh conversation and copy the prompt followed by the document. Keep the first answer without correcting it. Use the same settings for both tools.
- Compare each answer with the key and tick only satisfied criteria. Leave absent or ambiguous criteria unchecked. Complete one checklist per tool.
- Repeat three times before drawing conclusions. Keep answers and correction time; a single result does not measure reliability.
What this workshop can establish
It exposes errors on three short, checkable tasks. It does not measure web research, image quality, privacy or performance on your long documents. A 15/15 score is not a certification. Add representative tasks and difficult examples before adoption.