How We Research and Verify Every Question
A trivia game is only worth playing if the answers are right. This page describes exactly how questions get into GuessWho, what we throw out, and what to do when we make a mistake. It is written to be checkable rather than reassuring.
What is in this guide
Four principles
- Real material only. Every quote, tweet, lyric fragment, event and capital city in the games refers to something that genuinely exists or happened. We do not invent plausible material to fill a dataset.
- One defensible answer. If a question could reasonably be answered two ways, it is rewritten or dropped. Ambiguity is our error, not the player's.
- Verification before volume. When we cannot verify enough material to reach a target number, we ship the smaller number. Our Who Tweeted deck stopped at 165 entries rather than padding to 250 with tweets we could not document, and it will stay at 165 until we verify more.
- Corrections are published, not buried. Dataset fixes are listed in our public release notes with the reason.
What counts as a source
Sources are ranked, and the higher the rank the less corroboration a fact needs.
| Tier | Examples | Treatment |
|---|---|---|
| Primary | The original tweet, the published text of a speech, a film's own dialogue, a government's own designation of its capital, an official record or archive | Sufficient on its own. |
| Contemporaneous reporting | News coverage published at the time, quoting directly | Sufficient where the primary source is unavailable, preferably with a second independent report. |
| Reference works | Established encyclopaedias, quotation dictionaries with citations, standard reference books | Acceptable as corroboration, and as a route to the primary source they cite. |
| Aggregators | Quotation sites, lyric sites, listicles, screenshots, social posts about other social posts | Never accepted as evidence, at any tier. These are the main vector by which errors propagate. |
The general rule is that we chase a claim back to where it starts. If the trail ends at a page copying another page, the claim does not go in.
Rules for each game
Who Tweeted
Each tweet has to be traceable to the original account or to reporting that quotes it directly, with a date. Screenshots are not evidence: fabricated screenshots are trivial to make and circulate widely. Deleted tweets are allowed only where reputable contemporaneous coverage quotes them, since deletion is often itself the newsworthy part. Engagement figures are recorded as reported at the time and are indicative rather than live.
Who Said It
Attribution is the entire question, so it gets the strictest treatment. A quotation must trace to the speaker's own writing, a recorded speech, a contemporaneous report, or a reference work that cites one of those. Where a famous misattribution exists, we use the correct attribution and often place the popular wrong name among the incorrect options. Quotations we cannot pin down stay out, however good they are. Our guide to misattributed quotations explains the method in detail.
Guess the Movie
Lines are checked against the film's actual dialogue rather than quotation sites, which recycle each other's misquotes. Where a well known misquote exists, the real line is the question and the misquote frequently becomes a wrong answer. We use short, identifiable fragments only, always attributed to the film.
Guess the Song
Fragments are short, always attributed to artist and title, and used only as far as identifying a song requires. We do not reproduce lyrics beyond that, in guides or anywhere else, which is why our article on lyric writing quotes none at all. Artist and title are checked against release credits rather than lyric aggregators.
Guess the Year
Questions name the specific step being dated, because invention is a sequence rather than a moment: conception, demonstration, patent, release and mass adoption can span decades. Where the year is genuinely contested among historians, the question is dropped rather than adjudicated by us. The reasoning is set out in our historical timeline guide.
Guess the Capital
We use the capital as officially designated by the country itself, cross checked against United Nations member state listings. Where a country designates more than one capital, the question says so. Where a capital is mid relocation, we use the city currently functioning as the seat of government and revisit it as the situation changes. Territories that are not sovereign states are included and labelled as such.
What we refuse to publish
- Invented material. No fabricated tweets, quotes or events, in any game, for any reason. An earlier version of the Who Tweeted dataset contained four entries that did not survive verification; they were removed and replaced with documented ones.
- Unsourced claims presented as fact. If we cannot show where it comes from, it does not go in.
- Full song lyrics or extended copyrighted text. Short identifying fragments with attribution only.
- Questions about private individuals. Public figures in their public capacity, public statements, published work and documented events only.
- Trick questions. Deliberately misleading phrasing, technically true but misleading options, and near identical decoys are all bad quiz design. Difficulty should come from knowledge, not from parsing.
- Padding. Duplicate questions with reworded stems, or filler entries added to reach a round number.
How a fair four option question is built
The wrong answers matter as much as the right one. A question with three implausible decoys tests nothing, and one with a second defensible answer is simply broken.
- Decoys come from the same category as the answer: real capitals for a capital question, real films of comparable era for a film quote, real people who plausibly might have said it for a quotation.
- Decoys must be clearly wrong once you know the answer. If a decoy requires a judgement call to rule out, it is replaced.
- The correct answer is never distinguishable by form. Not longer, not more specific, not more carefully worded than the others. Answer position is randomised per player session.
- No overlapping options. Two options cannot both be true under different readings of the question.
- Difficulty is spread across the deck rather than tuned per player. Everyone gets the same five questions on a given day, so the mix has to be fair for a newcomer and interesting for a regular.
The automated checks
Every dataset is validated before release. The checks are unglamorous and catch real problems:
- Identifiers are sequential and unique, and existing ones are never reordered or reused.
- Every entry has exactly four options, and the correct answer is one of them.
- No duplicate questions, and no duplicate answer sets within a dataset.
- No entry appears twice under different identifiers, which is how a 2026 audit found 62 duplicated countries in the capitals dataset and replaced them with genuine, distinct ones.
- Metadata fields such as artist, year, continent or source are present and consistent.
Automated checks cannot tell you whether a fact is true. They tell you whether the dataset is structurally sound, which is a precondition rather than a substitute for verification.
Corrections
We get things wrong. When you find one, use the contact page and tell us which game and which question. What happens next:
- We check the claim against sources at the tiers described above.
- If you are right, the dataset is corrected and the fix ships with the next release.
- If the question was ambiguous rather than wrong, it gets rewritten or removed, because ambiguity is a defect either way.
- Substantive corrections are listed in the release notes, with the reason.
This is not hypothetical. Past audits and reader reports have already corrected official capitals recorded wrongly, television dialogue presented as film dialogue, tweets that could not be documented, and duplicate entries that made a dataset smaller than it claimed to be. Every one of those changes is listed in the public release notes.
Independence and funding
GuessWho is an independent project, free to play, funded by advertising. Nobody pays for a question, an answer or a mention, and no commercial relationship influences what appears in a game or a guide. Where a guide names a product, a company or a platform, it is because the subject required it.
Advertising is served by a third party, which sets its own cookies and collects its own data; that is described in the privacy policy. The games themselves need no account, no login and no personal data, and your scores, streaks and settings never leave your browser.
Frequently asked questions
Who writes GuessWho's questions and articles?
A small independent team. There is no newsroom and no crowd submission process: every question is researched, written and checked in house, and every guide on this site is original work written for this site.
Do you use AI to generate questions?
We use software to help draft candidate questions and to check datasets for duplicates and formatting errors. Nothing enters a game on that basis. Every fact has to be verified against a real source by a person before it is published, and material we cannot verify is discarded rather than published with a hedge.
What happens if I find a wrong answer?
Tell us through the contact page with the game and the question. We check it against sources, correct the dataset if you are right, and the fix goes out with the next release. Corrections are listed in our public release notes.
Are the daily questions the same for everyone?
Yes. The day's date seeds the question selection, so every player worldwide gets an identical set. Nothing about your history, device or location changes which questions you see.
Do you collect personal data to run the games?
No. There are no accounts and no logins. Scores, streaks and preferences are stored in your own browser, which is why they do not follow you to another device. The privacy policy sets out what third party services such as analytics and advertising do collect.