Skip to content
Purity Test Kit

Is the Rice Purity Test accurate?

Accurate at counting which of 100 fixed statements you tick, and invalid as a measure of innocence: 60 of those statements are sexual, and every one costs the same point.

5 min readPurity Test Kit

The Rice Purity Test is accurate at exactly one thing: counting how many of 100 fixed statements you are willing to tick. It is not accurate at the thing its name promises, because "purity" is not what the list measures. Read the questions and the mismatch is obvious — 60 of the 100 statements are about sexual experience. Thirteen are about romance, eight about trouble with the law, seven about alcohol and drugs, five about parties, four about mischief, and three about explicit messages and photographs.

So a Rice Purity score is, mostly, a sexual-experience count with a few other things attached. Once you know that, the number becomes useful in a narrow way and stops being useful in the broad way people want.

Reliable and invalid are different failures

Two words get mixed together whenever someone asks whether a test "works".

Reliability asks whether the same person gets the same result twice. The Rice Purity Test does well here, and for a dull reason: it is a memory checklist, not a judgement. "Been arrested?" has an answer you either know or do not. Take it twice a week apart, answer honestly both times, and the score should barely move — the only items that can drift are the vague ones.

Validity asks whether the number measures what it claims. Here the test fails, and not marginally. A 100-item list where 60 items are sexual cannot report on innocence, character or morality; it reports on one region of a life, heavily. A person who has never had sex but has been arrested twice, sold drugs and crashed a car scores higher than a person in a long marriage who has done none of those things. If your intuition says that ordering is wrong, your intuition is correct, and the arithmetic is what produced it.

Every point weighs the same, which is the second problem

Item 1 is holding hands romantically. Item 66 is being arrested. Both cost one point.

That flat weighting is why the score compresses things that are not comparable. Ticking twelve mild items in the romance block costs exactly what ticking twelve items from the sexual block costs, and both cost what twelve arrests would cost if the list had twelve arrests in it. The test has no way to say that one of those runs describes an ordinary adolescence and another describes a decade of consequences.

You can see the effect on any individual score page: the same number is reachable through completely different sets of items. Two people who both score 64 have each ticked 36 statements, and there is no guarantee they share more than a handful.

Nothing is verified, and the errors run both ways

There is no source of truth behind a self-graded checklist. That produces two opposite distortions, and they do not cancel out.

Downward pressure on honesty. Someone taking the test in a shared room, or on a partner's laptop, or fifteen minutes after being asked about their number by a friend, has reasons to leave boxes unticked. Social-desirability bias is the well-documented tendency to answer sensitive questions in the direction that looks better to whoever might see. On a list where most items are sexual, that pressure is not evenly distributed — it lands hardest on the items people are least likely to volunteer.

Upward pressure on drama. The test's second life is a screenshot. A low number is the interesting number, and clicking down a column takes less time than reading it. Any dataset of purity scores contains deliberate zeroes, and they pull an average further than they pull a median — which is the whole reason we withhold figures until 30 results are in and why the percentile explainer spends its time on selection effects rather than on arithmetic.

Neither distortion is measurable from inside the data. A 12 and a joke 12 look identical in a database.

The wording leaves room to disagree with yourself

Some items are precise. "Been arrested?" is precise. Others are not:

  • "Been in a relationship?" — for how long, and does the person you texted for two months count?
  • "Danced without leaving room for Jesus?" — a joke item from a church-dance idiom, which two people will read differently.
  • "Undressed or been undressed by a MPS (Member of Preferred Sex)?" — the abbreviation is glossed once and then reused bare, and copies of the list disagree about the gloss.

Where a statement is ambiguous, the score becomes partly a report on how strictly the taker reads English. That is real measurement error, and it is larger for readers who take instructions literally.

The version you took may not be the version someone else took

Sites copy the list from each other, and copies drift. We sourced ours from four independent live copies and reconciled the six items where they disagreed — including item 69, which some sites print as a bare question mark because the sheet they copied censored it.

That means a comparison between your score here and a friend's score from another site is a comparison between two instruments that may differ by a handful of items. Our four other tests — teen, couples, college and innocence — are deliberately different question sets, and their scores are not comparable with the classic list at all. The comparison page says which is which.

What the number can honestly be used for

Three things, all narrow:

  1. A private inventory. The list is a prompt. Reading 100 specific statements about your own life is a different exercise from thinking about it in the abstract, and the score is a side effect of having done it.
  2. A conversation opener with a specific person. Two people comparing which items they ticked learn something. Two people comparing only their numbers learn almost nothing, because of the compression problem above.
  3. A number to hold loosely. Your score is not a grade, and the ranges are not report cards — they describe how much of a fixed list applied to you.

What it cannot do: measure innocence, predict how anyone will behave, diagnose anything, or rank two lives against each other. Anyone presenting purity scores as research is presenting self-selected internet submissions as research, which is a different claim entirely.

Get your own number

The classic 100-question list scores instantly and shows you where your result sits in the live distribution. No sign-up, nothing stored against your name.