How the numbers work
Every ranking on this site is one line of arithmetic, printed on the page that uses it. This is that line, and the reasoning behind it.
The problem this site exists to fix
Consider two strains and one effect:
Sorted by percentage, Strain A wins. It should not. Seven people out of eleven is what noise looks like -- 64% there is not a measurement of Strain A, it is barely a measurement at all. Sort by effect on any strain site and watch the top of the list fill with strains nobody has heard of, each with a handful of reviews.
What we do instead
Each rate is pulled toward the database-wide average for that effect, by an amount that depends on how little evidence stands behind it:
With a database average of 30%, the two strains above become
36% and 52%, and the order is right. Strain B's
estimate barely moved, because 8,431 reports are plenty. Strain A's fell
most of the way back to average, because eleven reports are not.
Two things this buys, for free
A strain nobody has reviewed scores as average, not as zero.
Put n = 0 into that formula and it returns exactly
p0. This matters most for the effects people sort to avoid: in
a naive model, a strain with no data has "0% anxiety" and wins the
low-anxiety sort outright. Absence of evidence must never read as good
news, and here it cannot.
There is no special case. One report and ten thousand go through the same line, so there is no threshold to get wrong.
Turning that into an order
Your weights and the adjusted rates make the score:
Each term measures a strain against a typical strain rather than against zero. A score of 0 means average on everything you asked about; positive means better than typical for your weights. Because an effect with no data returns exactly the average, its term is exactly zero -- so a strain can never climb your ranking by being under-researched.
There is no 0-100 match percentage, because there is no ceiling to measure against: the maximum depends on weights you picked seconds ago. The table ranks, shows each term, and lets you check it.
Confidence
Every pooled rate carries a tier, derived from the data rather than assigned:
- High -- at least 500 reports from at least two independent sources.
- Moderate -- at least 100 reports.
- Limited -- at least 25 reports.
- Insufficient -- fewer than that. The rate is never shown as a standalone figure.
Two independent sources are required for the top tier because sample size alone cannot catch a methodology problem. Forty thousand reports from one site with one checkbox layout are forty thousand reports of that checkbox layout.
What this does not fix
Bias. Shrinking estimates fixes variance, not slant. If a source's reviewers skew toward people who liked a strain enough to write about it, more reports buy a more precise estimate of a biased quantity. Nothing in the arithmetic detects that, and this site does not pretend otherwise.
Double counting. Sources are pooled by summing their counts, which weights each by its sample size. If two sources syndicate the same underlying reviews, that counts them twice, and we cannot see it from the outside. So the number of distinct sources is shown beside every pooled figure, and each source's own numbers stay visible on the strain page.
Correlated effects. Energetic and sleepy are close to opposites. Weighting both is expressing roughly one preference twice, and it counts twice. Correcting for that properly would need a covariance matrix and would make the score impossible to print. Printing the score is worth more.
The name on the jar. This is the largest limitation and no statistic touches it. Strain names are not standardised, not certified, and not enforced. Two jars labelled with the same name, from different growers, can differ more from each other than from a third strain entirely. The chemistry section shows ranges across samples rather than a single figure for exactly this reason, and a range is the honest shape of that data.
Ranking on chemistry
Laboratory measurements are not percentages of people, so they get their own version of the same idea. Each strain's median is pulled toward the typical strain by how little evidence supports it:
Here m is not chosen. It is calculated, separately for each
ingredient, as the ratio of how much samples of one strain scatter
to how much strains differ from each other. That ratio is the
honest answer to "how many lab results before I believe this strain is
really different from average".
What it returned is worth stating, because it was not what we expected.
For THC it is 0.73; for limonene 0.66; for
caryophyllene 1.02. All far below the figure the same
calculation gives for self-reported effects. In plain terms: a
laboratory measuring the same strain twice agrees far more than two people
describing it do, so chemistry needs comparatively little
correction, and a handful of lab results genuinely does pin a strain down.
The correction still matters at the extreme, which is where it was always
needed: three samples no longer outrank five hundred.
Why every ingredient is scored in standard deviations
THC runs around 18%, limonene around 0.2%. A weighted total over raw percentages would be a THC ranking with a rounding error attached, and a request for more limonene would never visibly change the order. So each contribution is expressed as how far the strain sits from typical, measured in the spread between strains:
A term then reads as "this strain sits 0.8 standard deviations above a typical strain on limonene", which is comparable across ingredients and printable on the page. An ingredient never measured for a strain returns exactly the typical value, so its term is exactly zero and no strain can climb a ranking by being under-tested.
The published range is the middle 80%, not the extremes
Raw minimum and maximum are useless here, and the reason is instructive. Blue Dream's 1,793 results run from 0.4% to 32.7% THC — but six of those samples are under 5%, which is not flower anyone sold as Blue Dream, and one is over 30%. Seven bad rows out of 1,793 would define both ends of the published figure. The 10th to 90th percentile for the same strain is 14.4% to 22.2%. Raw extremes are kept as a way to spot bad data, never as the range shown to a reader.
What the research says about strain names
The largest limitation on this site is not statistical, and no amount of arithmetic touches it: a strain name is not a regulated, certified or standardised thing. Varietal names cannot be trademarked in the United States, federal trademark protection is unavailable for cannabis goods anyway, and plant variety protection reaches only hemp. There is no registry and no authority.
The published research is consistent about what follows from that:
- Indica and sativa labels do not track genetics. Watts et al. (2021) genotyped 137 samples at 116,296 markers and concluded that "Sativa–Indica labels thus do not accurately reflect genetic relatedness". Nature Plants, doi:10.1038/s41477-021-01003-y
- The same name is often not the same plant. The same study found samples sharing a name were frequently as chemically and genetically distant from each other as samples with different names. Schwabe & McGlaughlin (2019) found the same across 30 strain names in three states — and also found samples with different names that were genetically identical. doi:10.1186/s42238-019-0001-1
- Name variety overstates real variety. Reimann-Philipp et al. (2020) analysed 2,662 samples and found 396 breeder-reported strain names collapsing onto three distinct chemical profiles. doi:10.1089/can.2018.0063
- Even the laboratory is a variable. Jikomes & Zoorob (2020), across 215,285 Washington State test results, found the same chemotype reported at a median of 17.6% to 23.1% THC depending on which lab ran it. doi:10.1038/s41598-020-69680-x
This site shows ranges rather than single numbers because of that last point, and marks its classification labels as "how it is marketed" because of the first. None of it makes strain names useless — they are how people shop, and comparing them is the point of this tool. It does mean a number here describes what was measured in samples carrying that name, which is a weaker and more honest claim than it first appears.
Where the numbers come from
Every effect rate, every lab figure and every claim about lineage is stored with a source, a URL, a sample size and the date it was verified. Anything missing one of those fails the build and does not appear. The status page reports how much of the database is thinly evidenced, because a site that hides its own weak spots is asking to be taken on faith.