Jaccard Similarity Index Calculator

Written by Thierno Sadou Diallo, formula verified per our methodology • Last checked on 9/9/2026

The Jaccard index is calculated with intersection ÷ union, where union = size A + size B − intersection. For two samples of 10 and 15 species sharing 5 common species, the union is 20 and the Jaccard index is 0.25, or 25% similarity.

Explanation

The Jaccard similarity index measures the degree of overlap between two sets — typically, in ecology, two community surveys (the species present at two different sites, or at the same site on two different dates). It's calculated by dividing the size of the intersection of the two sets (the number of elements common to both) by the size of their union (the total number of distinct elements present in either one, without double-counting the common elements). The result ranges from 0 (no common elements, the two samples are completely disjoint) to 1 (the two samples are identical). This index differs from the diversity indices already covered on this site, like the Shannon diversity index or the reasoning behind Hardy-Weinberg equilibrium: those measure diversity or allele balance WITHIN a single population, while the Jaccard index measures similarity BETWEEN two distinct samples — a complementary question, not a competing one. In practice, the Jaccard index is used to compare the species composition of two study sites, to track how a single community changes over time, or more broadly in data science to compare two sets of categories (keywords, genetic motifs, user preferences) beyond the ecological context alone.

Example: two samples of 10 and 15 species, 5 in common

Inputs

Sample A: 10 species. Sample B: 15 species. Common species: 5.

Calculation

Union = 10 + 15 − 5 = 20. Jaccard index = 5 ÷ 20 = 0.25, or 25%.

Result

The two samples share 25% similarity according to the Jaccard index.

Frequently asked questions

Why divide by the union rather than the size of a single sample?

Dividing by the union guarantees a symmetric result (the order of the two samples doesn't matter) bounded between 0 and 1, regardless of the two samples' relative sizes. Dividing by the size of a single sample would give a different result depending on which of the two samples serves as the reference, making the comparison ambiguous.

What is the difference with the Sørensen-Dice index, a related measure?

The Sørensen-Dice index (2 × intersection ÷ (size A + size B)) measures a very similar idea but gives double weight to common elements in its calculation, which systematically produces a slightly higher value than the Jaccard index for the same data. Both indices are used in ecology, with the choice between them often depending on field convention or the reference publication being followed.

Does the Jaccard index apply only to ecology?

No, it's a set-theory measure used in many fields beyond ecology: comparing text documents (sets of words), recommendation systems (sets of user preferences), bioinformatics (sets of genetic motifs), or duplicate detection in large datasets.

Related resources

Similar calculators