Surface similarity is how closely a candidate phrase resembles the primary keyword at the level of the actual words used — an overlap of spelling and tokens, not of meaning. Two phrases can score high on surface similarity while pointing at different intents, and low while meaning nearly the same thing.
| Term | Surface Similarity |
|---|---|
| Category | Content and Topical Relevance |
| Also known as | Lexical Similarity |
| Where it appears | Variation evidence score |
What it means in RankGear
Surface similarity feeds the Variation evidence score, where RankGear weighs whether a candidate variation belongs to the same target as your primary keyword. It looks only at the literal words: how much the candidate’s tokens overlap the keyword’s tokens. “Best running shoes” against “best running shoe” registers high — the strings differ by a single plural. “Best running shoes” against “top trainers for jogging” registers low, because the words diverge even though the topic is close. This is deliberately the wording view of a variation, kept separate from meaning-based signals so the score can tell lexical closeness apart from semantic closeness rather than blending the two.
How to interpret it
Read surface similarity as a comparative closeness measure, not a quality verdict. A high value means the candidate is nearly the same string as the keyword — the metric is good at catching plurals, stopword swaps, and reordered words, but it says nothing about whether the variation adds any coverage. A low value is not a failure signal; a synonym or paraphrase can carry the same search intent while sharing almost no words. Its real use is to separate cosmetic variants from genuinely different phrasings, so always read it next to an intent signal — otherwise a same-words, different-intent pair looks equivalent when it is not.
| Surface similarity | What it tells you |
|---|---|
| High | Near-duplicate wording — plurals, stopwords, or word order differ. Cosmetic; rarely adds coverage on its own. |
| Mid-range | Shared core terms with real additions or substitutions; worth checking whether intent still matches. |
| Low | Few shared words. Often a paraphrase or synonym that can broaden a page without repeating the keyword. |
Example
Take a primary keyword of “commercial espresso machine.” The candidate “commercial espresso machines” scores high surface similarity — only a plural separates them — so it is a cosmetic variant that adds little. “Restaurant coffee equipment” scores low despite overlapping buying intent, because it shares no words with the keyword. “Commercial espresso machine reviews” scores high as well, since it keeps the whole phrase and appends one token; yet its intent has shifted toward comparison rather than purchase. Reading surface similarity beside intent is what stops RankGear from quietly folding that review query into the buying query on wording alone.
Important considerations
- Surface similarity counts shared wording only. It cannot distinguish a genuine match from a same-words, different-intent pair, so pair it with an intent signal before treating two phrases as interchangeable.
- Near-duplicates that differ only by a plural, a stopword, or word order score high but seldom add real topical coverage — don’t chase them for the number’s sake.
- Low surface similarity often marks the most useful variations: synonyms and paraphrases that widen a page’s coverage without stuffing the exact keyword. Preserve intent when adding them; repetition and forced wording do not build topical authority.
- It is a comparative wording indicator inside RankGear’s variation evidence, not a Google ranking score. A higher figure does not make a page rank — statistical association is not causation, and SERP-derived samples can be noisy.
Related terms
Part of the RankGear glossary · how RankGear measures · the 870 factors.