Word Count in Top 30KB

Word Count in Top 30KB counts qualifying word tokens inside RankGear’s compatibility view of the start of a page — the first 30,000 UTF-16 units of stripped source, with every collected image alt value appended. It is not a visible-prose word count, and some markup tokens survive, so read it only against other factors built from the same corpus.

Factor IDRG-KWD-099
FamilyKeyword Usage & Density
MeasurementCompatibility word count
Measured zoneGoogle-density top-30KB corpus plus image alt text

What it measures

This factor counts qualifying word-pattern tokens in RankGear’s compatibility representation of the beginning of the page. It approximates how much word-like material sits in the early portion of the source that a search engine is most likely to weigh, rather than the total word count of the rendered article.

How RankGear measures it

RankGear collects alt text from the full HTML, strips selected noncontent elements, truncates the remaining source to 30,000 UTF-16 units, appends all collected alt values, and counts tokens matching its 3-to-64-character word pattern. The token count of that assembled corpus is the value RankGear reports.

StageWhat RankGear does
Collect alt textPulls image alt values from the full HTML, including images past the truncation point.
StripRemoves selected noncontent elements from the source.
TruncateKeeps the first 30,000 UTF-16 units of the remaining source.
AppendAdds all collected alt values to the truncated corpus.
CountCounts tokens matching a 3-to-64-character word pattern.

How to optimize it

Provide substantial useful content early in the document and accurate alt text for meaningful images, without rearranging source solely to move this number. Treat the value as an observation about where word-like material falls in the source, not a target to game — padding the early source or stuffing alt attributes does not reliably relate to better position.

Important considerations

  • This is not a visible-prose word count — it counts word-pattern tokens in a compatibility corpus, not words a reader sees.
  • Some markup tokens remain in the corpus after stripping, so the count can include non-prose material.
  • Alt text from beyond the truncation point is appended, so images late in the page still contribute.
  • Compare it only with factors that use the same corpus construction; the number is not interchangeable with a plain content word count.
  • This is a prioritization signal that tracks with position within a measured result set. It does not, on its own, cause a page to rank.

Related factors

Part of the Factors reference · how RankGear measures · glossary.