Skip to content

Documentation

Lexical Diversity

Measure vocabulary variation and compare drafts with similar reference text.

Open Lexical Diversity

Lexical Diversity, the PirateSERP vocabulary-variation tool, scores how varied the words in a text are. It needs no API key. The scores describe vocabulary range only; they do not measure factual quality or predict rankings.

Analyze text

  1. Paste your draft into Text. The word count beside the field updates as you type.
  2. Choose Language. The project's default language is selected when you open the tool.
  3. Set Lemmatize and the MATTR window slider.
  4. Click Analyze.

Lexical Diversity accepts up to 200,000 characters of text. MTLD needs about 100 words or more for a stable score; a shorter text shows a warning that the score is indicative, not comparable.

Read the scores

The result opens with four counts: Tokens (total words), Types (distinct words), Sentences, and the counting mode, Lemmatized or Raw word forms.

The measure of textual lexical diversity (MTLD) is the headline score. A gauge places it in one of four bands: Low (under 40), Moderate (40 to 70), High (70 to 100), or Very high (over 100).

GroupMeasureHow to read it
Length-robustHD-D (hypergeometric distribution diversity)The probability that each word appears in a fixed-size random sample of the text. Very short texts show no HD-D score.
Length-robustMATTR (moving-average type-token ratio)The average share of distinct words across a sliding window. The label shows the window size, for example MATTR (w=50).
Length-sensitiveTTR (type-token ratio)Types divided by tokens.
Length-sensitiveRoot TTR and Log TTRGuiraud's R and Herdan's C, two length-adjusted forms of TTR.
Length-sensitiveMaasLower values indicate greater diversity.
Length-sensitiveMSTTR (mean segmental type-token ratio)The average TTR across equal text segments.

Length-sensitive scores fall as a text gets longer, so compare them only between texts of similar length. Top content words lists the most frequent content words and their counts.

Compare passages in the same language, topic, and format. Keep precise terminology; replacing a correct term with a synonym to raise a score weakens the text.

Choose word counting

For English text, Lemmatize counts inflected forms of a word as one word: "run," "running," and "ran" count once, and so do forms of "be" and "have" such as "is," "was," "has," and "had." Turn Lemmatize off to count each surface form separately.

Lemmatize is English only. Other languages always count surface forms, so a lemmatized English score and a score in another language are not directly comparable. Read Languages for other language limits.

The MATTR window slider sets the window from 20 to 100 words in steps of 5, with 50 as the default. Use the same window when you compare two texts.

Compare a Content Editor draft

Content Editor runs the same scores against your reference pages.

  1. Open a draft in Content Editor.
  2. Open Write & Review, then Diversity.
  3. Click Score my draft.

Reference median MTLD shows the median MTLD of your reference articles. With an even number of references, the median is the average of the two middle scores. The panel states whether your draft's MTLD is at or above, or below, that median, then lists each reference article with its word count and MTLD.

If reference scores are missing, the references were analyzed before diversity scoring existed; analyze the reference pages again to score them. A higher score alone does not make one draft better than another.