Reading Vocabulary Profiler
Compare the words in two texts. Count each word, find words they share and see each match in the original text.
How can I compare the words in two texts?
Add two texts and choose how to count their words. Run the comparison to find shared words, unique words and word counts. Select a word to see it in the original text. Download the full report when you are done.
Two passages. A closer reading.
See which words recur, which are shared and which make each passage distinctive. Counts always link back to the original text.
Passage A
153 / 100,000 characters
Passage B
174 / 100,000 characters
Word-counting rules
Everything runs on your device. Nothing is uploaded or automatically saved. Keep a project backup before leaving.
Three simple steps
How to use Reading Vocabulary Profiler
Add two passages and choose how words should be counted.
Compare vocabulary, filter the word list and inspect original occurrences.
Download the full comparison or save a project to continue later.
Common questions
Good to know
Understand the result.
Keep the original.
Is this tool free, and are my files uploaded?
This tool is free with no account required. The work happens in your browser. Your input and files are not sent to a processing server.
What are the limits and details?
Up to 100,000 UTF-16 characters and 10,000 word-like segments per passage, including ignored words and numbers, with at most 512 characters per segment; UTF-8 passage files up to 400 KB and project backups up to 2 MB. Word boundaries follow the language rules built into your browser. Segmentation and locale support can differ between browser versions. Words are normalized to NFC; optional lowercasing follows the resolved locale. There is no stemming, synonym matching or translation. Ignore up to 200 individual words. When numbers are excluded, a retained segment must contain a Unicode letter. Rates use words remaining after these rules; display filters do not change their denominators. Jaccard compares word-type sets and cosine compares raw word counts. These are vocabulary overlaps, not reading-level or comprehension measures. Processing has a 30-second deadline and can be cancelled. CSV guards formula-like text; exact original text and zero-based UTF-16 source positions remain in JSON/TXT.