Free online tool

Document Similarity Finder

Compare the words in up to 60 short documents. Rank similar pairs, see what they share and download every comparison.

Runs on a separate serverNo account neededCSV / JSON / ZIPOpen workspace ↓

How do I compare several documents for similar wording?

Add 2 to 60 short texts and choose a comparison method. Run it to rank pairs by shared words or character patterns. Open a pair to inspect both texts, then download all scores. Similar wording does not prove copying or shared meaning.

Find the passages that overlap.

Compare a collection by the words or character fragments it shares. Inspect each pair and keep the complete comparison matrix.

Import a collection from CSV

Use exactly these headers: id,text. Quoted multiline text is supported. Review five sample rows before replacing the collection.

Your passages

4 matching passages · Page 1 of 1

IDText previewEdit
garden-guidePlant tomato seeds in warm garden soil. Water the garden plants regularly.
growing-tomatoesWater tomato plants regularly. Plant seeds in warm soil in the garden.
bread-recipeMix flour and butter, then bake the bread in a hot kitchen oven.
train-journeyBook train tickets and pack a suitcase for a journey through the mountains.

Choose how to compare text

Words use Unicode letters, digits and underscores, including single-character tokens. Word pairs also include adjacent two-word phrases. Character fragments use 3–5 characters inside whitespace-separated words, with boundary padding. There is no stemming or semantic language model. English stop words only match lower-case forms when Match case is selected.

Running sends all passages and settings to Bookify’s isolated processing service. Temporary job files are cleared after processing; nothing is automatically saved in this page.

Three simple steps

How to use Document Similarity Finder

  1. Edit the examples or import a CSV and review its rows.

  2. Choose a method and send the experiment to the processing service.

  3. Inspect the results, then download the full analysis and reusable project.

Common questions

Good to know

Understand the result.
Keep the original.

Is this tool free, and are my files uploaded?

This tool is free with no account required. When you run the tool, your input goes to a separate processing service. Temporary files are deleted when the job ends. You can download the result to your device.

What are the limits and details?

Text is uploaded to an isolated processing worker using pinned scikit-learn 1.7.2. Up to 200,000 characters across all passages, 8,000 characters per passage, 80-character unique IDs and 10,000 retained vocabulary features. Projects must fit within 500,000 encoded JSON characters; imports are UTF-8 files up to 2 MB. Processing stops after two minutes. Words use Unicode letters, digits and underscores; character fragments use 3–5 characters inside whitespace-separated words with boundary padding. No stemming, translation or semantic model is used. The vocabulary cap retains the most frequent features. English stop words are optional and match only lower-case forms when case matching is enabled. Nothing is automatically saved. JSON preserves exact source text; CSV prefixes formula-like strings and spreadsheets can reinterpret numeric-looking text. All pairs among 2–60 documents are compared, up to 1,770 pairs. Features and TF-IDF weights are fitted across the whole collection; adding documents can change scores. Similarity ranges from 0 to 1 and distance is 1 minus similarity. A document with no usable features has undefined scores, including against itself. An entirely empty vocabulary is rejected. Filters never remove pairs from the export.