Free online tool

Text Classification Sandbox

Teach a small model to label text using your examples. Check it on separate test text and download the predictions and results.

Runs on a separate serverNo account neededCSV / JSON / ZIPOpen workspace ↓

How can I train a model to label text?

Add examples with labels, and mark which ones to use for training or testing. Train the model, then check how it handles the separate test examples and new text. Review the mistakes before using its labels. Download the report and examples together.

Teach it with your examples.

Label a small collection, reserve separate examples for testing, then try new passages. See the model’s scores and the mistakes it makes.

Import a collection from CSV

Use exactly these headers: id,label,role,text. Set role to train or test. Test labels must also appear in training. Quoted multiline text is supported. Review five sample rows before replacing the collection.

Your passages

18 matching passages · Page 1 of 1

IDLabelRoleText previewEdit
garden-1gardentrainPlant tomato seeds in warm garden soil.
garden-2gardentrainWater the flowers and pull weeds from the garden.
garden-3gardentrainThe gardener grows fresh vegetables near the greenhouse.
garden-4gardentrainAdd compost to improve soil around the plants.
garden-5gardentrainRose bushes need sunlight and regular watering.
garden-6gardentestGarden tools help plant seeds and trim flowers.
cooking-1cookingtrainBake the bread in a hot kitchen oven.
cooking-2cookingtrainCook pasta in boiling water and add tomato sauce.
cooking-3cookingtrainThis recipe combines fresh vegetables and olive oil.
cooking-4cookingtrainThe chef prepares soup with salt and herbs.
cooking-5cookingtrainMix flour and butter to make pastry dough.
cooking-6cookingtestKitchen recipes explain how to roast and bake food.
travel-1traveltrainBook a flight and pack a suitcase for the trip.
travel-2traveltrainThe train journey crosses mountains and small towns.
travel-3traveltrainTourists can reserve a hotel near the airport.
travel-4traveltrainA passport is needed for international travel.
travel-5traveltrainThe bus station has tickets for the coastal route.
travel-6traveltestA travel guide helps visitors plan their holiday.

Choose how to compare text

Words use Unicode letters, digits and underscores, including single-character tokens. Word pairs also include adjacent two-word phrases. Character fragments use 3–5 characters inside whitespace-separated words, with boundary padding. There is no stemming or semantic language model. English stop words only match lower-case forms when Match case is selected.

Running sends all passages and settings to Bookify’s isolated processing service. Temporary job files are cleared after processing; nothing is automatically saved in this page.

Three simple steps

How to use Text Classification Sandbox

  1. Edit the examples or import a CSV and review its rows.

  2. Choose a method and send the experiment to the processing service.

  3. Inspect the results, then download the full analysis and reusable project.

Common questions

Good to know

Understand the result.
Keep the original.

Is this tool free, and are my files uploaded?

This tool is free with no account required. When you run the tool, your input goes to a separate processing service. Temporary files are deleted when the job ends. You can download the result to your device.

What are the limits and details?

Text is uploaded to an isolated processing worker using pinned scikit-learn 1.7.2. Up to 200,000 characters across all passages, 8,000 characters per passage, 80-character unique IDs and 10,000 retained vocabulary features. Projects must fit within 500,000 encoded JSON characters; imports are UTF-8 files up to 2 MB. Processing stops after two minutes. Words use Unicode letters, digits and underscores; character fragments use 3–5 characters inside whitespace-separated words with boundary padding. No stemming, translation or semantic model is used. The vocabulary cap retains the most frequent features. English stop words are optional and match only lower-case forms when case matching is enabled. Nothing is automatically saved. JSON preserves exact source text; CSV prefixes formula-like strings and spreadsheets can reinterpret numeric-looking text. Training requires 2–20 labels. Test labels must appear in training; repeated example text is rejected after Unicode, case and whitespace normalization. Related paraphrases can still bias the evaluation. No automatic split or cross-validation is performed. Scores are uncalibrated model outputs, not verified probabilities. Text matching no training features remains unmatched. Logistic regression uses C=1 and at most 500 iterations. Projects preserve source and settings; models are rebuilt rather than exported as executable files.