Snehal Patel

Snehal Patel

I love to build things ✨

From Bag-of-Words to Jev: How Text Classification Got Here

The first question I got after the Jev post was a fair one: “Isn’t this just a glorified classifier?” Yes, and that’s the interesting part. Text classification is one of the oldest problems in NLP. Spam filters were doing it in 1998. So why did a classifier with a nice API become the most talked-about model release of the month?

The answer only makes sense if you know what each generation of classifier got wrong. Every method below was the best available answer for a while. Each one hit a wall that was obvious only in hindsight, and the next method was usually a direct fix for that wall: word order, then synonyms, then context, then the need for labeled data, then the cost of generating text just to deliver a label. Jev is the latest step in that chain, and it inherits the whole lineage.

Think of this post as a field notebook. Follow each method through how it worked, where it broke, and the idea that came next. Jump to an era in the map above, or explore all twelve methods in the index. The panel on the right (or the strip at the top, on a phone) keeps your place as you go.

A note on the numbers: I use the IMDb movie review test set (25,000 reviews, half positive) as a common yardstick because nearly every method in this history has been measured on it. Still, the setups differ: some numbers are tuned research results, some are quick tutorial baselines, and some aren’t directly comparable at all. All of them come from existing published benchmarks. Read the ladder as a rough shape, not a leaderboard.

Every breakthrough had a blind spot.

Follow the limitation.
It leads to the next idea.

Explore all 12 methods Open the index
Counting words
1998 to 2011
#

Bag-of-Words + Naive Bayes / Logistic Regression

Count the words, ignore the order.

Sahami et al., 1998 · Maas et al., 2011
How it works

Build a vocabulary from every word in the training set. Each document becomes a fixed-size vector with one slot per vocabulary word, holding a count (or a TF-IDF weight). A Naive Bayes or logistic regression model then learns one weight per word: refund pushes toward spam, meeting pulls away.

Used for

Email spam filters, news topic routing, the first wave of review sentiment.

IMDb
89.9% logistic regression on BoW features
Drawback

Word order is gone. “the dog bites the man” and “the man bites the dog” produce identical vectors, and “not good” is just one not plus one good.

same vector, two stories!

The next idea ↓

n-grams count adjacent word pairs, so short phrases like not good become features of their own.

Counting words
2012
#

n-grams + NBSVM

Count short phrases, not just words.

Wang & Manning, 2012
How it works

Add bigrams (and sometimes trigrams) to the vocabulary, so not good gets its own slot. NBSVM scales each feature by its Naive Bayes log-count ratio and then trains a linear SVM on top, which combines Naive Bayes’s strength on short text with the SVM’s strength on long text.

Used for

The baseline to beat for sentiment and topic classification for years. Still a great first model.

IMDb
91.22% NBSVM with bigrams · Wang & Manning, 2012
Drawback

The vocabulary explodes: every word pair is a new, sparse feature. And great, excellent, and superb are still three unrelated columns that share nothing.

phrases count now

The next idea ↓

Word embeddings give every word a dense vector, so similar words sit near each other.

Learned vectors
2013 to 2014
#

Word Embeddings and Paragraph Vectors

Words become points in space.

Mikolov et al., 2013 · Pennington et al., 2014 · Le & Mikolov, 2014
How it works

word2vec and GloVe learn a dense vector (typically 100 to 300 numbers) for every word from the contexts it appears in. Similar words end up close together. For classification you average the word vectors, or train a document vector directly (Paragraph Vector), and feed it to a linear classifier.

Used for

Cheap semantic features, similarity search, and the input layer of nearly every neural text model that followed.

IMDb
92.58% Paragraph Vector, as reported · Le & Mikolov, 2014
Drawback

One vector per word, whatever the sentence: bank means the same in “river bank” and “bank account”. Averaging still throws order away, and a later reproduction effort (Mesnil et al., 2014) could not match the reported Paragraph Vector number.

one 'bank' for every bank

The next idea ↓

Recurrent networks read the embeddings in order, carrying a memory from word to word.

Learned vectors
1997, popular by 2014
#

Recurrent Networks (LSTM, GRU)

Read one word at a time and keep a running memory.

Hochreiter & Schmidhuber, 1997 · Cho et al., 2014
How it works

An RNN reads word embeddings left to right. At each step it combines the current word with a fixed-size hidden state that summarises everything so far. Gates in the LSTM and GRU variants decide what to keep and what to forget. The final hidden state goes to a classifier head.

Used for

Sentiment, intent detection, and anything where order clearly matters, until about 2018.

IMDb
85.66% LSTM trained from scratch
Drawback

Strictly sequential, so it’s slow to train and can’t parallelise across a sentence. Long reviews squeeze through one small state vector, and trained from scratch on 25k reviews it overfits: worse than the bag-of-words baseline.

order matters... but slowly

The next idea ↓

Text CNNs look at all positions in parallel with small phrase detectors.

Learned vectors
2014
#

Text CNN

Slide learned phrase detectors over the sentence.

Kim, 2014
How it works

Borrowed from vision. Learned filters slide over windows of 3 to 5 word embeddings, each one becoming a detector for a kind of phrase (surprisingly good, waste of time). Global max-pooling keeps each filter’s strongest hit, so the output size doesn’t depend on review length.

Used for

Short-text sentiment, question classification, and fast production models.

IMDb
90.07% 1D CNN on IMDb
Drawback

The field of view is the window. A not thirty words before good never meets it. And like every model so far, it learns language from scratch from your 25k labeled reviews.

only 3 words at a time!

The next idea ↓

ULMFiT pretrains a language model on Wikipedia first, then fine-tunes it on your task.

Learned vectors
January 2018
#

ULMFiT: Transfer Learning Arrives

Learn English first, then learn your task.

Howard & Ruder, 2018
How it works

Three stages on one LSTM (AWD-LSTM): pretrain a language model on 103 million words of Wikipedia, fine-tune that language model on unlabeled text from your domain, then add a classifier head and fine-tune again on your labels, unfreezing layers gradually so it doesn’t forget what it learned.

Used for

Proved transfer learning works for text. With 100 labeled examples it matched models trained on 100x more.

IMDb
95.4% 4.6% test error · Howard & Ruder, 2018
Drawback

It’s still an RNN: sequential, a single state vector, and each word only sees what came before it (ULMFiT patches this by training a second, backward model).

pretrain → fine-tune → classify

The next idea ↓

Transformers replace recurrence with attention: every word looks at every other word, in both directions, in parallel.

Attention
October 2018
#

BERT and Encoder Transformers

Every word looks at every other word.

Devlin et al., 2018 · Sun et al., 2019 · ModernBERT, 2024
How it works

A bidirectional transformer is pretrained by hiding words and predicting them from both sides. Self-attention gives each word a vector that depends on its context, so bank near river and bank near loan finally differ. For classification you put a small head on the [CLS] token and fine-tune.

Used for

The default production classifier for moderation, intent, triage, and search ranking. ModernBERT (2024) still gets ~95% on IMDb with little tuning.

IMDb
95.79% BERT-large + in-task pretraining · Sun et al., 2019
Drawback

It’s brilliant at one task at a time. Every new label set needs a labeled dataset, a fine-tuning run, and a model to deploy.

'good' reads through 'not'

The next idea ↓

More pretraining (RoBERTa, XLNet) pushes accuracy further on the same recipe.

Attention
2019
#

RoBERTa and XLNet: Pretrain Harder

Same idea, much more data and compute.

Liu et al., 2019 · Yang et al., 2019
How it works

RoBERTa kept BERT’s architecture and trained it on roughly 10x the text (160GB vs 16GB), for longer, with bigger batches. XLNet changed the pretraining objective to permutation language modeling. Both are still fine-tuned with a classification head.

Used for

Topped leaderboards (GLUE, IMDb) and became the strong baseline for any supervised text task.

IMDb
96.21% XLNet, 3.79% test error · Yang et al., 2019
Drawback

IMDb is basically solved at this point, but each task still needs its own labels and its own fine-tune. The labels are frozen into the weights at training time.

10x the data, +0.4%

The next idea ↓

Zero-shot entailment and few-shot methods let you supply the labels at prediction time.

Attention
2019 to 2022
#

Zero-Shot Entailment and SetFit

Turn each label into a sentence and ask: does the text imply it?

Yin, Hay & Roth, 2019 · Tunstall et al., 2022
How it works

Train a model on natural language inference (premise → does it entail the hypothesis?). At prediction time, write each label as a hypothesis (“This review is positive.”) and score entailment for every one. SetFit goes the other way: fine-tune a sentence-embedding model on just 8 to 64 labeled examples per class.

Used for

Classifying with labels you invented this morning: ticket routing, tagging, quick prototypes.

IMDb
n/a No comparable full-IMDb figure. The win here was needing almost no labeled data.
Drawback

Sensitive to how you phrase each hypothesis, one forward pass per label, and the probabilities are rarely calibrated. Complex decisions (“should this invoice be held?”) stretch the entailment framing.

labels at request time

The next idea ↓

Generative LLMs can follow any instruction and pick a label in plain language.

Generative LLMs
2019 to now
#

GPT-Style LLMs: Prompt It or Fine-Tune a Head

Ask in English, read the answer.

Radford et al., 2019 · Brown et al., 2020
How it works

Two routes. Prompting: describe the task, maybe add a few examples, and let the model generate positive or negative. Fine-tuning: swap the vocabulary output layer for a small classification head and train it on the last token, which (thanks to the causal mask) is the only one that has seen the whole review.

Used for

Every one-off classification task since 2022: “just throw an LLM at it”.

IMDb
92.0% GPT-2 124M with a classification head (approx.)
Drawback

Prompting a big model is general but slow and expensive, and the answer arrives as text: sometimes “Positive.”, sometimes “The sentiment is mostly positive”, sometimes a paragraph. You end up writing a parser.

great answer, now parse it

The next idea ↓

Structured output constrains the generation to a JSON schema.

Generative LLMs
2023 to 2025
#

LLMs with Structured Output

Force the text into a schema.

Function calling, JSON mode, constrained decoding
How it works

The API constrains decoding so the model can only produce tokens that keep the output valid against a JSON schema, often with an enum for the label. Your code gets {"sentiment": "positive"} and parses it.

Used for

Most production GenAI workflows: extraction, routing, agents choosing tools.

IMDb
n/a Same accuracy as the underlying LLM. Structure fixes the format, not the decision.
Drawback

It’s still generating the answer one token at a time, so you pay latency and output tokens for braces and field names. “Confidence” is either missing or a number the model wrote, which is not a calibrated probability.

~10 tokens to say 1 bit

The next idea ↓

Jev skips generation: it scores the allowed answers directly and returns typed probabilities.

Typed decisions
September 2026
#

Jev: The Decision Without the Paragraph

One general model, any labels, a probability for each.

TypeSafe AI, 2026
How it works

You send state (the review) and typed questions (Choice, Score, or Noul) with plain-English descriptions of each option. Jev returns the chosen option and a probability distribution over all of them, with no answer text generated. TypeSafe says it’s trained with a calibration objective (RLCD); the architecture isn’t public.

Used for

Routing, triage, agent decisions, and the IMDb test set without any fine-tuning.

IMDb
96.47% Choice API, zero fine-tuning, ~$0.65 for 25k reviews
Drawback

Closed weights and an undisclosed method. Calibration is claimed, not independently measured, and a specialised fine-tuned model can still win on a single high-volume task.

no tokens generated for the answer

What comes next

Next: open-weight Jev-likes, and OpenAI’s Decision API (announced September 29, 2026).