The first question I got after the Jev post was a fair one: “Isn’t this just a glorified classifier?” Yes, and that’s the interesting part. Text classification is one of the oldest problems in NLP. Spam filters were doing it in 1998. So why did a classifier with a nice API become the most talked-about model release of the month?
The answer only makes sense if you know what each generation of classifier got wrong. Every method below was the best available answer for a while. Each one hit a wall that was obvious only in hindsight, and the next method was usually a direct fix for that wall: word order, then synonyms, then context, then the need for labeled data, then the cost of generating text just to deliver a label. Jev is the latest step in that chain, and it inherits the whole lineage.
Think of this post as a field notebook. Follow each method through how it worked, where it broke, and the idea that came next. Jump to an era in the map above, or explore all twelve methods in the index. The panel on the right (or the strip at the top, on a phone) keeps your place as you go.
A note on the numbers: I use the IMDb movie review test set (25,000 reviews, half positive) as a common yardstick because nearly every method in this history has been measured on it. Still, the setups differ: some numbers are tuned research results, some are quick tutorial baselines, and some aren’t directly comparable at all. All of them come from existing published benchmarks. Read the ladder as a rough shape, not a leaderboard.
Every breakthrough had a blind spot.
Follow the limitation.
It leads to the next idea.
Explore all 12 methods Open the index
Bag-of-Words + Naive Bayes / Logistic Regression
Count the words, ignore the order.
- How it works
Build a vocabulary from every word in the training set. Each document becomes a fixed-size vector with one slot per vocabulary word, holding a count (or a TF-IDF weight). A Naive Bayes or logistic regression model then learns one weight per word:
refundpushes toward spam,meetingpulls away.- Used for
Email spam filters, news topic routing, the first wave of review sentiment.
- IMDb
- 89.9% logistic regression on BoW features
- Drawback
Word order is gone. “the dog bites the man” and “the man bites the dog” produce identical vectors, and “not good” is just one
notplus onegood.
same vector, two stories!
n-grams count adjacent word pairs, so short phrases like not good become features of their own.
n-grams + NBSVM
Count short phrases, not just words.
- How it works
Add bigrams (and sometimes trigrams) to the vocabulary, so
not goodgets its own slot. NBSVM scales each feature by its Naive Bayes log-count ratio and then trains a linear SVM on top, which combines Naive Bayes’s strength on short text with the SVM’s strength on long text.- Used for
The baseline to beat for sentiment and topic classification for years. Still a great first model.
- IMDb
- 91.22% NBSVM with bigrams · Wang & Manning, 2012
- Drawback
The vocabulary explodes: every word pair is a new, sparse feature. And
great,excellent, andsuperbare still three unrelated columns that share nothing.
phrases count now
Word embeddings give every word a dense vector, so similar words sit near each other.
Word Embeddings and Paragraph Vectors
Words become points in space.
- How it works
word2vec and GloVe learn a dense vector (typically 100 to 300 numbers) for every word from the contexts it appears in. Similar words end up close together. For classification you average the word vectors, or train a document vector directly (Paragraph Vector), and feed it to a linear classifier.
- Used for
Cheap semantic features, similarity search, and the input layer of nearly every neural text model that followed.
- IMDb
- 92.58% Paragraph Vector, as reported · Le & Mikolov, 2014
- Drawback
One vector per word, whatever the sentence:
bankmeans the same in “river bank” and “bank account”. Averaging still throws order away, and a later reproduction effort (Mesnil et al., 2014) could not match the reported Paragraph Vector number.
one 'bank' for every bank
Recurrent networks read the embeddings in order, carrying a memory from word to word.
Recurrent Networks (LSTM, GRU)
Read one word at a time and keep a running memory.
- How it works
An RNN reads word embeddings left to right. At each step it combines the current word with a fixed-size hidden state that summarises everything so far. Gates in the LSTM and GRU variants decide what to keep and what to forget. The final hidden state goes to a classifier head.
- Used for
Sentiment, intent detection, and anything where order clearly matters, until about 2018.
- IMDb
- 85.66% LSTM trained from scratch
- Drawback
Strictly sequential, so it’s slow to train and can’t parallelise across a sentence. Long reviews squeeze through one small state vector, and trained from scratch on 25k reviews it overfits: worse than the bag-of-words baseline.
order matters... but slowly
Text CNNs look at all positions in parallel with small phrase detectors.
Text CNN
Slide learned phrase detectors over the sentence.
- How it works
Borrowed from vision. Learned filters slide over windows of 3 to 5 word embeddings, each one becoming a detector for a kind of phrase (
surprisingly good,waste of time). Global max-pooling keeps each filter’s strongest hit, so the output size doesn’t depend on review length.- Used for
Short-text sentiment, question classification, and fast production models.
- IMDb
- 90.07% 1D CNN on IMDb
- Drawback
The field of view is the window. A
notthirty words beforegoodnever meets it. And like every model so far, it learns language from scratch from your 25k labeled reviews.
only 3 words at a time!
ULMFiT pretrains a language model on Wikipedia first, then fine-tunes it on your task.
ULMFiT: Transfer Learning Arrives
Learn English first, then learn your task.
- How it works
Three stages on one LSTM (AWD-LSTM): pretrain a language model on 103 million words of Wikipedia, fine-tune that language model on unlabeled text from your domain, then add a classifier head and fine-tune again on your labels, unfreezing layers gradually so it doesn’t forget what it learned.
- Used for
Proved transfer learning works for text. With 100 labeled examples it matched models trained on 100x more.
- IMDb
- 95.4% 4.6% test error · Howard & Ruder, 2018
- Drawback
It’s still an RNN: sequential, a single state vector, and each word only sees what came before it (ULMFiT patches this by training a second, backward model).
pretrain → fine-tune → classify
Transformers replace recurrence with attention: every word looks at every other word, in both directions, in parallel.
BERT and Encoder Transformers
Every word looks at every other word.
- How it works
A bidirectional transformer is pretrained by hiding words and predicting them from both sides. Self-attention gives each word a vector that depends on its context, so
banknearriverandbanknearloanfinally differ. For classification you put a small head on the[CLS]token and fine-tune.- Used for
The default production classifier for moderation, intent, triage, and search ranking. ModernBERT (2024) still gets ~95% on IMDb with little tuning.
- IMDb
- 95.79% BERT-large + in-task pretraining · Sun et al., 2019
- Drawback
It’s brilliant at one task at a time. Every new label set needs a labeled dataset, a fine-tuning run, and a model to deploy.
'good' reads through 'not'
More pretraining (RoBERTa, XLNet) pushes accuracy further on the same recipe.
RoBERTa and XLNet: Pretrain Harder
Same idea, much more data and compute.
- How it works
RoBERTa kept BERT’s architecture and trained it on roughly 10x the text (160GB vs 16GB), for longer, with bigger batches. XLNet changed the pretraining objective to permutation language modeling. Both are still fine-tuned with a classification head.
- Used for
Topped leaderboards (GLUE, IMDb) and became the strong baseline for any supervised text task.
- IMDb
- 96.21% XLNet, 3.79% test error · Yang et al., 2019
- Drawback
IMDb is basically solved at this point, but each task still needs its own labels and its own fine-tune. The labels are frozen into the weights at training time.
10x the data, +0.4%
Zero-shot entailment and few-shot methods let you supply the labels at prediction time.
Zero-Shot Entailment and SetFit
Turn each label into a sentence and ask: does the text imply it?
- How it works
Train a model on natural language inference (premise → does it entail the hypothesis?). At prediction time, write each label as a hypothesis (“This review is positive.”) and score entailment for every one. SetFit goes the other way: fine-tune a sentence-embedding model on just 8 to 64 labeled examples per class.
- Used for
Classifying with labels you invented this morning: ticket routing, tagging, quick prototypes.
- IMDb
- n/a No comparable full-IMDb figure. The win here was needing almost no labeled data.
- Drawback
Sensitive to how you phrase each hypothesis, one forward pass per label, and the probabilities are rarely calibrated. Complex decisions (“should this invoice be held?”) stretch the entailment framing.
labels at request time
Generative LLMs can follow any instruction and pick a label in plain language.
GPT-Style LLMs: Prompt It or Fine-Tune a Head
Ask in English, read the answer.
- How it works
Two routes. Prompting: describe the task, maybe add a few examples, and let the model generate
positiveornegative. Fine-tuning: swap the vocabulary output layer for a small classification head and train it on the last token, which (thanks to the causal mask) is the only one that has seen the whole review.- Used for
Every one-off classification task since 2022: “just throw an LLM at it”.
- IMDb
- 92.0% GPT-2 124M with a classification head (approx.)
- Drawback
Prompting a big model is general but slow and expensive, and the answer arrives as text: sometimes “Positive.”, sometimes “The sentiment is mostly positive”, sometimes a paragraph. You end up writing a parser.
great answer, now parse it
Structured output constrains the generation to a JSON schema.
LLMs with Structured Output
Force the text into a schema.
- How it works
The API constrains decoding so the model can only produce tokens that keep the output valid against a JSON schema, often with an enum for the label. Your code gets
{"sentiment": "positive"}and parses it.- Used for
Most production GenAI workflows: extraction, routing, agents choosing tools.
- IMDb
- n/a Same accuracy as the underlying LLM. Structure fixes the format, not the decision.
- Drawback
It’s still generating the answer one token at a time, so you pay latency and output tokens for braces and field names. “Confidence” is either missing or a number the model wrote, which is not a calibrated probability.
~10 tokens to say 1 bit
Jev skips generation: it scores the allowed answers directly and returns typed probabilities.
Jev: The Decision Without the Paragraph
One general model, any labels, a probability for each.
- How it works
You send state (the review) and typed questions (Choice, Score, or Noul) with plain-English descriptions of each option. Jev returns the chosen option and a probability distribution over all of them, with no answer text generated. TypeSafe says it’s trained with a calibration objective (RLCD); the architecture isn’t public.
- Used for
Routing, triage, agent decisions, and the IMDb test set without any fine-tuning.
- IMDb
- 96.47% Choice API, zero fine-tuning, ~$0.65 for 25k reviews
- Drawback
Closed weights and an undisclosed method. Calibration is claimed, not independently measured, and a specialised fine-tuned model can still win on a single high-volume task.
no tokens generated for the answer
Next: open-weight Jev-likes, and OpenAI’s Decision API (announced September 29, 2026).
References
- Sahami, Dumais, Heckerman & Horvitz, A Bayesian Approach to Filtering Junk E-Mail (1998)
- Maas et al., Learning Word Vectors for Sentiment Analysis (2011), the IMDb dataset
- Wang & Manning, Baselines and Bigrams: Simple, Good Sentiment and Topic Classification (2012)
- Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013)
- Le & Mikolov, Distributed Representations of Sentences and Documents (2014)
- Mesnil et al., Ensemble of Generative and Discriminative Techniques for Sentiment Analysis of Movie Reviews (2014)
- Kim, Convolutional Neural Networks for Sentence Classification (2014)
- Howard & Ruder, Universal Language Model Fine-tuning for Text Classification (2018)
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers (2018)
- Sun et al., How to Fine-Tune BERT for Text Classification? (2019)
- Liu et al., RoBERTa: A Robustly Optimized BERT Pretraining Approach (2019)
- Yang et al., XLNet: Generalized Autoregressive Pretraining (2019)
- Yin, Hay & Roth, Benchmarking Zero-shot Text Classification (2019)
- Tunstall et al., Efficient Few-Shot Learning Without Prompts (SetFit) (2022)
- TypeSafe AI, Introducing System One Models and Jev (2026)
