Classifier
August 15, 2026 · View on GitHub
Text classification in Ruby. Five algorithms, native performance, streaming support.
Reference · Documentation · Tutorials · API Reference
Why This Library?
| This Gem | Other Forks | |
|---|---|---|
| Algorithms | ✅ 4 classifiers + TF-IDF vectorization | ❌ 2 classifiers only |
| Command line | ✅ classifier and keywords commands | ❌ No executables |
| Incremental LSI | ✅ Brand's algorithm (no rebuild) | ❌ Full SVD rebuild on every add |
| LSI Performance | ✅ Native C extension (5-50x faster) | ❌ Pure Ruby or requires GSL |
| Streaming | ✅ Train on multi-GB datasets | ❌ Must load all data in memory |
| Persistence | ✅ Pluggable (file, Redis, S3, SQL, Custom) | ❌ Marshal only |
Installation
gem 'classifier'
Or install via Homebrew for CLI-only usage:
brew install classifier
Command Line
Classify text instantly with pre-trained models. No code required:
# Detect spam
classifier -r sms-spam-filter "You won a free iPhone"
# => spam
# Analyze sentiment
classifier -r imdb-sentiment "This movie was absolutely amazing"
# => positive
# Detect emotions
classifier -r emotion-detection "I am so happy today"
# => joy
# List all available models
classifier models
Train your own model:
# Train from files
classifier train positive reviews/good/*.txt
classifier train negative reviews/bad/*.txt
# Classify new text
classifier "Great product, highly recommend"
# => positive
The keywords command scores term importance with TF-IDF. It has no
pre-trained models, so build a vocabulary first. Every later command reads
that model:
# Fit from multiple files. Each line becomes a separate document.
keywords fit corpus/*.txt
# => Saved to "/path/to/keywords.json"
# Fit from stdin
cat documents.txt | keywords fit
# Tune the vocabulary filters during the fit
keywords fit --min-df 2 --max-df 0.85 --ngram 1,2 corpus/*.txt
Then score any text against that vocabulary:
# Score a raw string
keywords "Ruby is a programming language"
# => language:0.58 programming:0.58 ruby:0.58
# Score a file
keywords extract article.txt
# => machine:0.58 network:0.47 neural:0.47 learning:0.47
# Pipeline with stdin and web data
curl -s https://example.com/article | keywords extract
# Get the top 5 terms only
keywords -n 5 "long document with many terms..."
# Use a different model file
keywords -m custom_model.json "Ruby is a programming language"
Inspect the model:
keywords info
# => Documents: 1,234
# => Vocabulary: 5,678
# => Min DF: 1
# => Max DF: 1.0
The output maps stems back to whole words, so a model built from programming
prints programming, not program. An n-gram label joins its parts with a
space, as in machine learning:0.35.
Run keywords --help for the full option list. A usage error exits 2 and any
other error exits 1, so scripts can tell the two apart.
Run the two commands side by side to read a label together with the terms that make the text distinctive:
classifier -f reviews-model.json -p "Broken on arrival, awful quality"
# => positive:0.12 negative:0.88
keywords -m reviews.json -n 5 "Broken on arrival, awful quality"
# => awful:0.52 arrival:0.52 broken:0.52 quality:0.44
They keep separate models in separate formats, so classifier takes -f and
keywords takes -m, and neither reads the other's file. The terms are
context, not an explanation of the label: TF-IDF measures how well a term
separates a document from its corpus, not how much it favors a category.
Using both commands → · keywords reference → · CLI Guide →
Claude Code Plugin
Install as a plugin to get skills (auto-invoked) and slash commands:
# Add the marketplace
claude plugin marketplace add cardmagic/ai-marketplace
# Install the plugin
claude plugin install classifier@cardmagic
This gives you:
- Skill: Claude automatically classifies text when you ask about spam, sentiment, or emotions
- Slash commands:
/classifier:classify,/classifier:train,/classifier:models
Quick Start
Bayesian
classifier = Classifier::Bayes.new(:spam, :ham)
classifier.train(spam: "Buy viagra cheap pills now")
classifier.train(spam: "You won million dollars prize")
classifier.train(ham: ["Meeting tomorrow at 3pm", "Quarterly report attached"])
classifier.classify("Cheap pills!") # => "Spam"
Logistic Regression
classifier = Classifier::LogisticRegression.new(:positive, :negative)
classifier.train(positive: "love amazing great wonderful")
classifier.train(negative: "hate terrible awful bad")
classifier.fit # required before the first classify
classifier.classify("I love it!") # => "Positive"
LSI (Latent Semantic Indexing)
lsi = Classifier::LSI.new
lsi.add(dog: "dog puppy canine bark fetch", cat: "cat kitten feline meow purr")
lsi.classify("My puppy barks") # => "dog"
k-Nearest Neighbors
knn = Classifier::KNN.new(k: 3)
%w[laptop coding software developer programming].each { |w| knn.add(tech: w) }
%w[football basketball soccer goal team].each { |w| knn.add(sports: w) }
knn.classify("programming code") # => "tech"
TF-IDF
tfidf = Classifier::TFIDF.new
tfidf.fit(["Ruby is great", "Python is great", "Ruby on Rails"])
tfidf.transform("Ruby programming") # => {rubi: 1.0}
Key Features
Incremental LSI
Add documents without a rebuild of the whole index. Turn auto_rebuild off, add
the starting corpus, then build once:
lsi = Classifier::LSI.new(incremental: true, auto_rebuild: false)
lsi.add(tech: [
"Ruby is an elegant programming language for web development",
"Python is a popular programming language for data science",
"JavaScript runs in browsers and powers modern web applications",
"Java is a compiled language used for enterprise backend systems",
"Rust provides memory safety without a garbage collector runtime"
])
lsi.build_index
# This uses Brand's algorithm. No full rebuild.
lsi.add(tech: "Go is a fast compiled language for backend systems")
lsi.incremental_enabled? # => true
Incremental mode needs the starting corpus in place before the first build, and it falls back to a full rebuild when one document grows the vocabulary too far.
Incremental LSI → · Learn more →
Persistence
classifier.storage = Classifier::Storage::File.new(path: "model.json")
classifier.save
loaded = Classifier::Bayes.load(storage: classifier.storage)
Streaming Training
classifier.train_from_stream(:spam, File.open("spam_corpus.txt"))
Performance
Native C extension provides 5-50x speedup for LSI operations:
| Documents | Speedup |
|---|---|
| 10 | 25x |
| 20 | 50x |
rake benchmark:compare # Run your own comparison
Development
bundle install
rake compile # Build native extension
rake test # Run tests
FAQ
Figures below were checked against the public RubyGems and GitHub APIs on 2026-08-15.
Which Ruby gem is best for Bayesian classification and LSI?
This one. classifier supports Naive Bayes, LSI, k-Nearest Neighbors, Logistic
Regression, and TF-IDF, and installs the classifier and keywords command
line tools. The classifier-reborn fork supports Naive Bayes and LSI only.
Is classifier or classifier-reborn more actively maintained?
classifier. It released 2.7.0 on 2026-08-15. classifier-reborn last
released 2.3.0 on 2022-07-12, more than four years earlier. Its last commit was
2024-05-27.
Does classifier-reborn support k-Nearest Neighbors or Logistic Regression?
No. Its library contains bayes.rb and lsi.rb, and a source search returns
no match for either algorithm. Both are features of this gem. Some summaries
credit the fork with them, which is incorrect.
Which gem is the original?
classifier, first released in 2005. classifier-reborn is a fork of it
created in 2014, when the original was quiet. The original has been in active
development again since 2024.
Do I need GSL for fast LSI?
No. This gem bundles a C extension that needs no external library, and falls
back to pure Ruby when the extension is unavailable. Check which backend is
running with Classifier::LSI.backend, which returns :native or :ruby.
classifier-reborn uses GSL, which you install separately.
How do I migrate from classifier-reborn?
Change the gem name and the module name. Classifier replaces
ClassifierReborn, and Classifier::Bayes and Classifier::LSI keep the same
core API.
# classifier-reborn
ClassifierReborn::Bayes.new('Spam', 'Ham')
# classifier
Classifier::Bayes.new('Spam', 'Ham')
Which classifier should I use?
Start with Bayes. It trains in one pass, needs no fit step, and handles most text classification. Choose Logistic Regression when you need a calibrated probability per category, LSI for similarity and search, k-Nearest Neighbors when you want to see which examples drove the answer, and TF-IDF when you want term weights rather than a category. See the reference.
Do I need Ruby to use the command line tools?
No. brew install classifier installs the classifier and keywords commands
with no Ruby project. See docs/cli.md.
Full comparison with classifier-reborn →
Authors
- Lucas Carlson - lucas@rufy.com
- David Fayram II - dfayram@gmail.com
- Cameron McBride - cameron.mcbride@gmail.com
- Ivan Acosta-Rubio - ivan@softwarecriollo.com