Configuration

August 15, 2026 ยท View on GitHub

Global settings

Classifier.configure sets the defaults for every classifier:

require "classifier"

Classifier.configure do |config|
  config.min_word_length = 2
end

Classifier.config.min_word_length
# => 2
SettingDefaultEffect
min_word_length3The tokenizer drops any word shorter than this

Set the configuration once at startup. The lazy setup is not thread-safe, so do not first touch it from several threads at once.

Every classifier also takes min_word_length on its own, which overrides the global value:

Classifier::Bayes.new(:spam, :ham, min_word_length: 2)

Tokenization

The tokenizer downcases the text, strips punctuation, drops the stop words in CORPUS_SKIP_WORDS, drops words shorter than min_word_length, and reduces each remaining word to its Porter stem.

"Ruby programming is elegant".word_hash
# => {rubi: 1, program: 1, eleg: 1}

clean_word_hash skips the punctuation strip when the text is already clean. stem_to_word_hash maps each stem back to the most frequent original word:

"Ruby programming is elegant and programming rocks".stem_to_word_hash
# => {rubi: "ruby", program: "programming", eleg: "elegant", rock: "rocks"}

Native extension

LSI uses a C extension for its linear algebra. It has no external dependency and builds during gem install. Pure Ruby runs when the extension is absent, with the same results and less speed.

Classifier::LSI.backend
# => :native

The value is :native or :ruby.

Force pure Ruby with an environment variable, which is useful to compare the two:

NATIVE_VECTOR=true bundle exec rake test

Build the extension from a checkout:

bundle exec rake compile

Silence the startup notice about the missing extension:

SUPPRESS_LSI_WARNING=true

Errors

Every error inherits from Classifier::Error.

ErrorRaised when
Classifier::NotFittedErrorA model is used before its fit. Logistic regression and TF-IDF
Classifier::UnsavedChangesErrorreload! would discard unsaved changes
Classifier::StorageErrorA storage backend operation fails
begin
  classifier.classify("text")
rescue Classifier::NotFittedError
  classifier.fit
  retry
end