Skip to content
VG/Tech

EngineeringBy Alexander KhomenkoOct 20268 min read

Classifiers: from rules to LLMs and beyond

Many software tasks only need a category or score. Here’s how to choose a classifier, from simple rules to decision models such as Jev.

You may have heard the buzz around Jev, a new model from TypeSafe AI. Over the last few years, most of the conversation about AI has focused on generative AI: LLMs and the products built on top of them. Jev is a classifier, not an LLM: it makes decisions rather than generating text. It's refreshing to see other kinds of AI tools entering the conversation.

Many parts of our software only need a category or score to make a decision. We already have several ways to provide that, so where do the new models fit?

Why do you need a classifier?

Classification is already common in software. An if/else statement can be a simple classifier, checking an input and choosing a category.

A support ticket may need a single category to route it to the right team. But a message about both payment and login problems may need multiple labels. Take a simple example:

“I got an email asking me to reset my password, but I didn't request it. Should I click the link?”

“Password reset” suggests technical support, but the customer didn't request it, which could indicate a security problem. We want security here. Another company might organise its teams differently, so we need to define the categories before choosing a classifier.

Let's start with the approaches we've been using for years.

Rule-based classifiers

An if/else is the simplest place to start. We can route requests based on an account's support plan, for example. The rules are explicit: we know which fields to check and which queue to choose.

Tools such as Open Policy Agent (OPA) let us manage rules separately from application code. OPA checks incoming data against those rules to decide what’s allowed.

Deterministic behaviour makes decisions easier to trace: we can see which rule matched and why. But rules can contain mistakes or become outdated, so they still need testing and maintenance.

Free text makes rules harder to maintain. Routing every message mentioning “password reset” to technical support gets our example wrong, and adding exceptions soon becomes difficult to manage.

Traditional ML classifiers

As rules accumulate, a rule-based approach becomes difficult to keep up to date. Spam filters and fraud detection are classic examples. Adversaries keep trying new patterns, so the ability to learn from new examples matters.

We train the model on labelled examples so it can classify new inputs. Approaches range from logistic regression to neural networks. A simple text classifier can learn which words and phrases are associated with each category. This is a useful starting point before trying larger models. Scikit-learn provides an example comparing several methods and their prediction costs.

Mistakes can be harder to explain. A small decision tree is easier to inspect than a large neural network, but either can learn misleading patterns from its training data. A model that works on old tickets may fail on new ones, so we need to track results and check whether retraining improves them.

Small language encoders belong here too

Language encoders offer another option: they turn text into numbers that represent its meaning. A small classifier then uses those numbers to choose a label.

SetFit is one example of an open-source tool for this. It trains an encoder and classifier on labelled examples and is designed to work with relatively few, though the amount needed depends on the task.

For our support example, we could train a classifier on previous login and security tickets. This suits a high volume of similar requests with stable categories. Changing the categories may require new examples and retraining.

Managed services such as Amazon Comprehend Custom Classification can train on your labelled documents and classify new ones individually or in batches. Each document can receive one or several labels.

LLMs as classifiers

LLMs let us define categories in ordinary language. With a few examples to clarify the differences, we can start classifying without training our own model.

For our ticket, the instructions could be: “Choose security for account activity the customer didn't initiate, including unexpected password-reset emails. Choose technical support for problems signing in or using the product.” When the policy changes, we update the instructions and test again.

Features such as OpenAI's Structured Outputs let us restrict answers to the labels our code expects. The model can still choose the wrong label or fail to return a usable answer. We can ask an LLM why it chose a label, but its response may not reflect how it made the decision, as Anthropic's research shows. To investigate mistakes, we still need the original inputs and results, along with the model and prompt versions.

Using a large model for every small decision can slow down a workflow and increase its cost. A smaller model may be enough, but we need to check that its decisions agree with human reviewers. Specialised services such as OpenAI's Moderation API classify text and images into predefined safety categories. We can't define our own categories, so it won't handle our support queues.

Decision models: Jev and alternatives

TypeSafe describes Jev as a System One model for fast decisions. We send it the application’s current state and our questions, and it returns answers in a format our code can use directly.

Jev supports three question types:

  • Choice: choose from a set of options, such as support queues.
  • Noul: estimate how likely a statement is to be true, such as “the customer wants to cancel.”
  • Score: rate something on a scale you define, such as how severe an issue is.

Choice and Score return probabilities and a confidence value; Noul returns only a probability. We can ask several independent questions in one request, such as where to route a message and how urgent it is.

Here’s an example request for our support ticket, using Jev’s documented API format:

json
{
  "model": "jev-1.13.0",
  "state": "I got an email asking me to reset my password, but I didn't request it. Should I click the link?",
  "questions": {
    "support_queue": {
      "type": "choice",
      "instructions": "Which queue should handle this message? Apply the supplied queue definitions.",
      "criteria": {
        "security": "Account activity the customer didn't start, such as unexpected password-reset emails, unknown logins, or suspected phishing.",
        "technical": "Product faults or access problems the customer is trying to fix, such as a forgotten password.",
        "billing": "Charge disputes, invoices, and payment failures.",
        "other": "No queue fits, or the requested action is unclear."
      }
    }
  }
}

We expect security here. The other option covers messages that don't fit any queue.

Price and speed

Jev can handle classification at significantly lower cost than generative LLMs, according to TypeSafe’s comparisons. The savings depend on the model and workload.

TypeSafe reports response times of 70–500 ms, while acknowledging that the short input in its demonstration favours Jev. Test it on your own workload to see how that compares.

A local classifier may still be cheaper for a stable task, while a generative model may handle difficult cases better. Compare accuracy as well as speed and cost.

What about accuracy?

A correctly formatted answer can still be wrong. Jev calculates Choice confidence from the leading option’s probability and the number of options. A confidence of 0.9 doesn’t mean the model is correct 90% of the time.

Check on your own data whether answers assigned an 80% probability are correct about 80% of the time. This is calibration, and it helps you set thresholds for automatic decisions.

TypeSafe says Jev performs best in English. If your product serves APAC, test it in your customers’ languages, including messages that mix languages.

Other decision models

Several other vendors announced decision models within weeks of Jev’s launch.

OpenAI announced a Decisions API at DevDay, running on GPT-6 Luna. It takes text or images and chooses from predefined answers. It’s in limited preview, and we haven’t seen public pricing or a request format yet.

Cloudflare’s Clef and Liquid AI’s d1 accept Jev’s request format, making it easier to compare models or switch between them. OpenRouter also offers an experimental Decisions endpoint that routes requests in this format to Jev or other decision models.

Perplexity also offers a Decisions API supporting the same three question types as Jev.

Some models, including Cloudflare’s Clef and Perplexity’s model, can run on your own infrastructure. This lets you keep data in-house and adapt the model to your needs, but you take on deployment and maintenance.

Vendors’ own benchmarks suggest their models compare well with Jev. Test them on your own data before choosing.

Combining approaches

These approaches can work together. Rules handle clear cases, while a classifier interprets the message. Difficult requests can go to a stronger model or a person.

Stripe Radar uses machine learning to assess payment fraud risk. Teams can add rules to allow or block payments and send suspicious cases for human review.

For our support example, a rule could send every security result to a person, regardless of confidence. Missing an account takeover costs more than misrouting a ticket. Uncertain results would also go to review. A security label can route a ticket. Locking an account or resetting credentials requires checking account activity, the customer’s identity, and permissions.

How to choose a classifier

Here’s where I’d start:

What you needWhere to startKeep in mind
Rules based on known fieldsApplication code or OPATest and maintain the rules
Stable categories with labelled examplesScikit-learn or SetFitYou manage training and deployment
A managed classifier for your categoriesAmazon ComprehendTraining examples must reflect real inputs
Flexible instructions and substantial contextAn LLM with structured outputsCompare accuracy, cost, and speed
Frequent classification decisionsJev or similar decision modelsTest confidence thresholds on your data
Image classificationClef, Perplexity’s model, or OpenAI Decisions APITest vendor claims
Content moderationOpenAI Moderation APICheck that its predefined categories fit

Are decision models production-ready?

At VG Tech Consulting, we’ve used Jev in some of our agent workflows and seen faster runs at lower cost. These are observations from our integrations, not benchmark results.

Whether it’s ready for your workflow depends on the consequences of mistakes. Mislabelled feedback is less serious than a customer complaint closed by mistake.

Start with one repeated decision and a golden dataset. Compare approaches on the same inputs, including unfamiliar or ambiguous requests in your customers’ languages. Keep the final test examples separate from those used for training or adjusting prompts.

Check which categories the model confuses and how often results need human review. Test long messages and instructions embedded in the input. Measure the full workflow’s speed and cost, including the work needed to correct mistakes.

Before relying on the classifier, run it alongside your existing process and compare results. Record the model version and instructions so you can repeat the tests. With Jev, use a fixed version if you’ve tuned confidence thresholds; jev-latest can change.

If your workflows make many small decisions, these models are worth testing. At VG Tech Consulting, we can help you choose and evaluate an approach for your product.

Sources

Sources checked in October 2026. The API request is illustrative.

Ready to put this into practice?