Jev, Laya, and Nimble: The Rise of Decision Models

AA
Atif AliFull Stack Engineer
https://res.cloudinary.com/dbonkjhet/image/upload/v1791194073/groooh/blog-images/ypqwcdwdvdqulykk0mxo.jpg

A technical comparison of Jev and Laya covering execution orchestration versus application composition, when to use each, and a step-by-step Laya setup guide.

If you have spent any time building AI agents recently you have likely run into a frustrating bottleneck: using a massive Large Language Model (LLM) just to make a simple choice. Asking an LLM to read a support ticket and output exactly "Billing", "Tech Support", or "Refund" is notoriously slow, expensive, and prone to formatting errors. Enter a new paradigm released in late 2026: Decision Models. Let us dive into exactly what these models are and how they are changing the landscape of artificial intelligence.

What are Decision Models?

Decision models like Jev and Laya flip the traditional AI architecture entirely on its head. Instead of generating open-ended unpredictable text tokens one by one, you feed them text alongside a strict set of typed questions such as multiple-choice options, scores, or simple yes/no queries. The model runs a single forward pass and returns a strict mathematical probability distribution. While they cannot write a poem or brainstorm a marketing campaign, they can tell you with near certainty which department should handle an incoming email, saving immense amounts of time and computational power.

What is Jev?

Jev is a hosted proprietary AI decision model developed by the team at TypeSafe AI. Operating as a managed API which is often referred to as a "System One" model, Jev is specifically designed for enterprise teams that want plug-and-play accuracy without the headache of managing their own infrastructure or server instances. Out of the box, Jev acts like an incredibly knowledgeable oracle. It boasts a massive context window of over 4000 tokens and highly calibrated confidence probabilities. If Jev states it is 90 percent sure about a classification, you can generally trust that metric, making it exceptionally reliable for business-critical automated routing where mistakes carry a high cost.

When weighing the pros and cons, Jev offers high out-of-the-box accuracy reaching 92.9 percent in independent benchmarks and requires zero maintenance since it is a fully hosted API. Its large context window handles extensive documents with ease and its well-calibrated confidence scores are highly reliable for programmatic decision-making. However, it operates as a closed-source system requiring data to be sent externally to TypeSafe AI's servers, which raises strict privacy concerns for healthcare and finance sectors. Furthermore, network round-trips introduce an unavoidable latency averaging around 136 milliseconds, and operational costs scale indefinitely with every API request made by your system.

What is Laya?

Laya on the other hand is an open-source non-autoregressive decision model created by Nandakishor M at Convai Innovations. Released freely under the Apache 2.0 license, Laya is built on a 421-million-parameter ModernBERT-large encoder equipped with a highly specialized typed decision head. Laya’s absolute superpower is its deployment flexibility. Because the weights are fully open and relatively small compared to massive LLMs, you can run Laya on your own local servers, deploy it in a private cloud environment, or astonishingly run it entirely inside a user's web browser tab using WebAssembly. This makes Laya incredibly fast, often clocking sub-40-millisecond latency, and completely private.

In terms of advantages, Laya ensures absolute data privacy and requires zero per-token fees, making it highly cost-effective at scale. Its local execution drastically outpaces API calls and it is built from the ground up to be specialized via fine-tuning or Mixture of Experts architectures. On the downside, Laya suffers from low zero-shot accuracy out of the box meaning it requires dedicated fine-tuning to become truly useful. It also possesses a limited 512-token context window for English inputs, struggles significantly when asked to choose from more than twenty categories simultaneously, and tends to be mathematically overconfident necessitating manual temperature refitting by developers to trust its probability scores.

What is Nimble?

Nimble is an open-source 9-billion-parameter decision model released by Bespoke Labs, fine-tuned from the Qwen3.5-9B checkpoint. Licensed under Apache 2.0, Nimble is designed as an open reconstruction of the "System One" approach popularized by Jev. Instead of generating reasoning or conversational text, Nimble reads your context prompt once and scores answer tokens directly, delivering an immediate probability distribution over allowed choices.

One of Nimble's major advantages is that it runs natively on Ollama (v0.35.0 and later) via a dedicated /v1/systemone endpoint. It offers an 8,194-token context window and handles up to 64 simultaneous multiple-choice, true/false, or rubric questions in a single request. Because it entirely bypasses autoregressive text generation, it is exceptionally fast—clocking in at under 100 milliseconds for inference on modern hardware like an Apple M5 Max. It effectively bridges the gap between Jev’s robust zero-shot capabilities and Laya’s local privacy, offering strong out-of-the-box performance without requiring developers to collect datasets and train their own classification heads.

Why Use This Over an LLM?

Imagine a standard customer support pipeline. An expensive billion-parameter LLM writes a custom reply to a frustrated user. Then that exact same massive model is invoked just to tag the ticket as urgent or normal, determine if a human manager needs to intervene, and route it to the Billing, Tech, or Sales department. Using a generative LLM for these simple classification tags is wildly inefficient. You are paying high prices for generated tokens, suffering immense latency, and constantly risking system failures when the LLM accidentally replies with conversational filler instead of a strict JSON format. Decision models completely eliminate this waste by reading the state and immediately providing structured outputs.

Architecture and Use Case Comparison

When comparing the models side-by-side, the differences in architecture, speed, and privacy are stark and define their ideal use cases. Jev is a proprietary hosted model that offers excellent accuracy without any training making it perfect for teams that need immediate results. However it operates significantly slower due to network transit and requires your private data to leave your localized network. Conversely, Laya is built on an open-source ModernBERT encoder that mandates fine-tuning to achieve comparable reliability. Yet once tuned, Laya operates at lightning-fast speeds natively on your hardware ensuring absolute privacy since the data never leaves the host machine. Nimble occupies a powerful middle ground: it provides excellent zero-shot accuracy natively on your own hardware via Ollama, giving you privacy and speed without the immediate need for manual fine-tuning.

Jev thrives in zero-maintenance enterprise routing scenarios involving large categorization sets and deep historical context logs. Laya excels in high-throughput privacy-first agents, localized expert pipelines, and browser-based automation where low latency is the absolute highest priority. Nimble is the go-to for developers already using Ollama who need a robust, private, zero-shot router that works seamlessly right out of the box.

At a Glance: Model Comparison

FeatureJev (TypeSafe AI)Laya (Convai Innovations)Nimble (Bespoke Labs)
ArchitectureProprietary Hosted API421M ModernBERT-large9B Qwen3.5-9B based
LicenseClosed-sourceApache 2.0Apache 2.0
Context Window>4,000 tokens512 tokens8,194 tokens
Input & Output TokensText input; JSON probability outputText input; Raw logit probabilities outputText input; Single letter code + probability score output
Token Cost & PricingPaid API tier; costs scale indefinitely per requestZero per-token fees (run locally)Zero per-token fees (run locally via Ollama)

Deploying Jev

Because Jev is a fully managed hosted solution by TypeSafe AI integration is straightforward. You do not need to provision GPUs manage container orchestrations or handle model weights. Here is the comprehensive step-by-step process for getting Jev into your production environment using Python.

Step 1: Environment Setup & Authentication

First navigate to the TypeSafe AI developer portal and generate a production API key. Set this securely in your environment variables to ensure it is not hardcoded into your application.

Pro Tip: Always use a dedicated staging API key during development. Jev charges per API request and runaway development loops can quickly drain your budget.

Step 2: Executing the Decision API Call (Jev)

Unlike generative LLMs where you write a conversational natural language prompt, Jev requires a strictly typed JSON payload. You provide the state (the raw text context to analyze) and an array of structured questions. Below is a complete, production-ready Python implementation using the standard requests library.

import os
import requests

JEV_API_KEY = os.getenv("JEV_API_KEY")
ENDPOINT = "https://api.typesafe.ai/v1/decide"

headers = {
    "Authorization": f"Bearer {JEV_API_KEY}",
    "Content-Type": "application/json"
}

payload = {
    "state": "The user clicked 'Cancel Subscription' but has 14 days left...",
    "questions": [
        {
            "id": "user_intent",
            "type": "choice",
            "options": ["churn_risk", "downgrade", "inquiry"]
        },
        {
            "id": "offer_discount",
            "type": "yes_no"
        }
    ]
}

response = requests.post(ENDPOINT, headers=headers, json=payload)
response.raise_for_status()
print(response.json())

Deploying Laya Locally

Laya demands a bit more hands-on engineering but the payoff is absolute privacy and zero latency costs. Because Laya is built on the ModernBERT architecture it integrates seamlessly with the open-source Hugging Face ecosystem. Here is how to run Laya entirely on your own hardware.

Step 1: Environment and Dependencies

You will need a Python environment with PyTorch and Transformers installed. Because Laya is highly optimized it runs perfectly fine on a standard CPU though a GPU will push inference times down to the ~10ms range.

Step 2: Loading and Initializing the Model

Pull the model weights directly from the Convai Innovations repository. We load both the tokenizer and the specialized sequence classification model into memory.

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

model_id = "convai/laya-v1-large"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

Step 3: Formatting and Inference

Laya expects the state and the question to be formatted using special tokens separating context from query. The model outputs raw logits mapped through a softmax function for confidence scoring.

state = "The user clicked 'Cancel Subscription'."
question = "Is the user churning? Yes or No."

inputs = tokenizer(state, question, return_tensors="pt")

with torch.no_grad():
    outputs = model(**inputs)
    logits = outputs.logits
    probabilities = torch.nn.functional.softmax(logits, dim=-1)

print("Computed Probabilities:", probabilities)

Deploying Nimble via Ollama

Deploying Nimble provides local privacy like Laya but with the ease of structured API-style HTTP requests. The model runs locally via Ollama (v0.35.0+) utilizing its optimized decision endpoints.

Step 1: Pulling the Model

First, ensure you are running Ollama version 0.35.0 or newer. Download the Q8_0 quantized model directly through your terminal CLI.

ollama pull nimble

Step 2: Executing the Local Decision API Call

Nimble utilizes Ollama's specialized /v1/systemone endpoint. You pass the context state along with up to 64 structured questions to receive an instantaneous probability matrix.

import requests
import json

ENDPOINT = "http://localhost:11434/v1/systemone"

payload = {
    "model": "nimble",
    "state": "The user clicked 'Cancel Subscription' but has 14 days left.",
    "questions": [
        {
            "id": "user_intent",
            "type": "choice",
            "options": ["churn_risk", "downgrade", "inquiry"]
        }
    ]
}

response = requests.post(ENDPOINT, json=payload)
response.raise_for_status()
print(json.dumps(response.json(), indent=2))

Frequently Asked Questions (FAQ)

Q: Are decision models going to replace LLMs? A: Not at all! They are designed to work alongside them. You should still use LLMs (like GPT-4 or Claude) for generating creative text, summarizing, and chatting. Use decision models for the "plumbing" of your app: routing tickets, classifying data, and triggering specific functions.

Q: Can I run Nimble or Laya without an expensive GPU? A: Yes! Laya is a smaller BERT-based model that can run exceptionally well on standard CPUs. Nimble, while a 9B parameter model, can be heavily quantized using Ollama to run smoothly on standard consumer hardware (like a Mac M-series chip) without a dedicated cloud GPU.

Q: Do I always have to fine-tune Laya? A: While Laya does have some basic zero-shot capabilities out of the box, its real power is unlocked when you fine-tune it on your specific domain data. If you want high accuracy without fine-tuning, you are better off using Jev (cloud) or Nimble (local).

Q: How do these models guarantee structured output compared to an LLM? A: LLMs generate text token-by-token, which means they can hallucinate or break formatting (like generating invalid JSON). Decision models do not generate text; they read the prompt and output a mathematical probability score across predefined categories. You get a guaranteed, parseable data structure every single time.

Final Verdict

Stop using massive conversational LLMs for simple routing logic. If you have the budget and need immediate enterprise reliability without managing infrastructure, integrate Jev today. If you are building a privacy-first agent operating at a massive unmetered scale, need sub-40ms latency, or want to embed intelligence directly into a browser app via WebAssembly, take the weekend to fine-tune Laya. If you want the best of both worlds—strong out-of-the-box zero-shot capabilities coupled with the absolute privacy of running locally on your own hardware—spin up Nimble through Ollama. All three models represent a massive leap forward in making AI agents faster, cheaper, and vastly more deterministic.

AA
Written byCore Contributor

Atif Ali

Full Stack Engineer

Full-Stack Engineer at Groooh with extensive expertise in cross-platform mobile development and cloud systems. Specializes in building unified digital ecosystems using React Native, Next.js, and PostgreSQL. Known for writing maintainable, test-driven code and optimizing app performance from database query tuning down to 60fps mobile UI interactions.

Have a Project in Mind?

Ready to build your next breakthrough product?

Let’s collaborate on your architecture roadmap, MVP sprint, or full product build with our senior team.

Start a Project

Relevant Blogs