Most LLMs I use write things. Paragraphs, SQL, commit messages. Then I came across Jev, a model from TypeSafe on OpenRouter that doesn't write anything at all. You give it some text plus a list of typed questions, and it gives you back probabilities. That's it. No prose, no chat, no "Certainly! Here's…".

That sounded boring until I realized it's exactly what most apps actually need. So I built Jev Playground to find out what a classifier model feels like when you put it right under your fingers.

What Jev Actually Is

Jev answers three kinds of questions about a piece of text (its "state"):

  • noul — a yes/no, returned as a probability from 0 to 1. "Is this a video call?" → 0.97.
  • choice — pick one of the options you define, with the full probability spread. "Which card is this?" → event 0.94, reminder 0.04, note 0.02.
  • score — a position on an ordered rubric. "How urgent is this?" on a scale from "not urgent" to "needs attention now".

Your code owns the workflow. The model just answers narrow questions, and it answers all of them in parallel in a single call. For routing, ranking, verification and gating, that's a much better fit than asking a chat model to "respond in JSON, please" and hoping it listens.

The Playground

The app is one text box. As you type, it turns into whatever you meant:

  • team standup monday 10am on meet → an event card with Monday, 10 AM and a video-call badge
  • buy coffee, batteries and limes → a shopping checklist
  • split $186 dinner between 4 → $46.50 each
  • 35000 feet in meters → 10,668 m, because I spend my days around aviation data

There are 19 card types in total: timers, habits, unit conversions, time zones, polls, countdowns, dice rolls and more. Every keystroke (debounced) sends the text to Jev with 14 questions at once: which card, how complete the input is, is it a question, is it recurring, how urgent, the tone, is it in person or on a call, flight or train, what category the expense is, and so on.

I didn't come up with the idea. It's based on Shapeshift by anishfn, an open-source Next.js app that uses TypeSafe's own SDK. I wanted it running on my own site with OpenRouter, which meant reworking how it talks to the model and where it runs.

Jev Decides, Code Computes

This is the part I liked most about the design. Jev never extracts a value. It never parses a date, adds up a bill or converts miles to kilometers. It only answers what kind of thing is this? and a handful of yes/no signals. Everything else is plain deterministic code: chrono-node for dates, a unit converter, a small math evaluator.

That split is what makes it trustworthy. A classifier being 94% sure something is an event is fine; the worst case is the wrong card, and you can pick the right one from a "did you mean" chip. An LLM being 94% sure that $186 ÷ 4 is $46.50 is not fine. As a data engineer this felt very familiar: let the fuzzy component route, let the exact component calculate.

There's also a small state machine in between so the UI doesn't flicker. Raw probabilities jump around as you type half-words, so a card only switches when a challenger wins twice in a row (or is very confident), and signal badges have an on/off hysteresis band. It feels calm even though the model is being hit on every pause.

Hosting It on a Static Site

My site is GitHub Pages, which is just files. No server, which means nowhere to hide an API key, and any key shipped in JavaScript gets scraped within minutes. Shapeshift solves this with a Next.js server route, which GitHub Pages can't run.

So the app is split in two:

  • The page is a Next.js static export, served from /jevplayground on this site like any other folder.
  • A Cloudflare Worker (about a hundred lines) holds the OpenRouter key as a secret. The browser sends it { text }; the Worker adds the key, asks Jev all 14 questions, and returns a clean result.

The Worker also does the boring safety work: it only accepts requests from my domain, rate-limits each visitor to 30 calls a minute using Cloudflare's built-in rate limiter, and caches identical inputs at the edge so retyping the same sentence is free. If the Worker is ever down or rate-limited, the app quietly falls back to an offline keyword classifier, so the page never breaks.

Cost-wise it's basically nothing. A full 14-question call costs around $0.00002, and Cloudflare's free plan covers 100k requests a day. I put a spending cap on the OpenRouter key anyway.

What I Learned

Not every AI feature needs a generative model. A huge share of "AI" in real products is really classification: route this ticket, flag this transaction, pick this template. Getting calibrated probabilities back instead of text you have to parse removes a whole class of bugs.

Ask many small questions, not one big one. Asking "is it a video call?" as its own yes/no is more reliable than asking the model to fill in an object with a mode field. And because Jev answers everything in parallel, adding a question barely changes the latency.

Static hosting plus a tiny edge function goes a long way. I didn't need Vercel or a backend. One Worker, one secret, one rate limit, and the static site stays exactly as simple as it was.

Read the response, don't assume it. OpenRouter returns score as an expected index on the rubric (0 to 2 for a three-step scale), not a 0–1 number. Checking one real response before wiring up the UI saved me from urgency badges that would never have turned on.

Try It

Open the Jev Playground and type anything. Try flight to tokyo in march, 9am cst in tokyo or run 100 miles this month, 34 done. Press / to see every card type, or add ?debug=1 to the URL to watch the raw probabilities move as you type.

The code, including the Worker, is on GitHub.