Input
model not loadedTyped decisions jev-web · one forward pass
—Load the model, then press Decide. Editing the state re-runs automatically.
Raw text parsing regex / keyword rules
—The other way to do this: a hand-written rule per question.
The baseline is a fair keyword matcher — it wins on easy text and collapses on negation, contrast and mixed-topic states. It is also N separate rule sets to maintain, and it can only ever return a hard label: no distribution, no confidence, no way to threshold it.
Network panel proof, not a promise
onlineoutbound hosts observed by the Performance API
jev-web vs. calling a hosted LLM API with fetch()
| jev-web (this demo) | hosted LLM API via fetch() | |
|---|---|---|
| Cost per decision | $0.00 — weights are static files | per-token, forever; a busy form multiplies it by every visitor |
| Latency | measured above — all questions in one pass | typically 0.4–3 s, plus a network round trip |
| Output | calibrated distribution per question, every question in one pass | prose you then have to regex back into a distribution |
| Determinism | same input → same probabilities, always | temperature 0 + structured outputs helps, still not guaranteed |
| Offline | works in airplane mode once cached | dead |
| Privacy | state text never leaves the device | your users' text is sent to a third party on every call |
| Ops burden | static files on a CDN, no API key, no rate limit | key management, rate limits, quotas, uptime, spend alerts |
| Ceiling | only picks among options you supply — never writes, never explains | general purpose; will happily hallucinate an option you never offered |
The honest trade: jev-web is a classifier, not a chatbot. If you need free-form answers, you need a generator. If you need a routing decision, a score, or a yes/no with a confidence you can threshold on — this is the cheaper, faster, private one.
Install
npm i jev-web @huggingface/transformers
@huggingface/transformers is an optional peer dependency (imported lazily); hosts that
already ship their own copy can inject it via createDecider({ transformers }).
This page runs on GitHub Pages, which cannot set COOP/COEP, so the ONNX runtime stays
single-threaded; run npm run dev locally for the threaded WASM build.