Jev by TypeSafe AI, explained: what it does, what it costs, and where it fails
Jev answers yes/no, choice and score questions in milliseconds for $0.042 per million tokens. Real tests, independent benchmarks, best uses and limits.

What Jev is, and what it is not
Jev is an AI model built by TypeSafe AI, a company that came out of stealth in September 2026 with a $40 million seed round led by DCVC. One of its co-founders, Diogo Almeida, worked at OpenAI on the research that turned GPT models into instruction-following assistants.
Most AI models you know are built to write: answers, emails, code. Jev is built to decide. You give it some data and a list of questions, and it returns an answer to every question at once, with a probability attached. TypeSafe calls this a "System One" model, after the fast, intuitive kind of thinking people do without deliberating.
That design has two consequences you should know before anything else:
- It cannot write anything. Not a reply, not a caption, not a summary. If the job needs new text, you still need a normal language model.
- It only reads text. It cannot look at an image, a video or a design. Anything visual has to be described in words first.
The three kinds of questions it answers
Everything Jev does fits into three question types:
| Type | What you ask | What you get back |
|---|---|---|
| Yes / no ("noul") | "This email needs a reply." | A probability from 0 to 1, for example 0.93 |
| Choice | "Which team should handle this ticket?" with a list of options | The chosen option, a probability for every option, and a confidence score |
| Score | "How urgent is this?" on a scale you define | A score on that scale, with a confidence score |
You can ask many questions about the same piece of data in one request, and they are answered together. That is where most of the speed and cost advantage comes from.
What it costs
- Input: $0.042 per million tokens.
- Output: free.
- Limits: up to 64,000 tokens per request and 1,200 requests per minute on the standard tier. TypeSafe says rate limits can change without notice.
Two details change the real price:
- Jev counts tokens differently. An independent benchmark sent the same English sentences to Jev and to an OpenAI model. Jev reported 340 tokens where OpenAI counted 162. So the effective price is closer to $0.088 per million of the tokens you are used to.
- Other languages cost more. In a separate audit, Russian text used about 3 times as many tokens as the same text in English. Arabic has not been measured publicly, but expect something similar. On Arabic text, Jev's price ends up close to the cheapest general-purpose models.
Our test: 1,000 emails, 7 questions each
We asked Jev seven questions about each of 1,000 emails (needs a reply, urgent, has a deadline, is a sales lead, is spam, category, sentiment) and ran the same job on GPT-5.6 Luna, OpenAI's cheapest model at the time.
| Jev | GPT-5.6 Luna | |
|---|---|---|
| Time | 7 seconds | 5 minutes |
| Cost | $0.09 | $0.62 |
That is roughly 7 times cheaper and 43 times faster. One fairness note: we ran Luna one request after another. Running it in parallel would shrink the speed gap, though not the price gap.
What independent testers found
Jev launched to a wave of excitement, so it is worth separating the company's claims from outside measurements.
- Speed and cost. TypeSafe advertises 193x faster and 444x cheaper. The one public benchmark that ran both models on the same tasks and the same clock measured 5x faster and 41 to 50x cheaper per call against a comparable OpenAI model. Still impressive, but a different order of magnitude.
- Accuracy. On simple classification it lands close to large models. On a harder banking-intent task, a small model trained on examples of that task beat it (93% against 83%).
- Other languages. In a pre-registered English-versus-Russian audit, accuracy on a reasoning task fell from 88% to 77%. Confidence became less reliable too: at 90% confidence or higher, Jev was right 97% of the time in English and 89% of the time in Russian.
- Repeatability. The same question with the same data changed its answer between 0.5% and 3.3% of the time.
- Out-of-scope inputs. Without a "none of these" option, Jev picks something anyway. One evaluation had it classify a cake recipe as a technical support issue with 94% confidence.
None of this makes Jev bad. It makes it a tool with a specific shape, and you should use it with that shape in mind.
Where Jev is genuinely useful
For anyone with an overflowing inbox
- Sorting email: needs a reply, urgent, invoice, lead, newsletter, spam. This is exactly the job it was built for.
- Filtering a feed: flag posts in your niche, breaking news, or text that reads as AI-generated, and treat the flag as a hint, not a verdict.
For content creators
- Sorting hundreds of comments or DMs by topic so you can spot the questions your audience keeps asking.
- Checking English copy against a list of rules before you post. One open-source linter built on Jev, Sniff Test, raised 1 false alarm on 54 clean paragraphs where a general model raised 37. It also missed more real problems than the strongest models did, so use it as a first pass.
For businesses
- Tagging public competitor ads. Open Meta's Ad Library yourself, collect the ad text you see, and let Jev tag each ad by hook, offer, call to action and customer awareness stage. Ads that have run for 30 days or more are usually the profitable ones. Two rules: Meta's official Ad Library API only returns political ads and ads shown in the EU, and Meta's terms forbid automated collection from the website, so read the pages the way a person would.
- Routing support tickets or form submissions to the right person, as long as you test it on your own messages first.
For founders building products
- A fast first filter in front of an expensive model: judge every item with Jev, then send only the few that matter to the model that writes.
Where not to use it
- Anything that needs writing. It does not write.
- Images, video and design. It cannot see, and it is documented to struggle even with color codes written as text.
- Numbers, prices and math. TypeSafe's own documentation says it is not a calculator.
- Sensitive decisions in Arabic or other non-English languages, until you have measured it on your own content.
- Anything that must give the same answer every time. Answers can flip on identical inputs.
- Trading with real money. Demos of Jev "trading Bitcoin" are demos. Do not put money on a model that is weeks old.
- Compressing your AI coding assistant's memory to save tokens. The popular plugins for this were measured independently and did worse than not compressing at all, and some upload your whole session to a third party.
- Private or customer data, unless you have read the terms (next section).
If you are building a product with it
- Put it in front of the expensive model, not in place of it. Nearly every "90% cheaper" demo works because a cheap judge removes most of the items before an expensive model runs. That structure is the saving, and a cheap conventional model often captures most of it. Compare both before you commit.
- Ask many questions about one item per request, not one question about many items. TypeSafe measured 13 questions in one request at 12x cheaper and 10x faster than 13 separate requests. Packing many items into one request is the documented way to lose accuracy.
- Always offer an escape option such as "none of these" or "unclear".
- Fail open. There is no SLA. Wrap calls in a short timeout and have a fallback path. The public status page shows about 99.86% uptime so far.
- Pin the model version (for example
jev-1.13.0) instead of the movingjev-latestalias, so answers do not change under you. - Read the terms before sending customer data. TypeSafe's Master Customer Agreement, updated on 19 September 2026, says it will not train models on your data without consent, but it also lets TypeSafe "process telemetry without restriction". There is no general zero-retention option when you go direct. Per-request zero data retention is available through Vercel's AI Gateway.
- Consider self-hosting. Open models with the same yes/no, choice and score interface already exist and can run on your own servers, which removes the data question entirely.
How to try it
- Request access at typesafe.ai. Access is still early, so there may be a wait.
- Open the console and try the Playground. You can paste data and questions without writing code.
- Create an API key and keep it in an environment variable, never in your code.
A minimal request looks like this:
curl https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-1.13.0",
"state": { "email": "Hi, the invoice for order 3312 failed again. Can you check today?" },
"questions": {
"needs_reply": { "type": "noul", "instructions": "This email needs a reply." },
"category": {
"type": "choice",
"instructions": "What is this email mainly about?",
"criteria": {
"billing": "Payments, invoices or refunds",
"support": "A problem using the product",
"sales": "Wants to buy or get a quote",
"none_of_these": "None of the above"
}
}
}
}'How we see it at Raqi
We evaluated Jev over three rounds for Raqi's AI agents, which reply to customers on WhatsApp, Instagram and Messenger in many Arabic dialects. On a small set of Iraqi, Egyptian and Arabizi customer messages, it answered all ten questions correctly: whether the customer wanted a human, what they wanted, and their mood.
We are still not putting it in front of customer conversations. Ten messages are not a benchmark, Arabic accuracy has not been measured publicly, and customer messages are exactly the data we will not send to a new provider under today's terms. Where we do see a fit is fast tagging of public information, such as competitor ads, and we will share what we learn when we have real results.
Frequently asked questions
Is Jev free?
No, but it is cheap. You pay $0.042 per million input tokens and nothing for output. Our 1,000-email test cost 9 cents.
Can Jev write replies or content?
No. It only makes decisions about text you give it. Pair it with a normal language model when you need writing.
Does Jev work in Arabic?
It reads Arabic, and it did well on our small test, but there are no public Arabic benchmarks. Expect lower accuracy and higher token counts than in English, and test on your own data before relying on it.
Is my data safe with Jev?
TypeSafe says it will not train on your data without consent, but its agreement allows unrestricted use of telemetry derived from your requests, and there is no zero-retention option when you use it directly. Keep sensitive and customer data out unless you route through a zero-retention gateway or self-host an alternative.
Is Jev really 400 times cheaper than GPT?
That is the company's figure. An independent head-to-head measured about 41 to 50 times cheaper per call and about 5 times faster. On non-English text the gap shrinks further.
How do I get access?
Sign up at typesafe.ai. Access is in early release, and the console includes a Playground where you can test it without code.
Sources
- TypeSafe AI and the Jev documentation
- TypeSafe Master Customer Agreement and status page
- jev-eval: speed, cost and tokenizer benchmark
- English versus Russian pre-registered audit
- PriorBench evaluation (out-of-scope inputs, thresholds)
- Baselines evaluation against small and supervised models
- Sniff Test prose linter
- Meta Ad Library API
- Vercel AI Gateway zero data retention
Run all of your social media from one place
Raqi is an AI social media management environment. Your conversations, customers, content and ads live together, and AI works across all of it.
Every conversation in one inbox
Instagram, Messenger, WhatsApp, Telegram, SMS, website chat and email, depending on your plan.
AI agents that speak your customers' dialect
From Starter. They reply in 90 dialects, answer from your knowledge sources and hand the chat to your team when it needs a person.
Free Auto-DMs
Basic Auto-DMs on Instagram and Messenger are free and unlimited.
Workflows
From Starter: questions, follow-ups, payment links and bookings, tested with Simulate before you publish.
Link page and guides
Publish a link page and guides on raqi.link, then send a guide from an Instagram or Messenger Auto-DM. Free keeps 3 guides live.
Content Hub
See which posts perform best and, from Starter, get AI post ideas based on your own results.
Ads Studio
Create images and videos for ads and posts with credits, then schedule and publish them from Raqi.
MCP for Claude and ChatGPT
From Starter, connect Raqi to Claude or ChatGPT and work with your conversations, contacts and content from the chat.
Creative Bridge
A free plugin for After Effects and Premiere Pro. Generating inside it uses your Studio credits.