Blog ·AI Engineering·

Taking Jev for a Test Drive

Nine models played 27 seeded games of 2048. TypeSafe AI's Jev never beat a 512 tile, and still made every small chat model look like the wrong tool: 2.6x the score of Haiku 4.5 at a 17th of the cost and 274 ms a move.

Jev's first game of 2048 finished in 160 seconds. It made 481 moves, built a 512 tile, and the whole thing cost 1.6 cents. Claude Haiku 4.5, playing the same board with the same random tiles, was on move 78 when Jev's game ended. Haiku died at move 151 with a 128.

That was the moment this stopped being a toy.

What Jev is

Jev is TypeSafe AI's "System One" model. It does not chat. You hand it a piece of state (a JSON object, a string, an array) and a set of typed questions: a choice among named options, a score on a rubric, a yes-or-no. It answers every question in one pass with probabilities attached, usually in 100 to 300 milliseconds. Input tokens cost $0.042 per million. Output is free, because there is almost none.

It also can't do arithmetic, and it won't write you a sentence. TypeSafe is upfront about both. The pitch is a model for the decision layer of software: routing, classification, rubric grading, verification. The fast, cheap judgment calls that currently get bounced to a chat model with a JSON schema bolted on.

I wanted to know how good that judgment actually is, so I gave it a game.

Why 2048

Nobody needs an AI to play 2048. But a game is a clean decision loop. Every move is a choice among at most four options, the consequences are computable, and a single game produces 200 to 1,000 decisions you can score. It is also the same shape as a lot of real agent work: here is the state, here are the legal actions, pick one.

Your turn

Play the game the models played

Same rules, no previews, no model. See if you can beat Jev's 512.

Score

0

Best

0

Click to play
Arrow keys or WASD, or drag across the board.

So I built a small Nuxt app where a model plays 2048 and explains itself, then a harness that plays without the UI, as fast as the gateway answers. Nine players through the Vercel AI Gateway: Claude Opus 5.5, Sonnet 5, Haiku 4.5, and Fable 5.1; GPT-6 Astra, Sol, and Luna; Gemini 3.1 Pro; and Jev. Every model played the same three seeded games, so the random tiles landed in the same cells for everyone and the boards only diverged through the models' own choices. A 1,000-move cap. Twenty-seven games. About $64 in gateway spend.

The chat models got the board as text, the exact board that would result from each legal move, a short strategy brief (keep the biggest tile in a corner, keep the edge ordered, keep room), and a JSON schema asking for a move and one sentence of reasoning. Reasoning effort was set low across the board to keep moves in the two-to-five second range, and that choice matters later.

Wiring Jev differently

Jev can't read a board. Give it a grid of numbers and ask which slide is best and it will guess, because "is 128 bigger than 64" is exactly the kind of question it is bad at. So the code does the arithmetic and Jev does the judging.

For each legal move, the engine simulates the slide and writes down what happened in words:

Slide left: one merge, some room (6 empty cells), largest tile stays anchored in the top-right corner, the edge next to it stays ordered large to small, 2 pairs lined up to merge next.

Then it asks one choice question with those descriptions as the options and a one-paragraph instruction about what a good move looks like. Jev returns a pick and a probability for each option: "down 71%, left 29%". That is the entire integration. About 40 lines.

Max tile per game

How far each game got

One dot per seeded game, bar marks the median. Same seed, same spawns.

6412825651210242048GPT-6 AstraClaude Opus 5.5GPT-6 SolJevClaude Fable 5.1Gemini 3.1 ProGPT-6 LunaClaude Sonnet 5Claude Haiku 4.5
Jev (evaluation model) Chat models

What happened

Nobody reached 2048. Three models built a 1024 reliably: GPT-6 Astra, Claude Opus 5.5, and GPT-6 Sol, and Astra put up the two highest scores of the run, 16,800 and 16,796. The best of these are genuinely good at the game, given the previews.

Jev topped out at 512. On every seed. Its three games scored 7,120, 5,556, and 6,828, and it was the only player with no variance to speak of. Every chat model had at least one collapse: Opus fell to a 256 on seed 1, Fable 5.1 to a 128 on seed 3, Haiku to 128 twice.

Below the 1024 club, the picture inverts. Sonnet 5, Haiku 4.5, GPT-6 Luna, and Gemini 3.1 Pro all landed at a median of 256 or 128. Jev beat every one of them on every seed.

Decision latency

How long a move takes

Median per move with the 95th percentile whisker, log scale.

100 ms1 s10 s20 sJevfastest · 274 msClaude Haiku 4.5Claude Sonnet 5Claude Opus 5.5GPT-6 AstraClaude Fable 5.1GPT-6 LunaGPT-6 SolGemini 3.1 Proslowest · 8.5 s
Jev (evaluation model) Chat models

Speed

Jev's median move took 274 milliseconds. The fastest chat model, Haiku, took 1.5 seconds. Gemini 3.1 Pro took 8.5. Under nine-way concurrency the chat models' 95th percentile ran 8 to 19 seconds a move regardless of size, which is the number that actually decides how long a game takes. Jev's 95th percentile was 479 milliseconds.

A full Jev game averaged 2.7 minutes. Haiku averaged 8.3 minutes for a game less than half as long, and Astra's 825-move games averaged 88.

Cost vs. result

What a game costs and what it buys

Mean gateway cost per game against the median max tile, both log scales. Up and left is the bargain.

6412825651210242048$0.010$0.100$1.00$10.00GPT-6 AstraClaude Opus 5.5GPT-6 SolJevClaude Fable 5.1Gemini 3.1 ProClaude Sonnet 5GPT-6 LunaClaude Haiku 4.5
Jev (evaluation model) Chat models

Cost

Jev played 1,334 moves for 4.5 cents. Per thousand moves: Jev 3 cents, Luna 19 cents, Haiku $1.21, Sol $2.62, Opus $6.01. Astra, the top scorer, ran $8.50 a game.

The comparison that matters

If you are choosing a model for a decision loop, the honest comparison is Jev against the small, fast chat models people reach for when latency and cost matter.

Jev against the small chat models

Same three seeds, same previews

Model Max tile per seed Mean score Median move Per 1,000 moves
Jevevaluation model5125125126,501274 ms3¢
Claude Haiku 4.51281285122,5321.5 s$1.21
GPT-6 Luna2562561282,5455.4 s19¢
Claude Sonnet 52561282562,4091.9 s$2.78
Same three seeds, same previews. Chat models at low reasoning effort.

Against Haiku 4.5, Jev scored 2.6 times higher, moved 5.6 times faster, and cost 17 times less. Against Sonnet 5 at low effort, 2.7 times the score at a 38th of the price. Against GPT-6 Luna, which is the cheapest chat model here, Jev was still 2.8 times cheaper and 20 times faster per move, with 2.5 times the score.

Jev wins on the division of labor. Compute the consequences in code, then ask for a judgment, and you beat handing a language model the raw state and hoping. Small chat models spend most of their tokens re-deriving what a for-loop already knows, and at low reasoning effort they re-derive it badly. Jev never has to derive anything. It is the right tool for the ranking step, and it gets to be cheap and fast because it does nothing else.

Where it stops

The 512 ceiling is real and it is structural. Jev only ever sees one move ahead, described in words. Building a 1024 in 2048 means planning several merges out, and a greedy policy over one-ply features will not do it. The chat models that won were doing something more like planning, slowly and expensively.

Two things would probably push Jev past 512, and neither is hard. Describe two-ply outcomes (the best reachable board after this move and the next) instead of one. Or ask a score question per candidate against a rubric instead of a single choice, so the harness can combine the answers with a heuristic in code. Both are afternoon projects.

Three caveats on the numbers. Three seeds is barely a benchmark. Every chat model ran at low reasoning effort, and Sonnet 5 and Fable 5.1 in particular would likely climb at medium or high, at several times the price. And the latency figures come from a run where nine models were hammering the gateway at once, so absolute values are noisier than the rankings.

What I'd use it for

Anywhere code can enumerate the options and a model needs to pick one: routing a support ticket, choosing the next tool in an agent step, grading a draft against a rubric, deciding whether a generated answer passes a guardrail. Monument Labs products make a few of those calls today on chat models with structured output, at a second or three each. Jev makes the same call in a quarter second for a rounding error.

The game, the harness, and every per-move log are public at github.com/GKjohns/2048-ai. The results page at /bench has the full charts, and bench/run.ts will reproduce the run for about $60.