Language model
PLANS, TALKS, WRITES
Interprets the conversation, plans richer tasks, generates missing text, synthesises research and performs bounded visual setup or recovery.
Browse, research and work through supported websites in plain language. jevry shows every action it takes, grounds every answer in the page it actually read, and hands control straight back to you when something matters.
Actions run in a visible, persistent browser — not a headless farm.
Answers cite what the page actually showed, with retained receipts.
Stop, redirect or take over mid-task and keep the context.
Bring a local CLI login or your own API provider for text and planning.
You state the task. A text model plans it, Jev chooses each supported action, and the runtime decides what is actually allowed to happen.
PLANS, TALKS, WRITES
Interprets the conversation, plans richer tasks, generates missing text, synthesises research and performs bounded visual setup or recovery.
DECIDES THE NEXT ACTION
TypeSafe's System One model returns typed judgments: Choice selects an option, Score rates against a rubric, Noul answers yes/no. Jevry's action loop is built on Choice.
EXECUTES AND VERIFIES
Builds the action space, validates responses, rechecks page freshness, dispatches input, retains receipts and enforces stop conditions.
Jev turns the controls a page actually offers into a finite set of typed choices. The selected choice is resolved by code that this project owns — so the decision can never become a script.
Jev chooses from the browser's actual capabilities. Its output never becomes arbitrary JavaScript, a generated selector or a shell command. A valid choice can still be the wrong choice: structured output is an interface guarantee, not proof of task success.
Anything outside that list is a boundary, not a surprise: the run stops and tells you rather than guessing at a control it cannot name.
The scenarios below use the app's own action vocabulary. Run one and the feed fills in as it goes; step, stop or redirect it at any point.
Cancellation policies — two sources
“Free cancellation up to 14 days before arrival. Two changes are included; after that the fare difference applies.”
“Cancel up to 7 days before check-in for a full refund to your wallet. Inside 48 hours, the first night is charged.”
| Question | What the pages said |
|---|---|
| Free window | reading… |
| Refund method | reading… |
| Both agree on | reading… |
Booking form · 5 fields, 4 sections
Weekend in Lisbon · Fri–Sun
2,048 tile · no powerups
“Compare the cancellation policies on these two pages and cite them.”
Two sources, one table, every sentence traceable to the page it came from.
Plan“Fill out the web form at example.org/booking and submit it.”
Four sections, five fields, then a follow-up that keeps every value you already gave.
Plan“Plan a weekend in Lisbon. Give me options, then book the best one.”
The search runs; the purchase stops and hands control back before anything is paid for.
Plan“Play this game and win it.”
Board reading happens locally, so the move loop stays fast and does not burn vision calls.
PlanEvery dispatched action, in order, with the choice that produced it.
Finished. The receipts above are what the app keeps: which node was used, what was dispatched, and what the page showed afterwards.
Half of these were measured on fixtures or a fixed corpus. The other half are directions to try.
“Search for Paris for two guests.” → “Now London, same guests.”
Exercised with live providers on controlled travel fixtures.
measured on fixtures“What result is currently shown?”
Read-only follow-ups tested without extra browser actions.
measured on fixtures“Compare the cancellation policies on these two pages and cite them.”
Live fixed-corpus study; native source discovery and citation checks.
measured on fixtures“Find the records matching these filters.”
Included in the WebArena-Verified development subset; multi-page completeness remains a known failure mode.
measured on fixtures“Find this library's getting-started guide.”
Supported link and search controls on accessible public sites. No separate success-rate claim.
explore“Play this game and win it.”
Recorded public 2048 victory plus a separate instrumented acceptance run.
measured on fixtures“Win this board game.” / “Reach the other side.”
Three-in-a-row and crossing-game fixtures reached their observed victory states.
measured on fixtures“Stop that search; compare these sources instead.”
Redirect cancellation, retained context and restart persistence tested in native workflows.
measured on fixturesExample prompts are starting points, not promises that every website is supported. 8 scenarios are documented here; the evidence column says which were measured and which are there to explore.
One prompt: “Play this game and win it.” A vision model calibrates the board once; after that local pixel matching and OCR read the tiles, and bounded lookahead ranks the legal moves. Jev selects every dispatched move.
The widget replays a recorded run: the same seeded board and the same move list, re-simulated in your browser. It reaches the 2048 tile in 1,333 moves. The instrumented acceptance run quoted further down is a different game with its own totals.
Move list: recorded locally by the same bounded-lookahead selection the app uses.
Three kinds of evidence, each from a named build. Nothing here is a leaderboard score and nothing is averaged into one.
The speed-006 candidate ran the packaged beta.5 app with live providers through Jevry's normal task interface. No task-solving script and no reference answers were supplied to the acting model.
| Same 12-task development subset | speed-005 | speed-006 |
|---|---|---|
| Official passes | 7/12 | 10/12 |
| Total actor time | 456.707 s | 542.831 s |
| Median task time | 23.367 s | 44.169 s |
| Median Jev request | 443 ms | 392 ms |
| Measurement | Recorded result |
|---|---|
| Points | 20,840 |
| Game-counted moves | 986 |
| Median Jev decision | 343 ms |
| Median key interval | 652 ms |
| 95th percentile key interval | 770 ms |
| General-purpose visual reviews | 0 |
| Total run, setup included | 11 min 47 s |
beta.1 study: three five-turn sessions per implementation, same fixtures, same resolved Claude model
A historical study, kept separate from the numbers above. Jevry also used Jev for browser decisions while the adapted Firecrawl source graph did not use Jev.
The rule we hold to: These are three different kinds of evidence and must not be combined into one score. Measurements belong to the build and protocol they were recorded under.
Development evidence. A repeatedly used development subset of an 812-task benchmark — not a held-out or full-suite result, and not a leaderboard submission. Accuracy improved while overall execution became slower: faster individual decisions did not remove repeated navigation and broad completion checks. Tasks 47 and 102 still failed.
Game evidence. One successful instrumented run. Initial setup still takes tens of seconds. The launch film is a separate recording with its own totals, and the replay on this page is a third run — same move selection, different game.
The loop is deliberately boring: the interesting decisions are made where they can be checked, and every step leaves something behind.
Page text and compatible controls are captured together, keeping real DOM-node references in an isolated browser world. The action refers to the observed node, never a selector invented afterwards.
Eligible pages become a Choice over complete operation/target pairs. Other states use an operation question plus speculative target questions in the same request; only the selected branch is consumed.
Routine action selection costs one inference round trip. Explicit user literals can be offered as field-value choices; missing text can call the text helper. Completion and recovery can add calls.
The offered choice set and probability distribution are checked, then the document, target, current value, visibility and occlusion are rechecked. A stale decision is discarded, not replayed.
Input receipts survive navigation failures. Once a mutation may have started, uncertainty stops the run instead of firing the action again.
Completion considers goal evidence and coverage. Seeing a matching row does not establish that every requested page was read — model-assessed completion stays distinct from independent verification.
Choices are compiled from what the page currently offers, then resolved to an operation this project owns. Nothing in the model's output is interpreted as code.
Fan-out lets independent questions travel in one request. A target question cannot read another question's answer, so a speculative branch carries its own assumptions instead of borrowing state it never saw.
This is a source preview: the app is real, the packaging is early, and the first-success path is deliberately small.
Download the developer build for your platform, or run the source preview from a clone.
npm installnpm run devYou need both connections. Setup validates them before the first task.
Jev: jev-latest @ api.typesafe.ai/v1/systemoneText: Codex CLI · Claude Code CLI · API keyOpen a small public page, ask one page-grounded question, then compare the answer with the page you can see.
What result is currently shown?22.12 or newer, with npm
A TypeSafe API key — the System One endpoint, not Chat Completions
A locally authenticated Codex or Claude Code CLI, or your own API provider
A provider/model that accepts images, for games and other visual flows
| Connection | How it is set up | What to know |
|---|---|---|
| Jev — required for browser decisions | TypeSafe API key, default model jev-latest | Requests use your TypeSafe account. A saved CLI login does not establish quota or model access. |
| Text model — required for conversation and planning | Codex CLI, Claude Code CLI, or an OpenAI-compatible / Anthropic API connection | The app can privately install a missing supported CLI through its Connect flow; npm is required. |
| Vision — for supported visual tasks | Any connected provider/model that accepts images | Used for game calibration and bounded recovery; local perception keeps the move loop out of the vision loop. |
Model usage can cost money and there is no fixed per-task cost: it depends on the models, the context, the number of actions and any retries. Check both accounts' usage before running a long task. Connections and conversation archives are encrypted locally — but local storage does not imply offline inference.
A short list, published on purpose. Everything here has been verified to be unsupported or unverified rather than assumed to work.
Sensitive actions such as purchases and sending messages hand control back to you. Label-based guards are not a complete defence against hostile pages, so keep an eye on the tab you handed over.
Named work, in the order it is being done — not a promise of dates.
Wider held-out runs with retained failure traces.
More reliable coverage checks and multi-page reading.
Repeat the provider matrix on clean machines and current models.
Signing, notarisation and automatic updates for the packaged apps.
Fresh-install testing as part of the release gate, not after it.
$JEVRY settles directly between wallets on Solana: no invoice, no middleman, no account. The token is a single public address published here first, and this page is the only one that speaks for it.
Every field above is verifiable from the public address; compare it before you buy.
Coming soon — the contract address is published on this page first.
Coming soon — the contract address is published on this page first.
A desktop browser with an agent inside it. You state a task in plain language; it observes the page, chooses one supported action at a time, executes it in a real Chromium tab and shows you what happened.
Jev is TypeSafe's System One model: it evaluates supplied context and returns typed judgments and probabilities. jevry is an independent project built around that model and does not train or host it.
Yes. Jev is required for browser decisions; a text model is required for conversation and planning. Vision is only needed for supported visual tasks such as games.
No. Website support varies. HTML and ARIA controls, native selects, open shadow roots, nested scrolling and bounded same-origin iframes are supported. Closed shadow roots, some cross-origin surfaces, uploads and extensions are not.
It stops and hands control to you for sensitive actions such as purchases and sending messages. Guards are label-based, which is a useful barrier, not a complete defence against a hostile page.
The run was recorded with the board state read back after every dispatched key, and the page you are reading shows a separate deterministic replay you can watch end to end — including its score and move count.
No. They come from a development subset of a much larger benchmark, scored by the official evaluator. They are reported as development evidence with their build boundaries, not as a held-out or certified score.
The app is open source and free. $JEVRY is the project's own token rail on Solana: a single public address, published on this page first, with no invoice and no middleman.
No. This is a static site with no analytics, no cookies and no third-party scripts. Fonts and assets are served from this domain.
Hand it one task. If it cannot do it, it will tell you where it stopped — and hand the tab back.