Skip to content
Desktop browser agent · source preview 0.4.0-beta.13

Give your browser a task.
Watch it work.

Browse, research and work through supported websites in plain language. jevry shows every action it takes, grounds every answer in the page it actually read, and hands control straight back to you when something matters.

MIT licensed Runs in a persistent profile Node 22.12+
Runtime
macOS 12+ · Windows 10+
Engine
Persistent Chromium profile
Decisions
Jev · System One, per action
License
MIT · developer preview
Action feed
CLICK:17 · Cancellation policy
REAL CHROMIUM TAB

Real Chromium tab

Actions run in a visible, persistent browser — not a headless farm.

Page-grounded

Answers cite what the page actually showed, with retained receipts.

Interruptible

Stop, redirect or take over mid-task and keep the context.

Model-agnostic

Bring a local CLI login or your own API provider for text and planning.

How it works

You say it.
Jev does it.

You state the task. A text model plans it, Jev chooses each supported action, and the runtime decides what is actually allowed to happen.

YouYou state the task in plain language — an instruction, not a script.
PlanA text model turns it into a goal with observable success conditions.
JevJev chooses the next supported action from the controls the page actually offers.
ChromiumDeterministic code validates the choice, dispatches native input and keeps a receipt.

Language model

PLANS, TALKS, WRITES

Interprets the conversation, plans richer tasks, generates missing text, synthesises research and performs bounded visual setup or recovery.

Jev

DECIDES THE NEXT ACTION

TypeSafe's System One model returns typed judgments: Choice selects an option, Score rates against a rubric, Noul answers yes/no. Jevry's action loop is built on Choice.

Deterministic code

EXECUTES AND VERIFIES

Builds the action space, validates responses, rechecks page freshness, dispatches input, retains receipts and enforces stop conditions.

One action at a time, from the page's real controls

Jev turns the controls a page actually offers into a finite set of typed choices. The selected choice is resolved by code that this project owns — so the decision can never become a script.

CLICK:17CLICK:42TYPE_TEXT:4SELECT_OPTION:2SCROLL_DOWNSCROLL_UPPRESS_KEY:EnterDONEBLOCKED

Jev chooses from the browser's actual capabilities. Its output never becomes arbitrary JavaScript, a generated selector or a shell command. A valid choice can still be the wrong choice: structured output is an interface guarantee, not proof of task success.

What it can drive today

HTML and ARIA controlsNative selectsContenteditable fieldsOpen shadow rootsNested scrollingBounded same-origin iframesMulti-tab follow-upsKeyboard-only controls

Anything outside that list is a boundary, not a surprise: the run stops and tells you rather than guessing at a control it cannot name.

See it in action

Watch every step.
Redirect or stop it anytime.

The scenarios below use the app's own action vocabulary. Run one and the feed fills in as it goes; step, stop or redirect it at any point.

example.org — opened by jevry in a real Chromium tab
Ready

Cancellation policies — two sources

Cancellation policy · source one14-day window · refreshed today
Cancellation policy · source two7-day window · refund to wallet
Results table appears here once both receipts are in.
Task

“Compare the cancellation policies on these two pages and cite them.”

Two sources, one table, every sentence traceable to the page it came from.

Plan
2 SOURCES READ6 CLAIMS CHECKED1 TABLE WRITTEN
0Actions
0Decisions
0Receipts
392 msMedian decision
Action feed

Every dispatched action, in order, with the choice that produced it.

Use cases

8 things people ask it for.

Half of these were measured on fixtures or a fixed corpus. The other half are directions to try.

Forms with follow-up context

“Search for Paris for two guests.” → “Now London, same guests.”

Exercised with live providers on controlled travel fixtures.

measured on fixtures

Page-grounded questions

“What result is currently shown?”

Read-only follow-ups tested without extra browser actions.

measured on fixtures

Research and comparison

“Compare the cancellation policies on these two pages and cite them.”

Live fixed-corpus study; native source discovery and citation checks.

measured on fixtures

Catalogs and admin tables

“Find the records matching these filters.”

Included in the WebArena-Verified development subset; multi-page completeness remains a known failure mode.

measured on fixtures

Documentation navigation

“Find this library's getting-started guide.”

Supported link and search controls on accessible public sites. No separate success-rate claim.

explore

Numeric puzzle games

“Play this game and win it.”

Recorded public 2048 victory plus a separate instrumented acceptance run.

measured on fixtures

Small discrete games

“Win this board game.” / “Reach the other side.”

Three-in-a-row and crossing-game fixtures reached their observed victory states.

measured on fixtures

Work you can interrupt

“Stop that search; compare these sources instead.”

Redirect cancellation, retained context and restart persistence tested in native workflows.

measured on fixtures

Example prompts are starting points, not promises that every website is supported. 8 scenarios are documented here; the evidence column says which were measured and which are there to explore.

Games

Plays 2048 to win.

One prompt: “Play this game and win it.” A vision model calibrates the board once; after that local pixel matching and OCR read the tiles, and bounded lookahead ranks the legal moves. Jev selects every dispatched move.

2048 TILE NO POWERUPS 0 VISUAL REVIEWS IN THE MOVE LOOP

The widget replays a recorded run: the same seeded board and the same move list, re-simulated in your browser. It reaches the 2048 tile in 1,333 moves. The instrumented acceptance run quoted further down is a different game with its own totals.

2048 — a real canvas game, driven by native keys
READY
Score0
Moves0

Move list: recorded locally by the same bounded-lookahead selection the app uses.

Results

What has actually been measured.

Three kinds of evidence, each from a named build. Nothing here is a leaderboard score and nothing is averaged into one.

10 / 12selected WebArena-Verified development tasks, scored by the unchanged official evaluator

The speed-006 candidate ran the packaged beta.5 app with live providers through Jevry's normal task interface. No task-solving script and no reference answers were supplied to the acting model.

Same 12-task development subsetspeed-005speed-006
Official passes7/1210/12
Total actor time456.707 s542.831 s
Median task time23.367 s44.169 s
Median Jev request443 ms392 ms

Development evidence. A repeatedly used development subset of an 812-task benchmark — not a held-out or full-suite result, and not a leaderboard submission. Accuracy improved while overall execution became slower: faster individual decisions did not remove repeated navigation and broad completion checks. Tasks 47 and 102 still failed.

Game evidence. One successful instrumented run. Initial setup still takes tens of seconds. The launch film is a separate recording with its own totals, and the replay on this page is a third run — same move selection, different game.

Architecture

Observe → choose → execute → verify

The loop is deliberately boring: the interesting decisions are made where they can be checked, and every step leaves something behind.

  1. Observe atomically

    Page text and compatible controls are captured together, keeping real DOM-node references in an isolated browser world. The action refers to the observed node, never a selector invented afterwards.

  2. Compile the decision

    Eligible pages become a Choice over complete operation/target pairs. Other states use an operation question plus speculative target questions in the same request; only the selected branch is consumed.

  3. Let Jev choose

    Routine action selection costs one inference round trip. Explicit user literals can be offered as field-value choices; missing text can call the text helper. Completion and recovery can add calls.

  4. Validate before input

    The offered choice set and probability distribution are checked, then the document, target, current value, visibility and occlusion are rechecked. A stale decision is discarded, not replayed.

  5. Record what happened

    Input receipts survive navigation failures. Once a mutation may have started, uncertainty stops the run instead of firing the action again.

  6. Check the whole objective

    Completion considers goal evidence and coverage. Seeing a matching row does not establish that every requested page was read — model-assessed completion stays distinct from independent verification.

The interface Jev answers on

Choice
Selects one option from the compiled action space — the judgment the loop is built on.
Score
Rates a candidate against a rubric, used where a ranked comparison is the honest answer.
Noul
Answers yes or no, for bounded checks such as “has the page settled?”.

The action space is finite by design

Choices are compiled from what the page currently offers, then resolved to an operation this project owns. Nothing in the model's output is interpreted as code.

CLICK:17CLICK:42TYPE_TEXT:4SELECT_OPTION:2SCROLL_DOWNSCROLL_UPPRESS_KEY:EnterDONEBLOCKED

Fan-out lets independent questions travel in one request. A target question cannot read another question's answer, so a speculative branch carries its own assumptions instead of borrowing state it never saw.

Install

Three steps to a first result.

This is a source preview: the app is real, the packaging is early, and the first-success path is deliberately small.

STEP 1

Get the app

Download the developer build for your platform, or run the source preview from a clone.

npm install
npm run dev
STEP 2

Connect your models

You need both connections. Setup validates them before the first task.

Jev: jev-latest @ api.typesafe.ai/v1/systemone
Text: Codex CLI · Claude Code CLI · API key
STEP 3

Give it a first task

Open a small public page, ask one page-grounded question, then compare the answer with the page you can see.

What result is currently shown?

What you need first

Node.js

22.12 or newer, with npm

Jev model key

A TypeSafe API key — the System One endpoint, not Chat Completions

Text model

A locally authenticated Codex or Claude Code CLI, or your own API provider

Vision

A provider/model that accepts images, for games and other visual flows

Connections, and what each one costs

ConnectionHow it is set upWhat to know
Jev — required for browser decisionsTypeSafe API key, default model jev-latestRequests use your TypeSafe account. A saved CLI login does not establish quota or model access.
Text model — required for conversation and planningCodex CLI, Claude Code CLI, or an OpenAI-compatible / Anthropic API connectionThe app can privately install a missing supported CLI through its Connect flow; npm is required.
Vision — for supported visual tasksAny connected provider/model that accepts imagesUsed for game calibration and bounded recovery; local perception keeps the move loop out of the vision loop.

Model usage can cost money and there is no fixed per-task cost: it depends on the models, the context, the number of actions and any retries. Check both accounts' usage before running a long task. Connections and conversation archives are encrypted locally — but local storage does not imply offline inference.

When something does not work

Install or build fails
node --version must be at least 22.12; use the checked-in lockfile with npm ci.
Interface opens but cannot browse
Launch the desktop app, not the web-only interface preview.
Jev connection fails
Use a TypeSafe key against the System One endpoint — a different API from Chat Completions.
Local CLI cannot connect
Finish its authentication flow, or configure an API provider instead.
Visual task fails
Confirm the selected provider and model accept image input.
Boundaries

What it does not do yet.

A short list, published on purpose. Everything here has been verified to be unsupported or unverified rather than assumed to work.

  • Uploads, drag-and-drop and download management
  • Browser extensions
  • Password-manager workflows
  • Some cross-origin and closed-shadow interactions
  • Multi-page completeness in long catalog tasks

Sensitive actions such as purchases and sending messages hand control back to you. Label-based guards are not a complete defence against hostile pages, so keep an eye on the tab you handed over.

Roadmap

What comes next

Named work, in the order it is being done — not a promise of dates.

  • Broader frozen evaluation

    Wider held-out runs with retained failure traces.

  • Completion and pagination

    More reliable coverage checks and multi-page reading.

  • Fresh-provider acceptance

    Repeat the provider matrix on clean machines and current models.

  • Signed distribution

    Signing, notarisation and automatic updates for the packaged apps.

  • Distribution hardening

    Fresh-install testing as part of the release gate, not after it.

$JEVRY · SOLANA

One wallet-to-wallet rail for the agent economy.

$JEVRY settles directly between wallets on Solana: no invoice, no middleman, no account. The token is a single public address published here first, and this page is the only one that speaks for it.

Every field above is verifiable from the public address; compare it before you buy.

NetworkSolana
StandardSPL
Supply1,000,000,000
Decimals6
Launchpadpump.fun
StatusPRE-LAUNCH
Contract address
Coming soon — the contract address is published on this page first.

Coming soon — the contract address is published on this page first.

FAQ

9 straight answers.

What is jevry, exactly?

A desktop browser with an agent inside it. You state a task in plain language; it observes the page, chooses one supported action at a time, executes it in a real Chromium tab and shows you what happened.

What is Jev?

Jev is TypeSafe's System One model: it evaluates supplied context and returns typed judgments and probabilities. jevry is an independent project built around that model and does not train or host it.

Do I need two connections to use it?

Yes. Jev is required for browser decisions; a text model is required for conversation and planning. Vision is only needed for supported visual tasks such as games.

Does it work on every website?

No. Website support varies. HTML and ARIA controls, native selects, open shadow roots, nested scrolling and bounded same-origin iframes are supported. Closed shadow roots, some cross-origin surfaces, uploads and extensions are not.

Can it spend my money or send messages?

It stops and hands control to you for sensitive actions such as purchases and sending messages. Guards are label-based, which is a useful barrier, not a complete defence against a hostile page.

How do you know the 2048 result is real?

The run was recorded with the board state read back after every dispatched key, and the page you are reading shows a separate deterministic replay you can watch end to end — including its score and move count.

Are the benchmark numbers a leaderboard result?

No. They come from a development subset of a much larger benchmark, scored by the official evaluator. They are reported as development evidence with their build boundaries, not as a held-out or certified score.

What does $JEVRY have to do with the app?

The app is open source and free. $JEVRY is the project's own token rail on Solana: a single public address, published on this page first, with no invoice and no middleman.

Is there anything tracking me on this page?

No. This is a static site with no analytics, no cookies and no third-party scripts. Fonts and assets are served from this domain.

Your browser. Ready to act.

Hand it one task. If it cannot do it, it will tell you where it stopped — and hand the tab back.