RADAR by Grit In early access

Agents say done.
RADAR checks.

RADAR runs AI coding agents such as Claude Code and Codex. Each task gets a test before work starts; the task counts as done only when RADAR runs that test itself and it passes.

  1. You give RADAR a task and its test. The test is written before any work starts.
  2. An agent does the work. In its own copy of the project.
  3. The agent says it is done. RADAR records it as a claim, not a result.
  4. RADAR runs the test itself. It fails. The task reopens with the failure attached.
  5. The agent fixes it. RADAR runs the test again. It passes.
  6. The task closes. Its proof is kept.
radar app
example session

radar appexample session

Scroll or swipe to move around the screen.

radar app
example

Fig. 1 · An example session, with sample data

Runs the agents you already use

  • Claude Code
  • Codex CLI
  • OMP
  • Hermes Agent

Works with

  • Claude
  • ChatGPT
  • Git
  • Python
  • Obsidian
  • Windows
  • Linux
Does RADAR replace my agent?

In plain words

What RADAR does, with one example.

An AI agent is a program that writes and changes software for you. It works fast and says when it has finished. RADAR checks whether it really has.

AI agent
A program that does software work on its own, such as fixing a bug. Claude Code and Codex CLI are two examples.
The test
A small, exact check that proves the work is right, written before the work starts. Engineers call it an acceptance test.
The record
Where every task is written down: who took it, what they claimed, and what the test showed. RADAR calls it the ledger.
  1. The test is written first

    Before anyone starts, the task gets its test: a wrong password must be refused, and the right one must let you in.

  2. An agent does the work

    RADAR hands the task to an AI agent. The agent works in its own copy of the project, so nothing else is touched.

  3. The agent says it is done

    The agent reports: done. RADAR writes that down as a claim, not as a result.

  4. RADAR runs the test itself

    RADAR runs the test on its own, on the latest version of the work, without taking the agent's word for it.

  5. If it fails, the task opens again

    Here a wrong password still lets you in. The task opens again with the failure attached, and the agent tries again. After three tries another agent takes over; after three more, the task waits for you.

  6. A passing test closes the task

    The agent fixes it, and RADAR runs the test again. It passes, so the task closes, and the record keeps the proof.

Why this matters

Agents can be confidently wrong. Without a check, a wrong “done” reaches your product and your customers. With RADAR, a task is finished when the proof says so.

Fig. 2 · The same example, close up

The chart room

Every agent, every task, on one screen.

This is RADAR's home screen. Five of its parts are explained below.

radar
sample data

radarsample data

Scroll or swipe to move around the screen.

Tap a number to zoom in.

Fig. 3 · The home screen, with sample data
  1. Top right: CHART

    The map

    Each island is a computer: yours, or a server you run. The dashed lines are the agents working on it. A machine that stops responding is marked AGROUND and stays on the map.

  2. Left: LOG

    The record

    Every task is counted by state: waiting, in progress or awaiting its test. All agents share one record, so nothing is lost between them.

  3. Bottom left: NEXT

    Needs you

    Only the decisions that need a person reach you, with the next command to run. An agent gets three attempts, then another agent takes over; after three more, the task waits here for you.

  4. Bottom left: ARRIVALS

    Checked, then landed

    Work joins your project only after its test passes, and the test runs once more after it joins. The bars show how many tasks landed each hour.

  5. Bottom right: COMMANDS

    Commands and memory

    RADAR is driven by a short list of commands. One of them searches the memory: notes of lessons learned, which RADAR shows the next agent before it starts. You can open them in Obsidian.

Fig. 3 · The home screen, with sample data

From Grit's own work

Observed, not a controlled comparison

In Grit's own work (20-28 Sep 2026), workers running one coding agent reported a task as finished 1,608 times (1,457 tasks); Grit's immediate re-run of the acceptance commands on the same server failed on 74 of those reports (4.6%). A claim alone does not close a task.

Which count to show

1,608 reports said done, so every one counts as finished.The tests were run again: 74 of those reports failed.

One square: one report that a task was finished Grit's re-run failed

Grit internal ledger, snapshot 28 Sep 2026 06:00 UTC

Editions

An open edition for individuals.
A company edition in development.

Open source (Apache-2.0) at public release

RADAR

For individual developers, on your own computer.

  • The record, the tests, the memory and agents working in parallel
  • Runs Claude Code, Codex CLI, OMP and Hermes Agent
  • Python 3.12 or newer; no account or cloud needed

In development

RADAR for companies

For teams, on your own infrastructure.

  • Runs on your team's own infrastructure, against your own code
  • Set up and supported by Grit
  • Autonomous systems built for your work, under the same rule: a task closes only when its test passes again
Is RADAR available today?

RADAR is in early access. To join, write to us through the contact page, or on X or Instagram. Once you have access, give your coding agent the repository link and say “install this”. The repository includes install steps written for agents: the agent checks your computer, asks before installing anything missing, sets RADAR up and reports back.

Does RADAR replace my coding agent?

No. RADAR runs the agents you already use and checks their work. The agents still do the work. RADAR gives them one record, one check they cannot skip, and one memory.

What counts as done?

A task is done when RADAR has run its test on the latest version of the work and the test has passed. The agent's own “done” is recorded as a claim; RADAR will not close the task until the test passes. A person can override this, and the override is recorded.

Which agents does it run?

Claude Code, Codex CLI, OMP and Hermes Agent. Each one works in its own copy of the project, on your computer or on Linux servers of your own.

Is RADAR better than Claude Code or Codex?

It is not a replacement, so it does not compete with them. RADAR runs them and checks what they deliver.

What if the test itself is wrong?

Then the check is wrong too. That is why the test is written before the work starts, where a person can read it, and any later change to it is recorded.

Does my code leave my computer?

Not through RADAR. It needs no account and no cloud, and every outside service is optional and off until you turn it on. Your agents still talk to their own providers, as they do today.

What does it cost to run?

RADAR adds no bill of its own. The agents it runs keep using the accounts you sign them in with, such as a Claude or ChatGPT plan, under those providers' terms and usage limits. RADAR uses a paid API only if you name the provider and a budget.

Your agents write the code. RADAR makes sure it's done.