RADAR runs AI coding agents such as Claude Code and Codex. Each task gets a test before work starts; the task counts as done only when RADAR runs that test itself and it passes.
An AI agent is a program that writes and changes software for you. It works fast and says when it has finished. RADAR checks whether it really has.
AI agent
A program that does software work on its own, such as fixing a bug. Claude Code and Codex CLI are two examples.
The test
A small, exact check that proves the work is right, written before the work starts. Engineers call it an acceptance test.
The record
Where every task is written down: who took it, what they claimed, and what the test showed. RADAR calls it the ledger.
1
The test is written first
Before anyone starts, the task gets its test: a wrong password must be refused, and the right one must let you in.
radar app · t-4821example
○youthe login accepts a wrong password; fix it
◆radartaskt-4821created
┊the test, written before any work:
┊
testwritten before the worknot run yet
…a wrong password→refused
…the right password→signed in
┊the working agent cannot change the test command
2
An agent does the work
RADAR hands the task to an AI agent. The agent works in its own copy of the project, so nothing else is touched.
radar app · t-4821example
○youthe login accepts a wrong password; fix it
◆radartaskt-4821created·test written first
◇agentWorking in my own copy of the project.
·▸ Read login_app/auth.py13 lines
·▸ Grep wrong_password in tests/1 match
·▸ Edit login_app/auth.py+1 −1
┊leaseclaimed attempt1/3
3
The agent says it is done
The agent reports: done. RADAR writes that down as a claim, not as a result.
radar app · t-4821example
◇agentWorking in my own copy of the project.
·▸ Read login_app/auth.py13 lines
·▸ Grep wrong_password in tests/1 match
·▸ Edit login_app/auth.py+1 −1
◇agentdone:wrong passwords are refused now
┊radar:claim, not result·it runs the test itself
4
RADAR runs the test itself
RADAR runs the test on its own, on the latest version of the work, without taking the agent's word for it.
radar app · t-4821example
◇agentdone:wrong passwords are refused now
┊radar:claim, not result·it runs the test itself
◆radarverifyt-4821·running the test itselfunder way
┊
testrun by radarrunning
a wrong password→…
the right password→…
5
If it fails, the task opens again
Here a wrong password still lets you in. The task opens again with the failure attached, and the agent tries again. After three tries another agent takes over; after three more, the task waits for you.
radar app · t-4821example
◇agentdone:wrong passwords are refused now
◆radarverifyt-4821·1 failedexit 1
┊
testrun by radar1 failed · exit 1
a wrong password→signed inshould be refused
the right password→signed in
┊check_login("ada", "wrong password") returned True
┊t-4821reopened· failure attached · attempt 2/3
6
A passing test closes the task
The agent fixes it, and RADAR runs the test again. It passes, so the task closes, and the record keeps the proof.
radar app · t-4821example
◇agentThe length check was not the bug: nothing compared the password with the stored hash.
·▸ Edit login_app/auth.py+2 −1
◆radarverifyt-4821·passed·judged by radar2 passed
┊
testrun by radar2 passed
a wrong password→refused
the right password→signed in
■t-4821closed·proof kept in the ledger
Why this matters
Agents can be confidently wrong. Without a check, a wrong “done” reaches your product and your customers. With RADAR, a task is finished when the proof says so.
radar app · t-4821example
○youthe login accepts a wrong password; fix it
◆radartaskt-4821created
┊the test, written before any work:
┊
testwritten before the worknot run yet
…a wrong password→refused
…the right password→signed in
┊the working agent cannot change the test command
○youthe login accepts a wrong password; fix it
◆radartaskt-4821created·test written first
◇agentWorking in my own copy of the project.
·▸ Read login_app/auth.py13 lines
·▸ Grep wrong_password in tests/1 match
·▸ Edit login_app/auth.py+1 −1
┊leaseclaimed attempt1/3
◇agentWorking in my own copy of the project.
·▸ Read login_app/auth.py13 lines
·▸ Grep wrong_password in tests/1 match
·▸ Edit login_app/auth.py+1 −1
◇agentdone:wrong passwords are refused now
┊radar:claim, not result·it runs the test itself
◇agentdone:wrong passwords are refused now
┊radar:claim, not result·it runs the test itself
◆radarverifyt-4821·running the test itselfunder way
┊
testrun by radarrunning
a wrong password→…
the right password→…
◇agentdone:wrong passwords are refused now
◆radarverifyt-4821·1 failedexit 1
┊
testrun by radar1 failed · exit 1
a wrong password→signed inshould be refused
the right password→signed in
┊check_login("ada", "wrong password") returned True
┊t-4821reopened· failure attached · attempt 2/3
◇agentThe length check was not the bug: nothing compared the password with the stored hash.
·▸ Edit login_app/auth.py+2 −1
◆radarverifyt-4821·passed·judged by radar2 passed
┊
testrun by radar2 passed
a wrong password→refused
the right password→signed in
■t-4821closed·proof kept in the ledger
Fig. 2 · The same example, close up
The chart room
Every agent, every task, on one screen.
This is RADAR's home screen. Five of its parts are explained below.
radarsample data120x36
1Tap a number to zoom in.
Fig. 3 · The home screen, with sample data
01
Top right: CHART
The map
Each island is a computer: yours, or a server you run. The dashed lines are the agents working on it. A machine that stops responding is marked AGROUND and stays on the map.
02
Left: LOG
The record
Every task is counted by state: waiting, in progress or awaiting its test. All agents share one record, so nothing is lost between them.
03
Bottom left: NEXT
Needs you
Only the decisions that need a person reach you, with the next command to run. An agent gets three attempts, then another agent takes over; after three more, the task waits here for you.
04
Bottom left: ARRIVALS
Checked, then landed
Work joins your project only after its test passes, and the test runs once more after it joins. The bars show how many tasks landed each hour.
05
Bottom right: COMMANDS
Commands and memory
RADAR is driven by a short list of commands. One of them searches the memory: notes of lessons learned, which RADAR shows the next agent before it starts. You can open them in Obsidian.
From Grit's own work
Observed, not a controlled comparison
In Grit's own work (20-28 Sep 2026), workers running one coding agent reported a task as finished 1,608 times (1,457 tasks); Grit's immediate re-run of the acceptance commands on the same server failed on 74 of those reports (4.6%). A claim alone does not close a task.
1,608 reports said done, so every one counts as finished.The tests were run again: 74 of those reports failed.
One square: one report that a task was finishedGrit's re-run failed
Grit internal ledger, snapshot 28 Sep 2026 06:00 UTC
Editions
An open edition for individuals. A company edition in development.
Open source (Apache-2.0) at public release
RADAR
For individual developers, on your own computer.
The record, the tests, the memory and agents working in parallel
RADAR is in early access. To join, write to us through the contact page, or on X or Instagram. Once you have access, give your coding agent the repository link and say “install this”. The repository includes install steps written for agents: the agent checks your computer, asks before installing anything missing, sets RADAR up and reports back.
Does RADAR replace my coding agent?
No. RADAR runs the agents you already use and checks their work. The agents still do the work. RADAR gives them one record, one check they cannot skip, and one memory.
What counts as done?
A task is done when RADAR has run its test on the latest version of the work and the test has passed. The agent's own “done” is recorded as a claim; RADAR will not close the task until the test passes. A person can override this, and the override is recorded.
Which agents does it run?
Claude Code, Codex CLI, OMP and Hermes Agent. Each one works in its own copy of the project, on your computer or on Linux servers of your own.
Is RADAR better than Claude Code or Codex?
It is not a replacement, so it does not compete with them. RADAR runs them and checks what they deliver.
What if the test itself is wrong?
Then the check is wrong too. That is why the test is written before the work starts, where a person can read it, and any later change to it is recorded.
Does my code leave my computer?
Not through RADAR. It needs no account and no cloud, and every outside service is optional and off until you turn it on. Your agents still talk to their own providers, as they do today.
What does it cost to run?
RADAR adds no bill of its own. The agents it runs keep using the accounts you sign them in with, such as a Claude or ChatGPT plan, under those providers' terms and usage limits. RADAR uses a paid API only if you name the provider and a budget.
Your agents write the code. RADAR makes sure it's done.