Cerevisor 3.2.0
· v3.2.0
Cerevisor 3.2.0 lets you continue an interrupted run where it stopped, have a second agent review any agent's work, see the evidence behind every note, answer an agent's question mid-run, and check a run at a glance from the new Overview on desktop and phone.
3.1 gave the agents on a run a place to talk to each other. 3.2 gives you a seat at it. A run window now opens on a plain summary of where the run stands. An agent’s work can be checked by a second agent before it counts. A claim on the board can carry what it is based on. An agent that is genuinely stuck can stop and ask you. When Cerevisor cannot tell whether something happened, it says so and lets you settle it. And a run the app was closed in the middle of no longer sits there claiming to be running: it is closed out, and offered back to you with one button.
1. Continue this run
If Cerevisor is closed, crashes, or is killed while a run is going, that run used to stay “running” forever. Now, on its next start, Cerevisor closes out any run from the last seven days that begins and never ends, and marks it plainly: Cerevisor closed before this run finished. That wording keeps it from reading as a failure the agents caused. Nothing already recorded is rewritten, and a record that cannot be read is left exactly as it is rather than given an ending it never had.
Beside each one is Continue this run. It starts a fresh run seeded from the interrupted one: the agents that already finished keep their work and cost nothing, and only what was left runs. The pre-run summary (the same one the Run button shows) lists what will be kept, what will run fresh, and what will pick up its own session, plus Changed while interrupted: the files that moved while the app was closed. If Cerevisor could not check, it says that instead of showing an empty list.
Agents on Codex, Claude Code and Cursor (cloud runs) resume their own sessions. Agents on Antigravity and Grok Build run their part from scratch, because those tools keep nothing to rejoin, and the summary says so before you confirm. One thing a resumed Cursor cloud agent cannot do is forget: it keeps the secrets it was created with until that agent is deleted, including one you have revoked since. That is said on the summary rather than papered over.
Continuing is always something a person does. Nothing unattended (no schedule, no Operator, no headless run) can decide on its own to pick up work the app died in the middle of.
2. Needs review before it counts
An agent can now have its work checked by a second agent. Switch on Needs review before it counts on an agent’s card and choose Always or Only when the agent is unsure; a reviewer card appears right after it on the canvas: a real card you can open, read and delete. A workflow-wide default is available too, and the chat builder understands the request in plain words.
A reviewer only reads. Its tools are the read-only set: no writing, no editing, no shell, no connected servers, no powers, no vault secrets, no output file. Cerevisor refuses to save a reviewer card that has been given any of those by hand, and narrows its tools again when the run starts.
Its answer is a note on the board, and you see where it stands on the agent’s own tile: Under review, Reviewed, Sent back, or Nobody checked this. A review result with no evidence attached does not count, and silence is never read as approval. A supported Sent back buys the producer exactly one revision, with the reviewer’s own sentence going into its next turn; a second one marks the work as Sent back by its reviewer: … and halts the run by default, which the workflow’s existing halt setting can turn off. The producer’s work is never thrown away; only its opening line is marked, and everything downstream reads why.
3. Why believe this
A Claim, a Decision and a Review result can carry what they are based on: up to eight pieces of evidence, shown under Why believe this. Each one says how it checked out: checked, not found, or not checked.
Evidence points; it never fetches. Cerevisor confirms a file exists and is inside the run’s own folder, following the same boundary rules the rest of the app uses (including after links and junctions are resolved), but never opens it to see whether it says what the agent claims. An agent cannot mark its own evidence as checked. And a web address on a note is shown as plain text everywhere: turning a page an agent read into something you click is the one thing this deliberately does not do.
4. Question for you
An agent that needs a decision from you can now stop and ask. The question is an ordinary board note addressed to you, called out as Question for you in the run window and on your phone, sometimes with choices to pick from. Your answer comes back to the agent as information, and it grants that agent nothing.
Only you can answer. Another agent replying on the board does not end the wait. The wait does not time out; pressing Stop is the other way it ends. A per-workflow setting decides whether the question holds just the agent that asked (the default) or the whole run.
Agents running on outside tools post their question and keep going, because Cerevisor cannot hold those tools mid-turn; your answer reaches the agents that run after it, and the launcher says so for those agents. A run with nobody watching never posts a question at all.
5. Cerevisor noticed
When the same thing fails over and over, Cerevisor now says so out loud, under Cerevisor noticed: read_file failed 5 times in a row, [credential] was switched away from 2 times, [agent] hit the same error twice in a row. These are observations, not interventions: Cerevisor tells you and keeps going. It never stops a run because of one.
6. Check by hand
Some actions end without an answer. A command killed by its own timeout after it had started; a web request that changes something, cut off after it went out; a connected server that stopped responding mid-call. Cerevisor genuinely does not know whether the thing happened, and it has stopped guessing.
The agent is told in plain words that the call may have gone through and must not be repeated, and within that turn Cerevisor makes sure of it: the very next identical call is refused before it can run. The step then appears in the run window under Cerevisor could not tell whether these ran. Only you can say, with two buttons: It happened and It did not happen. Your answer is written into the run’s own history as your decision.
The answer is permanent. A step takes one answer only; there is no way to correct a mis-press in this version, so read the row before you press.
Alongside this, every web page an agent fetches with Cerevisor’s own web tool is now recorded in the run’s history: the address, the method, the result and the size, never the contents. A page fetched some other way, whether inside a shell command or by an outside tool running its own agent, is not recorded this way. Secret-looking values in an address are replaced before it is written down, and an address that cannot be read has its query stripped entirely.
7. The Overview
A run window now opens on an Overview. A strip of counts covers running, waiting for you, under review, done, sent back, failed and check by hand, plus what the run has spent so far (an amount under a cent is shown to four decimals rather than rounded to zero). Then one tile per agent: where it stands in one word, what it spent, how much evidence it put on the board, how many of its questions are waiting on you, and where its review stands. Click a tile to see only that agent’s lines in the log.
A run that has not started says so. A run recorded by an older Cerevisor says This run finished before Cerevisor kept an Overview. Its log, files and board are still here, and its other tabs still hold all of it.
8. On your phone
The Companion shows the same picture: a run’s Overview with its counts, cost, tiles and notices; the Why believe this evidence behind a note; the run’s check-by-hand rows with the same two buttons; and any Question for you an agent addressed to you, which you can answer from there. Reading needs nothing extra; marking a step by hand needs the same remote-control permission every other phone action needs, and a run outside the workspace your phone can see is not shown at all.
An older Companion app keeps working exactly as it did: it simply does not offer the new screens. Nothing on your phone needs updating.
9. What an exported run carries
Exporting a run now brings the new material with it. The readable half gains the Overview as plain text (the counts, the cost and one line per agent), and every board note ends with how much evidence it carried ([evidence: 3], or [no evidence attached]) and, on a review result, whether the work was Reviewed or Sent back. The evidence itself, each pointer, its excerpt and how it checked out, rides in the structured half of the same file. As before, the export is a file on your machine and goes nowhere else.
10. Housekeeping
Two recording fixes and one prompt fix worth naming:
- An agent’s completion is recorded once, and says it completed. A duplicate internal row was being filed with the wrong stage, which could walk a finished agent’s tile back to “running”. That row now carries its real stage and the duplicate is no longer written, so a run’s history has one fewer activity line per agent than a 3.1 recording of the same work. It is a change in what is recorded, not only in what is shown.
- A locally built package now bundles this checkout’s own companion protocol. A build could pick up a stale copy from a neighbouring folder, which would ship a desktop that did not actually speak the new phone capability.
- Agents that run on Codex, Cursor or Grok Build are now told their board file is one of the files they may write. Those three are given a list of the files they may read or write, and the board file was missing from it while the same instructions asked them to post to the board, a contradiction that left agents on Codex almost never posting. The list now names the board file as well. Nothing else was added to it, and no other access changed.
Nothing new leaves your machine. 3.2 added no network access of any kind: a review result, an excerpt, a question and an Overview tile go to the reading agent’s own model and to your own paired phone, bounded and redacted, and nowhere else.
Known limits
- A native agent’s question holds it between turns. An agent that asks you something stops at the boundary of its own turn; a question raised in the middle of a chain of tool calls is put to you at the next break, not instantly.
- Agents on outside tools cannot wait for your answer. They post the question and continue. Your answer reaches the agents that run after them.
- “Cerevisor noticed” only reports. No streak stops a run, switches a credential or changes what happens next.
- The repeated-tool-failure notice covers Cerevisor’s own agents only. An agent running on Codex, Cursor, Claude Code, Antigravity or Grok Build produces no “failed N times in a row” line, because those tools do not report their calls in the form the count needs. Runs recorded before 3.0 report no tool streak either.
- The refusal to repeat an uncertain call lasts for that turn. If the agent is retried from the start, whether after an error or after switching to another credential, the count starts clean and the call could be tried again.
- A Check-by-hand answer cannot be changed. One answer per step, from any surface.
- A run whose history was torn by a crash cannot be checked by hand at all, which is precisely the kind of run that most invites the question. If the app was cut off part-way through writing the record, which step is which would be a guess, and Cerevisor refuses rather than settling the wrong one. This is permanent for that run, and the message says so.
- A question set to hold the whole run holds every agent’s place in the queue for as long as it stays open, with no time limit. An overnight question keeps those places overnight and other work waits behind them.
- The Overview does not yet show a note saying a run was continued from an interrupted one.
- Reviews do not apply to a single-agent re-run or to a workflow invoked as a Power by another agent.
- A reviewer reads roughly the first 3,000 characters of the work it checks. On a long output, “approved” means the opening looked right.
- A workflow-wide review default of Always, on a workflow with no reviewers, blocks Run until reviewers are added, and there is no Settings screen to reset it from.
- Turning review off and on again inserts a new reviewer card; the old one stays on the canvas as an ordinary agent until deleted.
- Everything listed under 3.1.0’s and 3.1.1’s known limits still applies.
The short version
| Feature | In one line |
|---|---|
| Continue this run | A run the app died in the middle of is closed out automatically and offered back with one button, picking up only what did not finish. |
| Needs review before it counts | Turn on a second, read-only agent to check any agent’s work before it counts, with a result of Reviewed or Sent back. |
| Why believe this | A claim, decision or review result can carry up to eight pieces of evidence, each marked checked, not found, or not checked. |
| Question for you | An agent that is genuinely stuck can stop and ask you directly, on the run window or your phone, and wait for your answer. |
| Cerevisor noticed | Repeated tool failures or credential switches are called out plainly, as an observation that never stops the run on its own. |
| Check by hand | When Cerevisor cannot tell whether an interrupted action went through, it refuses to guess and asks you to mark it happened or not. |
| The Overview | A run window now opens on a strip of counts and spend, plus one tile per agent showing its status, cost and evidence. |
| On your phone | The Companion mirrors the Overview, evidence and Question for you, and lets you answer or mark a step from your phone. |
| What an exported run carries | An exported run now includes the Overview as plain text, plus how much evidence each note carried. |
| Housekeeping | Two recording fixes and one prompt fix, including giving Codex, Cursor and Grok Build agents write access to the board file. |
Platforms
Available for Windows, macOS (Intel and Apple silicon) and Linux (AppImage and Debian package). macOS and Linux stayed on 3.1.0 while 3.1.1 shipped for Windows only, so this update brings them everything from 3.1.1 as well.
Updating
Existing installations receive the update through Cerevisor’s built-in updater, and the installers are available from cerevisor.com/download and the private GitHub release. Your workflows, run history and settings carry forward untouched; no file format changed.