The decisions in every turn that are not about the code go to a judge. This page lists all 38 decision points: what each one asks, what it changes, and who answers.
How it works
One decision point, three steps
A decision point is a small, routine choice inside a turn that is not about the code itself. mu takes it away from the big model and asks a judge instead.
01
Read a small state
Never the whole conversation: the message you just sent, one chunk of tool output, the command about to run.
You send "Run the tests and fix the failing ones".
02
Ask one bounded question
Yes or no, one of a few named answers, or a score, each with a probability. A warm question takes about 0.3 seconds.
◆ multi-step task · heavy gear · 752 ms
03
Change the next step
A one-line hint, a passage left out, a call stopped, a nudge. Never an extra question to you. Without an answer, mu does what it would do without a judge.
The model starts with a one-line hint about the kind of task, and the verdict shows in the conversation, as in the recording on the home page.
Each point has a mode
Activeactive
The verdict takes effect.
Shadowshadow
Asked and recorded, but nothing changes. For watching how well it judges first.
Offoff
Not asked; this decision is off.
Rules come first. A dangerous-looking command is flagged by rules; the judge only confirms that you asked for it.
Where they sit
Five moments in a turn
The points are grouped by when they are asked, the same way the desktop app's settings group them. Pick one to jump to it.
1
Input
3
When your message arrives: what kind it is, whether it changes the task, whether it should interrupt.
Halfway through, you add: "Don't touch the migrations folder."
Verdict
◆ a new hard constraint
Result
The task frame records it in your own words, with where you said it, and every later check reads it. A plain "thanks" is no change, and the frame stays as it is.
◆ Jev asksAfter going in circles, did the way out deserve a lesson?
When the agent went in circles or off course and the turn still ended with a passing check or the goal met, judges whether what finally worked was a different approach; if so, keeps the trap and the way around it. Asked at most once a turn.
◆ Jev asksA lesson the model or a sub-agent proposes: useful again, a one-off, or known already?
For a lesson the model keeps with the remember tool, or a Lesson: line in a sub-agent's report: useful again later, a one-off, or known already from the prompt or the project files. Only the first kind is kept.
◆ Jev asksThe same as an existing lesson, more precise, or contradicting it?
Before a new lesson is stored, compares it with the most similar kept ones: the same lesson is not kept twice, a more precise one replaces the old, and on a contradiction your latest word wins.
◆ Jev asksIn the Jev approves mode: does the task clearly need this command, this change outside the project, this outside action, this sub-agent?
In Jev-approves mode, commands, changes outside the project, outside actions and sub-agents go to Jev first: only what it is sure the task needs, done as you would expect, runs without you; anything else asks you in the status bar.
◆ Jev asksBefore a call that changes something: does it cross a constraint you stated?
What you ruled out is kept word for word in the task frame; before a call that changes something, each constraint is checked, and only a confident violation is stopped, in your own words.
◆ Jev asksA web page, a search result or an MCP server's output, passage by passage: does it carry instructions aimed at the AI?
Before the model reads a web page, search results or an MCP server's answer: which passages carry instructions aimed at an AI (ignore rules, hand over data, run commands, or carry the conversation out through a link)? Those are withheld and replaced by a one-line note. When Jev gives no answer, only plain injection phrases are withheld.
◆ Jev asksThe model's own yes/no question, about each of many items: files, log lines, findings
When the model has to sort hundreds of files, log lines or findings, it hands Jev one yes/no question to answer item by item, with a probability each, instead of reading them all. It appears only when a task needs it. The model asked, so it answers in shadow too; off turns it away.
◆ Jev asksFor each finding of /review: does it change behaviour, and is it about this change?
After the reviewer of /review reports, each finding gets two questions: does it change how the program behaves, and is it about this change; with the reviewer's own severity that orders them P0 to P3. None is dropped, P3 is collapsed, and a finding the reviewer insisted on never lands in P3.
◆ Jev asksNew language-server diagnostics after an edit: tell now, at the next pause, or never?
New language-server diagnostics after an edit: tell now, when the model pauses, or never (style warnings). New errors still there when the turn ends are always told.
◆ Jev asksThe same failure again and again: is this approach a dead end?
When the monitor sees the agent going in circles, or one command failing again and again, judges whether the approach is a dead end; only a confident dead end with no progress proposes going back to the turn's checkpoint. It never rewinds by itself.
◆ Jev asksThe run ends on "Let me run the tests next", or on asking for a go-ahead on work you asked for: did it stop short?
A run that ends on "next I'll run the tests" without doing it, or asks for a go-ahead on work you already asked for, is sent back to it. A step that is hard to undo or reaches beyond this machine (push, publish, delete, pay) is never pushed. At most twice per message.
◆ Jev asksWhile the model writes: does the tail of its output cross your constraints?
While the model writes, the judge reads the tail of its output against your hard constraints and the configured rules every few hundred characters; on a confident violation the output is cut, the rule is named, and the model carries on from there. Runs only with the feature switched on.
◆ Jev asksIn goal mode, when the big model gives no answer: is the condition met?
Goal mode is checked by a model by default; this decision is used when Jev is chosen or the model gives no answer: it reads the closing message for whether the goal holds and whether you are needed. An open acceptance item or an unverified edit always means not yet.
◆ Jev asksWhere do things stand, in multiple choice?
In a project with the board on, whenever the agent says something, after a check or a ticked acceptance item, every few tool calls, and whenever the agent stops, multiple-choice questions read its phase, the acceptance item it works on and whether it waits for you, and sort what happened since the last board into news and routine. Only something new has a plain-speaking model write the board again and retell the news on the running account; when a run ends, the news of the whole run is picked again for a summing up.
◆ Jev asksDid the sub-agent's patch stay within its task?
When an isolated sub-agent hands back a patch, judges from the task, the changed paths and the line counts whether it stays within the task and which files look unrelated. One line of advice for the main agent; it never blocks.
◆ Jev asksDoes a new finding replace, contradict or support an earlier one?
What a new finding does to an earlier one on the board: replaces it, contradicts it, supports it, or nothing. A replaced finding goes down and whoever heard it is told; a contradiction keeps both sides for a bee to settle.
Exact repeats. A failing run often prints the same diff, DOM dump or stack once per failed test. mu keeps the first copy and replaces each later copy with one line that names the lines it repeats. No model is called. The markers expand to the original byte for byte, and the full log stays on disk, behind a pointer at the end of the output.
Goal-aware selection. With a verbose reporter, what to keep depends on what you asked for: passing tests are noise when you debug a failure, and evidence when you ask which tests ran. In one request, Jev is asked about each block of passing tests and each block of test output: does the goal still need it? The summary and every failure are never asked about. A block is left out only when Jev gives "not needed" a probability of 0.9 or more.
Perfect judge reads the labels: the most a correct judge could cut. Keep failures only is a judge that always answers "leave it out", which is what a filter that ignores the goal does. All 282 live Jev requests of the study cost about $0.017 at list price; with the default wording, a request took 345 ms at the median.
Both are off by default. "features": { "admission": { "testLog": "rules" } } in ~/.mu/agent/mu.json, or Test log trimming in the desktop app's settings, folds repeats. "jev" also asks the judge which of the remaining parts the current goal needs; /mu mode tool.admission.test-log shadow only records that choice.
What these numbers are not:
The selection cases are real Vitest, node:test and pytest output of synthetic projects, plus 13 hand-written edge cases; the goals and labels are the authors'. The held-out goals were labeled first and run once, and nothing was changed afterwards.
They measure what reaches the model and what is lost, not whether the model then finishes the task.
The repeats come from two days of one developer's sessions: 15 test logs, all Vitest, of which the 7 over 4,000 characters are charted. Other runners are not measured.
There is no end-to-end comparison with pi, Claude Code or Codex on the same tasks yet.
node kyrn/spikes/judge-bench/test-log-replay.ts reruns every arm but Jev's in seconds, with no key; on the current code they come out 0.6 to 1.1 points above the chart, which was measured on 2026-09-21. The Jev arm needs TYPESAFE_API_KEY.
The charts as tables
Goal-aware selection
Tuned goals (29): cut
Required lines lost
Held-out goals (9): cut
Required lines lost
mu · Jev
40.2%
0 of 72
46.4%
0 of 19
Perfect judge
52.5%
0 of 72
46.5%
0 of 19
Keep failures only
61.9%
9 of 72
59.4%
6 of 19
Real failing test log
Characters
Folded
5 failures, one diff each
37,819
86%
DOM test, 4 failures
34,115
44%
The same run, seen by a sub-agent
34,115
44%
Shared stderr stack
15,565
47%
2 failures
9,249
16%
7 suites fail to parse
4,953
0%: one repeat, too little to fold
5 different failures
4,004
0%: no repeats
All 7
139,820
51.0%
If mu is useful to you, star it on GitHub
A star helps more people find it. The code, the discussions and every release live in the repository.