Debug
Hand it the error. ML.ai finds the real cause first.
Give ML.ai a failing test, an error message, or a stack trace and it reads the execution path, finds the actual cause, and proposes a fix as a diff you approve.
Anyone staring at a failing test, an exception in production logs, or behavior that does not match what the code appears to say.
This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.
- Reading src/auth/session.ts
- Reading src/lib/tokenStore.js
- Reading test/auth.test.ts
The cause was a token missing its expiry field slipping past a falsy check instead of raising, so refresh() silently returned null under a race. The fix and a regression assertion land together.
This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.
The line the trace points at is rarely the line that is wrong
Finding the actual cause is the hard part of debugging. These four things are what keep that part honest.
The line it points at vs. the line that's wrong
It reads the call stack, not just the top frame
A stack trace names where the error surfaced, not where it started. ML.ai walks the frames underneath it and follows the data through them, so the fix lands on the frame that actually caused the failure.
States the cause, then -- only then -- the fix
token missing expiry, falsy check hid it
+ throw new Error('expired')
It states the cause before it writes anything
A plain-language diagnosis comes first, and the fix only appears after. If the stated cause is wrong, that is obvious immediately, before any code is written on top of it.
A diagnosis you can check, not just trust
race on token refresh
The diagnosis is a claim you can check, not a black box
You see the reasoning in words, not just a patched file. Push back on it directly and it revises the diagnosis, the same way you would redirect a wrong plan before a wrong fix gets written.
Re-run against the failure it was meant to fix
auth.test.ts
passing
The fix is checked against the failure it was meant to close
Running the originally-failing test is a shell command like any other. Confirming red-to-green is part of the loop, not a separate step you have to remember to do yourself.
Doing it by hand vs. handing it off
You read the stack trace, open the file it names, and start guessing, because the line the trace points at is usually a symptom, not the actual cause.
ML.ai follows the failure back through its callers, states the real cause in plain language, then proposes the fix, so the diagnosis is checkable before a fix is written on top of it.
Which agent actually does this
general
Traces the cause and writes the fix
Reads the function that failed and its callers, states what is actually wrong, then proposes a fix as a diff behind the same approval gate as any other edit.
@explore
Optional, to map the area first
Read-only search across the codebase, useful when the failing code touches an unfamiliar part of the system before a fix is attempted.
How it actually runs
A bug report is rarely the same thing as the bug. ML.ai follows the failure back through the code that produced it, files and callers, not just the line the stack trace points at, before it proposes a change.
Hand over the symptom
Paste the error, the failing test name, or describe what is happening versus what should happen.
It traces the real path
The agent reads the function that failed and walks its callers and the data flowing through them, rather than guessing from the error text alone.
It states the cause before fixing
You get a short explanation of what is actually wrong, so a wrong diagnosis is obvious before a fix is written on top of it.
Review the fix as a diff
The proposed change is held behind the same Allow/Always/Deny gate as any other edit. Approve it, or push back if the cause looks wrong.
This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.
- Reading src/auth/session.ts
- Reading src/lib/tokenStore.js
- Reading test/auth.test.ts
The cause was a token missing its expiry field slipping past a falsy check instead of raising, so refresh() silently returned null under a race. The fix and a regression assertion land together.
This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.
What actually goes wrong, without it
Three specific ways this kind of task breaks, and the part of the mechanism above that closes each one.
01 · Problem
The stack trace names a symptom, not a cause
The line an error is thrown from is often just where a bad value finally broke something, three function calls away from where that value actually went wrong.
How Debug closes it
general reads the function that failed and walks its callers and the data flowing through them, rather than treating the top frame of the trace as the answer.
02 · Problem
A plausible-looking fix that treats the symptom
Patching the line the trace points at often makes the immediate error go away while leaving the actual defect in place, ready to resurface somewhere else.
How Debug closes it
A diagnosis is stated in plain language before any fix is proposed, so a fix aimed at the wrong cause is visible as wrong before it is written, not after it ships.
03 · Problem
A fix nobody re-ran against the original failure
A change can look correct on inspection and still not actually close the bug it was written for, because the one test that matters was never run again.
How Debug closes it
Running the originally-failing test is a normal shell command in the same loop, so red-to-green is confirmed as part of the fix, not left for you to remember separately.
Where this stops
The real edges of what general can do here, stated plainly instead of left for you to discover.
It can only trace what is in the repo. A bug caused by a stale cache or bad prod config needs you to describe that part yourself.
A flaky bug may need you to reproduce it once first, so there is an actual failure to read instead of just a description.
Running a test to confirm a fix still triggers a permission prompt, the same as any other command.
The same agent, applied elsewhere
Build
Most agents bill every step of a feature at frontier rates. ML.ai reads and plans on a lighter model, spends the frontier tier only on writing and checking the diff, and hands you the same reviewed change for about 30% less.
See how it worksQuestions worth asking
Try ML.ai today, or talk to us about what is next.
Install the editor agent on your own machine, or book a call to talk through your team's workloads.
