Debug

Hand it the error. ML.ai finds the real cause first.

Give ML.ai a failing test, an error message, or a stack trace and it reads the execution path, finds the actual cause, and proposes a fix as a diff you approve.

Anyone staring at a failing test, an exception in production logs, or behavior that does not match what the code appears to say.

ML.ai

This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.

3 of 3 files read
  • Reading src/auth/session.ts
  • Reading src/lib/tokenStore.js
  • Reading test/auth.test.ts
Wants to write 5 lines
Allowed
42 async function refresh(token) {
43- if (!token) return null
43+ if (!token?.exp) throw new Error('expired')
44+ return await rotate(token)
45 }

The cause was a token missing its expiry field slipping past a falsy check instead of raising, so refresh() silently returned null under a race. The fix and a regression assertion land together.

This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.

DebugML.ai Standard
Ready

The line the trace points at is rarely the line that is wrong

Finding the actual cause is the hard part of debugging. These four things are what keep that part honest.

The line it points at vs. the line that's wrong

auth.test.ts:12
session.ts:43cause
tokenStore.js:8

It reads the call stack, not just the top frame

A stack trace names where the error surfaced, not where it started. ML.ai walks the frames underneath it and follows the data through them, so the fix lands on the frame that actually caused the failure.

States the cause, then -- only then -- the fix

token missing expiry, falsy check hid it

+ throw new Error('expired')

It states the cause before it writes anything

A plain-language diagnosis comes first, and the fix only appears after. If the stated cause is wrong, that is obvious immediately, before any code is written on top of it.

A diagnosis you can check, not just trust

diagnosis

race on token refresh

The diagnosis is a claim you can check, not a black box

You see the reasoning in words, not just a patched file. Push back on it directly and it revises the diagnosis, the same way you would redirect a wrong plan before a wrong fix gets written.

Re-run against the failure it was meant to fix

auth.test.ts

passing

The fix is checked against the failure it was meant to close

Running the originally-failing test is a shell command like any other. Confirming red-to-green is part of the loop, not a separate step you have to remember to do yourself.

Doing it by hand vs. handing it off

Without it

You read the stack trace, open the file it names, and start guessing, because the line the trace points at is usually a symptom, not the actual cause.

WITH ML.ai CODE

ML.ai follows the failure back through its callers, states the real cause in plain language, then proposes the fix, so the diagnosis is checkable before a fix is written on top of it.

Which agent actually does this

general

Traces the cause and writes the fix

Reads the function that failed and its callers, states what is actually wrong, then proposes a fix as a diff behind the same approval gate as any other edit.

@explore

Optional, to map the area first

Read-only search across the codebase, useful when the failing code touches an unfamiliar part of the system before a fix is attempted.

How it actually runs

A bug report is rarely the same thing as the bug. ML.ai follows the failure back through the code that produced it, files and callers, not just the line the stack trace points at, before it proposes a change.

    1

    Hand over the symptom

    Paste the error, the failing test name, or describe what is happening versus what should happen.

    2

    It traces the real path

    The agent reads the function that failed and walks its callers and the data flowing through them, rather than guessing from the error text alone.

    3

    It states the cause before fixing

    You get a short explanation of what is actually wrong, so a wrong diagnosis is obvious before a fix is written on top of it.

    4

    Review the fix as a diff

    The proposed change is held behind the same Allow/Always/Deny gate as any other edit. Approve it, or push back if the cause looks wrong.

ML.ai

This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.

3 of 3 files read
  • Reading src/auth/session.ts
  • Reading src/lib/tokenStore.js
  • Reading test/auth.test.ts
Wants to write 5 lines
Allowed
42 async function refresh(token) {
43- if (!token) return null
43+ if (!token?.exp) throw new Error('expired')
44+ return await rotate(token)
45 }

The cause was a token missing its expiry field slipping past a falsy check instead of raising, so refresh() silently returned null under a race. The fix and a regression assertion land together.

This test fails intermittently: `expect(session.isValid).toBe(true)` in auth.test.ts.

DebugML.ai Standard
Ready

What actually goes wrong, without it

Three specific ways this kind of task breaks, and the part of the mechanism above that closes each one.

01 · Problem

The stack trace names a symptom, not a cause

The line an error is thrown from is often just where a bad value finally broke something, three function calls away from where that value actually went wrong.

How Debug closes it

general reads the function that failed and walks its callers and the data flowing through them, rather than treating the top frame of the trace as the answer.

02 · Problem

A plausible-looking fix that treats the symptom

Patching the line the trace points at often makes the immediate error go away while leaving the actual defect in place, ready to resurface somewhere else.

How Debug closes it

A diagnosis is stated in plain language before any fix is proposed, so a fix aimed at the wrong cause is visible as wrong before it is written, not after it ships.

03 · Problem

A fix nobody re-ran against the original failure

A change can look correct on inspection and still not actually close the bug it was written for, because the one test that matters was never run again.

How Debug closes it

Running the originally-failing test is a normal shell command in the same loop, so red-to-green is confirmed as part of the fix, not left for you to remember separately.

Where this stops

The real edges of what general can do here, stated plainly instead of left for you to discover.

01

It can only trace what is in the repo. A bug caused by a stale cache or bad prod config needs you to describe that part yourself.

02

A flaky bug may need you to reproduce it once first, so there is an actual failure to read instead of just a description.

03

Running a test to confirm a fix still triggers a permission prompt, the same as any other command.

The same agent, applied elsewhere

Build

Most agents bill every step of a feature at frontier rates. ML.ai reads and plans on a lighter model, spends the frontier tier only on writing and checking the diff, and hands you the same reviewed change for about 30% less.

See how it works

Questions worth asking

Try ML.ai today, or talk to us about what is next.

Install the editor agent on your own machine, or book a call to talk through your team's workloads.