Factory’s Droids can act without asking.ML.ai Code never ships a change you haven’t seen.

Factory is built to delegate a well-defined task end-to-end, unattended if you configure it that way. ML.ai Code is built around one rule: nothing changes on disk until you’ve seen the actual diff. That rule holds even at ML.ai Code’s most autonomous setting, and it’s backed by a measured benchmark result, not a target.

What each one actually is

ML.ai

ML.ai Code is an in-editor agent built around one default: nothing writes to disk until you approve it. Four scoped subagents (general, explore, architect, planSubAgent) each have a fixed, different write scope, and a command-safety classifier judges every shell command by effect before it runs, with a short list of absolute refusals that hold even with auto-approve on. On the official SWE-bench Verified harness, it solved 86% of a 50-instance slice, 28 points ahead of a leading frontier model given one guess.

Factory.ai

Factory’s "Droid" is a configurable agent usable from a desktop app, CLI, browser, JetBrains/VS Code/Vim, Slack, or Teams. Four autonomy tiers, set per org, range from read-only up to High, where a Droid can run Docker operations and push to git on its own, gated only by org-configured safety checks rather than a per-action approval. Factory has stopped publishing SWE-bench scores at all, saying the benchmark is Python-only and doesn’t match real enterprise work.

One fixed position, watch how far Factory’s ceiling can move past it.

Factory lets an org configure the autonomy ceiling anywhere from read-only up to High, where a Droid can push to git and run Docker on its own. ML.ai Code never moves off its one position: every edit and command gated, always.

Autonomy, gated to unattended

Factory.ai’s four tiers

ML.ai

Off

Read-only

Low

Edit + low-risk

Medium

Commits

High

Docker, git push

Factory sets the ceiling per org, anywhere from read-only to High. ML.ai Code has one position: every edit and command is gated, and no setting moves that pin.

Which one do you actually need?

ML.ai

Use this if

You want every edit and command to require a look at the real diff before it happens, backed by a measured benchmark result, with no setting that removes that gate.

Factory.ai

Use this if

You want to delegate a well-defined task end-to-end and are willing to configure autonomy ceilings and blocklists to run it unattended, including at High autonomy where it can push to git on its own.

Every claim, side by side

The same categories, compared row by row. Every line on both sides traces back to a real doc, pricing page, or published benchmark.

ML.aivsFactory.ai

What it is

What it actually is

A VS Code/Cursor extension with four scoped subagents (general, explore, architect, planSubAgent), each with a different, fixed write scope.

A configurable agent ("Droid") usable via desktop app, CLI, browser, JetBrains/VS Code/Vim, Slack, or Teams, with four autonomy tiers from read-only up to git push and Docker operations.

Autonomy & safety

Default autonomy

Every edit and command is held behind an Allow/Always/Deny gate; a native diff opens before you decide. Auto-approve cannot widen what Plan mode restricts.

Configurable per org: Off (read-only), Low (edit + low-risk commands), Medium (reversible workspace changes, commits), or High (Docker, git push, gated by safety checks).

Command safety model

Every shell command is lexed and classified by effect before running; a fixed list of absolute refusals (deletion outside the workspace, privilege escalation, fork bombs) holds regardless of mode or auto-approve setting.

Risk-tiered commands compared against your Autonomy Level; a separate blocklist is rejected outright at every level, with no approval prompt possible even at High.

Multi-step task handling

Background runs delegate a job to its own child session that survives closing the panel; finished output lands as a draft in the composer, never auto-sent.

"Missions" decompose a large task into parallel tracks across a fleet of Droids; "Spec Mode" restricts a session to read-only and low-risk commands only, no edits, while a spec is being reviewed.

Track record

Published benchmarks

A real measured result: 86% on a 50-instance slice of SWE-bench Verified, 28 points ahead of a leading frontier model’s single-shot baseline of 58%. The full 500-instance run and SWE-bench Pro are in progress.

SWE-bench Lite: 31.67% pass@1, dated June 2024 in their own technical report, over two years stale. No current SWE-bench Verified score is published; Terminal-Bench 58.75% is current (on Claude Opus 4.1).

Compliance & reach

Compliance certifications

SOC 2 Type II, ISO 27001, GDPR + CCPA, HIPAA on request, PCI DSS on dedicated capacity.

SOC 2 Type I (a point-in-time attestation, not the over-time Type II) and ISO 42001; GDPR and CCPA compliant, confirmed on their own security page.

Platform support

macOS Apple Silicon and Windows x64 only, inside VS Code or Cursor.

Desktop app (Mac Apple Silicon/Intel, Windows x64/ARM64), CLI, browser, mobile, IDE plugins (VS Code, Cursor, Windsurf, the full JetBrains family, Vim), Slack, Teams, Linear.

Pricing

Five tiers, $0 (30-day trial) to custom Enterprise; usage-based completion fee only past the plan’s included execution credit.

Six tiers: Pro $20/mo, Plus $100/mo, Max $200/mo, Teams $60/mo team fee + $40/mo per seat (up to 10 seats), Business and Enterprise custom. No confirmed free trial.

Where each one actually falls short

ML.ai’s own limits get the same weight as theirs, not a footnote after the sales pitch.

Where Factory.ai is limited

  • SOC 2 Type I is a point-in-time attestation, not the over-time operating-effectiveness attestation (Type II) that some enterprise security teams specifically require.

  • Their most recent published SWE-bench figure is from June 2024, on an unspecified model backend. No current SWE-bench Verified score is published; Factory’s own founders have said in interview they stopped running the benchmark because it’s Python-only and doesn’t reflect real enterprise work.

  • The original named-Droid lineup (Code/Review/Knowledge/Tutorial/Reliability Droid) was replaced within roughly a year by a different model (Missions, Spec Mode, custom droids), per the company’s own June 2026 "Factory 2.0" announcement. Older indexed content may not reflect the current product.

  • No confirmed free trial on the public pricing page.

Where ML.ai is limited

  • ML.ai Code has no equivalent to Factory’s configurable High-autonomy tier. By design, there is no setting that lets an agent push to git or run Docker operations without a human approving each one.

  • The 86% SWE-bench Verified result is measured on a 50-instance slice, not the full 500-instance public set; that fuller run and SWE-bench Pro are still in progress.

  • No headless/CI execution mode equivalent to Factory’s Droid Exec is documented for ML.ai Code today.

  • Supported platforms are narrower: macOS Apple Silicon and Windows x64 only, versus Factory’s desktop/CLI/browser/mobile/IDE/Slack/Teams spread.

Questions worth asking

Try ML.ai Code today, or talk to us about what is next.

Install the editor agent on your own machine, or book a call to talk through your team's workloads.