Autonomous project executionLocal-first / model-independent

Great worktakes time.Keep yours.

Ten minutes can make a prototype. Complex software takes planning, testing, failure, and refinement. Crinkle stays for the whole process, so you don’t have to.

Follow the work
A different measure

The goal isn’t to generate faster. It’s to finish without you managing every step.

01 / 04

Crinkle is an autonomous software team for multi-hour and multi-day projects. Give it the brief, the boundaries, and the definition of done. It decomposes the work, assigns specialists, tests the real application, repairs failures, and keeps going.

Benchmarked in the open

Same model.
8× more requirements met.

We benchmarked one mid-size open model three ways on the same complex tasks: one-shot generation, a naive retry loop, and inside Crinkle. Graded by external browser checks that click the real app. No model grades itself.

8×

more hard-task runs met every coding requirement

Multi-feature builds, cross-file refactors, symptom-only bug hunts: on its own, the model met the full requirements in 1 of 15 runs, and the retry loop in 2. Inside Crinkle, the same model met them in 8, including a monolith-to-modules refactor at 30/30 checks, three runs straight.

0

false “done” claims in 66 runs

The bare model declared success on software that failed its requirements in 67% of hard runs, and a naive retry loop still did it 60% of the time. Crinkle’s verified-done shipped zero false completions across the entire cycle.

100%

of requirements met on the standard suite

Every one of 51 Crinkle runs produced software that passed every independent check: 168 of 168. The bare model’s best arm met every requirement in 90% of its runs.

qwen3.6-35B running locally · 20 tasks × 3 runs per arm · graded by Playwright driving the built apps · full methodology + raw scorecards

The long run

It works while you don’t.

Models time out. Tests fail. Designs need another pass. Crinkle recovers, reroutes, and continues toward verified milestones instead of handing the management job back to you.

09:42Elapsed / unattended
Mira / planningForge / repairVerity / browserSentinel / review
Project telemetrySimulated run
22:14
Scope locked12 acceptance criteria · 4 milestones
23:08
Core flow implementedBuild green · browser test started
01:46
Failure foundRace condition reproduced and routed
03:31
Regression repaired11 of 12 criteria now evidenced
07:56
Independent reviewSecurity and dependency scan clear
Human interventions00
The team

Specialists that prompt each other.

Most tools make you the go-between: approve this, re-prompt that, ferry one AI's output into another. In Crinkle the agents brief each other. Failures route straight back to the Builder with the evidence attached, and nobody waits for you to press send.

Team channelSimulated run
Proof, not promises

Done has evidence.

Crinkle does not confuse generated code with working software. Completion has to survive the actual toolchain, the running application, adversarial input, and an independent review.

Verified build / CR-1842Illustrative · 07:58:42 local
12/12

Acceptance criteria proven against the running application. Full evidence retained with the run.

Toolchain

Typecheck, build, unit and integration suite

Passed

Browser acceptance

Real clicks, inputs, navigation, and rendered output

Passed

Adversarial pass

Invalid states, edge cases, and regression replay

Passed

Security review

Code, commands, dependencies, and exposed secrets

Clear
Built for endurance

Autonomy with discipline.

01 / Persistent

Keeps the goal intact

Milestones, project memory, and acceptance criteria survive across long runs. Context is managed without forgetting what matters.

02 / Recoverable

Failure isn’t the end

Retries, fallback models, checkpoints, and restore points keep a provider outage or bad edit from killing the project.

03 / Accountable

Agents check agents

Managers coordinate. Engineers build. Testers try to break it. Security reviews last. No single model grades its own homework.

04 / Bounded

You set the limits

Time, spend, tool, and approval boundaries are yours. Pause instantly, steer when useful, or let the system continue unattended.

05 / Observable

See convergence

Follow milestones completed, failures shrinking, checks passing, current work, cost, and the reason behind every stop.

06 / Local-first

Your machine. Your code.

Projects and credentials remain local. Use cloud providers, subscriptions, or private local models without rebuilding the workflow.

You set the terms

Autonomy you can interrupt.

Unattended never means unaccountable. Every run operates inside boundaries you define, and it leaves a plain-language record of every decision, command, and dollar.

Windows betaWorks offline with local modelsProjects never leave your machine
Run boundariesYours to change mid-run
Spend capThe run stops at the cap, never past it
$7.44 / $12.00
Command consentAnything outside your allowlist waits for you
$ npm install playwright
Restore pointsCheckpoint at every milestone · one-click rollback, no git required
14 saved
Steer anytimeType mid-run; the team folds it in at the next step without stopping
Never required
Instant pauseStop takes effect now, and the run report explains where it stood
Model-independent

The system is the constant. Models are the crew.

Use the best model for each job. Crinkle routes work by capability, cost, context, and current health, then walks the fallback chain when a provider fails at 3 AM.

SubscriptionsCodex · Claude Code · Google
API providersOpenAI · Anthropic · Gemini · NIM + more
Private + localOllama · LM Studio · Open WebUI
ControlRouting · budgets · fallbacks
Crinkle beta

Give it the hard project.

Set the goal. Set the boundaries. Come back to working software, and the evidence that it works.

Download for Windows
Free beta · runs locally · bring your own AI
$ npx crinkle