Keeps the goal intact
Milestones, project memory, and acceptance criteria survive across long runs. Context is managed without forgetting what matters.
Ten minutes can make a prototype. Complex software takes planning, testing, failure, and refinement. Crinkle stays for the whole process, so you don’t have to.
The goal isn’t to generate faster. It’s to finish without you managing every step.
Crinkle is an autonomous software team for multi-hour and multi-day projects. Give it the brief, the boundaries, and the definition of done. It decomposes the work, assigns specialists, tests the real application, repairs failures, and keeps going.
We benchmarked one mid-size open model three ways on the same complex tasks: one-shot generation, a naive retry loop, and inside Crinkle. Graded by external browser checks that click the real app. No model grades itself.
Multi-feature builds, cross-file refactors, symptom-only bug hunts: on its own, the model met the full requirements in 1 of 15 runs, and the retry loop in 2. Inside Crinkle, the same model met them in 8, including a monolith-to-modules refactor at 30/30 checks, three runs straight.
The bare model declared success on software that failed its requirements in 67% of hard runs, and a naive retry loop still did it 60% of the time. Crinkle’s verified-done shipped zero false completions across the entire cycle.
Every one of 51 Crinkle runs produced software that passed every independent check: 168 of 168. The bare model’s best arm met every requirement in 90% of its runs.
qwen3.6-35B running locally · 20 tasks × 3 runs per arm · graded by Playwright driving the built apps · full methodology + raw scorecards
Models time out. Tests fail. Designs need another pass. Crinkle recovers, reroutes, and continues toward verified milestones instead of handing the management job back to you.
Most tools make you the go-between: approve this, re-prompt that, ferry one AI's output into another. In Crinkle the agents brief each other. Failures route straight back to the Builder with the evidence attached, and nobody waits for you to press send.
Crinkle does not confuse generated code with working software. Completion has to survive the actual toolchain, the running application, adversarial input, and an independent review.
Acceptance criteria proven against the running application. Full evidence retained with the run.
Typecheck, build, unit and integration suite
Real clicks, inputs, navigation, and rendered output
Invalid states, edge cases, and regression replay
Code, commands, dependencies, and exposed secrets
Milestones, project memory, and acceptance criteria survive across long runs. Context is managed without forgetting what matters.
Retries, fallback models, checkpoints, and restore points keep a provider outage or bad edit from killing the project.
Managers coordinate. Engineers build. Testers try to break it. Security reviews last. No single model grades its own homework.
Time, spend, tool, and approval boundaries are yours. Pause instantly, steer when useful, or let the system continue unattended.
Follow milestones completed, failures shrinking, checks passing, current work, cost, and the reason behind every stop.
Projects and credentials remain local. Use cloud providers, subscriptions, or private local models without rebuilding the workflow.
Unattended never means unaccountable. Every run operates inside boundaries you define, and it leaves a plain-language record of every decision, command, and dollar.
$ npm install playwrightUse the best model for each job. Crinkle routes work by capability, cost, context, and current health, then walks the fallback chain when a provider fails at 3 AM.
Set the goal. Set the boundaries. Come back to working software, and the evidence that it works.
$ npx crinkleWe’ll send a verification code. Use the same email when you sign in to Crinkle.