OpenFactory

Kalmantic / Outcome-led software delivery

Ship more customer-ready software.

We commit to a production outcome, operate the redesigned workflow, and transfer the resulting system to your team.

One workflowOne metricOne operating windowCustomer ownership

The current state

You have already bought AI speed.

The investment already exists across coding agents, engineering talent, QA, infrastructure and executive attention. The question now: did that investment change customer-ready release?

01Coding agents

Claude Code, Codex, Cursor and Devin

02Engineering system

Repositories, CI/CD, cloud and test infrastructure

03Operating attention

Training, prompt libraries and process experiments

04Leadership time

CEOs and CTOs redesigning delivery by hand

The return gap

Your developers got faster. Did production?

More code and more PRs create value only when they become customer-ready releases.

Observed across software teams

The same pattern appears at different scales.

11engineering + QA

80–90% on Claude Code. 3–5× POC speed. Review and QA remain manual.

30engineering + QA

About $10K monthly Claude spend. Merged PRs doubled. Safe autonomy remains unresolved.

187dev + testing

$36K monthly Claude spend. 3× reported output. Regression and test data constrain delivery.

Observed customer-discovery patterns. These are not Kalmantic-produced outcomes.

The engagement starts here

A committed production outcome.

A production outcome measures customer-ready work reaching production. It does not measure tokens, prompts, code generated or agent activity.

Throughput

Production-ready stories per month

Lead time

Approved requirement to production

Release flow

Safe releases per week

Autonomy

Stories reaching production with defined human intervention

One primary metric becomes the commitment.
CEOCustomer commitments shipped

Can the company deliver its roadmap and new-market promises faster?

CFOProduction output from the existing budget

Does the current engineering investment produce more finished software?

CTORequirement-to-production flow

Where does work wait, fail or return for rework?

One shared scoreboard

Before the commitment

The baseline makes the outcome credible.

We trace one workflow from approved requirement to production and measure where work waits, returns or requires scarce human judgment.

Baseline = throughput + lead time + rework + human intervention
  1. 01ApprovedInput accepted
  2. 02PlannedGaps resolved
  3. 03BuiltChange complete
  4. 04VerifiedCorrectness proved
  5. 05ProductionCustomer-ready
Example commitmentIncrease production-ready stories from 20 to 30 per monthwithin 90 days

Or: approved requirement to production from 20 days to 12. The target is set with you after baseline. The metric is production. Never tokens, PRs or activity.

The operating commitment

Diagnose. Redesign. Operate. Transfer.

01

Diagnose

Weeks 1–2

Trace the workflow, set the baseline, and find the constraint.

02

Redesign

Weeks 3–6

Change the workflow. Build only what moves the metric.

03

Operate

Weeks 7–12

Run it with your team until the number holds. Showcase every week.

04

Transfer

Week 12

Source, runbooks, metrics and a named internal owner. Yours to keep.

The guarantee

If we miss the number, you do not pay for it.

Every pilot on the market charges for activity and hopes for a result. We charge for the result. The risk of being wrong sits with us.

Outcome fee: zero

Not earned unless the agreed metric hits the agreed target.

We keep operating

Past 90 days at no charge, until the number holds or you call it.

You keep everything

Source, agents, evals and runbooks. Hit or miss, it stays with you.

Tokens on us

Model and compute cost through the window is ours, not a pass-through.

The price, printed

$190K for 90 days.

A third of it only if the number hits.
Foundation$40K

Baseline, workflow redesign, first build. Invoiced when the operating foundation is delivered, week 6.

Operating$30K / month

Three months. Founders embedded, weekly showcase. Invoiced monthly while Kalmantic operates the workflow.

Outcome$60K

The committed production result. Earned only when the agreed metric reaches the agreed target.

Tokens and compute through the window included. Larger teams and second workflows priced from this base after baseline.

No predetermined tool

The intervention follows the constraint.

We do not sell a generic automation recipe. We change the part of the production system that limits customer-ready output.

Planning

Requirements and gap resolution consume the cycle

Context

Corrections disappear after each agent session

Review

Generated code overwhelms senior reviewers

Test data

Developers seed fixtures and environments manually

Regression

QA queues absorb the gains from faster coding

Release

Deployment and service coordination remain human-bound

What remains with the customer

A customer-owned agentic software factory system.

Workflow

The path from approved work to production

Agents

Customer-specific builders, reviewers and operators

Evaluations

Evidence that defines and proves correct

Learning

Corrections that persist into the next run

Control

Human approval where judgment still matters

Runs on OpenFactory, the workspace your platform team keeps. Models will change. The system stays under your control.

Proof

Already built, already running, already paid for by someone else.

NuveproCustomer-specific sandbox production

Factory stations, maker-checker agents, a swappable harness and adaptive documentation. 11% measured experience uplift. Paying design partner since August 2026.

OpenFactoryThe agent workspace your team inherits

Built and shipped by the founders. 53 builders, 3 design partners. Every engagement leaves it installed.

TokenTopperOpen-source token and PR visibility

Installed at a 20-person team running 70–80 PRs a week on 9–10B tokens a month. Their first usage number.

Case study / Nuvepro

Faster production of customer-specific sandboxes.

Before

An engineer studied the technology, built the sandbox, tested it and prepared the supporting material.

1–2 days per sandbox
Production constraint

New technologies arrive faster than a manual, customer-specific production process can absorb.

70–80% AI-generation ambition
Kalmantic intervention

Factory stations, maker-checker agents, swappable components, terminal re-engineering and adaptive documentation.

11% measured experience uplift

The first architecture experiment outperformed the previous RAG approach. A second experiment is underway.

Low regret by design

Your team needs us less.

One workflow, your code, 90 days. If it fails there is little to unwind. If it works you hold the system and the evidence to fund the rollout.

Source

Customer-owned code and deployment configuration

Operations

Workflow definitions, runbooks and production metrics

Correctness

Evaluations, controls and evidence requirements

Ownership

Training and a named internal platform owner

Start small enough to be reversible. Prove enough to fund the rollout.

Why Kalmantic

Product research, embedded execution, production accountability.

Kalmantic builds inference infrastructure, coding agents and self-improving harnesses. Founder-run, three engagements at a time. The people here are the people in your repo.

RajanCEO & co-founder

Business outcome, executive alignment and commercial accountability.

KashiCo-founder & CTO

Technical architecture, system construction and embedded operation.

AnanyaCo-founder, GTM

Customer workflow, adoption and operating change.

The first decision

Which production outcome matters now?

Choose one workflow. In the first working session we baseline it, find the constraint, set the 90-day number and print the price.

Next working sessionWorkflow • Baseline • Target • Owner • PriceBook the session ↗