10/10
Main development gate, 5 runs per repo
12/12
Four smaller gates, 3 runs each
0/9
Frozen held-out set of then-unseen repos
3 of 3
Held-out repos that later passed once in development
Regression closure in progress
Give it a Flask repo and a recipe; it plans the target architecture, rewrites existing files and creates the new modules the migration requires, verifies against the repo's own tests in a network-off Docker sandbox, recovers from failures under bounded budgets, and reports honestly, including when it fails.
Portage is a durable, measured code-migration agent. Its current recipe migrates Flask applications to FastAPI.
It surveys a repository, builds a structural graph, and plans the target architecture, including new modules when the migration needs them. Cross-file interfaces and capability ownership freeze before generation. Portage executes dependency-aware groups of changes, rejects invalid code through mechanical AST and topology checks, and runs repository tests in an ephemeral Docker sandbox with networking disabled.
LangGraph execution checkpoints in Postgres. Recovery runs under bounded attempt and cost budgets, and a failed targeted repair restores the last coherent cut. A run is green only when every task is complete, the full suite passes, the test oracle is intact, and the measured tree is migrated. Passing the original tests after rollback still counts as red.
The CLI drives autonomous migrations. MCP exposes repository graphs, blast-radius queries, and proposed-patch verification in an isolated copy. The frontend is the observability and evidence surface.
Development results are strong, but unseen-repository reliability is not yet proven. Frozen R5 v1 remains 0/9. All three former R5 repositories later reached one strict autonomous K1 green as development inputs. Regression closure is still in progress, and another generalization claim requires a genuinely fresh frozen held-out set after that work closes.
Below is the core durability claim, live: the worker is killed mid-migration, and a restarted worker resumes from the Postgres checkpoint instead of starting over.
reproduce: bash scripts/demo_kill_resume.sh · stricter: scripts/dod_check.sh
Most "AI migration" demos are single-shot prompts with no verification story. Portage is built around the opposite claim: a migration is only real if the repo's own tests still pass, every planned file was actually migrated, and recovery cannot game the score by giving up. The governing principle is narrow + measured beats broad + unproven. One hard migration (Flask → FastAPI) instead of a catalogue of half-working recipes. An eval harness that runs the real queue/worker path, not a mocked agent loop. A failure taxonomy with SOLVED / PARTIAL / OPEN statuses and evidence, not an all-green sheet. The autonomous + eval core is the credibility engine for the MCP product: if the verify/recover loop is measured, another agent can trust verify_patch_in_sandbox over a raw sandbox.
A submitted job runs this LangGraph graph, checkpointing to Postgres after every node: kill the worker mid-run and a restarted worker resumes from the last completed node. When Verify fails, Recover classifies the failure and routes back: targeted rollback + regenerate to Execute, replan to Plan for planner misses, or give up to Integrate once budgets are exhausted, reporting an honest red rather than a gamed green.
v1 ships one recipe: Flask → FastAPI. That target is deliberate. Routing decorators, request/response handling, blueprints → routers, error handlers, app factories, and ambient request context (g, session) need understanding, not mechanical rewriting; deterministic codemods cannot do this reliably. The architecture is recipe-pluggable; the evidence is recipe-specific by design. The capability that unlocked the hard repos: some migrations are unreachable by rewriting existing files. Flask's g/session have no FastAPI equivalent. A correct port needs a new request-context module, a test-compatibility surface, a rendering layer, and every consumer wired to them coherently. Portage plans those artifacts with a bounded architect call, freezes their contracts before generation, compiles the deterministic parts itself, and enforces that a framework-shaped capability is only valid when the plan owns and implements it, so a model can't reference a helper it wishes existed. A second wave, coherent-cut preservation, closed the gap the first one left open. One bad file inside an otherwise-correct migration used to trigger a full rollback of every file in its verification cut, so a single local mistake could sink a ten-file run. Recover now checkpoints the last coherent state before a targeted repair and restores that on failure instead of the whole migration, and one shared gate (caller, capability, import-direction, cycle, and contract checks) runs identically across every generation path: first draft, contract repair, and targeted repair alike. That is what took watchlist, a Flask-SQLAlchemy app that had never gone green, to autonomous 15/15, and pushed flaskr to a 5-for-5 reliability gate. Then the first frozen held-out evaluation supplied the correction. On three repositories that had never been migrated during development, Portage scored 0/9 strict green. The engine failed honestly: five trees restored coherently, four stayed migrated-but-red, zero were hybrid. But the recipe did not generalize. The project now has both halves of a credible result, strong development convergence and a measured unseen-repository gap, and publishes them together. A later forensic audit then corrected part of that failure explanation. The protected ws-example test file was byte-identical; a truncated read of a long protected test file caused a false oracle-integrity alarm. That correction changes the explanation, not the 0/9, because each sample was independently red for migration reasons. Development results are strong, but unseen-repository reliability is not yet proven. Frozen R5 v1 remains 0/9. All three former R5 repositories later reached one strict autonomous K1 green as development inputs. Regression closure is still in progress, and another generalization claim requires a genuinely fresh frozen held-out set after that work closes. One core engine, two interfaces. Autonomous mode: `portage migrate <repo> --watch` drives the full graph. Co-pilot mode: Claude Code / Cursor call verify_patch_in_sandbox, repo_graph, and blast_radius over MCP, the same verified primitives the eval numbers were measured on. The dashboard is the observability and proof surface, not the front door.
The portage console script is a thin httpx client over the REST API; it never touches the DB or queue directly, the same boundary the dashboard respects. migrate --watch streams live task transitions; status, jobs, report and diff <job-id> (add --stat for a summary) cover inspection.
The exit-code contract belongs to the migration, not to every command that can print something about it. A completed status reflects the strict result: 0 only for an honest green, 1 for finished but not complete-and-green, 2 for usage or infrastructure. On a job that is still running, status can return 0 because the query succeeded, and jobs listing rows means the listing worked, not that a migration passed. Read the inspection commands as inspection: the verdict is the strict outcome they report, not their exit status.
The MCP server exposes the verified core so another AI agent can test its own work before writing to the caller's tree: verify_patch_in_sandbox copies the repo, applies a unified diff, runs the tests network-off, and returns structured pass/fail with failing test names, never mutating the caller's files. repo_graph and blast_radius give it structural awareness. Here is what a blast-radius query actually computes:
The blast_radius primitive in action: when db.py changes, Portage walks the structural code graph outward, direct callers first (hop 1) and then their dependents (hop 2), and selects only the tests that cover the impacted set. Verify uses this to iterate fast; the final honesty bar still runs the full suite. The same query is exposed to co-pilot agents as the blast_radius MCP tool.
Next.js App Router, REST only; the frontend never owns schema. Jobs list with launch form, job detail with live pipeline route, per-file diffs, and attempt tier/model timelines, and a public /eval leaderboard rendering per repo×scenario green rates, mean±variance, cost, and recovery straight from the runs/metrics tables.
A bounded architect call proposes new target-architecture modules; a deterministic contract compiler fills in what the engine already derives; contracts freeze before generation and bind every retry, escalation, replan, and resume. Created files get the same ordering, diffs, rollback, and cost accounting as rewrites.
A Flask-shaped capability (test_client, app_context, g, session) is accepted only when a frozen plan artifact owns and implements it, checked receiver-aware. "The model referenced a module it wished existed" becomes a pre-sandbox rejection, not a silent runtime failure.
LangGraph Postgres checkpointer after every node; worker lease with heartbeat. Kill the worker mid-run and a restarted one resumes: Ingest runs once, Execute skips already-applied files via content hashes.
Uniquely attributable failures repair the single owning artifact (measured: a stray .decode() fixed for $0.011 without touching its ten-file cut). Otherwise: targeted rollback + regenerate, widen-on-repeat, replan, skip-and-continue as last resort, all budget-bounded.
A failed targeted repair restores the last known-coherent checkpoint, not the original sources, so one bad file can no longer roll back the nine correct ones beside it. The single highest-leverage fix in the project: it converted watchlist and flaskr from occasional greens into repeatable ones.
Test files are protected artifacts: names, assertions, raises/parametrize/skip structure and fixture lifecycles are frozen at Plan; only sanctioned plumbing may differ. The guard also had to survive an audit of itself. A held-out 0.75 integrity score looked like deleted tests, but the files were byte-identical and both readers had truncated a long protected test file. Fixed, regression-covered, and the run stayed red on its own merits.
Green cannot be gamed by skip-and-continue, empty diffs, or all-skipped suites. Report reloads task truth from Postgres, Integrate recomputes the diff, Verify requires passed > 0, and engine errors count against the score.
First N attempts use the driver model tier; later attempts escalate. Every attempt lands in attempts_log with tier, model, tokens, and USD, so "how often does escalation rescue?" is a SQL query.
Every LLM call's tokens and USD recorded per attempt, summed per job, averaged per eval cell, with retries, escalations, and architect calls included. Cost scales with recovery, and that relationship is part of the result.
Three repositories were frozen, baseline-vetted, and unseen at the time R5 v1 ran. It ran once from a pinned commit and scored 0/9. No failed sample was renamed, replaced, or rerun, and the result is published beside the development gates rather than behind them. All three became development inputs once those findings shaped the work, so their later greens cannot supply held-out evidence and the next generalization test needs a genuinely fresh frozen corpus.
Development performance and unseen-repository performance answer different questions. These are disclosed development milestones and one frozen held-out evaluation, not a claim that every regression gate is green on the current tree.
Green requires the full suite passing, every planned task done, zero skips, oracle integrity 1.0, and a tree_state of migrated: a run that recovery rolls back to original sources passes the original suite and still scores red.
| Evidence set | Result | What it means |
|---|---|---|
| Flaskr + Watchlist · disclosed K=5 v4 development gate | 10/10 green | Flaskr 24/24 tests and Watchlist 15/15 per successful run. Earlier gate generations stay disclosed. |
| Items, RESTX, Structural, Minimal · their K=3 development gates | 12/12 green | Four independently scoped gates, not the entire general-preservation suite. |
| Seven-repository autonomous development confirmation | 6/7 green | One sample per repository. Microblog was red. |
| Microblog accepted-plan replay | 4/4 tests · 26/26 tasks | Historical replay diagnostic. Excluded from autonomous rates. |
| Frozen R5 v1 · three then-unseen repositories, K=3 each | 0/9 green | Ran exactly once as the frozen held-out evaluation. The recipe had not generalized. |
| Post-R5 ws-example · development K1 | 42/42 tests · 5/5 tasks | One strict autonomous green after becoming a development input. |
| Post-R5 Silicon · development K1 | 34/34 tests · 14/14 tasks | One strict autonomous green after becoming a development input. |
| Post-R5 flask-email-login · development K1 | 18/18 tests · 15/15 tasks | One strict autonomous green after becoming a development input. |
Each of the three post-R5 development K1 greens has tree_state=migrated and oracle integrity 1.0. They do not replace R5 v1 and do not count as held-out evidence. A later accepted-plan Microblog preservation replay is red, so current regression closure remains open.
The held-out set, in full
R5 v1 ran exactly once as the frozen held-out evaluation, from commit 3b25ee9 against corpus/heldout.toml, with the frozen offline sandbox and GPT-4o on both model tiers. It scored 0/9 strict autonomous green. Five runs restored coherently, four stayed migrated but red, and zero produced hybrid trees. Those failures showed that the development performance had not generalized.
| Unseen repo | Baseline | K=3 | Dominant failure |
|---|---|---|---|
| ws-example | 42/42 | 0/3 | generated test-client facade shadowed FastAPI route decorators; two samples stalled at 13/42 |
| silicon | 34/34 | 0/3 | invalid generated signatures; a raw FastAPI object constructed instead of the frozen facade |
| flask-email-login | 18/18 | 0/3 | architect missed the required context owner; the fallback left CSRF and mail providers as None |
The scoring machinery held even though the recipe failed
This is the part worth reading. Rejected cuts restored the original suite, and those restored passes contributed exactly zero migration score. Trees came back 4 migrated / 5 restored-coherent / 0 hybrid. All nine jobs produced durable reports with no missing run rows: 119 LLM calls, 19 recovery visits, $3.8643, architect acceptance 6/9.
A later forensic audit corrected part of the failure explanation. The protected ws-example test file was byte-identical; a truncated read of a long protected test file caused a false oracle-integrity alarm. That correction changes the explanation, not the 0/9. Each sample was independently red for migration reasons: two stalled at 13/42 behind the shadowed route decorators, and one restored the original tree with tasks still incomplete.
Once those findings influenced development, all three repositories became development inputs permanently. Their later successful runs cannot supply held-out evidence. The original result stays visible. The next meaningful generalization test requires a genuinely fresh, frozen corpus after current regression closure.
What convergence looks like when it works
flaskr, the canonical Flask tutorial app (templates + factory + auth + SQLite + Click CLI), went from never green in any grid to 24/24 tests, 12/12 tasks and zero recovery, and now holds 5/5 at K=5. One stored flaskr job cost $0.154 including retries and planning; that is a single run, not a rate across the full history. watchlist, a Flask-SQLAlchemy app that had never gone green, holds 5/5 at 15/15 tests. Both needed new modules to exist; the engine designed and wired them. What made them repeatable rather than occasional was coherent-cut preservation.
Where it stands now
All three former R5 repositories have now reached one strict autonomous development K1 green: ws-example at 42/42 tests and 5/5 tasks; Silicon at 34/34 and 14/14; flask-email-login at 18/18 and 15/15. All three have migrated trees and oracle integrity 1.0.
The current R5 remediation goal is not fully closed. A new accepted-plan Microblog preservation replay exposed a general callback-retention bug involving a source-defined initializer callback. That is a replay diagnostic, not autonomous evidence. Current regression closure is still in progress.
Portage has strong development results. It is not production-ready, generally solved, or held-out validated. Unseen-repository reliability remains unproven. After the current goal closes, the next meaningful generalization claim requires a newly frozen held-out corpus.
Evidence reviewed September 8, 2026. Development gate history is from July 2026; frozen R5 v1 ran in July 2026; the three post-R5 development K1 greens are from August 2026. These stored runs were inspected again for this page; they were not rerun on the current tree.
Every implementation choice, numbered, from the queue claim to the sandbox runtime.
Friction
Takeaways
portage migrate --watch · exit 0 = honest green
verify_patch_in_sandbox · repo_graph · blast_radius
jobs · recovery timelines · public /eval
Every moving part explained: the animated architecture, graph-node lifecycle, checkpoint and lease mechanics, sandbox anti-gaming predicates, artifact-producing plans, the eight recovery strategies, K-run eval methodology with non-claims, and the ten-category failure taxonomy with evidence.