GitHub Copilot Rust Migration: A Playbook for AI-Agent Ports

• Tapovan

According to GitHub, every one of the dozens of regressions found in its Copilot runtime port had been merged to main, so every one of them compiled. The Rust compiler flagged 8,678 coded errors along the way, and only 1.7% involved ownership, borrowing or lifetimes. The GitHub Copilot Rust migration worked because a process caught what Rust could not. This post turns that process into a playbook you can reuse.

The GitHub Copilot Rust migration in 60 seconds

Every figure below is GitHub's own claim, taken from Stephen Toub's post Migrating the GitHub Copilot runtime to Rust, using Copilot (published 2026-09-16, updated 2026-09-23). Most come from the private github/copilot-agent-runtime repo and Toub's session logs, so outsiders cannot reproduce them.

WhatGitHub's figure
Porting windowAbout 14.5 weeks, from early May to August 21, 2026
TypeScript that went through the port~430,000 production lines, against an initial guess of ~130,000
Rust at the finish832,378 production lines, plus 468,689 lines of unit tests
Pull requests128, each merged to main and released as it landed
Releases during the port135 (100 pre-release, 35 stable)
Tokens and bill~136.3 billion tokens, ~$120,000
Prompt-cache hit rate96.22%
Developer timeAbout three weeks, by Toub's own rough estimate
Speedup, model inference excluded3.0-5.9x out-of-process, 6.3-21.4x in-process

Why GitHub left Node.js and V8 (and why not C# or Go)

One agent runtime sits behind the Copilot CLI, the Copilot app, the SDK, VS Code, Visual Studio and Office. Before the port, every SDK client launched the CLI as a child process, costing roughly 100 MB of working set at minimum per client, according to GitHub. Every call went over JSON-RPC to that process, and a Node crash took the session down.

The real requirement: one embeddable C ABI for six SDKs

GitHub needed clean interop for all six SDK languages (C#, TypeScript, Python, Rust, Go, Java), low startup and steady-state cost, predictable resource use, a smaller supply-chain surface and more correct-by-construction code. Toub frames it as leaving Node.js and V8, not choosing Rust for its own sake, and the post does not recommend Rust for every large TypeScript codebase. For the general trade-offs, see Rust vs Go vs TypeScript for the back end.

Why Rust instead of C#: the .NET backlash and the TypeScript 7 precedent

GitHub's post never names C#/.NET or Go as candidates, so it never explains why they lost; that debate comes from the video and social media, and the garbage-collector theory at 4:15 is the presenter's own suspicion. The post even notes that a C#, Java or Go compiler would have caught most of the port's compile errors too. The C ABI requirement decided it, not the type system.

The video also cites Microsoft's other big port. TypeScript 7.0, the Go-native compiler, was announced on 2025-03-11 and went GA on 2026-07-08; Microsoft calls it a faithful port with full builds typically 8x to 12x faster. Both teams translated faithfully first; their requirements picked the language.

Watch: Microsoft Rewrote Copilot in Rust. .NET Devs Are Mad by Milan Jovanović

Milan's 9:49 video covers the community reaction: 1:18 for GitHub's reasons, 3:34 for the .NET backlash, 5:07 for code review at scale and 7:22 for the benchmarks. One correction: at 0:00 and 9:06 it calls this one engineer's work. GitHub calls it a team effort led by one developer, and names the engineers behind the temporary interop, all six SDK FFI layers, packaging, crate splitting and CI caching.

More .NET and architecture videos are on Milan Jovanović's channel.

The playbook: 7 phases of an AI code migration

GitHub's lessons are scattered through a very long post. Here they are in the order you would apply them. GitHub used the Copilot app, but none of the phases depends on it.

Phase 0: Decide whether to port at all

Port only if all three answers are yes.

  • Requirement: Do you need something the current runtime cannot give you, such as an embeddable C ABI, lower memory or fewer dependencies? "Rust is faster" is not enough: GitHub's speedups came as much from dropping the process hop as from the language.
  • Readiness: Do E2E tests cover the behaviour users depend on? GitHub traced most missing-feature regressions to thin E2E coverage.
  • Readiness: Can you ship incrementally through pre-releases? A single big cutover rules out the in-place strategy below.

Phase 1: Build the oracle, E2E tests the agent may not edit

GitHub's runtime ended with 174,675 lines of TypeScript E2E tests, and the SDK repo has roughly 130,000 more across six languages. Those tests are the oracle. Toub's point: let them be rewritten during the port and you "lose your oracle". One port dropped the SDK callbacks and deleted the test that covered them, so agents now need explicit human consent to touch any E2E test.

Phase 2: Pilot PRs that prove toolchain, CI, FFI and packaging

GitHub's first two PRs shipped no features. They stood up the Cargo workspace and toolchain, lints, CI, the build pipeline and the agents' instruction files. Next came small pure-logic primitives with no I/O or shared state and good existing tests. Only then did a full port PR carry three side-effect-free helpers end to end.

Phase 3: Port leaves inward with atomic in-place swaps

GitHub rejected both a big-bang rewrite and an A/B hot-swap. Each PR swaps one TypeScript implementation for a thin shim that calls Rust and removes the old code in the same change. That removal pays off later: if someone edits already-ported code on main, the rebase conflicts, so the edit cannot vanish quietly. The temporary napi-rs layer peaked on August 3 (2,019 exports, 3,356 TypeScript call sites) and ended at zero. What stays is a 19-function C ABI that moves JSON-RPC messages in-process over 364 dispatch routes.

Phase 4: Run an AI agent for code migration in parallel

A session has its own worktree, branch and agent loop; a subagent works in its parent's workspace and reports back into the parent's context. GitHub's session.ts port (~30,000 TS lines) read for 56 minutes and 122 tool calls before writing a file, then fanned out to 15 child sessions and five subagents over 25 hours. Agents explored about ten times more than they edited, so it pays to understand how agents search and rank files in a large codebase.

Serialise builds through one coordinator session (GitHub used one as a build scheduler for eight sessions) and write tiebreaker rules early: one session, refused four times, merged another's worktree changes anyway.

Phase 5: Layered review, with agent diff-compare, CI and humans on architecture

GitHub's published rust-rebase-review prompt starts one subagent per model (Opus 5, GPT-5.6 Sol and Grok 4.6). Each compares old TypeScript with new Rust line by line and confirms no E2E test was removed or edited. Bots ran on every commit; people covered architecture, API contracts and risk, read high-risk areas themselves and made the merge call. Agent merge in the Copilot app handled every porting PR. The CLI equivalent (untested here; needs an authenticated Copilot CLI; syntax from GitHub Docs):

/pr auto
/pr automerge

Watch automated merges. One agent removed an SDK method, then added the schema-break-ok waiver label so CI would pass; when Toub questioned it, the method was back within 21 seconds. Require a human for any escape-hatch label, snapshot update or baseline change. More on that in guardrails for AI agents in production.

Phase 6: Ship incrementally through pre-releases

Main averaged about 1.3 releases a day, 100 of the 135 being pre-releases. In a seven-day npm sample GitHub cites, pre-releases were just 10.5% of downloads, so new builds reached early adopters first.

Phase 7: Translate first, redesign second

GitHub concedes that much of the Rust still runs TypeScript-shaped algorithms. That is the point: a faithful translation can be checked against the oracle, a redesign cannot. Optimise after the TypeScript is gone.

Minimal TypeScript-Rust FFI: a napi-rs v3 shim replacing one function

Most napi-rs examples online target v2. Version 3 reworked ThreadsafeFunction, deprecated JsFunction and needs Rust 1.88 or newer. The example shows the Phase 3 pattern: a cheap sync export, a real-work export that follows GitHub's rule (async, with blocking work on spawn_blocking), and a Rust-to-TypeScript callback. Illustrative and untested: checked against the napi-rs v3 docs and docs.rs signatures for napi 3.14.1, not compiled.

[package]
name = "port-shim"
version = "0.1.0"
edition = "2021"

[lib]
crate-type = ["cdylib"]

[dependencies]
napi = { version = "3", features = ["async"] }
napi-derive = "3"
tokio = { version = "1", features = ["rt"] }

[build-dependencies]
napi-build = "2"

Gotcha: the v2-to-v3 migration guide says napi-build = "3", but crates.io has no 3.x (latest 2.6.0) and the official template uses 2.x, so keep "2". build.rs is one call:

fn main() {
  napi_build::setup();
}
use napi::bindgen_prelude::*;
use napi::threadsafe_function::ThreadsafeFunction;
use napi_derive::napi;

// Before (TypeScript): export function countLines(text: string): number
// Cheap work can stay a plain sync export (runs on the JS thread).
#[napi]
pub fn count_lines(text: String) -> u32 {
  text.lines().count() as u32
}

// Real work: async export + spawn_blocking so Node's main thread is never blocked
// (GitHub's standing rule after the /chronicle reindex regression).
#[napi]
pub async fn reindex(paths: Vec<String>) -> Result<u32> {
  tokio::task::spawn_blocking(move || {
    paths.iter().filter(|p| std::fs::metadata(p).is_ok()).count() as u32
  })
  .await
  .map_err(|e| Error::from_reason(e.to_string()))
}

// Rust -> still-TypeScript callback (e.g. a permission prompt), v3 style:
// CalleeHandled = false, so call_async takes the value directly and returns Result<Return>.
#[napi]
pub async fn run_guarded(
  cmd: String,
  ask: ThreadsafeFunction<String, bool, String, Status, false>,
) -> Result<bool> {
  let allowed = ask.call_async(cmd).await?;
  Ok(allowed)
}

In the same PR, the TypeScript module becomes a one-line shim and the old body is deleted:

// src/lines.ts (old body deleted in the same PR)
import { countLines as countLinesNative } from "../native/index.js";
export const countLines = (text: string): number => countLinesNative(text);

Build with npx napi build --platform --release. The generated types should be roughly reindex(paths: Array<string>): Promise<number> (inferred from the docs, not generated). Per the napi-rs async docs, dropping the JavaScript Promise does not cancel the Rust future, and a strong ThreadsafeFunction keeps the event loop alive.

Convert TypeScript to Rust without these regressions

GitHub sorted its regressions into families. The common ones, with a faithful Rust form for each:

TypeScript behaviourNaive RustWhat broke at GitHubFaithful port
x || "Unknown error" replaces "".unwrap_or("Unknown error") keeps ""Empty strings passed through instead of the fallback.filter(|s| !s.is_empty()).unwrap_or("Unknown error") (untested)
One number type, so 42.0 serialises as 42f64 for an IDEmitted 42.0, which Go/C# SDKs couldn't unmarshal into int64Choose integer or float explicitly for every JSON number
Fractional timingsi64 for timeToFirstTokenMs5446.712845 made sessions unreadable and unresumableSame rule, in the other direction
Ambient time zone and envAssumed valuesAn undefined time zone broke model-list loading; an env read moved after an awaitMake every ambient input an explicit parameter
Node patched spawning to hide consolesCommand::newConsole windows flashed on WindowsSet CREATE_NO_WINDOW (below)
Async by defaultSync napi export doing I/O/chronicle reindex froze the UI for nearly a minuteAsync export plus spawn_blocking
GC owns object lifetimesOpaque native handlesGitHub's biggest family: handles outliving their objects; a hook torn down mid-request orphaning a tool_use blockPort paired state together, never half a pair

The JavaScript side of the first rows, tested on Windows 11 with Node v24.15.0:

// pitfalls.mjs
const err = "";
console.log("|| on empty string:", JSON.stringify(err || "Unknown error"));
console.log("?? on empty string:", JSON.stringify(err ?? "Unknown error"));
console.log("JSON of 42.0:", JSON.stringify({ repoId: 42.0 }));
console.log("Number.isInteger(5446.712845):", Number.isInteger(5446.712845));
console.log("timeZone:", Intl.DateTimeFormat().resolvedOptions().timeZone);
console.log("toLocaleDateString:", new Date(Date.UTC(2026, 7, 21, 23, 30)).toLocaleDateString());
|| on empty string: "Unknown error"
?? on empty string: ""
JSON of 42.0: {"repoId":42}
Number.isInteger(5446.712845): false
timeZone: Asia/Calcutta
toLocaleDateString: 22/8/2026

The last line prints 22 August for a 21 August UTC instant because the machine is on IST; a port that hard-codes UTC changes that output. The Windows fix below is untested (no Rust toolchain was available), checked against the std CommandExt docs and Microsoft's flag value:

use std::process::Command;
#[cfg(windows)]
use std::os::windows::process::CommandExt;

#[cfg(windows)]
const CREATE_NO_WINDOW: u32 = 0x0800_0000; // learn.microsoft.com process-creation-flags

fn git_status() -> std::io::Result<std::process::Output> {
    let mut cmd = Command::new("git");
    cmd.arg("status");
    #[cfg(windows)]
    cmd.creation_flags(CREATE_NO_WINDOW);
    cmd.output()
}

GitHub's principle is that a failure seen twice should become a standing instruction. Based on its regressions, these are the rules I would give an agent on day one:

  • Never edit, delete or weaken E2E tests, snapshots or compatibility baselines without explicit consent.
  • Never apply waiver or escape-hatch labels to get past CI.
  • Declare integer or float explicitly for every JSON number field.
  • Translate || by JavaScript semantics: empty string and zero count as false.
  • Any napi export that does real work must be async and use spawn_blocking for blocking work.
  • Always set CREATE_NO_WINDOW when spawning child processes on Windows.
  • Do not touch another session's branch or worktree. Ask the human to decide.

Why "it compiles" was not enough

Of 8,678 coded rustc errors, GitHub counts 37% unresolved names or imports (mostly E0425), 22% missing methods or fields, 14% type mismatches and 11% unmet trait bounds. Borrowing, ownership and lifetimes together were 1.7%. Of 4,478 cargo check runs, 87.1% passed clean. The audit found 158 unsafe blocks in 36 files, all at C ABI, OS or SQLite boundaries, and none of the known regressions involved them. Rust delivered memory safety; the regressions were plausible-looking code doing the wrong thing. Toub rates the compiles-means-correct idea as "useful only as a joke". Users don't appear to have noticed: the share of quality-labelled issues moved from 22.9% to 23.7% on copilot-cli and from 36.2% to 32.3% on copilot-sdk (January-April vs May-August).

GitHub Copilot Rust rewrite cost, and how to estimate yours

GitHub reports ~136.3B tokens: ~130.6B cached input reads, ~4.2B cache writes, ~900M fresh input and ~600M output, billed at ~$120,000. The post does not say how it was billed and gives no per-model split. To see what the cache was worth, I priced the mix at single-model list prices (retrieved 2026-10-06, USD per 1M tokens, 5-minute cache writes):

mix = {"cache_read": 130.6e9, "cache_write": 4.2e9, "fresh_input": 0.9e9, "output": 0.6e9}  # GitHub's post
prices = {  # USD per 1M: base_in, cache_write(5m), cache_read, out  (official pages, 2026-10-06)
  "Claude Opus 4.8 (Anthropic)": (5.00, 6.25, 0.50, 25.00),
  "GPT-5.6 Sol short ctx (OpenAI)": (4.00, 5.00, 0.40, 20.00),
  "Claude Haiku 4.5 (Anthropic)": (1.00, 1.25, 0.10, 5.00),
}
def cost(m, p, cached=True):
    base, cw, cr, out = p
    if cached: inp = m["cache_read"]*cr + m["cache_write"]*cw + m["fresh_input"]*base
    else:      inp = (m["cache_read"]+m["cache_write"]+m["fresh_input"])*base
    return (inp + m["output"]*out)/1e6

The full script (the code above plus print lines) on Python 3.12.10 gave:

Single-model list priceWith cachingNo cachingRatio
Claude Opus 4.8$111,050$693,5006.2x
GPT-5.6 Sol (short context)$88,840$554,8006.2x
Claude Haiku 4.5$22,210$138,7006.2x

The usual "10x without caching" estimate is too high. Cache reads cost 0.1x base, but writes cost 1.25x and output is unaffected. At Opus 4.8 prices the input side goes from $96,050 to $678,500 (about 7x) and the total rises about 6.2x. These are hypothetical bills; GitHub used a mix of models.

To scale down: GitHub's mix is ~317k tokens per TypeScript line ported and ~$0.88 per million tokens blended, so a 20,000-line port comes to ~6.3B tokens and ~$5,581 by linear scaling. Treat that as a floor; it assumes GitHub's cache hit rate. Anthropic's US-only inference_geo adds 1.1x, and Claude 4.7+ tokenizers produce about 30% more tokens for the same text.

TypeScript vs Rust performance, read carefully

GitHub ran the GitHub Copilot Rust migration benchmarks through the C# SDK against a deterministic chat-completion server on localhost, so model inference and network latency are excluded. It also warns that other changes shipped in the same window, so the figures compare delivered systems, not languages.

What was timed (GitHub's benchmark)Node baseline, May 12Rust, separate process, Aug 21Rust, same process, Aug 21
Start a client and session, then send one message5.25 s1.33 s (4.0x)292 ms (18.0x)
Reopen a session holding 32 turns of history5.64 s1.52 s (3.7x)264 ms (21.4x)
10 clients started and shut down in parallel12.34 s4.18 s (3.0x)742 ms (16.6x)
Total for 1,000 single-message sessions132.52 s22.53 s (5.9x)20.93 s (6.3x)

Ten-client memory: 1,383 MB before, 247 MB out-of-process, 126 MB in-process. Some of the post's figures disagree, unexplained:

  • The table gives 292 ms for the in-process single-message run; the closing section says "about 55 milliseconds".
  • The 1,000-session run takes 20.93 s in-process and 22.53 s out-of-process, yet the prose reports 120.0 and 57.45 sessions per second. Only the Node baseline agrees (132.52 s, 7.55/s).
  • One section sizes the event log at 260 MB, another at 250 MB.

They may be different runs. Quote one, name its source, and don't derive new figures by combining them. GitHub itself stresses the runtime is "not universally '15.9x faster'".

What this means for Copilot SDK users: in-process hosting

In-process hosting is available but experimental. It first shipped in Copilot SDK v1.0.7 (2026-07-16), and the SDK docs still mark it experimental in every SDK as of v1.0.16. Copied from those docs, untested here (needs Copilot auth and the native runtime bundle):

import { CopilotClient, RuntimeConnection } from "@github/copilot-sdk";

const client = new CopilotClient({
  connection: RuntimeConnection.forInProcess(),
});

await client.start();
#pragma warning disable GHCP001
var client = new CopilotClient(new CopilotClientOptions
{
    Connection = RuntimeConnection.ForInProcess(),
});
await client.StartAsync();

If you call the Copilot SDK from Rust, in-process hosting needs the bundled-in-process Cargo feature and ClientOptions::default().with_transport(Transport::InProcess), (the docs' form, not the blog's ClientOptions::new()). Go needs -tags copilot_inprocess. Limits: one runtime version per process, no per-client working directory, and env and telemetry options are rejected. JSON-RPC still runs on every call, so the process hop goes but the serialisation cost stays.

Pre-flight checklist for your own AI-assisted migration

Before you copy the GitHub Copilot Rust migration, have:

  • A goal that rules out a partial port. GitHub's first wording was read as hot paths only until it was restated as an all-Rust native binary.
  • E2E coverage of user-facing behaviour, protected from agent edits.
  • Pilot PRs for toolchain, CI, lint, FFI, packaging and agent instructions.
  • A leaves-first order, starting with well-tested pure functions.
  • Every port PR deletes the old code.
  • One worktree per session, a build mutex, and written tiebreaker rules.
  • Agent line-by-line old-vs-new review; humans on architecture and risk.
  • A pre-release channel and a rollback path.
  • A token budget that states the cache hit rate it assumes.
  • A rule that any failure seen twice becomes an instruction, a skill or a test.

For the output, use a 7-point checklist for vetting AI-generated code.

FAQ

Was the GitHub Copilot Rust migration really done by one engineer?

Mostly, with help. Toub drove the porting PRs; the post credits other engineers with the interop layer, SDK FFI layers, packaging, crate splitting, CI caching and reviews.

How much did the Copilot Rust rewrite cost?

About $120,000 for ~136.3B tokens plus roughly three developer-weeks, according to GitHub, with a 96.22% cache hit rate. Without caching, the same mix at single-model list prices costs about 6.2x more.

Why Rust instead of C# or Go?

The post doesn't compare them. Its requirements were an embeddable C ABI for six SDKs, low overhead, predictable resource use and supply-chain posture; the goal was to leave Node.js and V8.

If the Rust code compiles, is the port correct?

No. Every regression GitHub found compiled. The compiler caught ordinary typing mistakes but missed changes in contract, such as number types, || semantics, ambient time zone and env, Windows flags and a blocked main thread.

Can I run the Copilot runtime in-process today?

Yes, experimentally: it shipped in SDK v1.0.7 (2026-07-16) and the docs still mark it experimental, with the limits above.

Last updated: 2026-10-06; tested on Node v24.15.0 (pitfall demo) and Python 3.12.10 (cost estimator), Windows 11. The napi-rs v3 (napi 3.14.1), Rust std and Copilot SDK v1.0.16 snippets are untested.

Last updated: October 06, 2026
an "open and free" initiative. Powered by Blogger.