arXiv:2610.04862October 2026

GitSwarm: Decentralized Compounding Inference

Many agents, one shared repository, and work that keeps building on itself.

Vedant Shah1,2,3, Ankur Samanta1,4, Paras Dahal1, Mikhail Plekhanov1, Carole-Jean Wu1, Scott Yih1, Remi Munos1, Rob Fergus1, Jakob Foerster1, Ruslan Salakhutdinov5,∗, Sanjeev Arora6,∗, Jason Weston1, Aaron Courville2,3, Anirudh Goyal1
1Meta Superintelligence Labs   2Mila – Québec AI Institute   3Université de Montréal   4Columbia University   5Carnegie Mellon University   6Princeton University   ∗Work done at Meta
worker episode (spawns, contributes, exits) draft repair verification refuted integration semantic inheritance final answer and its ancestry

Difficult long-horizon problems are rarely solved in one attempt. Partial solutions, experiments and even failed ideas often turn out to be useful later, but most ways of spending inference compute throw them away when an attempt ends. We call the alternative compounding inference: organizing inference-time computation so that intermediate work persists and can be inspected, extended, combined or challenged by later computation.

GitSwarm puts this idea into practice. A pool of identical agents works asynchronously on a shared, branchable Git repository. No one hands out tasks: each agent reads what exists, decides for itself what would help most, and publishes its work as a commit that records which earlier commits it built on.

Why compounding inference?

Most ways of giving a model more compute at inference time treat attempts as disposable. You can sample many answers and keep the best one, or let one agent reason for longer inside a single context. Either way, when an attempt ends, almost everything it worked out ends with it.

On long-horizon problems, though, the value of intermediate work is often unclear when it is produced. A failed experiment can point to a better direction. A partial proof can contain the lemma someone needs ten attempts later. Two half-working programs can be worth more combined than either one alone. If this work disappears, later attempts have to rediscover it. If it is kept, later attempts can build on it.

Figure 1Three ways to spend more inference compute

Independent attempts

Many samples in parallel. One is picked and the rest are discarded, including any useful pieces they contain.

One long trajectory

A single agent keeps going. Earlier work survives only as long as it fits, and stays useful, in one context.

Compounding (GitSwarm)

Every attempt persists as a commit. Later work extends, verifies, combines or refutes it, across branches.

How GitSwarm works

GitSwarm has three parts: a shared Git repository that acts as the swarm's memory, a pool of identical workers that each choose their own next move, and a light execution harness that schedules workers, isolates them, and validates what they publish. There is no planner, no assigned roles and no privileged final judge.

The repository is the memory

Each task starts as a repository containing the task specification. Every piece of work a worker publishes becomes one atomic, immutable commit: a draft solution, a test, an experiment and its result, a repair, or a negative finding. Branches let competing approaches grow side by side, so the swarm never has to commit to one direction too early.

Each commit records two kinds of history. Its Git parent is the state the worker started from. Its informed_by= field lists the other commits the worker relied on, including commits on other branches. We call this second link semantic inheritance: it lets a worker build on one branch while borrowing a discovery from another, without merging. Following both links backwards from any commit gives its ancestry: all the earlier work it builds on.

Figure 2A small GitSwarm repository. Click any commit to trace its ancestry.
Git parent semantic inheritance ancestry of the selected commit

Identical workers, no manager

Every worker runs the same model with the same instructions. Nobody is told to be a solver, a verifier or an integrator. Each worker looks at the repository and decides what would add the most value right now. Specialization, when it appears, comes from the state of the repository: early workers write first drafts because nothing exists yet, and later ones repair, test, combine and compare. We return to this below.

Asynchronous execution

The harness keeps up to C workers running at once (five in our benchmark runs) and launches a replacement whenever a slot frees up, until a budget of B worker episodes is spent. Workers never wait for each other, so a commit is visible to every worker that starts after it is published.

Figure 3Twenty worker episodes on five slots, with the same episode durations in both modes

Asynchronous

a new worker starts as soon as a slot frees up

Synchronized waves

every worker waits for the slowest member of its wave (dotted)

Each bar is one worker episode; dots mark when its commit becomes visible to later workers. On nine ProgramBench development tasks, asynchronous runs finished in 38.0 minutes on average versus 49.3 for synchronized waves, with similar scores.

Seeing the live swarm

Work that hasn't been committed yet is invisible in Git, so workers get two read-only views of the live run. Every worker receives a snapshot of the dashboard when it spawns, and can fetch an up-to-date one at any time by running gs_dashboard. It shows the remaining budget, how many workers are in flight, and the nomination tally, which is explicitly labelled as not being evidence of correctness; in research runs it shows shared GPU usage and the best validation score instead. To see what other workers are doing right now, a worker runs gs_inflight: it lists the workers in flight, and gs_inflight inspect shows one worker's public log of what it announced it would do, its progress updates, and its uncommitted changes. Each worker writes to this log itself, so others can avoid duplicating its work. Neither view assigns work.

Figure 4What a worker sees: dashboard and a peer's in-flight log (abridged from real traces)
gs_dashboardshared state
$ gs_dashboard=== Git-Swarm dashboard ===Total worker budget:          200Remaining worker budget:      101Workers currently in flight:  5Nomination workers in flight: 0Total valid nominations made: 7Nominated commits and vote counts  (selection state; not correctness evidence):  7ce10344e6db: 2 nominations  1e5a9adb44da: 1 nomination  ...--- research tasks show instead ---GPUs: 6/8 allocated (75.0%; 2 available)Jobs: 2 running, 0 starting, 11 queuedWARNING: Resource usage below targetBest protected dev: val_loss=3.9004  commit 1d9edbb
gs_inflight inspectpeer activity
$ gs_inflight inspect w0101=== In-flight worker detail ===peer w0101  head=9ac922c47c87  branch=agent/w0101/word-boundary-regex-escapes 06:06:07 announce  targets=9ac922c  Gap: regex candidates diverge on doubled  word-boundary escapes | Deliverable: 9ac922c-  based normalizer | Success: differential  checks | Abort: if this exact refinement is  already published06:08:39 update  Implemented a localized escape  normalizer; narrowing probes ...06:09:16 update  py_compile and compile.sh pass.working tree:  M ambr.py  (+36 lines)

Anatomy of one episode

A worker episode is one agent run from spawn to exit, inside its own isolated Git worktree. The prompt asks it to understand the existing work before acting, to scope one bounded contribution, and to say up front what would make it give up.

Figure 5One worker's episode, step by step (a constructed example; commit IDs and outputs are illustrative)
    worker w0101isolated worktree

    Three ways to end an episode

    Every episode ends with exactly one of three actions. Validation checks only that a commit follows the task's publication rules (required files, allowed paths), not that it is correct.

    Contribute

    Publish one validated, atomic commit with new work: a partial solution, an experiment, a test, a repair, or a synthesis of several branches.

    uses one unit of budget

    Nominate

    Cast the worker's vote for an existing commit to be the final output, by publishing a separate nomination record with a written rationale. The commit itself is not changed.

    uses one unit of budget

    Abstain

    Exit without publishing when nothing useful is left to add, or when equivalent work is already done or underway.

    replaced, does not use budget

    Decentralized readout

    Eventually the swarm has to decide which commit is its answer. Rather than handing this to a privileged judge, GitSwarm leaves it to the workers too. Nominations are stored outside the commit graph, so workers can evaluate candidates while others keep building new ones, without changing anyone's ancestry. When the budget runs out, the candidate with the most nominations is submitted, with ties broken deterministically. For the research tasks, which have a validation score, the best-scoring model is selected instead.

    Figure 6Nominations accumulate while the run continues; the most-nominated candidate is submitted

    Experimental setup

    We test GitSwarm in two settings: long-horizon problems with a well-defined, checkable answer, and open-ended research, where the workers must also decide which experiments are worth running.

    Long-horizon problem solving

    IMOProofBench-Advanced
    30 olympiad-level proof problems; we count fully correct proofs.
    ProgramBench
    Rebuild a program from its documentation and an execute-only reference binary, scored on hidden tests. 50-task stratified subset.
    Workers
    Codex (GPT-5.5-high), base GPT-5.5, Gemini 3.1 Pro. Five concurrent workers. Budgets of 100–500 episodes (programs) and 20–160 (proofs).
    Baseline
    Forced single agents with the same model, prompted to continue every time they stop, compared at similar token usage.

    Sustained research (AutoResearch)

    Tasks
    Improve three models under fixed rules and a fixed metric: the Residual Matrix Transformer and Looped Transformer (architecture changes only), and NanoChat (architecture, optimizer and training code).
    Workers
    12–20 workers sharing a fixed pool of GPUs through a SLURM scheduler, for 24–72 hours: Codex workers with GPT-5.6 Sol (RMT) or GPT-5.5 (Looped Transformer), and GPT-6 Astra or GPT-5.6 Sol workers on NanoChat.
    Selection
    Nomination is turned off; the best validation score is selected and, where available, checked on a held-out set.
    Baselines
    Single agents with the same model, GPUs and time, plus an evolutionary search baseline on NanoChat.

    Results on ProgramBench and IMOProofBench

    ProgramBench. With Codex workers, GitSwarm improves from 73.1% at 100 worker episodes to 79.4% at 500. The forced single Codex agent stays between 63.1% and 65.1% however long it is pushed. With base GPT-5.5 workers, GitSwarm improves from 64.6% to 71.2%. Gains shrink at larger budgets, and input tokens grow because each worker reads a larger repository.

    IMOProofBench-Advanced. With Gemini 3.1 Pro workers, GitSwarm goes from 23.3 to 29.0 fully correct proofs as the budget grows from 20 to 160 workers (with a local dip at 80). The best forced baseline reaches 26.5. With GPT-5.5 workers, a single run solves all 30 problems at budgets of 40 and 80. Best-of-N sampling with an oracle verifier also reaches this ceiling, so solving all 30 is not, by itself, an advantage unique to GitSwarm.

    Figure 7Performance against inference compute (input tokens)

    ProgramBench

    Mean test score, 50 tasks

    IMOProofBench-Advanced · Gemini 3.1 Pro

    Fully correct proofs, of 30

    Each GitSwarm point is one run at the labelled worker budget B. Baseline points are the natural stopping checkpoints of a single agent that is repeatedly told to continue. Tokens are summed over all requests per task (left) or per problem (right).

    Does the work actually compound?

    A higher score could simply come from running more agents. To check whether later work really builds on earlier work, we looked inside the repositories from the benchmark runs: what workers cite, which files they open and run, and what ends up in the selected solution.

    99.9% / 96.2%of eligible episodes cite earlier work (ProgramBench / proofs)
    94.7% / 44.4%of contributions are later built on by other workers
    82–93% / 21–33%of a run's commits lie in the selected solution's ancestry
    ~80% / 583 of 759voluntary artifacts actively reused / probe bundles opened by later workers

    The selected solution builds on earlier work

    In both benchmarks, the selected solution builds on work that earlier workers published, but how much of a run it draws on differs. On ProgramBench its ancestry covers 82–93% of the commits in a run, and only about a quarter of that comes in through Git parents; most of the rest arrives through semantic inheritance across branches. On IMOProofBench the selected proofs reach 21–33% of the graph: good proofs tend to appear early, and later workers mostly check and compare candidates rather than extend them. That checking still uses earlier work, and it decides which proof is selected, even though it is not part of the proof's ancestry.

    Figure 8Every commit in one ProgramBench run (FFmpeg, B = 100), over normalized runtime
    in the selected winner's ancestry outside it
    Figure 9Share of the contribution graph in the selected solution's ancestry, by worker budget
    via Git parents added via semantic inheritance outside the ancestry

    Workers autonomously adopt roles

    We labelled every worker episode with its main activity after the run. On ProgramBench, a short wave of workers creates first versions; after that, adding missing features, fixing behavior and combining candidates continue right to the end, and nominations pick up only near the end. The selected program usually first appears at 93–98% of the way through a run. Proof runs shift much earlier: once good candidates exist, more and more workers check and select among them instead of writing new proofs. About one in six ProgramBench workers abstains after finding that equivalent work already exists.

    Figure 10What workers spend their episodes on, over the course of a run

    ProgramBench

    Codex workers, one panel per worker budget

    IMOProofBench-Advanced

    Gemini 3.1 Pro workers, one panel per worker budget

    Share of worker episodes by primary activity against worker-spawn progress. The dark line marks the top of all commit-producing activities.

    Different problems, different workflows

    The same protocol produces quite different collaboration patterns depending on the problem. In the delta run, separate branches improve different parts of the program and are combined late into one strong integration that the winner builds on. In the cheat run, workers spend the first third on command-line and configuration behavior, then pivot to output renderers after a burst of renderer experiments, and the renderer line becomes the winner.

    Figure 11Commit graphs from two ProgramBench runs, over normalized runtime

    Late integration

    delta, B = 100

    Search and pivot

    cheat, B = 500

    AutoResearch: running a research program

    Benchmarks have fixed answers. Research does not: workers must choose experiments, run them on shared GPUs, interpret noisy results and decide what to try next. We gave GitSwarm three research tasks, built on the Residual Matrix Transformer (Mak & Flanigan, 2025), the Looped Transformer Huginn (Geiping et al., 2025) and NanoChat (Karpathy, 2025; we start from the autoresearch setup). In each one, workers start from a published model and must improve it under fixed rules: architecture changes only for the Residual Matrix Transformer and the Looped Transformer, and architecture, optimizer and training code together for NanoChat. They submit training jobs to a shared GPU queue much like a research group would. Code, results and notes all persist in the repository.

    Each task also illustrates a different way the research compounded.

    Figure 12Best-so-far validation loss, and how the final model came about

    How the final model came about

    lineage of the final model developed on a side branch failed / refuted doubted / inconclusive reuses an earlier failure code or evidence reused

    Research behaviors we did not program

    Reading the research histories turned up behaviors that no prompt asked for explicitly. A few that stood out, with details in Appendix E of the paper:

    Residual Matrix Transformer

    A failure, revived 49.5 hours later

    Letting each token see the previous token's memory state failed. Two days later, another worker cited that failure as evidence that longer look-backs were still unexplored, built a mixer over several earlier tokens, and it entered the submitted model.

    Looped Transformer

    Auditing their own yardstick

    After a 128-example check nearly killed the winning idea, a worker measured how often small evaluations misrank models: 5 of 12 pairs. The swarm then adopted a 128 → 512 → 4,873 example evaluation funnel, and shared proxy setups were reused by 257 commits.

    All three tasks

    Negative results are not thrown away

    Workers publish experiments that didn't work, labelled as such, and others cite them: 74% of negative architectural results in the Looped Transformer run and 70% in NanoChat were cited by later contributions, often to rule out a direction or build a control.

    Residual Matrix Transformer

    Breadth that single agents lacked

    The swarm explored 19 families of ideas and kept trying new ones throughout. The single Codex agent opened no new line of ideas after 26.5 hours, and the direct GPT agent spent 68% of its GPU time on a single family.

    NanoChat · Looped Transformer

    Checking each other's work

    A worker re-ran an identical configuration and turned an apparent win into a near-tie; the next worker cited that and chose a different change. In the Looped run, four workers independently ran parameter-matched controls, and a later worker overturned a small-sample win.

    Residual Matrix Transformer

    Compounding across whole runs

    Seeded with the best models from three earlier, independent runs (about 220 research-hours), a new swarm combined them, adding one run's head router and another's deeper stack, and lowered the inherited best loss from 2.909 to 2.898.

    Takeaways

    • Inference compute can accumulate. With persistent, branchable memory and explicit cross-branch links, independent agent episodes add up to a body of work that later episodes reuse, instead of a pile of disposable attempts.
    • Compounding looks different on different problems. Program reconstruction folds most of the run into the final program. Proof solving uses earlier work mainly to verify and select. Research mixes both, and adds the recovery of ideas that first looked like failures.
    • Coordination can stay decentralized. Without planners or roles, workers divided the work, avoided some duplication, and selected answers by themselves.

    Limitations. Our dependency and trace analyses show that workers read, reuse and build on earlier work, but they do not isolate how much each reused piece contributes to performance. ProgramBench results use a 50-task subset; the perfect proof score comes from a single run; and the research comparisons are not matched on every axis (for example, NanoChat starting points and GPU allocations differ). Persistent collaboration also has costs: input tokens grow as the repository grows, and redundant episodes still use compute.

    A natural next step is to train agents inside this kind of environment, rewarding contributions whose value only shows up through what other agents later build on them, so that agents learn to advance a shared process of discovery and not only to finish their own episode.

    Citation

    @article{shah2026gitswarm,
      title   = {GitSwarm: Decentralized Compounding Inference},
      author  = {Shah, Vedant and Samanta, Ankur and Dahal, Paras and Plekhanov, Mikhail
                 and Wu, Carole-Jean and Yih, Scott and Munos, Remi and Fergus, Rob
                 and Foerster, Jakob and Salakhutdinov, Ruslan and Arora, Sanjeev
                 and Weston, Jason and Courville, Aaron and Goyal, Anirudh},
      journal = {arXiv preprint arXiv:2610.04862},
      year    = {2026}
    }