Court Rules
Developer Guide• 7 min read

Meet Our Newest Team Member: A Laptop in a Drawer

Agents write most of Court Rules, and they write it faster than anything can check it. This is how our software factory works, and why its checking station lives on a laptop in a drawer.

A closed silver laptop wrapped in a folded blanket on a charcoal desk, tagged CI, with a queue of index cards waiting beside it and a tally counter

Court Rules is a database of the rules that decide whether a court filing gets accepted: each judge's standing orders, each court's local rules, the filing requirements, the holidays. On October 5, 2026 it lists 1,077 courts and holds 146,135 current rules. Lawyers read them on the website. Programs fetch them through an API, and AI assistants ask for them through an MCP server, the standard way an assistant like ChatGPT or Claude calls an outside tool.

Agents write most of the software behind it. A person decides what to build, and agents do the building. Legal experts check the quality of the rules it produces.

A software factory

In 2026 a few teams began describing what they call a software factory: coding agents that plan, write, test and ship software in parallel while people set direction. Dan Shapiro laid out the ladder in January, from autocomplete at level 0 to a "dark factory" at level 5, where nobody reads the code. In February StrongDM published a factory run on two rules: humans don't write the code, and humans don't review it. In September Gergely Orosz described OpenAI's own factory, where load on some internal systems grew about tenfold in six months.

Ours is smaller and younger. In Shapiro's terms it sits at level 4: people write the intent, and agents do the work. Here is how it runs.

  1. Step 1

    Intent

    People and monitors say what is needed: a court to add, a bug to fix, a finding from a review.

  2. Step 2

    Build

    Agents work in parallel, each on its own branch, and claim a shared file before they edit it.

  3. Step 3

    Check

    21 automatic checks guard every change. For bigger ones, a model from a different family reviews the work too.

    This post

  4. Step 4

    Ship

    An agent merges a pull request once its checks pass, and the merge deploys itself to Vercel and Trigger.dev.

  5. Step 5

    Learn

    Alerts, usage and questions we could not answer become the next piece of intent.

What the factory learns starts the next round of intent.

People: legal experts check rule quality and break apart conflicts. Others set direction, approve spending and run changes to the production database.

Our software factory, from intent to production. This post is about the third station.

The speed is the point. We have merged 236 pull requests since July, and 168 of them landed in the last 14 days. On October 4 alone, 58 merged. Many of the bursts are courts: that day, 17 agents worked on 14 county court scrapers at the same time.

September

October

Pull requests merged per day, September 22 to October 5, 2026. October 5 (lighter bar) was still under way.

Writing code stopped being the slow part. Checking it is. Every one of those pull requests has to pass the same gates before it can merge, and OpenAI's account says the same of its own pipeline: every part of build, test and deploy now carries far more load. This post is about our checking station.

What the checks are

A check is a program that fails when something is wrong. We have 21 of them: formatting, lint, types, dead code, duplicated code, and the test suites for each part of the code. A change runs the ones it can affect. Engineers call running them automatically on every proposed change continuous integration, or CI. A proposed change is a pull request, and a pull request can't merge until its checks pass.

Ten agents, one laptop

Our first limit was the laptop. A pre-push hook, a script that runs when an agent tries to push code, ran every check first, and it took 195 seconds. With several agents pushing at once, the checks competed for one machine. On October 4 ten agents ran tests at the same time. The load average, which counts how many programs are waiting for a core, reached 194 on a 14-core machine, and 7 GB of swap filled. No single agent did anything wrong. We had asked one laptop to do ten people's work at once.

So we made the checks cheaper before we moved them anywhere. The suite now runs only the checks a change can affect, skips a tree that has already passed, and runs the quick checks alongside the slow ones. The pre-push hook dropped from 195 seconds to 29. A tree that has already passed takes under a second.

Time to check one change before it is pushed. The bars are drawn to scale.

Why not rent more machines

One of Pulumi's seven rules for an AI-native factory is to run it in the cloud, not on a laptop. We agree about the agents, which run wherever there is room. The checks are the exception, for now. GitHub's hosted runners start each job on a fresh machine, so each job downloads its dependencies and rebuilds its caches from nothing, and GitHub rounds every job up to a full minute. Two of our jobs only decide which other jobs to run and report the final result. They last seven seconds and three seconds. A fresh machine per job suits an occasional pull request. It does not suit 127 of them in a week.

The new hire

We already owned an M1 Max MacBook Pro with ten cores and 32 GB of memory that stays on and logs itself in. We call it the sleeper, because it never does. We registered three GitHub runners on it. A runner is a program that takes a job from GitHub and runs it. Two take the heavy jobs, and one takes the jobs that only last seconds, so a short job never waits behind a build.

A repository variable decides where checks run. Unset, they run on GitHub's machines as before. Set, they run on the sleeper. A script on the sleeper clears the variable when fewer than two runners are online, so checks fall back to GitHub instead of waiting for a machine that is down.

A pull request opens

GitHub reads one repository variable

Variable set, two or more runners online

The sleeper, a Mac that stays on

  • Heavy runner 1builds and test suites
  • Heavy runner 2builds and test suites
  • Light runnerjobs that last seconds

Variable unset, under two runners online

GitHub's own machines

A fresh machine for every job

The result appears on the pull request

Merging opens when every check passes

Where the checks for a pull request run.

Keeping the work on one machine also made it faster. The runners share a package store on local disk, so an install after a lockfile change downloads only the new packages. The workspace stays between runs with its TypeScript and Next.js build caches intact, and a cleanup step removes everything else a previous run left behind. A hosted runner fetches those caches over the network onto a fresh machine every time. Reading a local disk is quicker.

What the laptop did not know about Linux

GitHub's machines run Ubuntu, and so does production. The sleeper runs macOS, and the differences showed up quickly. The default service configuration allows 256 open files per process, which is too few for ESLint and Next.js, so we install the runners from our own template that raises the limit to 65,536 and runs jobs at a lower priority than interactive work. macOS file systems ignore letter case and Linux doesn't, so a wrongly capitalized import compiles on the Mac and fails on Vercel. The type checker is set to report it. The schema tests expect a Postgres container, and the Mac has no Docker, so it runs two Postgres clusters on separate ports, one for each heavy runner.

The easiest thing to miss was a workflow nobody thinks of as CI. The job that deploys our background workers was still pinned to GitHub's machines while everything else followed the variable. We moved it, and now every job follows the same switch.

The fine print

A self-hosted runner runs the code in every pull request on a machine we own, so we only use it on a private repository where our own agents open the pull requests, and pull requests from forks never run. Before we add collaborators or open the repository, the runners need their own user account or a virtual machine. The runbook says so.

The deploy that ships our background workers, 199 tasks in all, now runs on the sleeper in about a minute.