Autopilot Pulse · #02 · 2026-08-17
The Bill: what it costs to run a company on AI

TL;DR
- →Polsia's model bill went $500K → $1.2M/month → ~$100K. This week the receipts got audited — both sides below.
- →Cofounder raised $87M for a "manager agent" that wields 80+ tools — and provisions a full GitHub + Vercel + Supabase stack for every customer.
- →Egbe runs a zero-employee company mostly on open-weights GLM — usage up ~8× in a month.
- →The experiments got real: agents ran actual shops and real-money deals; 48 AIs have now run a (simulated) food truck — 16 went bankrupt.
- →Stack alert: Stripe is reportedly buying OpenRouter for $7B+ — the "neutral" model gateway may soon live inside a payments giant.
Edition #1 asked whether businesses can run themselves. This one opens the engine rooms: what the agents actually run on, what it costs — and a story that broke this week and stress-tested the index's own rules.
1 — Polsia: the bill, and the bill come due

The most-watched company on the index is a one-founder, zero-employee "AI operating system" that builds and runs online businesses end-to-end — and it gave the category its defining unit-economics story: a monthly Anthropic bill that climbed from $500K to $1.2M, then dropped to ~$100K after routing routine workloads to open-source models on rented GPUs (Sciforium for inference, Sapiom routing every task to the cheapest capable model). The detail that makes it credible: Sapiom just raised $35M — with Anthropic among the backers. The model vendor is investing in the company that cuts its customers' model bills.
Then, this week, the other side of the ledger. An independent investigation alleged some Polsia-launched businesses are "hollow shells" — polished landing pages with no working product behind them. Trustpilot sits at 1.9/5 across 74 reviews, and an independent breakdown pegs subscription revenue at $6.96M of the $10.22M ARR headline. The counterpoints are real too: reviewers have documented genuinely fast autonomous output, and founder Ben Cera discloses numbers most founders hide — roughly 50% month-one churn, and that about 10% of customer companies have made at least a dollar. Our call: Polsia stays on the index, flagged DISPUTED, with both sides linked. Vibes don't chart — in either direction.
2 — Cofounder: the $87M manager agent
Cofounder's architecture is the clearest picture yet of what an "org chart as software" looks like: a GPT-5 + Claude "superoptimizer" manager agent holding 80+ tools in context, directing worker agents underneath. The detail CTOs will appreciate: it auto-provisions a complete GitHub + Vercel + Supabase stack for every customer company — infrastructure-as-onboarding, confirmed in Supabase's own customer story. $87M raised to bet that the org chart is now a routing diagram.
3 — Egbe: the contrarian model bet
Nikolay Vyahhi's zero-employee company runs the majority of its build workload on GLM — open-weights models from Z.ai — rather than a frontier lab API. Requests and tokens are up roughly 8× in a month (self-reported). And Z.ai just shipped GLM-5.3 this week, so the bet keeps compounding. The lesson: frontier-lab loyalty is now optional. The tradeoff: concentration risk in a single fast-moving vendor — which is true of every stack choice on this list.
4 — Atoms & Nanocorp: one prompt → one company
Atoms (built by the creator of MetaGPT) doesn't stop at building your product — its agents market and sell it too, which is the part most builders underestimate. Nanocorp's pitch is the purest distillation of the category: "One prompt. One company. Zero code." — with the honest footnote that its ARR claims are disputed and labeled as such on the index. Building is now table stakes; distribution is the moat.
5 — The experiments got real tills

FoodTruckBench deserves its own spotlight this week — not just as a benchmark, but as a solo-founder story. Nicholas S. built the whole thing alone — simulation engine, 12-factor demand model, the playable game, everything — and it has quietly become the most useful reality check in the category: 48 frontier models, each handed $2,000 and 34 tools to run an Austin food truck for 30 days. 16 went bankrupt. Claude Opus 5 currently leads at $75,264 net worth (+3,663% ROI); the best human run is still ahead at roughly $101K.
And because this is The Bill edition, the benchmark's newest lens fits perfectly: net worth per API dollar. GPT-5.6 Luna returned $144,501 of simulated net worth for every $1 of API spend; Claude Opus 5 — the outright winner — returned $2,800. Winning and being worth the bill are different metrics.
In December, Anthropic ran the experiment quietly and published it in April as Project Deal: 69 employees handed their buying and selling to custom Claude agents in a real Slack marketplace — 500+ items, real money, zero human intervention — and the agents closed 186 deals worth $4,000+. The finding that should keep every builder up at night isn't that it worked; it's that model quality was negotiation power — while the humans with weaker agents rated their outcomes exactly as fair as everyone else. The disadvantage was real, measurable, and completely invisible. Your stack choice is now your negotiating position.
And Andon Labs keeps going further — handing agents actual shops and a café to run: inventory, pricing, customers. Free due-diligence for your own agent-business idea. Run the experiment before you run the company.
Stack alert — Stripe ⇄ OpenRouter, $7B+
Bloomberg reports Stripe has finalized a deal to buy OpenRouter — the gateway many builders use as their "one API over 400+ models" neutrality layer — for more than $7B, barely three months after its $1.3B Series B. Stripe declined to comment, so treat it as reported, not confirmed. But if it closes, the neutral routing layer will live inside a payments giant — worth knowing if OpenRouter is your fallback plan.