Compound Labs
Get the newsletter
THE STANDUPAnthropic fixes Claude 5.1 error spikesERROR SPIKES, FIX DEPLOYEDCOMPOUND MASTODONSkillWorks records inactive files beside their roster status22 INACTIVE FILES STILL LOADCTXWINDOWContext Window tracks AI lab changes and filters their evidenceFILTERS KEEP EVIDENCE TOGETHERAGENTWIREmicrosoft/agent-governance-toolkit - AI Agent Governance ToolkitPOLICY ENFORCEMENT FOR AI AGENTSCOMPOUND FACEBOOKCompound Portfolio restores its sweep after a missing module stops desk dataSWEEP RESTORED, DESK DATA WRITTENBLOCKDEXBlockDex shows no destination for captured rows with rejected slash names31 ROWS, NO DESTINATIONCONTEXT WINDOWOpenAI adds 1024p image pricing: $0.50 and $0.251024P IMAGE, $0.50 AND $0.25COMPOUND BLUESKYCompound Labs renamed job labels so job-run finds breaker stateBREAKER STATE, NOW FINDABLECTXWINDOWContext Window tracks document changes but failed mobileMOBILE WORKFLOW, ONE COLUMNAGENTWIREMineDojo Voyager: Open-Ended Embodied AgentOPEN-ENDED EMBODIED AGENTFOUNDER BLUESKYdeploy.sh now fails false Vercel production landingsFALSE LANDINGS NOW FAILCOMPOUNDlast known fixes redirect tracking after migrationSTUDIO SOURCES NOW SAY COMPOUNDCONTEXT WINDOWOpenAI adds sora-2-pro pricing720P VIDEO, $0.15, $0.30FOUNDER LINKEDINdeploy.sh now stops failed promotes from marking production landedFAILED PROMOTES NO LONGER LANDUSINGITUPUsing It Up records a receipt fixing a wobbling wooden tableONE RECEIPT, TABLE STAYED LEVELCOMPOUNDpnpm v12.4.2POSIX BIN SHIMS REPLACEDFOUNDER PEERLISTCompound Labs' deploy runner rejects failed promotesFAILED PROMOTES NO LONGER LANDTRUSTDESKTrustDesk fixes host migration redirects while keeping /api/ outOLD HOST, 308; API STAYS LIVE
Independent product R&D labFounded and run by Isaiah Kim, @kyisaiah47Newest commit Sep 16, 2026, AgentwireNewest writing Sep 16, 2026Site changelog Sep 16, 2026
ToolproofSep 4, 20263 min read

675 claims on our boards said verified and 38 of them could re-check themselves, so Toolproof now prints both dates

A date next to a number is the only reason anybody believes the number. On our boards that date said a person had gone to the source, read it, and written down what it said, which is a real assertion and a good one. It's also an assertion about a day in the past, and nothing on the page told a reader whether the source still said the same thing.

We swept every claim register on the estate on 2026-08-15 and counted 705 claims, 675 of them stamped verified. 38 carried the one thing that lets a machine check a figure again: the literal string that has to still be sitting on the source page. All 38 were on a single index. The other 667 named an address and a day somebody read it, and nothing else.

What Toolproof is

Toolproof is the masthead over the studio's measurement indexes, SkillWorks, StillShipping, ToolDrift and BlockDex among them. It carries the four things none of them had on their own: one published methodology, one open dataset, one report, one byline. Nothing on it measures anything itself. Its own register carries two claims, both about which fields Claude Code's documentation marks required for a subagent and optional for a skill.

Toolproof feature card showing verified and recheck dates with statistics on 705 board claims, Three concrete outcomes: 675 verified stamps, 38 source literal matches, 667 address-only claims. Demonstrates scope and coverage of the feature across real data.
Toolproof feature card showing verified and recheck dates with statistics on 705 board claims, Three concrete outcomes: 675 verified stamps, 38 source literal matches, 667 address-only claims. Demonstrates scope and coverage of the feature across real data.

What an HTTP 200 actually proves

The warning was already written down beside the comparison tables, months before we went near the registers: a freshness job that re-fetches a page and checks it returns 200 proves the company still has a pricing page, and proves nothing whatsoever about the number in our table. Reachability and truth are two different measurements, and a board that shows one while a reader reads it as the other is worse than a board with no date at all.

So each checkable claim names a string. When that string is gone, the figure we print is very likely wrong now.

What the re-read comes back withWhat it proves about our figureHow the board says it
The string is still on the pageThe figure is still the one on the sourceThe machine date moves
The page loaded with real text, the string is goneThe figure is very likely wrong nowFails, never green
The page came back with almost no readable textNothing, neither confirmed nor refutedReported on its own line, never green
No string was ever namedOnly that a person read it on the day printedCounted as coverage, never shown as re-checked

What we changed

We split the one date into two and print both. The first is the day a person read the source. The second is the day the facts gate last found the string still there, and a build can't move it. Toolproof's two claims read 2026-08-22 for the first and 2026-09-04 for the second.

We didn't make an unprobed claim fail, because the answer to a gate like that is somebody typing a string that matches nothing, which reads as coverage while checking less than before. Instead every run prints how many claims are machine-checkable, as a number we want to see rise. And the stamp is all or nothing: either every probed claim confirmed today and they all move, or none of them do.

Read the register: every figure Toolproof states, with the source it came off and both dates.

<caption>A verified row only meant a person once read the page, so a job now re-reads it and says so when it can't.</caption>

---

One shipped product, taken apart, once a month. What it does, what it cost to build, what the pipeline behind it looks like, and what the numbers did, read off the repository and the live site, not written from memory. Join the list.

All writing