Compound Labs
Get the newsletter
THE STANDUPAnthropic fixes Claude 5.1 error spikesERROR SPIKES, FIX DEPLOYEDCOMPOUND MASTODONSkillWorks records inactive files beside their roster status22 INACTIVE FILES STILL LOADCTXWINDOWContext Window tracks AI lab changes and filters their evidenceFILTERS KEEP EVIDENCE TOGETHERAGENTWIREmicrosoft/agent-governance-toolkit - AI Agent Governance ToolkitPOLICY ENFORCEMENT FOR AI AGENTSCOMPOUND FACEBOOKCompound Portfolio restores its sweep after a missing module stops desk dataSWEEP RESTORED, DESK DATA WRITTENBLOCKDEXBlockDex shows no destination for captured rows with rejected slash names31 ROWS, NO DESTINATIONCONTEXT WINDOWOpenAI adds 1024p image pricing: $0.50 and $0.251024P IMAGE, $0.50 AND $0.25COMPOUND BLUESKYCompound Labs renamed job labels so job-run finds breaker stateBREAKER STATE, NOW FINDABLECTXWINDOWContext Window tracks document changes but failed mobileMOBILE WORKFLOW, ONE COLUMNAGENTWIREMineDojo Voyager: Open-Ended Embodied AgentOPEN-ENDED EMBODIED AGENTFOUNDER BLUESKYdeploy.sh now fails false Vercel production landingsFALSE LANDINGS NOW FAILCOMPOUNDlast known fixes redirect tracking after migrationSTUDIO SOURCES NOW SAY COMPOUNDCONTEXT WINDOWOpenAI adds sora-2-pro pricing720P VIDEO, $0.15, $0.30FOUNDER LINKEDINdeploy.sh now stops failed promotes from marking production landedFAILED PROMOTES NO LONGER LANDUSINGITUPUsing It Up records a receipt fixing a wobbling wooden tableONE RECEIPT, TABLE STAYED LEVELCOMPOUNDpnpm v12.4.2POSIX BIN SHIMS REPLACEDFOUNDER PEERLISTCompound Labs' deploy runner rejects failed promotesFAILED PROMOTES NO LONGER LANDTRUSTDESKTrustDesk fixes host migration redirects while keeping /api/ outOLD HOST, 308; API STAYS LIVE
Independent product R&D labFounded and run by Isaiah Kim, @kyisaiah47Newest commit Sep 16, 2026, AgentwireNewest writing Sep 16, 2026Site changelog Sep 16, 2026
Compound WireCOMPOUND31 August 2026

PRAXIST: an autonomous, gradable research loop

sapientinc published PRAXIST on GitHub, described by its authors as an autonomous research system for measurable, computer-executable research.

The two adjectives are the design. An unattended research loop lives or dies on grading, because a loop with no scorer generates plausible ideas forever and never closes. Restricting the domain to work a machine can run end to end and score with a number is what makes closing it possible at all. That takes in algorithms, training runs and simulation. It rules out anything needing a bench, an instrument, a human subject, or a judgment call about whether a result is interesting rather than merely significant.

The question a one-line description cannot answer is who owns the metric. A system that proposes the hypothesis, writes the code and defines the measurement is in a position to optimise the measurement, and that failure is not exotic, it is the ordinary history of every benchmark that got saturated by things tuned against it. So the thing to look for in the repo is whether the evaluation is fixed and external to the agent, or generated alongside the work it is grading.

https://github.com/sapientinc/PRAXIST

#dev

What each account said

sapientinc/PRAXIST is an autonomous research system scoped, in its own words, to "measurable, computer-executable research". That bound is the whole claim: it covers research whose result a machine can run and score, and the description stops there. #dev
Bluesky