Vol. III · Issue 14 · JUL 2026 · Established MMXXIV

AutoResearch: How We Use AI to Improve Our Own Platform

We built an autonomous experimentation pipeline that runs real experiments on our B2B dashboard: the Karpathy Loop with governance, cost tracking, and knowledge accumulation. Here's how it works and what we learned.

BB
Boris BarashBuilder of things with AI. Creator of curate-me.ai.
#autoresearch#karpathy-loop#openclaw
AI Collaboration
§ Colophon

AutoResearch Pipeline, Autonomous experimentation

blog-dev, Page implementation

Total AI cost: $2.16

Governed by curate-me.ai

Two weeks ago, Andrej Karpathy released AutoResearch, a 630-line Python script that ran 700 experiments in 2 days and found 20 optimizations. Shopify's CEO let it run overnight on internal data: 37 experiments, 19% performance gain.

So we pointed it at our own product.

The Experiment

We built an AutoResearch pipeline that runs real experiments on the Curate-Me B2B dashboard, a Next.js 15 app with 229 pages and 1,100+ source files. The actual product, not a toy benchmark.

Every experiment follows the Karpathy Loop:

  1. Read current state: check what the last experiment achieved
  2. Hypothesize: propose a single, targeted change
  3. Modify: make the code change
  4. Evaluate: run the eval command, extract the metric
  5. Decide: if the metric improved, commit. If not, revert.
  6. Repeat

Each experiment writes its learnings to a knowledge store. Future experiments read those learnings before starting. The system gets smarter with every run.

What We Actually Found

We ran experiments across 6 metrics. Here are the dashboard baselines we measured:

MetricBaselineWhat We Found
lucide-react imports4 files (vs 690 Phosphor files)99.4% of the codebase already used Phosphor. 4 holdout files from an early template.
@monaco-editor bundle1 static import, 380KB gzippedUsed in exactly 1 file. Dynamic import saves 180KB from initial load.
@xyflow/react bundle27 files, all static importsOnly used in workflow builder. Already behind a dynamic boundary at page level.
Test coverage4% line threshold162 test files for 1,106 source files, a 14.6% ratio.
Playwright workers1 (sequential)Config comments suggest previous OOM issues. Safe to increase to 2-3.
TypeScript any types198 in source70% are in test files. Only 15-20 non-test instances.

The first three experiments alone would reduce the initial bundle by 233KB. That's 500ms faster on 3G.

The Governance Layer

This is different from running a script locally: every LLM call in the experiment pipeline flows through our gateway.

The gateway adds 6 governance checks before each request reaches the LLM provider:

  1. Rate limit: 100 RPM per org
  2. Cost estimate: BPE tokenization predicts the cost before tokens are consumed
  3. PII scan: regex scan for secrets and personally identifiable information
  4. Security scanner: prompt injection and jailbreak detection
  5. Model allowlist: enforce which models each org can use
  6. HITL gate: flag high-cost requests for human approval

Total governance latency: 3.1ms. Every experiment is cost-tracked, auditable, and governed.

The Pipeline: Blog to PR

The full chain looks like this:

Blog trigger button
   Blog API route (/api/demos/autoresearch/trigger)
     Gateway autopilot endpoint (/api/v1/autopilot/run/dev_team)
       Docker container starts (openclaw-base image)
         Claude Code CLI runs the experiment
           Git changes committed
             GitHub PR created
               Results stored in MongoDB
                 Blog experiment archive updated

We triggered experiments directly from its-boris.com/demos/autoresearch. No B2B dashboard access needed. Click a button, watch a real experiment execute, see the PR appear.

4 PRs created in our first session:

  • PR #1413: E2E pipeline verification
  • PR #1414: Remove lucide-react imports
  • PR #1415: Blog file inventory
  • PR #1416: Add smoke tests for demo pages

Total cost: $1.35 for 20 experiments. Compare that to a senior engineer spending 4 hours on the same work.

Self-Improving Agents

The experiments page isn't the only new demo. We also built Night Owl, a daily AI news digest agent inspired by the self-improving-agent skill (1,100+ stars on ClawHub, 90K+ downloads).

Night Owl runs daily at 8 AM UTC. Before scanning news, it reads its knowledge base: past digests, topic engagement data, source quality scores, reader feedback. After writing the digest, it records what it learned.

The quality trend chart on the demo page shows the agent improving over time. Readers can rate each digest with a star rating that feeds directly back into the learning loop.

Governance and learning sit in the same platform, so the agents that get smarter over time are also the ones you can audit.

What We Built (Technical)

In one session, we shipped:

  • 3 new demo pages: AutoResearch, Night Owl, Runner Spotlight
  • 8 API routes: experiment archive, SSE stream proxy, on-demand triggers, feedback endpoints
  • 4 patent components: Cost Governance, MCP Self-Governance, HITL Auto-Replay, Audit Chain Viewer
  • Navigation overhaul: dropdown menus surfacing 9+ hidden pages
  • Learning dashboard: cross-agent knowledge metrics on the platform page
  • 7 infrastructure fixes: BYOVM executor attributes, agent token auth, model ID format, MongoDB client, internal URLs, stale containers

The blog at its-boris.com went from 4 nav links hiding 9 pages to a comprehensive showcase with 10 interactive demos, all powered by the Curate-Me gateway.

Try It

The AutoResearch demo is live at its-boris.com/demos/autoresearch. You can see the experiment archive with real data, trigger new experiments on demand, and watch the knowledge base grow.

The Night Owl cron agent is at its-boris.com/demos/cron. The Runner Spotlight is at its-boris.com/demos/runner.

Every experiment, every digest, every runner task flows through the governance gateway. Every LLM call is cost-tracked. Every result is auditable.

We're using our own platform to improve our own platform.

Comments

Loading comments...

Leave a comment

Comments are moderated by our AI agent and reviewed by a human.