Evidence and benchmarks¶
Looplet's primary claim is not that a loop library makes a model smarter. It is that an owned harness can be inspected, changed, and regression-tested. The evidence is organized accordingly: a controlled behavioral proof first, then model-controlled harness comparisons, then package-cost measurements.
Every study below names its variable and limitations. Historical numbers are snapshots, not timeless leaderboard claims.
1. Behavioral proof: failure → regression contract¶
The repository includes a network-free experiment with two harness versions:
- the scripted model makes the same
publish_reportanddonedecisions; - v1 contains one arithmetic bug and the required outcome grader fails;
- captured-response replay executes the fixed v2 tool in a fresh workspace;
- an independent collector reads
report.json; the same grader passes.
| Variable | Held fixed / changed |
|---|---|
| Model calls | Fixed captured responses; no provider or network |
| Tool decisions | Fixed: publish_report → done |
| Harness code | One line changes: addition → subtraction |
| Outcome | Profit changes from 200 → 40 |
| Required eval | 0.00 fail → 1.00 pass |
Run it with uv run python examples/regression_demo/run_demo.py. Read the
full proof and limitations before treating replay as an
experiment design: tools and side effects execute again, so this is
captured-response replay, not deterministic simulation.
2. Model-controlled harness comparison¶
The optional
coder harness benchmark
compares examples/coder.cartridge with GitHub Copilot CLI while requesting the
same model family (claude-sonnet-4.6 in the recorded snapshot). Looplet used a
proxy and Copilot used its own serving connection, so the study compares the
end-to-end harness and serving paths rather than isolating scaffolding alone.
| Recorded suite | Looplet coder | Copilot CLI | What the snapshot supports |
|---|---|---|---|
| 9 short coding/non-coding tasks with deterministic verifiers | 9/9 | 9/9 | Correctness parity on this small suite; Looplet averaged 18.3s vs 23.1s. |
| 4 hard coding tasks with hidden verifiers | 4/4 | 4/4 | No observed correctness separation. |
| 6 open-ended tasks, blind LLM judges | 18.0/20 after a one-line prompt change | 18.6/20 | Copilot retained a small prose-quality edge; the result is judge- and sample-sensitive. |
Reports preserve task-level results and methodology:
Limits of this comparison¶
- Same model family does not guarantee identical serving stacks, hidden provider policy, caching, or sampling implementation.
- Nine and four tasks are demonstration-sized samples, not population estimates. There are no confidence intervals.
- Looplet token counts are
chars / 4estimates because the proxy omitted usage; Copilot counts are self-reported. Compare direction, not precision. - Open-ended scores are single-generation LLM-judge results. Small differences may be noise even with blind, counterbalanced judging.
- The suites do not represent huge repositories, external services, or multi-hour autonomous operation.
The honest result is modest: a small, owned cartridge was competitive on these tasks. It does not establish that Looplet is universally faster, cheaper, or more capable than a turnkey coding agent.
3. Runtime footprint¶
Core Looplet 0.3.0 declares zero third-party runtime dependencies. Provider SDKs are optional extras. This is a current package property, not a benchmark inference.
The table below is an archived environment snapshot from 2026-04-21 (Python 3.11.13, Linux x86_64, fresh uv-managed environments, then-current PyPI versions). "Packages installed" includes the target package and excludes pip, setuptools, and wheel.
| Install in that snapshot | Packages installed |
|---|---|
pip install looplet |
1 (Looplet itself; zero runtime dependencies) |
pip install looplet[all] |
20 |
pip install claude-agent-sdk |
30 |
pip install langgraph |
31 |
pip install strands-agents |
49 |
pip install pydantic-ai |
144 |
In that archived run (Looplet 0.1.7, 2026-04-21), median cold import over nine fresh subprocesses was 289 ms. The comparison packages ranged from 1,885 ms to 3,975 ms. Those latency numbers are historical: hardware, OS caches, Python, wheel versions, and package releases all affect them. Re-run before using them in a current engineering decision.
Reproduce instead of repeating the headline¶
# Current cold-import snapshot (choose an explicit interpreter/environment).
python scripts/bench_cold_import.py --runs 9 --markdown
# Current dependency snapshot (creates throwaway environments).
python scripts/bench_dep_footprint.py --markdown
# Behavioral proof: no model, network, or external CLI.
uv run python examples/regression_demo/run_demo.py
The model-controlled coder comparison has external prerequisites and environment knobs documented in its benchmark README. Commit the raw environment metadata and task-level output when publishing a new snapshot; do not silently replace historical results.
What these results do not claim¶
They do not show that Looplet improves model reasoning, guarantees safer agents, makes arbitrary replays deterministic, or outperforms graph runtimes and turnkey agents on their intended workloads. They show three narrower properties:
- a harness failure can become an executable behavioral contract;
- a compact, owned harness can be competitive in a small same-model study;
- the core package imposes no third-party runtime dependency tree.