Files
partridge-cpp/BENCHMARKING.md
T
Codex instance ddf07e730a bench: record clean solver baseline
Record reproducible order-7 and order-8 measurements from the committed benchmark framework with a clean working tree. Include stable work counts, solve-time spread, validation and rendering cost, and interleaved instrumentation overhead.

Order 9 remains supported by the runner but retains its historical timing because each current run takes several minutes.

Benchmark: Apple arm64 Release, one warm-up, three repetitions

Refs: #8
2026-07-30 17:10:01 +01:00

89 lines
4.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Solver benchmarks
The benchmark suite records repeatable performance data independently of the
default correctness tests. It exercises the exhaustive infeasible order 7 and
the first-solution orders 8 and 9. The benchmark executable is opt-in:
```sh
cmake -S . -B build-benchmark -DCMAKE_BUILD_TYPE=Release \
-DPARTRIDGE_BUILD_BENCHMARKS=ON
cmake --build build-benchmark
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
> benchmark.json
```
The default policy is one unrecorded warm-up followed by five repetitions per
case, with a 600-second timeout for each process. Solver stdout is captured;
the probe renders into an in-memory stream so grids do not perturb terminal I/O.
Override the policy with `--orders`, `--warmup`, `--repetitions`, and
`--timeout`. Order 9 is intentionally supported but may be omitted during
local iteration because the current solver takes minutes:
```sh
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
--orders 7 8 --warmup 1 --repetitions 5 --timeout 60 > benchmark.json
```
The JSON contains every run and median, range, and median absolute deviation
for solve, construction, independent validation, and rendering. The current
direct-search solver reports zero construction time: its setup and allocation
remain part of solve time. The separate construction field is reserved for
future constructive solution paths. The document also records compiler,
flags, build type, commit, OS/CPU metadata, worker count, search policy, seed,
timeouts, errors, invalid outputs, and the stdout policy. Search counts must
be stable across repeated runs. Prune and task counters are zero for the
current unpruned, single-threaded solver and reserve stable schema fields for
later work.
The runner writes its JSON report before returning a failure status if any mode
has no completed runs or produces an error or invalid output. Timeouts are
reported but do not fail a case when another repetition completed.
Normal `partridge_cpp` calls instantiate a compile-time counter-free solver.
Use `--measure-overhead` to run both counter-free and counted variants and
report their median difference:
```sh
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
--orders 8 --warmup 2 --repetitions 7 --timeout 60 --measure-overhead \
> overhead.json
```
Counter-free and counted warm-ups and measurements are interleaved. The first
mode alternates on each repetition, limiting systematic bias from temperature,
frequency scaling, and run order. Reported overhead is the difference between
the two independently summarized medians.
Do not use wall-clock thresholds as correctness checks. Keep the generated
JSON outside version control unless it is being deliberately added as a named
comparison baseline.
## Post-correctness baseline
This framework starts from commit `ce39d0a` after the rendering assertion fix
in #7 and completion check fix in #14. The earlier Apple M1 Release results in
`results.md` are approximately 1.76 seconds for order 8 and 158.69 seconds
elapsed for order 9; they predate the structured runner and do not contain
search counters.
The first clean structured baseline used commit `598667b`, Apple Clang 21.0.0
with `-O3 -DNDEBUG`, Apple arm64, one worker, one warm-up, and three measured
repetitions. The runner reported a clean working tree:
| Order | Result | Counted solve median (range) | Nodes | Placements | Backtracks |
| --- | --- | --- | ---: | ---: | ---: |
| 7 | infeasible | 3.453 s (3.4443.455 s) | 110,483,315 | 110,483,314 | 110,483,314 |
| 8 | solution | 1.817 s (1.8141.817 s) | 60,485,176 | 60,485,176 | 60,485,140 |
Counts were stable across repetitions. Interleaved counter-free medians were
3.435 seconds for order 7 and 1.805 seconds for order 8, giving counted
overheads of 0.53% and 0.64% respectively. Construction time was zero; median
independent validation and rendering times were each below 0.02 milliseconds.
Order 9 was not rerun for this initial baseline because its documented runtime
is several minutes. The default suite includes it with a per-run timeout.
New optimization issues should quote the exact JSON
environment, policy, median/spread, stable counters, and counted overhead from
this runner for both before and after revisions.