Make candidate ordering an explicit deterministic policy and benchmark ascending, descending, and exact-width-first choices. Keep ascending as the default because it produces the best measured time to first solution despite best-fit's slightly smaller tree. Retain the benchmark v1 interface and document why randomized and duplicate-work portfolio policies are deferred. Tests: Release and Debug CTest (10 passed each) Refs: #12
223 lines
12 KiB
Markdown
223 lines
12 KiB
Markdown
# Solver benchmarks
|
||
|
||
The benchmark suite records repeatable performance data independently of the
|
||
default correctness tests. It exercises the exhaustive infeasible order 7 and
|
||
the first-solution orders 8 and 9. The benchmark executable is opt-in:
|
||
|
||
```sh
|
||
cmake -S . -B build-benchmark -DCMAKE_BUILD_TYPE=Release \
|
||
-DPARTRIDGE_BUILD_BENCHMARKS=ON
|
||
cmake --build build-benchmark
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
> benchmark.json
|
||
```
|
||
|
||
The default policy is one unrecorded warm-up followed by five repetitions per
|
||
case, with a 600-second timeout for each process. Solver stdout is captured;
|
||
the probe renders into an in-memory stream so grids do not perturb terminal I/O.
|
||
Override the policy with `--orders`, `--warmup`, `--repetitions`, `--timeout`,
|
||
and `--candidate-order`. The choices are `ascending`, `descending`, and
|
||
`best-fit`; ascending candidate sizes are the production default.
|
||
Order 9 uses the constructive odd-order path, searching order 8 and then tiling
|
||
the enlarged border, so it is suitable for normal local benchmarking:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 8 9 --warmup 1 --repetitions 5 --timeout 60 > benchmark.json
|
||
```
|
||
|
||
Use `--direct-search` when benchmarking the skyline core rather than the public
|
||
even-predecessor construction used for odd orders from 9 onwards:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 6 7 8 9 --warmup 1 --repetitions 5 --timeout 60 \
|
||
--candidate-order ascending --direct-search > direct.json
|
||
```
|
||
|
||
The JSON contains every run and median, range, and median absolute deviation
|
||
for solve, construction, independent validation, and rendering. Direct-search
|
||
cases report zero construction time: their setup and allocation remain part of
|
||
solve time. Constructed odd-order cases report predecessor search and
|
||
construction separately. The document also records compiler,
|
||
flags, build type, commit, OS/CPU metadata, worker count, search policy, seed,
|
||
timeouts, errors, invalid outputs, and the stdout policy. Search counts must
|
||
be stable across repeated runs. Prune and task counters are zero for the
|
||
current unpruned, single-threaded solver and reserve stable schema fields for
|
||
later work.
|
||
|
||
The runner writes its JSON report before returning a failure status if any mode
|
||
has no completed runs or produces an error or invalid output. Timeouts are
|
||
reported but do not fail a case when another repetition completed.
|
||
|
||
Normal `partridge_cpp` calls instantiate a compile-time counter-free solver.
|
||
Use `--measure-overhead` to run both counter-free and counted variants and
|
||
report their median difference:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 8 --warmup 2 --repetitions 7 --timeout 60 --measure-overhead \
|
||
> overhead.json
|
||
```
|
||
|
||
Counter-free and counted warm-ups and measurements are interleaved. The first
|
||
mode alternates on each repetition, limiting systematic bias from temperature,
|
||
frequency scaling, and run order. Reported overhead is the difference between
|
||
the two independently summarized medians.
|
||
|
||
Do not use wall-clock thresholds as correctness checks. Keep the generated
|
||
JSON outside version control unless it is being deliberately added as a named
|
||
comparison baseline.
|
||
|
||
## Smallest-valley skyline
|
||
|
||
The solver stores one filled height per board column instead of one value per
|
||
cell. Equal adjacent heights form conceptual vertical bars. A valley is a
|
||
maximal bar lower than both neighbours, with board edges treated as bars of
|
||
full board height. Each node scans for the smallest-width valley, breaking
|
||
ties by lower height and then leftmost position, and tries every available
|
||
square which fits at that valley's far-left edge.
|
||
|
||
This branching remains complete: the bottom-left cell of the selected valley
|
||
must be covered, a square covering it cannot begin to the left across the
|
||
taller neighbour, and it cannot extend beyond the equal-height run without
|
||
overlap or leaving an unreachable hole. Trying every fitting available size
|
||
therefore includes the placement used by every possible completion.
|
||
|
||
For board width `W` and order `n`, the skyline scan is `O(W)`. A node tries at
|
||
most `n` candidates and each placement or exact undo changes at most `n`
|
||
heights, giving `O(W + n^2)` local work. The skyline, multiplicities,
|
||
placements, and recursion stack use `O(W + n^2)` state, compared with the
|
||
former `O(W^2)` cell grid.
|
||
|
||
One-run exploratory measurements used the issue #4 dirty worktree at base
|
||
commit `0a7ce1e`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker,
|
||
no warm-up, and a 15-second timeout. Every completed result passed the
|
||
independent benchmark validator:
|
||
|
||
| Order | Result | Ascending time | Ascending nodes | Descending time | Descending nodes |
|
||
| --- | --- | ---: | ---: | ---: | ---: |
|
||
| 6 | infeasible | 0.040 s | 659,598 | 0.039 s | 659,598 |
|
||
| 7 | infeasible | 3.103 s | 43,604,507 | 3.071 s | 43,604,507 |
|
||
| 8 | solution | 0.585 s | 7,735,369 | 0.941 s | 12,186,125 |
|
||
| 9 direct | solution | 3.831 s | 45,840,266 | timeout | unavailable |
|
||
|
||
The infeasible orders exhaust the same tree in either direction. Ascending
|
||
was selected as the default because it reaches the first order-8 solution with
|
||
36% fewer nodes and also completed direct order 9 within the timeout;
|
||
descending direct order 9 did not.
|
||
|
||
A direct ascending order-10 probe exceeded 20 seconds. Public order 10 is
|
||
also a direct search, and public order 11 first searches order 10 before using
|
||
odd-predecessor construction. Consequently neither 10 nor 11 is in the
|
||
routine correctness suite: doing so would test the same unresolved order-10
|
||
search bottleneck, while the existing route-boundary test still verifies that
|
||
11 selects construction. Revisit both sizes when order 10 completes within a
|
||
practical test budget.
|
||
|
||
## Candidate policy selection
|
||
|
||
Candidate ordering is a deterministic search policy and does not alter the
|
||
smallest-valley selection or set of placements tried. `ascending` tries
|
||
smaller available squares first and `descending` tries larger ones first.
|
||
`best-fit` first tries a square exactly as wide as the selected valley, because
|
||
that placement closes the valley without leaving a shelf remainder, then tries
|
||
the other sizes in ascending order. If no exact-width square fits, best-fit
|
||
and ascending are identical at that node.
|
||
|
||
The policy comparison used the issue #12 working tree based on commit
|
||
`d751d1b`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, and one worker.
|
||
Order 8 used two warm-ups and seven sequential measured repetitions; direct
|
||
order 9 used one warm-up and three measured repetitions. All completed
|
||
results passed independent validation and node counts were stable:
|
||
|
||
| Order | Policy | Counted median (range) | Counter-free median | Nodes |
|
||
| --- | --- | --- | --- | ---: |
|
||
| 8 | ascending | 0.808 s (0.790–0.852 s) | 0.794 s | 7,735,369 |
|
||
| 8 | descending | 1.310 s (1.279–1.449 s) | 1.268 s | 12,186,125 |
|
||
| 8 | best-fit | 0.817 s (0.813–0.857 s) | 0.823 s | 7,679,349 |
|
||
| 9 direct | ascending | 5.522 s (5.499–5.830 s) | not measured | 45,840,266 |
|
||
| 9 direct | best-fit | 5.651 s (5.533–5.820 s) | not measured | 45,746,016 |
|
||
|
||
The earlier direct-order-9 descending probe exceeded its 15-second limit.
|
||
Ascending is retained as the stable single-threaded default because it had the
|
||
lowest measured median time to the first solution at both measured solvable
|
||
sizes. Best-fit's slightly smaller trees did not compensate for its policy
|
||
checks, while descending was substantially worse. Exhaustive infeasible
|
||
order-5 tests visit the same number of nodes under all three policies, which
|
||
checks that ordering does not affect completeness.
|
||
|
||
One policy therefore applies to the currently measured sizes 8 and 9. This
|
||
does not establish that ascending is optimal for order 10: a bounded best-fit
|
||
order-9 comparison changed the search tree by only 0.2%, so there was no
|
||
evidence that repeating the known long order-10/11 search would be useful.
|
||
Keep 10 and 11 as opt-in benchmark cases. Public order 11 is particularly
|
||
important to interpret correctly: it constructs from an order-10 search, so it
|
||
does not independently measure an odd-order candidate policy.
|
||
|
||
A worker portfolio was considered but not added. Running identical policies
|
||
duplicates the same deterministic traversal. Pairing ascending with best-fit
|
||
adds little diversity on the measured trees, and pairing ascending with
|
||
descending dedicates a worker to the consistently slower policy. Splitting a
|
||
shared frontier could avoid duplicated prefixes, but that is the parallel
|
||
frontier work tracked separately in issue #11. Seeded randomized ordering was
|
||
also rejected for now: the deterministic alternatives already select a clear
|
||
default, and there is no measurement showing that seed distributions would
|
||
improve time to first solution. The benchmark schema retains its nullable
|
||
seed field so a future evidence-backed randomized policy can report
|
||
reproducible runs without changing the format.
|
||
|
||
## Post-correctness baseline
|
||
|
||
This framework starts from commit `ce39d0a` after the rendering assertion fix
|
||
in #7 and completion check fix in #14. The earlier Apple M1 Release results in
|
||
`results.md` are approximately 1.76 seconds for order 8 and 158.69 seconds
|
||
elapsed for order 9; they predate the structured runner and do not contain
|
||
search counters.
|
||
|
||
The first clean structured baseline used commit `598667b`, Apple Clang 21.0.0
|
||
with `-O3 -DNDEBUG`, Apple arm64, one worker, one warm-up, and three measured
|
||
repetitions. The runner reported a clean working tree:
|
||
|
||
| Order | Result | Counted solve median (range) | Nodes | Placements | Backtracks |
|
||
| --- | --- | --- | ---: | ---: | ---: |
|
||
| 7 | infeasible | 3.453 s (3.444–3.455 s) | 110,483,315 | 110,483,314 | 110,483,314 |
|
||
| 8 | solution | 1.817 s (1.814–1.817 s) | 60,485,176 | 60,485,176 | 60,485,140 |
|
||
|
||
Counts were stable across repetitions. Interleaved counter-free medians were
|
||
3.435 seconds for order 7 and 1.805 seconds for order 8, giving counted
|
||
overheads of 0.53% and 0.64% respectively. Construction time was zero; median
|
||
independent validation and rendering times were each below 0.02 milliseconds.
|
||
|
||
Order 9 was not rerun for this initial baseline because the former direct
|
||
search took several minutes. The default suite includes it with a per-run
|
||
timeout.
|
||
|
||
## Odd construction comparison
|
||
|
||
The order-9 construction was measured from the issue 6 working tree based on
|
||
commit `ddf07e7`, using Apple Clang 21.0.0 with `-O3 -DNDEBUG`, macOS arm64,
|
||
one worker, one warm-up, and three measured repetitions. Counter-free and
|
||
counted runs were interleaved:
|
||
|
||
| Order | Mode | Search median (range) | Construction median | Nodes |
|
||
| --- | --- | --- | --- | ---: |
|
||
| 8 | counter-free | 1.773 s (1.770–1.775 s) | 0 | 0 |
|
||
| 8 | counted | 1.849 s (1.848–1.850 s) | 0 | 60,485,176 |
|
||
| 9 | counter-free | 1.778 s (1.773–1.779 s) | 0.458 us | 0 |
|
||
| 9 | counted | 1.852 s (1.845–1.912 s) | 0.416 us | 60,485,176 |
|
||
|
||
All runs completed with valid results and stable counters. The matching
|
||
order-8 and order-9 search counts demonstrate that the new path searches only
|
||
the predecessor. Compared with the recorded 158.69-second direct order-9
|
||
elapsed time in `results.md`, the 1.778-second counter-free median plus
|
||
construction is approximately 89 times faster. The benchmark working tree
|
||
was necessarily dirty with the issue 6 implementation.
|
||
|
||
New optimization issues should quote the exact JSON
|
||
environment, policy, median/spread, stable counters, and counted overhead from
|
||
this runner for both before and after revisions.
|
||
|
||
The separate optional CP-SAT reference benchmark and its model, memory, worker,
|
||
and timing report are documented in [CP_SAT_REFERENCE.md](./CP_SAT_REFERENCE.md).
|