A selected valley cannot admit a square wider than itself until it reaches the lower neighbouring rim. Reject states whose remaining narrow-square area cannot fill that strip, including width-one and width-two gaps. Keep an unpruned benchmark mode and dedicated counters so the rule remains independently measurable. Record the soundness argument and the measured default-on improvement. Tests: Debug CTest (12 passed) Tests: ASan+UBSan CTest (12 passed) Refs: #10
305 lines
16 KiB
Markdown
305 lines
16 KiB
Markdown
# Solver benchmarks
|
||
|
||
The benchmark suite records repeatable performance data independently of the
|
||
default correctness tests. It exercises the exhaustive infeasible order 7 and
|
||
the first-solution orders 8 and 9. The benchmark executable is opt-in:
|
||
|
||
```sh
|
||
cmake -S . -B build-benchmark -DCMAKE_BUILD_TYPE=Release \
|
||
-DPARTRIDGE_BUILD_BENCHMARKS=ON
|
||
cmake --build build-benchmark
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
> benchmark.json
|
||
```
|
||
|
||
The default policy is one unrecorded warm-up followed by five repetitions per
|
||
case, with a 600-second timeout for each process. Solver stdout is captured;
|
||
the probe renders into an in-memory stream so grids do not perturb terminal I/O.
|
||
Override the policy with `--orders`, `--warmup`, `--repetitions`, `--timeout`,
|
||
and `--candidate-order`. The choices are `ascending`, `descending`, and
|
||
`best-fit`; ascending candidate sizes are the production default.
|
||
The production default also removes equivalent D4 board orientations by
|
||
constraining the unique unit square. Pass `--no-symmetry` to obtain an
|
||
otherwise identical unconstrained baseline.
|
||
The production default also applies the valley-capacity rule described below.
|
||
Pass `--no-pruning` to obtain an otherwise identical unpruned search.
|
||
Order 9 uses the constructive odd-order path, searching order 8 and then tiling
|
||
the enlarged border, so it is suitable for normal local benchmarking:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 8 9 --warmup 1 --repetitions 5 --timeout 60 > benchmark.json
|
||
```
|
||
|
||
Use `--direct-search` when benchmarking the skyline core rather than the public
|
||
even-predecessor construction used for odd orders from 9 onwards:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 6 7 8 9 --warmup 1 --repetitions 5 --timeout 60 \
|
||
--candidate-order ascending --direct-search > direct.json
|
||
```
|
||
|
||
The JSON contains every run and median, range, and median absolute deviation
|
||
for solve, construction, independent validation, and rendering. Direct-search
|
||
cases report zero construction time: their setup and allocation remain part of
|
||
solve time. Constructed odd-order cases report predecessor search and
|
||
construction separately. The document also records compiler,
|
||
flags, build type, commit, OS/CPU metadata, worker count, search policy, seed,
|
||
timeouts, errors, invalid outputs, and the stdout policy. Search counts must
|
||
be stable across repeated runs. Prune counters report enabled search rules;
|
||
they are zero in the corresponding disabled modes. Task counters remain zero
|
||
for the current single-threaded solver and reserve stable schema fields for
|
||
later work.
|
||
|
||
The runner writes its JSON report before returning a failure status if any mode
|
||
has no completed runs or produces an error or invalid output. Timeouts are
|
||
reported but do not fail a case when another repetition completed.
|
||
|
||
Normal `partridge_cpp` calls instantiate a compile-time counter-free solver.
|
||
Use `--measure-overhead` to run both counter-free and counted variants and
|
||
report their median difference:
|
||
|
||
```sh
|
||
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
|
||
--orders 8 --warmup 2 --repetitions 7 --timeout 60 --measure-overhead \
|
||
> overhead.json
|
||
```
|
||
|
||
Counter-free and counted warm-ups and measurements are interleaved. The first
|
||
mode alternates on each repetition, limiting systematic bias from temperature,
|
||
frequency scaling, and run order. Reported overhead is the difference between
|
||
the two independently summarized medians.
|
||
|
||
Do not use wall-clock thresholds as correctness checks. Keep the generated
|
||
JSON outside version control unless it is being deliberately added as a named
|
||
comparison baseline.
|
||
|
||
## Valley-capacity pruning
|
||
|
||
For a selected valley, let its rim be the lower height of its two neighbours,
|
||
treating a board edge as full height. Until the valley reaches that rim, no
|
||
square wider than the valley can enter it. Therefore the combined area of all
|
||
remaining squares no wider than the valley must be at least the valley width
|
||
times its depth below the rim. Rejecting a state when that necessary
|
||
inequality fails is sound. At widths one and two it provides the usual
|
||
narrow-gap capacity checks without separate special cases.
|
||
|
||
The rule is compiled out of the recursive search in `--no-pruning` mode.
|
||
Instrumented output records `valley_capacity_checks` and
|
||
`valley_capacity_prunes` as well as the aggregate pruning counters, allowing
|
||
enabled and disabled runs to report nodes, checks, hits, and elapsed time.
|
||
|
||
Measurements used the issue #10 working tree based on commit `f37e087`, Apple
|
||
Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker, one warm-up, and five
|
||
measured repetitions. The production ascending policy and D4 constraint were
|
||
enabled. Every result passed the benchmark's independent validator and all
|
||
counters were stable:
|
||
|
||
| Order | Valley capacity | Median solve (range) | Nodes | Checks | Prunes |
|
||
| --- | --- | ---: | ---: | ---: | ---: |
|
||
| 7 exhaustive | enabled | 0.996 s (0.990-0.998 s) | 13,833,048 | 13,833,048 | 7,411,551 |
|
||
| 7 exhaustive | disabled | 1.110 s (1.109-1.111 s) | 14,997,603 | 0 | 0 |
|
||
| 8 first solution | enabled | 0.205 s (0.204-0.208 s) | 2,724,096 | 2,724,095 | 1,606,836 |
|
||
| 8 first solution | disabled | 0.228 s (0.228-0.233 s) | 2,931,203 | 0 | 0 |
|
||
|
||
The rule reduced nodes by 7.8% for order 7 and 7.1% for order 8. Its low
|
||
per-node cost also reduced median counted solve time by 10.3% and 9.9%,
|
||
respectively, so it remains enabled by default.
|
||
|
||
## D4 board symmetry
|
||
|
||
Every solution contains exactly one 1-by-1 square. Rotations and reflections
|
||
of the whole board preserve square sizes, multiplicities, and coverage, so the
|
||
unit square can select the orientation without assigning identities to any of
|
||
the repeated larger squares. For a board of width `W`, the solver accepts the
|
||
unit square only in the closed fundamental triangle
|
||
`x <= y <= floor((W - 1) / 2)`. Reflecting a cell toward the left edge,
|
||
swapping its coordinates if necessary, and reflecting toward the top edge
|
||
maps every D4 orbit into this triangle. Closed diagonal and midline
|
||
boundaries retain the smaller orbits of symmetric cells.
|
||
|
||
The check is made only when the skyline search is ready to place the unit
|
||
square, so it never rejects a partial state before the square's position is
|
||
decidable. The implementation remains compile-time counter-free in normal
|
||
solver calls and keeps the single-threaded deterministic search policy.
|
||
Instrumented runs count examined and rejected unit-square placements as prune
|
||
checks and hits.
|
||
|
||
Measurements used the issue #3 dirty working tree based on commit `74bde26`,
|
||
Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker, one warm-up, and
|
||
five sequential measured repetitions. All results passed the benchmark's
|
||
independent cell-coverage validator and counts were stable:
|
||
|
||
| Route | D4 constraint | Median solve (range) | Nodes | Prune hits |
|
||
| --- | --- | ---: | ---: | ---: |
|
||
| Order 8 public | enabled | 0.245 s (0.244-0.248 s) | 2,931,203 | 1,329,567 |
|
||
| Order 8 public | disabled | 0.679 s (0.659-0.709 s) | 7,735,369 | 0 |
|
||
| Order 9 direct | enabled | 1.705 s (1.671-1.740 s) | 16,231,918 | 6,755,171 |
|
||
| Order 9 direct | disabled | 4.542 s (4.482-4.662 s) | 45,840,266 | 0 |
|
||
|
||
Thus the constraint reduced nodes by 62% for order 8 and 65% for direct order
|
||
9; counted median solve time fell by 64% and 62%, respectively.
|
||
|
||
A public order-10 probe with the D4 constraint, ascending policy, and a
|
||
45-second per-process bound did not complete. Public order 11 first performs
|
||
that identical order-10 core search and only then adds its inexpensive odd
|
||
border, so adding either order to routine correctness tests would duplicate
|
||
the same unresolved bottleneck. They remain useful opt-in heavyweight
|
||
benchmark targets with explicit timeouts. A direct order-11 benchmark is a
|
||
different experiment: it bypasses the public odd construction and searches
|
||
the larger core itself, so it must not be presented as public order-11
|
||
performance.
|
||
|
||
## Smallest-valley skyline
|
||
|
||
The solver stores one filled height per board column instead of one value per
|
||
cell. Equal adjacent heights form conceptual vertical bars. A valley is a
|
||
maximal bar lower than both neighbours, with board edges treated as bars of
|
||
full board height. Each node scans for the smallest-width valley, breaking
|
||
ties by lower height and then leftmost position, and tries every available
|
||
square which fits at that valley's far-left edge.
|
||
|
||
This branching remains complete: the bottom-left cell of the selected valley
|
||
must be covered, a square covering it cannot begin to the left across the
|
||
taller neighbour, and it cannot extend beyond the equal-height run without
|
||
overlap or leaving an unreachable hole. Trying every fitting available size
|
||
therefore includes the placement used by every possible completion.
|
||
|
||
For board width `W` and order `n`, the skyline scan is `O(W)`. A node tries at
|
||
most `n` candidates and each placement or exact undo changes at most `n`
|
||
heights, giving `O(W + n^2)` local work. The skyline, multiplicities,
|
||
placements, and recursion stack use `O(W + n^2)` state, compared with the
|
||
former `O(W^2)` cell grid.
|
||
|
||
One-run exploratory measurements used the issue #4 dirty worktree at base
|
||
commit `0a7ce1e`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker,
|
||
no warm-up, and a 15-second timeout. Every completed result passed the
|
||
independent benchmark validator:
|
||
|
||
| Order | Result | Ascending time | Ascending nodes | Descending time | Descending nodes |
|
||
| --- | --- | ---: | ---: | ---: | ---: |
|
||
| 6 | infeasible | 0.040 s | 659,598 | 0.039 s | 659,598 |
|
||
| 7 | infeasible | 3.103 s | 43,604,507 | 3.071 s | 43,604,507 |
|
||
| 8 | solution | 0.585 s | 7,735,369 | 0.941 s | 12,186,125 |
|
||
| 9 direct | solution | 3.831 s | 45,840,266 | timeout | unavailable |
|
||
|
||
The infeasible orders exhaust the same tree in either direction. Ascending
|
||
was selected as the default because it reaches the first order-8 solution with
|
||
36% fewer nodes and also completed direct order 9 within the timeout;
|
||
descending direct order 9 did not.
|
||
|
||
A direct ascending order-10 probe exceeded 20 seconds. Public order 10 is
|
||
also a direct search, and public order 11 first searches order 10 before using
|
||
odd-predecessor construction. Consequently neither 10 nor 11 is in the
|
||
routine correctness suite: doing so would test the same unresolved order-10
|
||
search bottleneck, while the existing route-boundary test still verifies that
|
||
11 selects construction. Revisit both sizes when order 10 completes within a
|
||
practical test budget.
|
||
|
||
## Candidate policy selection
|
||
|
||
Candidate ordering is a deterministic search policy and does not alter the
|
||
smallest-valley selection or set of placements tried. `ascending` tries
|
||
smaller available squares first and `descending` tries larger ones first.
|
||
`best-fit` first tries a square exactly as wide as the selected valley, because
|
||
that placement closes the valley without leaving a shelf remainder, then tries
|
||
the other sizes in ascending order. If no exact-width square fits, best-fit
|
||
and ascending are identical at that node.
|
||
|
||
The policy comparison used the issue #12 working tree based on commit
|
||
`d751d1b`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, and one worker.
|
||
Order 8 used two warm-ups and seven sequential measured repetitions; direct
|
||
order 9 used one warm-up and three measured repetitions. All completed
|
||
results passed independent validation and node counts were stable:
|
||
|
||
| Order | Policy | Counted median (range) | Counter-free median | Nodes |
|
||
| --- | --- | --- | --- | ---: |
|
||
| 8 | ascending | 0.808 s (0.790–0.852 s) | 0.794 s | 7,735,369 |
|
||
| 8 | descending | 1.310 s (1.279–1.449 s) | 1.268 s | 12,186,125 |
|
||
| 8 | best-fit | 0.817 s (0.813–0.857 s) | 0.823 s | 7,679,349 |
|
||
| 9 direct | ascending | 5.522 s (5.499–5.830 s) | not measured | 45,840,266 |
|
||
| 9 direct | best-fit | 5.651 s (5.533–5.820 s) | not measured | 45,746,016 |
|
||
|
||
The earlier direct-order-9 descending probe exceeded its 15-second limit.
|
||
Ascending is retained as the stable single-threaded default because it had the
|
||
lowest measured median time to the first solution at both measured solvable
|
||
sizes. Best-fit's slightly smaller trees did not compensate for its policy
|
||
checks, while descending was substantially worse. Exhaustive infeasible
|
||
order-5 tests visit the same number of nodes under all three policies, which
|
||
checks that ordering does not affect completeness.
|
||
|
||
One policy therefore applies to the currently measured sizes 8 and 9. This
|
||
does not establish that ascending is optimal for order 10: a bounded best-fit
|
||
order-9 comparison changed the search tree by only 0.2%, so there was no
|
||
evidence that repeating the known long order-10/11 search would be useful.
|
||
Keep 10 and 11 as opt-in benchmark cases. Public order 11 is particularly
|
||
important to interpret correctly: it constructs from an order-10 search, so it
|
||
does not independently measure an odd-order candidate policy.
|
||
|
||
A worker portfolio was considered but not added. Running identical policies
|
||
duplicates the same deterministic traversal. Pairing ascending with best-fit
|
||
adds little diversity on the measured trees, and pairing ascending with
|
||
descending dedicates a worker to the consistently slower policy. Splitting a
|
||
shared frontier could avoid duplicated prefixes, but that is the parallel
|
||
frontier work tracked separately in issue #11. Seeded randomized ordering was
|
||
also rejected for now: the deterministic alternatives already select a clear
|
||
default, and there is no measurement showing that seed distributions would
|
||
improve time to first solution. The benchmark schema retains its nullable
|
||
seed field so a future evidence-backed randomized policy can report
|
||
reproducible runs without changing the format.
|
||
|
||
## Post-correctness baseline
|
||
|
||
This framework starts from commit `ce39d0a` after the rendering assertion fix
|
||
in #7 and completion check fix in #14. The earlier Apple M1 Release results in
|
||
`results.md` are approximately 1.76 seconds for order 8 and 158.69 seconds
|
||
elapsed for order 9; they predate the structured runner and do not contain
|
||
search counters.
|
||
|
||
The first clean structured baseline used commit `598667b`, Apple Clang 21.0.0
|
||
with `-O3 -DNDEBUG`, Apple arm64, one worker, one warm-up, and three measured
|
||
repetitions. The runner reported a clean working tree:
|
||
|
||
| Order | Result | Counted solve median (range) | Nodes | Placements | Backtracks |
|
||
| --- | --- | --- | ---: | ---: | ---: |
|
||
| 7 | infeasible | 3.453 s (3.444–3.455 s) | 110,483,315 | 110,483,314 | 110,483,314 |
|
||
| 8 | solution | 1.817 s (1.814–1.817 s) | 60,485,176 | 60,485,176 | 60,485,140 |
|
||
|
||
Counts were stable across repetitions. Interleaved counter-free medians were
|
||
3.435 seconds for order 7 and 1.805 seconds for order 8, giving counted
|
||
overheads of 0.53% and 0.64% respectively. Construction time was zero; median
|
||
independent validation and rendering times were each below 0.02 milliseconds.
|
||
|
||
Order 9 was not rerun for this initial baseline because the former direct
|
||
search took several minutes. The default suite includes it with a per-run
|
||
timeout.
|
||
|
||
## Odd construction comparison
|
||
|
||
The order-9 construction was measured from the issue 6 working tree based on
|
||
commit `ddf07e7`, using Apple Clang 21.0.0 with `-O3 -DNDEBUG`, macOS arm64,
|
||
one worker, one warm-up, and three measured repetitions. Counter-free and
|
||
counted runs were interleaved:
|
||
|
||
| Order | Mode | Search median (range) | Construction median | Nodes |
|
||
| --- | --- | --- | --- | ---: |
|
||
| 8 | counter-free | 1.773 s (1.770–1.775 s) | 0 | 0 |
|
||
| 8 | counted | 1.849 s (1.848–1.850 s) | 0 | 60,485,176 |
|
||
| 9 | counter-free | 1.778 s (1.773–1.779 s) | 0.458 us | 0 |
|
||
| 9 | counted | 1.852 s (1.845–1.912 s) | 0.416 us | 60,485,176 |
|
||
|
||
All runs completed with valid results and stable counters. The matching
|
||
order-8 and order-9 search counts demonstrate that the new path searches only
|
||
the predecessor. Compared with the recorded 158.69-second direct order-9
|
||
elapsed time in `results.md`, the 1.778-second counter-free median plus
|
||
construction is approximately 89 times faster. The benchmark working tree
|
||
was necessarily dirty with the issue 6 implementation.
|
||
|
||
New optimization issues should quote the exact JSON
|
||
environment, policy, median/spread, stable counters, and counted overhead from
|
||
this runner for both before and after revisions.
|
||
|
||
The separate optional CP-SAT reference benchmark and its model, memory, worker,
|
||
and timing report are documented in [CP_SAT_REFERENCE.md](./CP_SAT_REFERENCE.md).
|