Files
partridge-cpp/BENCHMARKING.md
T
Codex instance f37e08768d solver: break board dihedral symmetry
Constrain the unique unit square to a closed D4 fundamental region once its placement is known. This preserves one representative of every board-orientation orbit without assigning identities to repeated squares.

Keep a symmetry-disabled benchmark path, document the proof and measurements, and cover generic, diagonal, midline, corner, and centre orbits.

Tests: Release, Debug, ASan, and UBSan CTest (11 passed each)

Refs: #3
2026-07-30 18:18:45 +01:00

270 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Solver benchmarks
The benchmark suite records repeatable performance data independently of the
default correctness tests. It exercises the exhaustive infeasible order 7 and
the first-solution orders 8 and 9. The benchmark executable is opt-in:
```sh
cmake -S . -B build-benchmark -DCMAKE_BUILD_TYPE=Release \
-DPARTRIDGE_BUILD_BENCHMARKS=ON
cmake --build build-benchmark
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
> benchmark.json
```
The default policy is one unrecorded warm-up followed by five repetitions per
case, with a 600-second timeout for each process. Solver stdout is captured;
the probe renders into an in-memory stream so grids do not perturb terminal I/O.
Override the policy with `--orders`, `--warmup`, `--repetitions`, `--timeout`,
and `--candidate-order`. The choices are `ascending`, `descending`, and
`best-fit`; ascending candidate sizes are the production default.
The production default also removes equivalent D4 board orientations by
constraining the unique unit square. Pass `--no-symmetry` to obtain an
otherwise identical unconstrained baseline.
Order 9 uses the constructive odd-order path, searching order 8 and then tiling
the enlarged border, so it is suitable for normal local benchmarking:
```sh
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
--orders 8 9 --warmup 1 --repetitions 5 --timeout 60 > benchmark.json
```
Use `--direct-search` when benchmarking the skyline core rather than the public
even-predecessor construction used for odd orders from 9 onwards:
```sh
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
--orders 6 7 8 9 --warmup 1 --repetitions 5 --timeout 60 \
--candidate-order ascending --direct-search > direct.json
```
The JSON contains every run and median, range, and median absolute deviation
for solve, construction, independent validation, and rendering. Direct-search
cases report zero construction time: their setup and allocation remain part of
solve time. Constructed odd-order cases report predecessor search and
construction separately. The document also records compiler,
flags, build type, commit, OS/CPU metadata, worker count, search policy, seed,
timeouts, errors, invalid outputs, and the stdout policy. Search counts must
be stable across repeated runs. Prune and task counters are zero for the
current unpruned, single-threaded solver and reserve stable schema fields for
later work.
The runner writes its JSON report before returning a failure status if any mode
has no completed runs or produces an error or invalid output. Timeouts are
reported but do not fail a case when another repetition completed.
Normal `partridge_cpp` calls instantiate a compile-time counter-free solver.
Use `--measure-overhead` to run both counter-free and counted variants and
report their median difference:
```sh
python3 benchmarks/run.py --binary build-benchmark/partridge_benchmark \
--orders 8 --warmup 2 --repetitions 7 --timeout 60 --measure-overhead \
> overhead.json
```
Counter-free and counted warm-ups and measurements are interleaved. The first
mode alternates on each repetition, limiting systematic bias from temperature,
frequency scaling, and run order. Reported overhead is the difference between
the two independently summarized medians.
Do not use wall-clock thresholds as correctness checks. Keep the generated
JSON outside version control unless it is being deliberately added as a named
comparison baseline.
## D4 board symmetry
Every solution contains exactly one 1-by-1 square. Rotations and reflections
of the whole board preserve square sizes, multiplicities, and coverage, so the
unit square can select the orientation without assigning identities to any of
the repeated larger squares. For a board of width `W`, the solver accepts the
unit square only in the closed fundamental triangle
`x <= y <= floor((W - 1) / 2)`. Reflecting a cell toward the left edge,
swapping its coordinates if necessary, and reflecting toward the top edge
maps every D4 orbit into this triangle. Closed diagonal and midline
boundaries retain the smaller orbits of symmetric cells.
The check is made only when the skyline search is ready to place the unit
square, so it never rejects a partial state before the square's position is
decidable. The implementation remains compile-time counter-free in normal
solver calls and keeps the single-threaded deterministic search policy.
Instrumented runs count examined and rejected unit-square placements as prune
checks and hits.
Measurements used the issue #3 dirty working tree based on commit `74bde26`,
Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker, one warm-up, and
five sequential measured repetitions. All results passed the benchmark's
independent cell-coverage validator and counts were stable:
| Route | D4 constraint | Median solve (range) | Nodes | Prune hits |
| --- | --- | ---: | ---: | ---: |
| Order 8 public | enabled | 0.245 s (0.244-0.248 s) | 2,931,203 | 1,329,567 |
| Order 8 public | disabled | 0.679 s (0.659-0.709 s) | 7,735,369 | 0 |
| Order 9 direct | enabled | 1.705 s (1.671-1.740 s) | 16,231,918 | 6,755,171 |
| Order 9 direct | disabled | 4.542 s (4.482-4.662 s) | 45,840,266 | 0 |
Thus the constraint reduced nodes by 62% for order 8 and 65% for direct order
9; counted median solve time fell by 64% and 62%, respectively.
A public order-10 probe with the D4 constraint, ascending policy, and a
45-second per-process bound did not complete. Public order 11 first performs
that identical order-10 core search and only then adds its inexpensive odd
border, so adding either order to routine correctness tests would duplicate
the same unresolved bottleneck. They remain useful opt-in heavyweight
benchmark targets with explicit timeouts. A direct order-11 benchmark is a
different experiment: it bypasses the public odd construction and searches
the larger core itself, so it must not be presented as public order-11
performance.
## Smallest-valley skyline
The solver stores one filled height per board column instead of one value per
cell. Equal adjacent heights form conceptual vertical bars. A valley is a
maximal bar lower than both neighbours, with board edges treated as bars of
full board height. Each node scans for the smallest-width valley, breaking
ties by lower height and then leftmost position, and tries every available
square which fits at that valley's far-left edge.
This branching remains complete: the bottom-left cell of the selected valley
must be covered, a square covering it cannot begin to the left across the
taller neighbour, and it cannot extend beyond the equal-height run without
overlap or leaving an unreachable hole. Trying every fitting available size
therefore includes the placement used by every possible completion.
For board width `W` and order `n`, the skyline scan is `O(W)`. A node tries at
most `n` candidates and each placement or exact undo changes at most `n`
heights, giving `O(W + n^2)` local work. The skyline, multiplicities,
placements, and recursion stack use `O(W + n^2)` state, compared with the
former `O(W^2)` cell grid.
One-run exploratory measurements used the issue #4 dirty worktree at base
commit `0a7ce1e`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, one worker,
no warm-up, and a 15-second timeout. Every completed result passed the
independent benchmark validator:
| Order | Result | Ascending time | Ascending nodes | Descending time | Descending nodes |
| --- | --- | ---: | ---: | ---: | ---: |
| 6 | infeasible | 0.040 s | 659,598 | 0.039 s | 659,598 |
| 7 | infeasible | 3.103 s | 43,604,507 | 3.071 s | 43,604,507 |
| 8 | solution | 0.585 s | 7,735,369 | 0.941 s | 12,186,125 |
| 9 direct | solution | 3.831 s | 45,840,266 | timeout | unavailable |
The infeasible orders exhaust the same tree in either direction. Ascending
was selected as the default because it reaches the first order-8 solution with
36% fewer nodes and also completed direct order 9 within the timeout;
descending direct order 9 did not.
A direct ascending order-10 probe exceeded 20 seconds. Public order 10 is
also a direct search, and public order 11 first searches order 10 before using
odd-predecessor construction. Consequently neither 10 nor 11 is in the
routine correctness suite: doing so would test the same unresolved order-10
search bottleneck, while the existing route-boundary test still verifies that
11 selects construction. Revisit both sizes when order 10 completes within a
practical test budget.
## Candidate policy selection
Candidate ordering is a deterministic search policy and does not alter the
smallest-valley selection or set of placements tried. `ascending` tries
smaller available squares first and `descending` tries larger ones first.
`best-fit` first tries a square exactly as wide as the selected valley, because
that placement closes the valley without leaving a shelf remainder, then tries
the other sizes in ascending order. If no exact-width square fits, best-fit
and ascending are identical at that node.
The policy comparison used the issue #12 working tree based on commit
`d751d1b`, Apple Clang 21.0.0, `-O3 -DNDEBUG`, macOS arm64, and one worker.
Order 8 used two warm-ups and seven sequential measured repetitions; direct
order 9 used one warm-up and three measured repetitions. All completed
results passed independent validation and node counts were stable:
| Order | Policy | Counted median (range) | Counter-free median | Nodes |
| --- | --- | --- | --- | ---: |
| 8 | ascending | 0.808 s (0.7900.852 s) | 0.794 s | 7,735,369 |
| 8 | descending | 1.310 s (1.2791.449 s) | 1.268 s | 12,186,125 |
| 8 | best-fit | 0.817 s (0.8130.857 s) | 0.823 s | 7,679,349 |
| 9 direct | ascending | 5.522 s (5.4995.830 s) | not measured | 45,840,266 |
| 9 direct | best-fit | 5.651 s (5.5335.820 s) | not measured | 45,746,016 |
The earlier direct-order-9 descending probe exceeded its 15-second limit.
Ascending is retained as the stable single-threaded default because it had the
lowest measured median time to the first solution at both measured solvable
sizes. Best-fit's slightly smaller trees did not compensate for its policy
checks, while descending was substantially worse. Exhaustive infeasible
order-5 tests visit the same number of nodes under all three policies, which
checks that ordering does not affect completeness.
One policy therefore applies to the currently measured sizes 8 and 9. This
does not establish that ascending is optimal for order 10: a bounded best-fit
order-9 comparison changed the search tree by only 0.2%, so there was no
evidence that repeating the known long order-10/11 search would be useful.
Keep 10 and 11 as opt-in benchmark cases. Public order 11 is particularly
important to interpret correctly: it constructs from an order-10 search, so it
does not independently measure an odd-order candidate policy.
A worker portfolio was considered but not added. Running identical policies
duplicates the same deterministic traversal. Pairing ascending with best-fit
adds little diversity on the measured trees, and pairing ascending with
descending dedicates a worker to the consistently slower policy. Splitting a
shared frontier could avoid duplicated prefixes, but that is the parallel
frontier work tracked separately in issue #11. Seeded randomized ordering was
also rejected for now: the deterministic alternatives already select a clear
default, and there is no measurement showing that seed distributions would
improve time to first solution. The benchmark schema retains its nullable
seed field so a future evidence-backed randomized policy can report
reproducible runs without changing the format.
## Post-correctness baseline
This framework starts from commit `ce39d0a` after the rendering assertion fix
in #7 and completion check fix in #14. The earlier Apple M1 Release results in
`results.md` are approximately 1.76 seconds for order 8 and 158.69 seconds
elapsed for order 9; they predate the structured runner and do not contain
search counters.
The first clean structured baseline used commit `598667b`, Apple Clang 21.0.0
with `-O3 -DNDEBUG`, Apple arm64, one worker, one warm-up, and three measured
repetitions. The runner reported a clean working tree:
| Order | Result | Counted solve median (range) | Nodes | Placements | Backtracks |
| --- | --- | --- | ---: | ---: | ---: |
| 7 | infeasible | 3.453 s (3.4443.455 s) | 110,483,315 | 110,483,314 | 110,483,314 |
| 8 | solution | 1.817 s (1.8141.817 s) | 60,485,176 | 60,485,176 | 60,485,140 |
Counts were stable across repetitions. Interleaved counter-free medians were
3.435 seconds for order 7 and 1.805 seconds for order 8, giving counted
overheads of 0.53% and 0.64% respectively. Construction time was zero; median
independent validation and rendering times were each below 0.02 milliseconds.
Order 9 was not rerun for this initial baseline because the former direct
search took several minutes. The default suite includes it with a per-run
timeout.
## Odd construction comparison
The order-9 construction was measured from the issue 6 working tree based on
commit `ddf07e7`, using Apple Clang 21.0.0 with `-O3 -DNDEBUG`, macOS arm64,
one worker, one warm-up, and three measured repetitions. Counter-free and
counted runs were interleaved:
| Order | Mode | Search median (range) | Construction median | Nodes |
| --- | --- | --- | --- | ---: |
| 8 | counter-free | 1.773 s (1.7701.775 s) | 0 | 0 |
| 8 | counted | 1.849 s (1.8481.850 s) | 0 | 60,485,176 |
| 9 | counter-free | 1.778 s (1.7731.779 s) | 0.458 us | 0 |
| 9 | counted | 1.852 s (1.8451.912 s) | 0.416 us | 60,485,176 |
All runs completed with valid results and stable counters. The matching
order-8 and order-9 search counts demonstrate that the new path searches only
the predecessor. Compared with the recorded 158.69-second direct order-9
elapsed time in `results.md`, the 1.778-second counter-free median plus
construction is approximately 89 times faster. The benchmark working tree
was necessarily dirty with the issue 6 implementation.
New optimization issues should quote the exact JSON
environment, policy, median/spread, stable counters, and counted overhead from
this runner for both before and after revisions.
The separate optional CP-SAT reference benchmark and its model, memory, worker,
and timing report are documented in [CP_SAT_REFERENCE.md](./CP_SAT_REFERENCE.md).