[DO NOT MERGE] Add SIMD primitive comparison kernels - #9587
Conversation
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | compare_int_constant_left_avx2 |
3 µs | 3.4 µs | -10.88% |
| ⚡ | WallTime | compare_u8_avx2 |
3.4 µs | 1.7 µs | ×2 |
| ⚡ | WallTime | compare_f32_avx2 |
4.6 µs | 3.1 µs | +48.27% |
| ⚡ | WallTime | compare_u8_avx512 |
2.3 µs | 1.6 µs | +46.71% |
| ⚡ | WallTime | compare_f32_avx512 |
3.6 µs | 2.7 µs | +37.02% |
| ⚡ | Simulation | search_index_above_max_chunked |
608.7 µs | 549.4 µs | +10.79% |
| ⚡ | Simulation | search_index_in_range_chunked |
610.3 µs | 550.9 µs | +10.77% |
| 🆕 | WallTime | add_i32_nonnull_neon |
N/A | 7.7 µs | N/A |
| 🆕 | WallTime | add_i64_constant_neon |
N/A | 9.8 µs | N/A |
| 🆕 | WallTime | add_i64_nonnull_neon |
N/A | 9.8 µs | N/A |
| 🆕 | WallTime | add_i64_nullable_neon |
N/A | 11.4 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(128, ConstantPerRow)] |
N/A | 3.1 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(128, PerRowConstant)] |
N/A | 3.1 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(128, PerRowNullableConstant)] |
N/A | 3.7 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(128, PerRowPerRow)] |
N/A | 1.9 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(16384, ConstantPerRow)] |
N/A | 9.8 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(16384, PerRowConstant)] |
N/A | 9.9 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(16384, PerRowNullableConstant)] |
N/A | 10.5 µs | N/A |
| 🆕 | WallTime | add_shapes_neon[(16384, PerRowPerRow)] |
N/A | 9.8 µs | N/A |
| 🆕 | WallTime | add_u32_nonnull_neon |
N/A | 6.6 µs | N/A |
| ... | ... | ... | ... | ... | ... |
ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ct/primitive-comparison-simd (1163162) with develop (0ef496c)2
Footnotes
-
106 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
No successful run was found on
develop(17bd9e2) during the generation of this report, so 0ef496c was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩
3e39406 to
931772a
Compare
931772a to
c0148bd
Compare
## Summary The measurement half of #9587, split out so the numbers land before the kernels do. Bottom of a three-PR stack: this PR, then #9599, then #9587. Tagging these with `#[cpu_features]` first means the walltime legs record the portable lane kernel's throughput on `avx2`, `avx512`, and `neon` metal as a baseline series. The kernel PR then reports against it rather than introducing both the benchmark and the thing it measures in one diff. ## Changes - Adds the primitive comparison cases a hand-written SIMD kernel would have to beat: constant on the left, `u8`, `u64`, and `f32`. - Tags those four with `#[cpu_features]`, so each walltime leg measures them under its own build flags instead of in simulation. The existing cases are untouched and keep their simulation series. - `bench_compare` now carries an `ItemsCount`, so the report reads as throughput rather than a time that only means something next to another run over the same array length. - Adds module docs recording why these four are tagged and the rest are not. Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
c0148bd to
46ad4bb
Compare
…9599) ## Summary Sweeps `#[cpu_features]` across the microbenchmarks that clearly earn it: the binary numeric arithmetic and comparison kernels. Those are portable lane loops — the source is identical on every target and the vector width the compiler picks comes from the build flags — which is the case the attribute exists for. Measuring them in simulation under one fixed `+avx2` build hides the only variable that matters. Middle of a three-PR stack: #9598, then this PR, then #9587. It carries no kernel changes of its own — everything here is a benchmark attribute — so it can be reordered or rebased onto `develop` without touching the other two. ## Changes Tagged: - `binary_ops`: the primitive arithmetic cases (`add_*`, `subtract_*`, `multiply_*`, `mul_*`, `div_i64_*`, `sub_i64_constant`, and the three `*_shapes` matrices) and the two primitive comparison cases (`eq_i64_constant`, `lt_i64_nullable`). - `compare`: `compare_int`, `compare_int_nullable`, `compare_int_constant`, `compare_int_eq`, `compare_float`. - `scalar_subtract`. - `lane_kernels`: `lanezip_checked_add_u32` and its `arrow_checked_add_u32` baseline. The baseline is tagged too — comparing the two is only meaningful under the same build flags. Left in simulation, with the reasoning recorded in each file's module docs: - Decimal arithmetic and comparison: `i128` widening and per-lane rescaling, not something a wider vector register decides. - Boolean `and`/`or`: already word-at-a-time over a bitmap. - String and struct comparison: dominated by view chasing and per-field dispatch. - Casts in `lane_kernels`: vectorization-sensitive, but out of scope here. Also left alone: the `between` benchmarks in `vortex-fastlanes` (`new_raw_prim_test_between` is a raw-primitive comparison kernel and does qualify) and the bit-packed comparison matrices. Both are `types =`/`consts =` parameterized, which `#[cpu_features]` has no coverage for yet, and both would fan out to dozens of walltime series. Worth a follow-up rather than a guess in this PR. Note that tagging moves a benchmark out of the sharded simulation job, so these series restart on the walltime legs instead of continuing their simulation history. Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Moves the explicit AVX2 and AVX-512 primitive comparison kernels from #9547 into the handwritten comparison path, leaving the RowFn work in #9547 and #9548 untouched. Non-x86 targets and x86-64 CPUs without AVX2 keep the portable lane-kernel fallback. Full-word kernels cover every primitive type, comparison operator, and array/constant orientation, with scalar tail handling, split by lane width so each width reads on its own. Boundary and nullable coverage comes with them. Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
46ad4bb to
1163162
Compare
Summary
Moves the explicit AVX2 and AVX-512 primitive comparison kernels from #9547 into the handwritten comparison path. This leaves the RowFn work in #9547 and #9548 untouched while preserving the existing portable lane-kernel fallback on non-x86 targets and x86-64 CPUs without AVX2.
Top of a three-PR stack: #9598 adds the comparison benchmarks and their
#[cpu_features]tags, #9599 sweeps the same tag across the rest of the numeric arithmetic and comparison microbenchmarks, and this PR adds the kernels. The benchmarks land first so the walltime legs record the portable baseline on each feature set, and this PR's CodSpeed report compares kernels against it rather than introducing both at once.Changes
ymmmasks and AVX-512 useszmm/kmasks with direct bitmap stores. This environment cannot provide x86 runtime timings.