Skip to content

[DO NOT MERGE] Add SIMD primitive comparison kernels - #9587

Draft
connortsui20 wants to merge 1 commit into
developfrom
ct/primitive-comparison-simd
Draft

[DO NOT MERGE] Add SIMD primitive comparison kernels#9587
connortsui20 wants to merge 1 commit into
developfrom
ct/primitive-comparison-simd

Conversation

@connortsui20

@connortsui20 connortsui20 commented Aug 24, 2026

Copy link
Copy Markdown
Member

Summary

Moves the explicit AVX2 and AVX-512 primitive comparison kernels from #9547 into the handwritten comparison path. This leaves the RowFn work in #9547 and #9548 untouched while preserving the existing portable lane-kernel fallback on non-x86 targets and x86-64 CPUs without AVX2.

Top of a three-PR stack: #9598 adds the comparison benchmarks and their #[cpu_features] tags, #9599 sweeps the same tag across the rest of the numeric arithmetic and comparison microbenchmarks, and this PR adds the kernels. The benchmarks land first so the walltime legs record the portable baseline on each feature set, and this PR's CodSpeed report compares kernels against it rather than introducing both at once.

Changes

  • Adds full-word SIMD kernels for all primitive types, comparison operators, and array/constant orientations, with scalar tail handling, split by lane width.
  • Adds boundary and nullable coverage.
  • Confirms the intended x86-64 codegen by cross-compiling for macOS: AVX2 uses packed ymm masks and AVX-512 uses zmm/k masks with direct bitmap stores. This environment cannot provide x86 runtime timings.

@connortsui20 connortsui20 added the changelog/performance A performance improvement label Aug 24, 2026
@codspeed-hq

codspeed-hq Bot commented Aug 24, 2026

Copy link
Copy Markdown

Merging this PR will regress 1 benchmark

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

⚠️ Different runtime environments detected

Some benchmarks with significant performance changes were compared across different runtime environments,
which may affect the accuracy of the results.

Open the report in CodSpeed to investigate

⚡ 6 improved benchmarks
❌ 1 regressed benchmark
✅ 1934 untouched benchmarks
🆕 156 new benchmarks
⏩ 106 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
WallTime compare_int_constant_left_avx2 3 µs 3.4 µs -10.88%
WallTime compare_u8_avx2 3.4 µs 1.7 µs ×2
WallTime compare_f32_avx2 4.6 µs 3.1 µs +48.27%
WallTime compare_u8_avx512 2.3 µs 1.6 µs +46.71%
WallTime compare_f32_avx512 3.6 µs 2.7 µs +37.02%
Simulation search_index_above_max_chunked 608.7 µs 549.4 µs +10.79%
Simulation search_index_in_range_chunked 610.3 µs 550.9 µs +10.77%
🆕 WallTime add_i32_nonnull_neon N/A 7.7 µs N/A
🆕 WallTime add_i64_constant_neon N/A 9.8 µs N/A
🆕 WallTime add_i64_nonnull_neon N/A 9.8 µs N/A
🆕 WallTime add_i64_nullable_neon N/A 11.4 µs N/A
🆕 WallTime add_shapes_neon[(128, ConstantPerRow)] N/A 3.1 µs N/A
🆕 WallTime add_shapes_neon[(128, PerRowConstant)] N/A 3.1 µs N/A
🆕 WallTime add_shapes_neon[(128, PerRowNullableConstant)] N/A 3.7 µs N/A
🆕 WallTime add_shapes_neon[(128, PerRowPerRow)] N/A 1.9 µs N/A
🆕 WallTime add_shapes_neon[(16384, ConstantPerRow)] N/A 9.8 µs N/A
🆕 WallTime add_shapes_neon[(16384, PerRowConstant)] N/A 9.9 µs N/A
🆕 WallTime add_shapes_neon[(16384, PerRowNullableConstant)] N/A 10.5 µs N/A
🆕 WallTime add_shapes_neon[(16384, PerRowPerRow)] N/A 9.8 µs N/A
🆕 WallTime add_u32_nonnull_neon N/A 6.6 µs N/A
... ... ... ... ... ...

ℹ️ Only the first 20 benchmarks are displayed. Go to the app to view all benchmarks.

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing ct/primitive-comparison-simd (1163162) with develop (0ef496c)2

Open in CodSpeed

Footnotes

  1. 106 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports.

  2. No successful run was found on develop (17bd9e2) during the generation of this report, so 0ef496c was used instead as the comparison base. There might be some changes unrelated to this pull request in this report.

@connortsui20
connortsui20 force-pushed the ct/primitive-comparison-simd branch from 3e39406 to 931772a Compare August 24, 2026 21:34
@connortsui20
connortsui20 changed the base branch from develop to ct/compare-bench-cpu-features August 24, 2026 21:34
@connortsui20
connortsui20 force-pushed the ct/primitive-comparison-simd branch from 931772a to c0148bd Compare August 24, 2026 21:44
@connortsui20
connortsui20 changed the base branch from ct/compare-bench-cpu-features to ct/cpu-features-numeric-benches August 24, 2026 21:44
@connortsui20 connortsui20 changed the title Add SIMD primitive comparison kernels [DO NOT MERGE] Add SIMD primitive comparison kernels Aug 24, 2026
connortsui20 added a commit that referenced this pull request Aug 24, 2026
## Summary

The measurement half of #9587, split out so the numbers land before the
kernels do.

Bottom of a three-PR stack: this PR, then #9599, then #9587. Tagging
these with `#[cpu_features]` first means the walltime legs record the
portable lane kernel's throughput on `avx2`, `avx512`, and `neon` metal
as a baseline series. The kernel PR then reports against it rather than
introducing both the benchmark and the thing it measures in one diff.

## Changes

- Adds the primitive comparison cases a hand-written SIMD kernel would
have to beat: constant on the left, `u8`, `u64`, and `f32`.
- Tags those four with `#[cpu_features]`, so each walltime leg measures
them under its own build flags instead of in simulation. The existing
cases are untouched and keep their simulation series.
- `bench_compare` now carries an `ItemsCount`, so the report reads as
throughput rather than a time that only means something next to another
run over the same array length.
- Adds module docs recording why these four are tagged and the rest are
not.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
@connortsui20
connortsui20 force-pushed the ct/primitive-comparison-simd branch from c0148bd to 46ad4bb Compare August 24, 2026 23:30
Base automatically changed from ct/cpu-features-numeric-benches to develop August 25, 2026 01:35
connortsui20 added a commit that referenced this pull request Aug 25, 2026
…9599)

## Summary

Sweeps `#[cpu_features]` across the microbenchmarks that clearly earn
it: the binary numeric arithmetic and comparison kernels. Those are
portable lane loops — the source is identical on every target and the
vector width the compiler picks comes from the build flags — which is
the case the attribute exists for. Measuring them in simulation under
one fixed `+avx2` build hides the only variable that matters.

Middle of a three-PR stack: #9598, then this PR, then #9587. It carries
no kernel changes of its own — everything here is a benchmark attribute
— so it can be reordered or rebased onto `develop` without touching the
other two.

## Changes

Tagged:

- `binary_ops`: the primitive arithmetic cases (`add_*`, `subtract_*`,
`multiply_*`, `mul_*`, `div_i64_*`, `sub_i64_constant`, and the three
`*_shapes` matrices) and the two primitive comparison cases
(`eq_i64_constant`, `lt_i64_nullable`).
- `compare`: `compare_int`, `compare_int_nullable`,
`compare_int_constant`, `compare_int_eq`, `compare_float`.
- `scalar_subtract`.
- `lane_kernels`: `lanezip_checked_add_u32` and its
`arrow_checked_add_u32` baseline. The baseline is tagged too — comparing
the two is only meaningful under the same build flags.

Left in simulation, with the reasoning recorded in each file's module
docs:

- Decimal arithmetic and comparison: `i128` widening and per-lane
rescaling, not something a wider vector register decides.
- Boolean `and`/`or`: already word-at-a-time over a bitmap.
- String and struct comparison: dominated by view chasing and per-field
dispatch.
- Casts in `lane_kernels`: vectorization-sensitive, but out of scope
here.

Also left alone: the `between` benchmarks in `vortex-fastlanes`
(`new_raw_prim_test_between` is a raw-primitive comparison kernel and
does qualify) and the bit-packed comparison matrices. Both are `types
=`/`consts =` parameterized, which `#[cpu_features]` has no coverage for
yet, and both would fan out to dozens of walltime series. Worth a
follow-up rather than a guess in this PR.

Note that tagging moves a benchmark out of the sharded simulation job,
so these series restart on the walltime legs instead of continuing their
simulation history.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
Moves the explicit AVX2 and AVX-512 primitive comparison kernels from #9547 into
the handwritten comparison path, leaving the RowFn work in #9547 and #9548
untouched. Non-x86 targets and x86-64 CPUs without AVX2 keep the portable
lane-kernel fallback.

Full-word kernels cover every primitive type, comparison operator, and
array/constant orientation, with scalar tail handling, split by lane width so
each width reads on its own. Boundary and nullable coverage comes with them.

Signed-off-by: Connor Tsui <connor.tsui20@gmail.com>
@connortsui20
connortsui20 force-pushed the ct/primitive-comparison-simd branch from 46ad4bb to 1163162 Compare August 25, 2026 01:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

changelog/performance A performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant