randyzwitch
## Problem Summing 90 expressions over one Int16 column (ClickBench q29: `sum(ResolutionWidth + i)` for i in 0..89) takes 325 ms against 17 ms for Polars and 9 ms for DuckDB, about 3.6 ms per expression for 1M rows. ## Cause (from the profiles and the code) For each of the 90 expressions: - `col + i` on Int16 computes each row through `_int_binary` and `_narrow` (`dataframe/expr_kernels.mojo`), which can raise, so the loop does one checked operation per row and cannot be vectorised; - the result is materialized as a full column, with validity built as a `List[Bool]` and packed (`_pack_bits`); - `sum` then copies the column to Int64 with another `List[Bool]` of validity (`_canonical` in `dataframe/aggregate.mojo`) before reducing (`Reducer._update`); - `select_exprs` (`dataframe/frame.mojo`) evaluates the expressions one after another, so 90 expressions are 90 sequential passes. ## How Polars and DuckDB do it - **Polars** does integer arithmetic with wrapping SIMD kernels (`crates/polars-compute/src/arithmetic/`), sums with a SIMD widening reduction (`crates/polars-compute/src/sum.rs`), and evaluates a projection's expressions in parallel on its thread pool (`par_iter` in `crates/polars-mem-engine/src/executors/projection_utils.rs`). - **DuckDB** checks overflow per vector with typed operators and sums narrow integers into a wider accumulator directly. ## Proposed solution 1. Checked integer arithmetic that computes a block of rows in a wider type with SIMD and checks the block's range once, raising only if any lane overflowed (same error, no per-row branch). 2. Reductions over narrow integers and Float32 read the native width and accumulate wide, without copying to Int64 first. 3. Evaluate independent expressions of one `select`/`with_columns` in parallel across workers when each is too small to split by rows. 4. Fuse `sum(col + literal)` style aggregate-over-elementwise into one pass without materializing the intermediate column. ## Profiled queries where this is at least 10% of the samples Found by profiling every benchmark query with `perf` on `main` at d10ff69: quick tier (H2O 1M rows, PDS-H scale 0.1, ClickBench 1M rows), 8 threads, Threadripper 3970X, runners built with `-O3 -g1`. **Share** is the fraction of the query run's CPU samples spent in this cause's code (it includes the one-time table load, so it slightly understates). Times are the quick tier's medians in ms. | Query | Share | mojo ms | Polars ms | DuckDB ms | vs fastest | |---|---:|---:|---:|---:|---:| | clickbench q29 | 86% | 324.6 | 17.3 | 9.2 | 35.34x | Held-out queries (pdsh, clickbench) are listed as the evidence that found this cause, not as targets (docs/benchmarks.md, rule 3). The fix must be the general mechanism described here, measured on its own and on the development suites. ## Acceptance criteria - `col + 1` on 1M rows of Int8, Int16, Int32 and Int64 is within 1.2 times Polars; overflow still raises with the same message. - `sum` over 1M Int16 rows is within 1.2 times Polars; a `select` of 50 independent reductions over one 1M-row column is within 1.5 times Polars. - Arithmetic and aggregation tests and the oracle pass; no regression on the development suites. 🤖 Generated with [Claude Code](https://claude.com/claude-code)