map_with_hit / each_with_hit pair records with the wrong hits when a result's DB record is missing

#1099 · open · 1 comments

View on GitHub ↗

Alan-Marx

### Summary When iterating search results with `Response::Records#map_with_hit` (or `#each_with_hit`), a record can be silently paired with the **wrong** hit. The record ends up decorated with a *different* result's `highlight` / `_score` / `sort`, and the trailing hit is dropped. No exception is raised — the results just come back subtly wrong. ### Root cause `map_with_hit` / `each_with_hit` pair the records and the hits **positionally** (`elasticsearch-model/lib/elasticsearch/model/response/records.rb`): ```ruby def map_with_hit(&block) records.to_a.zip(results).map(&block) end ``` This assumes `records.length == results.length`. That assumption breaks whenever a hit's DB record is missing (eventual consistency: the document is still indexed but the row was deleted / a reindex gap): - **Multiple adapter** (`adapters/multiple.rb`): `#records` fetches via `where(id:)` and ends with `records.compact`, so a missing row is dropped from the array. - **ActiveRecord adapter** (`adapters/active_record.rb`): `where(klass.primary_key => ids)` simply returns fewer rows than there are hits. Either way `records` is now shorter than `results`, and `zip` (called on the shorter array) shifts every pair after the gap: ``` hits (results): [H0, H1(missing record), H2, H3] # length 4 records: [R0, R2, R3] # length 3 records.to_a.zip(results) => [[R0, H0], [R2, H1], [R3, H2]] # correct R2+H1! R3+H2! (H3 dropped entirely) ``` So `R2` is yielded with `H1`'s highlight/score, `R3` with `H2`'s, etc. ### Reproduction Any search returning >= 2 hits where one hit's DB record no longer exists: ```ruby # ES returns hits for ids [1, 2, 3], but row #2 has been deleted. response.records.map_with_hit { |record, hit| [record.id, hit._id] } # => [[1, "1"], [3, "2"]] # record 3 paired with hit "2"; hit "3" dropped # expected: [[1, "1"], [3, "3"]] (hit "2" skipped, since it has no record) ``` This surfaced in production as search rows showing the **wrong title** — the displayed highlight came from the hit, the link/metadata from the record, and the two belonged to different documents. ### Affected versions Observed on `elasticsearch-model` 8.0.1; the mismatch exists for any version since missing records became tolerated. ### Prior art / not a duplicate - #369 made the Multiple adapter *tolerate* missing records (`where(id:)` + `.compact`) so `find` would stop raising — but did not update `map_with_hit` / `each_with_hit` to account for the now-shorter array. - #99 fixed a related **ordering** mismatch (same set of records, wrong order) via the ActiveRecord adapter's re-sort to hit order. This issue is the orthogonal **length** mismatch (a record is *absent*), which re-sorting does not address — a shorter array still zips out of alignment. ### Suggested fixes 1. Pair by `hit._id` rather than positionally — look each hit's record up by id (and type, for multi-model), skipping hits whose record is absent. 2. Or stop compacting / keep `nil` placeholders so positions stay aligned with `results`, and have `map_with_hit` / `each_with_hit` skip the `nil`s.

Comments

Alan-Marx

### Example fix The core change is to pair each record with its hit **by document id** instead of positionally, in `Response::Records` (`lib/elasticsearch/model/response/records.rb`): ```ruby # before def map_with_hit(&block) records.to_a.zip(results).map(&block) end def each_with_hit(&block) records.to_a.zip(results).each(&block) end # after — pair by _id, so a missing record no longer shifts the rest def map_with_hit(&block) hits_by_id = results.index_by { |hit| hit._id.to_s } records.to_a.map { |record| block.call(record, hits_by_id[record.id.to_s]) } end def each_with_hit(&block) hits_by_id = results.index_by { |hit| hit._id.to_s } records.to_a.each { |record| block.call(record, hits_by_id[record.id.to_s]) } end ``` Notes: - Both adapters already return `records` in hit order (the ActiveRecord adapter re-sorts to hit order; the Multiple adapter maps over hits then `compact`s), so mapping over `records` preserves result order. - Iterating `records` (not `results`) means hits whose record is missing are simply not visited — no positional shift, and no need to special-case them. - `index_by` is from ActiveSupport, which `elasticsearch-model` already depends on; a plain-Ruby `each_with_object({}) { |hit, h| h[hit._id.to_s] = hit }` works too. **Multi-model caveat.** With the Multiple adapter, two models can share an `id` value (e.g. integer primary keys), so a bare-`_id` key can collide across types. For that case the key should include the type — e.g. `[__type_for_hit(hit), hit._id.to_s]` on the hit side and `[record.class, record.id.to_s]` on the record side (or `[hit._index, hit._id]`). Apps whose searchable models all use UUID primary keys don't hit this, since the id is already globally unique. For reference, this is the workaround we shipped in our app (UUID ids, so bare-id keying is safe): an additive `map_with_hit_by_id` rather than overriding the gem method, so nothing relies on patched gem behavior.