agea/cruciverba

โ˜… 0Forks 0JavaScriptGitHub โ†—Compare

Project website โ†—

README

Cruciverba ๐Ÿ‡ฎ๐Ÿ‡น

A generator and player for Italian crosswords, running entirely in the browser. No backend, no build step, no external dependencies โ€” just HTML and vanilla JavaScript. Installable as a PWA and fully playable offline, tuned for iPad and desktop.

โ–ถ๏ธ Play it: https://agea.github.io/cruciverba/


โœจ Features

  • Automatic generation of dense, Settimana Enigmisticaโ€“style grids with black squares, with square and landscape presets from 5ร—5 up to 25ร—13.
  • Interactive solving: pick a clue, type with the physical or on-screen keyboard, with highlighting of the active word and crossing cell.
  • Helpers: check wrong letters, reveal a cell or a whole word, clear.
  • Persistence: the game state is saved in localStorage, so you can pick up where you left off.
  • Game timer and automatic completion detection.
  • "Pencil-on-paper" design, touch-friendly and offline-first (PWA with a service worker).

๐Ÿ—๏ธ Architecture

Three decoupled components:

  1. Clue database โ€” cruciverba_db.json: pure data, kept separate from logic.
  2. Grid generator โ€” gen_dense.js: the grid-construction algorithm, run inside an adaptive pool of at most two Web Workers so the UI never blocks.
  3. UI โ€” the HTML app: grid rendering, input, helpers, persistence.

The main worker starts immediately. On 13ร—13 grids and rectangular grids from 17ร—11 upward, devices reporting at least four logical processors can start a second independent worker after an adaptive delay (1.6 s for 13ร—13, 350โ€“700 ms for large rectangles); the first valid result wins and the other worker is terminated. gen_dense.js is precached by the service worker together with the database for offline PWA use.


๐Ÿ—‚๏ธ The database

The source database lives in voci/: 26 CSV files, one per initial letter, with 36,738 definition rows. node builddb.js turns them into cruciverba_db.json, a compact JSON array of ["SOLUTION", clue] entries, where clue is either a string (one definition) or an array of strings (several definitions for the same solution). When a word has multiple clues, the generator picks one at random per puzzle, so the same answer can be asked differently from one grid to the next.

The clue database and definitions are licensed separately from the software: see LICENSE-CONTENT.md.

  • 22,157 solutions / 36,738 clues / 11,102 multi-clue solutions. Solutions are uppercase, letters Aโ€“Z only (accents and spaces stripped at build time).
  • Length distribution is deliberately skewed toward short words, which feed the dense crossings:
Letters 2 3 4 5 6 7 8 9 10 11 12 13 14+
Solutions 190 324 1337 3023 3993 4156 3246 2355 1603 956 511 260 203

The 14+ bucket is made of 111 words of length 14, 61 words of length 15, 22 words of length 16, 4 words of length 17, 3 words of length 18 and 2 words of length 19. Short slots (2โ€“3 letters) lean on the classic Italian-puzzle style: initialism, car plates, musical notes and chemical symbols โ€” and, being the most frequent, often carry several alternative clues.

Extending the database

  1. Add rows to the right file in voci/ โ€” one file per initial letter (voci/A.csv, voci/B.csv, โ€ฆ), each under the header soluzione,definizione. A word goes in the file of its first letter. Wrap a clue in double quotes if it contains a comma. To give a word more than one clue, add several rows with the same solution and different definitions.
  2. Run node builddb.js. It reads every voci/*.csv, normalizes (NFD โ†’ uppercase โ†’ Aโ€“Z only), groups every distinct definition under its solution (only exact (solution, clue) duplicates are dropped), drops invalid entries, sorts, and rewrites cruciverba_db.json. (A single legacy voci.csv is still accepted as a fallback.)

๐Ÿ’ก Solutions shorter than 2 letters or without a clue are discarded automatically. Duplicate solutions are kept as one JSON entry with all distinct clues attached.


โš™๏ธ The dense generator (gen_dense.js)

Produces dense grids in the Italian style: a filled rectangular grid with black squares, where every white run of length โ‰ฅ 2 (across or down) is a database word with its clue.

Pipeline:

  1. Word bank โ€” the DB is indexed by length and by (position, letter), to quickly fetch candidates for a partially filled slot.
  2. Black-square pattern โ€” randomly generated with controlled density, then carved around a near-central crossing between one long across answer and one long down answer. Over-long non-seed white runs are split (maxRun) to keep the fill tractable; isolated white cells are removed; black squares are normalized to avoid 2ร—2 black blocks and black runs longer than 3 cells. Large rectangles use lower preferred starting densities while retaining a 20โ€“21% hard ceiling as a safety net during fallback. Valid patterns are ranked using seed-domain breadth, density and short-slot penalties: batches contain 12 patterns on 13ร—13 and 24 on large rectangles, so the most promising are filled first without permanently discarding the others.
  3. Slot extraction โ€” all white runs โ‰ฅ 2, across and down, plus a cell โ†’ slot map.
  4. Filling (backtracking) โ€” the central crossing is pre-filled first, then most-constrained-slot selection (propagation from already-filled slots + a static seed on the most-crossed ones), forward-checking on crossings, no repeated words. Candidate domains are narrowed and restored incrementally along the search path instead of being rebuilt at every node. Large low-density grids may contain answers that differ by one letter; exact duplicates remain forbidden.
  5. Fallback cascade โ€” if a configuration can't be completed, it retries with different starting patterns and search budgets before falling back to a smaller grid. The 17ร—11 starts directly from the 10โ€“12% profiles that most often succeed; 21ร—13 and 25ร—13 prioritize maxRun=6, avoiding the long maxRun=7 dead ends found by profiling. Every plan still respects the density ceiling.

During generation the worker emits throttled progress updates by phase, attempted patterns and backtracking activity. The displayed percentage is intentionally conservative because backtracking progress is not linear.

Preset sizes

Shape Sizes
Square 5ร—5, 7ร—7, 9ร—9, 11ร—11, 13ร—13
Landscape 11ร—7, 13ร—9, 17ร—11, 21ร—13, 25ร—13

Smoke-test timings (Node, current DB, fixed seeds, representative preset-plan parameters, zero clueless words):

Requested Actual Words Black Longest 7+ words Time
5ร—5 5ร—5 10 1/25 (4.0%) 5 0 ~0.04 s
7ร—7 7ร—7 22 10/49 (20.4%) 5 0 ~0.02 s
9ร—9 9ร—9 29 15/81 (18.5%) 7 7 ~0.26 s
11ร—11 11ร—11 45 26/121 (21.5%) 7 7 ~3.75 s
13ร—13 13ร—13 60 27/169 (16.0%) 8 14 ~7.31 s
11ร—7 11ร—7 29 11/77 (14.3%) 7 8 ~4.45 s
13ร—9 13ร—9 38 27/117 (23.1%) 8 9 ~0.42 s
17ร—11 17ร—11 65 35/187 (18.7%) 12 17 ~5.14 s
21ร—13 21ร—13 94 53/273 (19.4%) 12 5 ~0.34 s
25ร—13 25ร—13 113 62/325 (19.1%) 12 5 ~7.68 s

Output structure

{
  width, height,
  cells: [[ null | { ch, num }, ... ], ...],  // null = black square
  across: [{ num, clue, answer, row, col, len }, ...],
  down:   [{ num, clue, answer, row, col, len }, ...],
  wordCount,
  ghosts  // words without a clue (expected: 0)
}

๐Ÿ“ Project structure

File Role
voci/ Database source, split by initial (A.csv โ€ฆ Z.csv)
builddb.js Build script: CSV โ†’ JSON
cruciverba_db.json Generated database, ignored by git locally and rebuilt in CI before deploy
gen_dense.js Dense generator (Web Worker)
index.html Playable PWA app
sw.js ยท manifest.webmanifest ยท icons/ PWA assets, offline cache and update flow
version.json Build metadata used by the deployed app to detect updates

๐Ÿš€ Development

No toolchain required.

# rebuild the database after editing any voci/*.csv
node builddb.js

# refresh README database statistics after editing any voci/*.csv
node scripts/update-readme-stats.js

# serve locally (a service worker needs an HTTP origin)
python3 -m http.server 8000
# then open http://localhost:8000

Enable the included pre-commit hook with:

git config core.hooksPath .githooks

The hook runs node scripts/update-readme-stats.js and stages README.md if the statistics changed.

Pushing to main triggers the GitHub Actions workflow, which rebuilds the database, writes deploy-time version.json metadata from the commit SHA and UTC build time, and deploys the static site to GitHub Pages.

Offline use

Full offline use requires opening the app at least once from GitHub Pages (or any HTTP origin), so the service worker can cache index.html, gen_dense.js, cruciverba_db.json, README.md and the other assets. The service worker also refreshes version.json and cruciverba_db.json from the network when available, then shows an in-app update prompt for a newly deployed version. Opening directly via file:// is not the main target, because browsers restrict fetch and service workers outside an HTTP/HTTPS origin.


License

  • Software code: GNU General Public License v3.0 only. See LICENSE.
  • Crossword entries, clue definitions and generated database content: Creative Commons Attribution-NonCommercial 4.0 International. See LICENSE-CONTENT.md.

Contributors

agea

Issues