Lua SDK for BenchGecko -- the data platform for comparing AI model benchmarks, estimating inference costs, and exploring performance across providers.
benchgecko provides clean, idiomatic Lua functions for working with LLM benchmark data. Models are plain tables, operations are pure functions, and everything works with standard Lua 5.1+ (including LuaJIT). Build comparison tools, cost calculators, model selectors, and leaderboard UIs for games, embedded systems, or any Lua environment.
The library provides:
- new_model() for constructing model tables with scores and pricing
- compare_models() for head-to-head analysis across shared categories
- estimate_cost() and estimate_monthly() for inference cost calculations
- model_tier() and filter_by_tier() for S/A/B/C/D tier classification
- rank_by_category() for leaderboard sorting across 9 benchmark dimensions
- best_value() for finding the most cost-effective model
- value_score() for computing performance-per-dollar ratios
- model_summary() for human-readable one-line descriptions
luarocks install benchgeckolocal bg = require("benchgecko")
-- Create models with builder-style chaining
local gpt4 = bg.new_model("gpt-4o", "OpenAI")
bg.set_context_window(gpt4, 128000)
bg.add_score(gpt4, "reasoning", 92.3)
bg.add_score(gpt4, "coding", 89.1)
bg.add_score(gpt4, "knowledge", 88.7)
bg.set_pricing(gpt4, 2.50, 10.00)
local claude = bg.new_model("claude-sonnet-4", "Anthropic")
bg.set_context_window(claude, 200000)
bg.add_score(claude, "reasoning", 94.1)
bg.add_score(claude, "coding", 93.7)
bg.add_score(claude, "knowledge", 91.2)
bg.set_pricing(claude, 3.00, 15.00)
-- Compare across shared categories
local result = bg.compare_models(gpt4, claude)
print("Winner: " .. result.winner.name)
print("GPT-4o wins: " .. #result.a_wins .. " categories")
print("Claude wins: " .. #result.b_wins .. " categories")
-- Estimate cost for a request
local cost = bg.estimate_cost(gpt4, 5000, 2000)
print(string.format("Request cost: $%.4f", cost))All setter functions return the model table, so you can chain them:
local bg = require("benchgecko")
local model = bg.set_pricing(
bg.add_score(
bg.add_score(
bg.set_context_window(
bg.new_model("gemini-2", "Google"), 1000000),
"reasoning", 91.5),
"coding", 88.3),
1.25, 5.00)
print(bg.model_summary(model))
-- gemini-2 (Google) [S-Tier] avg=89.9 value=14.4Models are classified into tiers based on average benchmark score:
| Tier | Average Score | Description |
|---|---|---|
| S | 90+ | Elite frontier models |
| A | 80-89 | Strong general-purpose models |
| B | 70-79 | Capable mid-range models |
| C | 60-69 | Budget or older generation |
| D | <60 | Entry-level or legacy |
print(bg.model_tier(gpt4)) -- "S"
print(bg.tier_description("S")) -- "Elite frontier models (90+)"
-- Filter by tier
local elite = bg.filter_by_tier(models, "S")
-- Rank by specific category
local leaderboard = bg.rank_by_category(models, "coding")
for _, entry in ipairs(leaderboard) do
print(entry.model.name, entry.score)
endAll cost functions return nil when pricing is missing, making them safe for pipelines:
-- Single request
local cost = bg.estimate_cost(gpt4, 10000, 5000)
print(string.format("Cost: $%.4f", cost))
-- Monthly projection
local monthly = bg.estimate_monthly(gpt4, 1000, 3000, 1000)
print(string.format("Monthly: $%.2f", monthly))
-- Value analysis
print(string.format("GPT-4o value: %.1f", bg.value_score(gpt4)))
print(string.format("Claude value: %.1f", bg.value_score(claude)))
local best = bg.best_value({gpt4, claude})
print("Best value: " .. best.name)The library tracks 9 benchmark dimensions aligned with BenchGecko:
| Category | Key | Typical Benchmarks |
|---|---|---|
| Reasoning | reasoning |
GSM8K, MATH, ARC |
| Coding | coding |
HumanEval, MBPP, SWE-bench |
| Knowledge | knowledge |
MMLU, HellaSwag |
| Instruction | instruction |
MT-Bench, AlpacaEval |
| Multilingual | multilingual |
MGSM, XLSum |
| Safety | safety |
TruthfulQA, BBQ |
| Long Context | long_context |
RULER, Needle-in-a-Haystack |
| Vision | vision |
MMMU, MathVista |
| Agentic | agentic |
WebArena, SWE-bench |
Works with Lua 5.1, 5.2, 5.3, 5.4, and LuaJIT. No external dependencies.
Benchmark data, model metadata, and pricing information are maintained by BenchGecko. Visit the platform for live leaderboards, interactive comparisons, and the full model database covering 300+ models across 50+ providers.
MIT