BenchGecko/benchgecko-lua

Official lua SDK for the BenchGecko API. Compare AI models, benchmarks, and pricing.

★ 2Forks 0LuaGitHub ↗Compare

README

benchgecko

Lua SDK for BenchGecko -- the data platform for comparing AI model benchmarks, estimating inference costs, and exploring performance across providers.

Overview

benchgecko provides clean, idiomatic Lua functions for working with LLM benchmark data. Models are plain tables, operations are pure functions, and everything works with standard Lua 5.1+ (including LuaJIT). Build comparison tools, cost calculators, model selectors, and leaderboard UIs for games, embedded systems, or any Lua environment.

The library provides:

  • new_model() for constructing model tables with scores and pricing
  • compare_models() for head-to-head analysis across shared categories
  • estimate_cost() and estimate_monthly() for inference cost calculations
  • model_tier() and filter_by_tier() for S/A/B/C/D tier classification
  • rank_by_category() for leaderboard sorting across 9 benchmark dimensions
  • best_value() for finding the most cost-effective model
  • value_score() for computing performance-per-dollar ratios
  • model_summary() for human-readable one-line descriptions

Installation

luarocks install benchgecko

Quick Start

local bg = require("benchgecko")

-- Create models with builder-style chaining
local gpt4 = bg.new_model("gpt-4o", "OpenAI")
bg.set_context_window(gpt4, 128000)
bg.add_score(gpt4, "reasoning", 92.3)
bg.add_score(gpt4, "coding", 89.1)
bg.add_score(gpt4, "knowledge", 88.7)
bg.set_pricing(gpt4, 2.50, 10.00)

local claude = bg.new_model("claude-sonnet-4", "Anthropic")
bg.set_context_window(claude, 200000)
bg.add_score(claude, "reasoning", 94.1)
bg.add_score(claude, "coding", 93.7)
bg.add_score(claude, "knowledge", 91.2)
bg.set_pricing(claude, 3.00, 15.00)

-- Compare across shared categories
local result = bg.compare_models(gpt4, claude)
print("Winner: " .. result.winner.name)
print("GPT-4o wins: " .. #result.a_wins .. " categories")
print("Claude wins: " .. #result.b_wins .. " categories")

-- Estimate cost for a request
local cost = bg.estimate_cost(gpt4, 5000, 2000)
print(string.format("Request cost: $%.4f", cost))

Chained Construction

All setter functions return the model table, so you can chain them:

local bg = require("benchgecko")

local model = bg.set_pricing(
    bg.add_score(
        bg.add_score(
            bg.set_context_window(
                bg.new_model("gemini-2", "Google"), 1000000),
            "reasoning", 91.5),
        "coding", 88.3),
    1.25, 5.00)

print(bg.model_summary(model))
-- gemini-2 (Google) [S-Tier] avg=89.9 value=14.4

Tier Classification

Models are classified into tiers based on average benchmark score:

Tier Average Score Description
S 90+ Elite frontier models
A 80-89 Strong general-purpose models
B 70-79 Capable mid-range models
C 60-69 Budget or older generation
D <60 Entry-level or legacy
print(bg.model_tier(gpt4))           -- "S"
print(bg.tier_description("S"))       -- "Elite frontier models (90+)"

-- Filter by tier
local elite = bg.filter_by_tier(models, "S")

-- Rank by specific category
local leaderboard = bg.rank_by_category(models, "coding")
for _, entry in ipairs(leaderboard) do
    print(entry.model.name, entry.score)
end

Cost Estimation

All cost functions return nil when pricing is missing, making them safe for pipelines:

-- Single request
local cost = bg.estimate_cost(gpt4, 10000, 5000)
print(string.format("Cost: $%.4f", cost))

-- Monthly projection
local monthly = bg.estimate_monthly(gpt4, 1000, 3000, 1000)
print(string.format("Monthly: $%.2f", monthly))

-- Value analysis
print(string.format("GPT-4o value: %.1f", bg.value_score(gpt4)))
print(string.format("Claude value: %.1f", bg.value_score(claude)))

local best = bg.best_value({gpt4, claude})
print("Best value: " .. best.name)

Benchmark Categories

The library tracks 9 benchmark dimensions aligned with BenchGecko:

Category Key Typical Benchmarks
Reasoning reasoning GSM8K, MATH, ARC
Coding coding HumanEval, MBPP, SWE-bench
Knowledge knowledge MMLU, HellaSwag
Instruction instruction MT-Bench, AlpacaEval
Multilingual multilingual MGSM, XLSum
Safety safety TruthfulQA, BBQ
Long Context long_context RULER, Needle-in-a-Haystack
Vision vision MMMU, MathVista
Agentic agentic WebArena, SWE-bench

Compatibility

Works with Lua 5.1, 5.2, 5.3, 5.4, and LuaJIT. No external dependencies.

Data Source

Benchmark data, model metadata, and pricing information are maintained by BenchGecko. Visit the platform for live leaderboards, interactive comparisons, and the full model database covering 300+ models across 50+ providers.

License

MIT

Contributors

BenchGecko

Issues