Base language model benchmark

BananaMindBench Leaderboard

Text-completion performance on 350 BananaMind Base Bench 1.1 examples, ranked by fixed-item Overall Elo.

Maximum parameter size for submissions: 250M.

350 examples 7 categories 4 continuations per example Mean conditional log-likelihood Higher Elo is better

Model advisor

Not sure where to start?

Choose your use case, capability priorities, and parameter range.

Organization
All organizations
Model size

Relative model ranking

Benchmark cards

Choose Overall or any benchmark category to rank the cards. Scores are normalized to the active size class. Models with the best selected-metric Elo in their size-adjusted neighborhood receive the frontier glow when their selected score is at least 2.0/10. Category neighborhoods are 80% wider than Overall, and category glow requires at least 600K parameters.

Rank cards by

Official results

Base Bench 1.1

Parameter efficiency

Overall Elo vs parameters

Higher Overall Elo is better. Fewer parameters are farther right. Models above the dashed 805 Elo line exceed random-choice chance. Frontier models are directly labeled; near-identical markers are offset slightly for visibility. Hover any point for exact details.

Model advisor

Which model should you use?

Answer a few questions to find the strongest benchmark fit.