HAAKKIM

حَكِّم

ḥakkim: "to judge, to arbitrate"

An open arena for evaluating Arabic language models by human preference, across eleven dialects, ranked with a Bradley-Terry model.

People vote on pairs of model responses. Votes cast in Ranked Arena matches feed the public leaderboard.

Try the live arena, check the leaderboard, or browse the dataset.

How evaluation works

Ranked Arena
Random model pairing, single-turn MSA, a matched system instruction. The only mode that feeds the leaderboard.
Side-by-Side
You choose both models yourself, in any dialect. Counted for win rate only, kept out of the ranking to avoid selection bias.
10 Questions
A fixed pool of Arabic prompts, any dialect, for consistent side-by-side comparison. Also win rate only.

The Haakkim-1.0v dataset

Battles1,273 total (1,130 voted, 143 skipped)
Ranked-eligible831
Models ranked67
Dialects11
LicenseCC BY 4.0
FormatParquet, PII-scrubbed

Full conversation transcripts, sampling weights, and category tags are included. View the dataset card on the Hub.

Citation

If you use Haakkim or this dataset in your research, please cite the IEEE Access paper.

@ARTICLE{11595788,
  author   = {Mars, Mourad and Barmandah, Hassan and Alassaf, Abdulrhman},
  journal  = {IEEE Access},
  title    = {Haakkim: An Arena-Style Platform for Human Preference Evaluation of LLMs in Arabic and Its Dialects},
  year     = {2026},
  volume   = {14},
  pages    = {112112-112137},
  keywords = {Modeling; Ranking (statistics); Generative Pre-trained transformer; Large language models;
               Voting; Labeling; Estimation; Clamps; Probability; Computational linguistics;
               Arabic LLMs; human preference; dialects; evaluation; leaderboard},
  doi      = {10.1109/ACCESS.2026.3710665}
}