HAAKKIM
ḥakkim: "to judge, to arbitrate"
An open arena for evaluating Arabic language models by human preference, across eleven dialects, ranked with a Bradley-Terry model.
People vote on pairs of model responses. Votes cast in Ranked Arena matches feed the public leaderboard.
Try the live arena, check the leaderboard, or browse the dataset.
| Battles | 1,273 total (1,130 voted, 143 skipped) |
| Ranked-eligible | 831 |
| Models ranked | 67 |
| Dialects | 11 |
| License | CC BY 4.0 |
| Format | Parquet, PII-scrubbed |
Full conversation transcripts, sampling weights, and category tags are included. View the dataset card on the Hub.
If you use Haakkim or this dataset in your research, please cite the IEEE Access paper.
@ARTICLE{11595788, author = {Mars, Mourad and Barmandah, Hassan and Alassaf, Abdulrhman}, journal = {IEEE Access}, title = {Haakkim: An Arena-Style Platform for Human Preference Evaluation of LLMs in Arabic and Its Dialects}, year = {2026}, volume = {14}, pages = {112112-112137}, keywords = {Modeling; Ranking (statistics); Generative Pre-trained transformer; Large language models; Voting; Labeling; Estimation; Clamps; Probability; Computational linguistics; Arabic LLMs; human preference; dialects; evaluation; leaderboard}, doi = {10.1109/ACCESS.2026.3710665} }