The leaderboard measures the task. It does not decide who you can trust in Czech.

The public leaderboard is a map of one task, one prompt population, and one scoring method. It is not a judgment about which model is best for Czech work.

The leaderboard measures the task. It does not decide who you can trust in Czech.
Editorial illustration created with AI assistance.The public leaderboard is a map of one task, one prompt population, and one scoring method. It is not a judgment about which model is best for Czech work.
Evidence record

Evidence passport

Sources and checks
22 sourcesSources checked:
Publication and updates
Published: Updated: Not stated
Correction history
0 corrections
Role of AI
Not stated
Reading time
27 min read
Listen
00:0000:00
Advanced controls
1.00 ×
Ready
Evidence record

Conclusion at a glance

What is established

The public leaderboard is a map of one task, one prompt population, and one scoring method. It is not a judgment about which model is best for Czech work.

What remains uncertain

A demonstrated technical capability alone does not establish real-world adoption, error rates in another setting or effects on particular people.

What would change the conclusion

An independent audit of the deployed system, reproducible measurement in a matching setting or new real-world impact data would change the conclusion.

Article contents
  1. Each ranking advocates a different task
  2. One "best AI" has at least five layers
  3. Length and markdown are not capabilities. They are a style factor
  4. MMLU is not general intelligence
  5. Price, Safeguard and Denial are not Elo
  6. Self-reported scores are not independent runs

Section: Technology & AI Author: IN Reading Length: ~27 min Sources and further reading: 22 items Topics: AI, benchmarks, ChatGPT, Gemini, Claude, Grok, Czech, model evaluation SEO / Working Title: ChatGPT vs. Gemini vs. Claude vs. Grok: Which AI is the best?

The question is: which AI is the best — ChatGPT, Gemini, Claude, or Grok? It sounds simple until we notice that we are not comparing one thing. Arena preference is not correctness. SWE-bench Verified is not your GitLab and five attempts at max effort is not one attempt. MMLU is not general intelligence. The English ranking is not the Czech usage. Only when these four things are separated, it is possible to read what the public tables aggregate as of September 2026 — and what of them for the Czechthe work does not flow.

1. Each ranking advocates a different task

Let's imagine four readers of the same word "best". The Product Manager will open LMArena and show the first lines of the public leaderboard. He's right that people's anonymous pairwise preferences capture which answer users like better. Paper Chatbot Arena reports over 240,000 votes from April 2023 to January 2024 and approximately 90,000 users. However, it does not measure factual error, Czech grammar or Czech jurisprudence [1].

The development manager pulls out the SWE-bench. He is right that the proportion resolved after passing the tests is tougher than the impression from the chat. But SWE-bench Verified is 500 human-verified instances from python GitHub issues. The original SWE-bench is based on 12 open-source python repositories. This is not your enterprise monorepo, your GitLab, SAP, internal library or the Czech Cybersecurity Act [6][7].

The teacher will show MMLU. He is correct that 57 subjects and 15,908 questions make up the named academic set. However, it does not measure open work with documents, Czech realities or whether the model admits uncertainty. And following the work of Gem et al. it is no longer possible to pretend that every difference on the raw MMLU is pure model ability: the authors of MMLU-Redux estimated 6.49 percent of questions with a key error [9][10].

Czech user adds BenCzechMark. He is right that Czech is not just translated English. BenCzechMark has 50 tasks, of which 14 are newly collected, in 8 categories, with mostly native Czech and about 10 percent machine translated tasks. It uses a duel win score from statistical tests with an alpha level of 5 percent, not a simple mean percentage. At the same time, it often lacks closed models, because authors deliberately do not send samples to paid APIs. "Not Measured"not "loser" [17][18].

At first glance, it's a brand dispute. In fact, each describes a different role. Therefore, they can all be partly right and the collective verdict "this AI is the best" still wrong. One table cannot decide chat, programming, school tests, Czech language, safety and corporate responsibility without reservations.

A leaderboard is a single task map. There is no judgment about who you can trust in Czech.

— Jiný Kontext

2. One "best AI" has at least five layers

When you say the best model, you often think of a single number. But the result of the public assessment is created in several layers. The first is the task. Are we asking about chat preferences, fixing a GitHub issue, a four-choice test, Czech NLI, school style or searching in a legal text? These are not different examples of the same thing. They are different problems.

The second layer is a population of prompts or instances. The arena collects self-selected prompts of real users and is mainly in English. SWE-bench works with issues from 12 python repositories. MMLU is an English exam bank. BenCzechMark adds historical Czech, pupil styles, spoken word and Czech tasks [1][6][9][17].

The third layer is scoring. The Bradley-Terry model of pairwise preferences is not the same as percentage of tests solved. Accuracy on four choices is not the same as a duel win score with statistical significance. Each scoring has its own question and its own blind spot.

The fourth layer is harness and disclosure policy. Agent scaffold, thinking budget, number of attempts, Docker, enabled tools and choice of variant can change the result. Paper Leaderboard Illusion highlights private variant testing and selective disclosure. This is not a small detail if the ranking is being made into a business argument [4].

The fifth layer is cost, latency, language, and denial. Token price, speed, Czech quality, data mode, safeguard and the possibility that the system will reject part of the task usually do not fit into one Elo. At the same time, they often decide for users more than the difference of one line on the ranking.

This layering is not a detail. It says if we are comparing the same work. If one model ran five times at max effort and the other ran once in a different month, they don't just share a benchmark brand. They do not share a protocol.

3. Preference in the Arena is not correctness

Chatbot Arena is useful precisely because it measures something that static tests often cannot: people's preferences over real prompts. The user asks a question, gets two anonymous answers and chooses the better one. A ranking is created from a large number of such duels. Paper from ICML 2024 describes data from April 2023 to January 2024, more than 240,000 votes and about 90,000 users [1].

However, it does not follow that a higher rank means a truer answer. Preference combines tone, length, formatting, responsiveness, politeness, certainty, and sometimes correctness. A user can vote for an answer that is more readable, even if it contains a minor factual error. Or, conversely, brief answers, because they are more suited to his goal.

The arena is not useless. It's a conversational behaviour map. It just has to be read as a preference, not a forensic truther. When asked which system wrote a nicer email, Arena comes closer to the task than MMLU. When we ask which system correctly interprets the Czech paragraph, Elo alone is not enough.

In addition, preference often arises without knowing the correct answer. The user compares two answers at a time when he himself may not have a primary source, test or legal support. That's fine if we're measuring user choice. It is a problem if we infer truth from this choice. A model that answers confidently, fluently, and with good articulation can get a vote even where a later review would find a misquotation. A model who answers with a shorter sentence and confessesuncertainty, may appear weaker, even though it is more useful for risky work.

At the date of the search, the live first place of the text Arena was not used, because without a stable primary snapshot it would be a number without an anchor. This is not a weakness of the argument. It is his condition. The leaderboard changes and must have a date, a URL, and a description of what it measured.

4. Length and markdown are not capabilities. They are a style factor

LMSYS does not fight the problem of style by denying it. Li, Angelopoulos, and Chiang's August 28-29, 2024 blog shows that response length was the dominant style factor in the Bradley-Terry regression. The coefficient of the normalized length difference states 0.249. Authors controlled length and markdown elements such as lists, headings, and bold [2].

This is important because many users don't just vote for the truth. He votes on the feeling that the answer is complete. A longer answer often looks more workmanlike. Markdown feels more organised. Headings and bullet points give a sense of structure. This may be a useful feature of the product, but it is not the same as substantive accuracy.

Style control separates some of this influence. It doesn't remove it perfectly. The authors themselves draw attention to possible confounders, for example chain-of-thought or other characteristics of the answer that are related to length. An observational control is not a randomized experiment. He does not say that after subtracting the style we have the pure truth [2].

Therefore, it is a mistake to read the style-controlled ranking as a universal verdict. It is better than a rough impression. It's still a corrected preference map, not a factual error meter. If your work is about whether the citation really supports the sentence, you need to measure the citation's support. Not just the ranking in the Arena.

5. A private option before publication is not a public record

The Paper Leaderboard Illusion draws attention to a practice that is unpleasant precisely because the leaderboards appear public and simple. The authors report 27 private LLM variants of one provider, Meta, tested before the release of Llama 4. They also estimate the asymmetry of access to Arena data: 20.4 percent for OpenAI, 19.2 percent for Google, and 29.7 percent for the 83 open-weight models combined [4].

These shares were disputed by LMSYS in its response. Therefore, they cannot be read as a closed audit of the entire Arena. They can be read as a warning that the public ranking is not just mathematics above a neutral world. Who is allowed to test privately, who downloads scores, who only publishes the best variation, and who has access to votes affects what the public sees [4].

This is an audit of one practice, not proof that rankings are worthless. Quite the contrary: if they have a value, they need a protocol. They need to say which variant ran, when, with what rules, and whether the result matches the production model. Without it, the user does not buy the winner. It buys a name that might have meant something slightly different in the rankings.

6. SWE-bench Verified is not your GitLab

SWE-bench is a powerful test because it asks practically: can the system solve a real GitHub issue so that the patch passes the tests? The original benchmark contains 2,294 instances from 12 open-source python repositories. It is evaluated whether the proposed patch passes FAIL_TO_PASS and PASS_TO_PASS tests in the given harness [6][8].

SWE-bench Verified was created because the original set had various impurities. In August 2024, OpenAI presented 500 human-verified instances from 1,699 randomly selected test samples. Reviewed by 93 python developers. In the verified set there are 196 easy instances with a human estimate of time under 15 minutes and 45 hard instances with an estimate of over 1 hour. These times are a human estimate, not an automated metric [7].

Verified is cleaner. It is not universal. It still measures python issues from specific repositories, in a specific harness, with specific tests. It does not measure code review, security, architecture, TypeScript monorepo, historical SAP or Czech documentation. And SWE-bench Pro is a different set: heavier, more files, less ground truth leakage. Verified is not Pro and Pro is not your GitLab.

Therefore, it makes sense to take SWE-bench as a candidate filter for programming. There is no point in taking it as proof that the model will fix your issue at the same speed and with the same risk. Your harness, your tests and your review rules are part of the performance.

In the company repository, what is decisive is what the public benchmark does not see. Code can pass local tests and still violate an internal architectural rule. A patch may be functional but unreadable by the team that will maintain it. The agent can fix the symptom and bypass where the error actually originates. Or, on the contrary, it fails only because the company did not give it the same context that a human has: runtime variables, historical decisions, an internal library andknowledge of the release process. Therefore, the programmer's benchmark must be supplemented with its own pilot. Not because the public set is bad, but because your work has a different threshold for correctness.

7. Five attempts at max effort is not one attempt

The biggest confusion arises when we put numbers with the same benchmark name and different protocol next to each other. Anthropic in Fable 5 and Mythos 5 system card from June 2026 states on SWE-bench Verified 95.5 percent resolved for Mythos 5 and 95 percent for Fable 5. These are self-reported numbers, average of 5 attempts, adaptive thinking and max effort. Fable has production safeguards including a fallback to Opus 4,8, which the card cites as a reason why it may be slightly below Mythos [11][12].

On SWE-bench Pro, the same card reports 80.3 percent for Mythos 5 and 80 percent for Fable 5. Google DeepMind in the 19 February 2026 model card Gemini 3,1 Pro reports 80.6 percent on SWE-bench Verified for Gemini 3,1 Pro Thinking High, but as a single attempt. On public split SWE-bench Pro reports 54.2 percent, also single attempt [11][14].

OpenAI lists 64.6 percent on SWE-bench Pro in the GPT-5,6 table of 9 July 2026. It does not list SWE-bench Verified in this table. The benchmark shows the Fable 5 as 80 percent Pro, but that's still a number in the OpenAI table and a name match, not proof of the same run with the same harness. Historically, OpenAI at SWE-bench Verified listed GPT-4o with 33.2 percent resolved as of August 2024, in a note to Agentless [7][13].

What needs to be tested before adding two percent

Type Verified or Pro. Number of attempts. Thinking. Card date. Self-report or independent run. Scaffold agent. Allowed tools. Only then compare. Without that, the number 95 next to 80.6 cannot be read as a battle of brands. It is a duel of two lines with a different protocol.

Self-reported scores share the name of the benchmark. It does not share the protocol.

— Jiný Kontext

8. MMLU is not general intelligence

MMLU has a good reason to be famous. Hendrycks et al. built an extensive four-choice academic test in English: 57 subjects and 15,908 questions in total, of which 14,079 are test and 1,540 are validation. Five few-shot examples per subject are used in the original paper [9].

Such a set helped measure breadth of knowledge. But breadth of academic questions is not general intelligence. MMLU does not measure Czech law, Czech realities, open-ended work with a document or the ability to admit that a source is missing. The model can be excellent on a four-choice question and weaker on a long Czech letter to a client. Or vice versa.

Also, the human anchor often reads stronger than the paper allows. Hendrycks et al. report 34.5 percent for unspecialized MTurk and about 89.8 percent as an expert 95th percentile estimate from source trials. That expert figure is a guess, partly from percentiles and partly educated guess. It is not a measured panel of experts on the entire MMLU [9].

MMLU-Redux and MMLU-Pro are other objects. Contamination review from GEM 2026 recalls MMLU as the archetype of a set susceptible to leakage into training and repeated use [20]. This is not to say that MMLU has no value. It means that it cannot be unreservedly made a judgment on work that the MMLU has never measured.

9. The error in the key is not a model difference

Gema et al. at work Are We Done with MMLU? re-annotated 5,700 questions from all 57 subjects to create MMLU-Redux. They estimated that 6.49 percent of questions have an error in the key. In virology, in the section analyzed, they report 57 percent of erroneous entries [10].

This is methodologically unpleasant. If two frontier models differ by one percentage point on raw MMLU, part of the difference may be label noise, not ability. A model can answer substantively correctly and be penalized with an incorrect key. Or he may hit a key that is itself questionable.

It is not fair to make a sentence out of this that MMLU is useless. It is fair to make this a rule: the smaller the difference, the greater the need to look at label quality, uncertainty interval, contamination and the nature of the task. The exact percentage can accurately measure the defective condition.

10. A cluster at the GPQA ceiling is not a brand verdict

GPQA Diamond is a different type of academic test. Rein et al. they constructed it as a graduate-level science, google-proof kit; diamond subset has 198 questions. High scores appear in the 2026 model cards: Anthropic lists Mythos 5 at 94.1 percent as an average of 5 attempts, Google lists Gemini 3,1 Pro at 94.3 percent without tools, and OpenAI lists GPT-5,6 Sol at 94.6 percent [11][13][14][22].

A difference of 94.1, 94.3 and 94.6 percent looks like a rank. But it is a cluster at the ceiling, across different cards and protocols. Without the same run in the same week, with the same setup and public logos, it's impossible to make a brand verdict.

The practical question sounds different again. If you are choosing a model for scientific Q&A, GPQA can be a useful map. If you are choosing a model for Czech internal legal analysis, it maps a different terrain. A ceiling score in English science is not an authorization to believe a Czech citation.

11. The English ranking is not a Czech usage

Czech is not just a dictionary. It is syntax, inflection, context, cultural references, schools, authorities, legal wording, diacritical errors and the ability not to translate an English phrase as a Czech sentence. That is why Czech benchmarks are important, even if they are not as visible as global arenas.

BenCzechMark from TACL 2025 lists 50 tasks, 14 newly collected and 8 categories. It is mostly native Czech; about 10 percent of jobs are machine translated. The score is not a simple accuracy average, but a duel win score from statistical tests with an alpha of 5 percent. This is a different rating language than MMLU or Arena Elo [17].

At the same time, a caveat is necessary. BenCzechMark authors deliberately do not send samples to paid APIs, so closed ChatGPT, Claude or Gemini on BCM are often missing. This cannot be read as a defeat. It's a measurement gap [17][18].

CzechBench is another Czech track, but even here it is not appropriate to pretend to be accurate where the sources diverge. The dossier states that the BenCzechMark paper mentions 17 tasks for CzechBench and 8 native ones, while the public leaderboard reports 15 tasks. Therefore, one official CzechBench number does not stand here as a fixed figure. The named benchmark Czeval or CzEval was not found in primary sources as of September 2026 [19].

Those who only read LMArena or MMLU do not see Czech duels. Anyone who reads only the Czech benchmark without closed models is also not seeing the entire market. The correct sentence is not "this model won the Czech language". The correct sentence often reads: in this Czech set, this is measured, that closed model is missing, and a separate test is necessary.

In addition, the Czech usage has several forms. Another problem is writing a natural business email. Another is to summarize the minutes of the meeting with Czech names and institutions. Another is to distinguish whether the quoted paragraph of the law applies, has been changed or belongs to another part of the legal system. And another is to correct the student's text so that the system does not transcribe the author's style into smooth general Czech. A single score can bring any of these jobs closer. It cannot replace them all.

12. Price, Safeguard and Denial are not Elo

The public leaderboard usually does not carry an account. In doing so, the user is not only buying the ability, but the received result. Anthropic lists Fable 5 as $10 per 1 million entry tokens and $50 per 1 million exit tokens in June 2026. OpenAI lists GPT-5,6 Sol at $5 per 1 million input and $30 per 1 million output tokens in July 2026 [11][12][13].

These numbers are not a verdict in themselves. A cheaper token can be more expensive if the model needs more trials, more review and more human correction. A more expensive model can be cheaper if it reduces the number of failed runs. The price list is the price of the raw material, not the price of the result.

Safeguard is another layer. Anthropic on Fable 5 introduces production safeguards and fallback to Opus 4,8, which launches in less than 5 percent of sessions on average [11]. This can be a safety advantage. But it can also mean that the Czech "I can't do it" in a specific task is not the inability of the model, but the policy of the product.

Rejection, latency, and data mode don't fit easily into Elo. Nevertheless, they make decisions at work. The best model in the ranking may not be the best system for a process where you need auditing, low latency, stable response, cheap control and predictable behaviour.

The same logic applies to damage. For a marketing proposal, rejection can only be a delay. With customer support, an unpredictable rejection can be a breach of service. For legal or health text, caution may be desirable if the system clearly states what it does not want to decide and why. A leaderboard that measures issues resolved or chat preference will not usually show this threshold. It must be shown by the pilot in a specific process.

13. Grok card is not a SWE chart

Grok appears in the debates as another brand in the same fight. However, the dossier states a simple limit: the public cards Grok 4 dated 20 August 2025 and Grok 4,20 dated 7 April 2026 are security and dual-use assessments. The comparable SWE-bench Verified xAI number in these cards has not been verified as of the date of the search [15][16].

The absence of a number is not proof of incompetence. It is proof that these cards cannot be put together in a fair fight on SWE-bench Verified. If one vendor publishes a coding score in a spreadsheet and another publishes mainly a security card, we don't have two measurements of the same thing.

This is a more general rule. Not every model card is a leaderboard. Not every safety document contains product performance. And not every marketing table has a log that allows comparison with another table. When the background is missing, the correct answer is "unverified", not "lost".

14. The best map is the one you can bear the mistake of

The following recommendations are not a product ranking. They are a signpost to your first own run. They follow the text Who do you allow to make mistakes?, which handles AI selection based on the type of error you can afford. This is not a second brand verdict. It's about reading the leaderboard like a map.

A kind of failure What does he look like? How to test What will limit the damage
Preference as correctness Higher Elo = truer Czech Ten tasks in the same language; factual error, not "like" Take style control as a caveat, not as a test of truth
Hidden harness 95% next to 80.6% as a battle of the brands Write down attempts, thinking, date, Verified vs Pro, self-report vs independent Do not add cards with different protocols
English map in Czech LMArena or MMLU as a verdict for Czech work BenCzechMark, CzechBench or custom set; "not measured" not "lost" Don't leave MMLU or raw Elo first column
Cherry-pick variant Private runs, download scores, just the best number out Asking which variant is in production Leaderboard Illusion as caveat to the table
A cluster at the ceiling 94.1 / 94.3 / 94.6 as mark order GPQA Diamond: 1 p.p. difference at ceiling Don't decide on a purchase by one percentage point

Illustrative diagram: Three columns — Arena (preferences), SWE-bench (harness tests), BenCzechMark (Czech duels). The "best" arrow does not lead between the columns. It only leads inside the column, with date and protocol.

A good map is not the one with the biggest number. A good map is one whose fallacy you know. If it measures preference, use it for work where the preference makes sense. If it measures the patch after the tests, use it to select candidates for programming. If it measures Czech, read it in Czech and with a coverage limit.

15. Ten custom tasks will tell more than first place

The practice test is not a brand battle. It is a check that the map fits your task. Compile ten quests in the same language, same genre, and with the same damage threshold. Ten Czech e-mails to a customer. Ten issues from your repository. Ten questions from the internal knowledge base. Ten short analyzes where you know the right support.

How to test two candidates on the same set

Run two candidates on the same day, with the same thinking setup, without switching prompts between runs, and with the same access to tools. Record factual error, hallucinated quote, rejection, length, cost and time to usable result. Do not attribute Elo. Add up which error would pass your inspection.

Three harness tests

If the task is Czech, do not leave MMLU or raw Elo Arenas as the first column. The first column is BenCzechMark, CzechBench or your set. If the closed model is not on BenCzechMark, write "not measured", not "lost".

If the task is code, specify the harness. "95 percent Verified" without trial count, thinking mode, and agent stack is a marketing headline, not a lab record. If the assignment is research, check the citations sentence by sentence. A citation that exists may not yet support a particular conclusion.

Your own set doesn't have to be big. It must be true. It's supposed to measure the work you actually do, in the language you do it in, and with the error you can imagine.

A good set also has correct negative examples. It is not enough to give the model ten tasks that can be solved smoothly. Add an entry in which the background is missing. Add a document with a conflict between an older and a newer version. Add a question where the correct answer is "can't tell from the available data". At such a moment, the difference between a model that wants to finish the sentence at all costs and a system that helps maintain control becomes apparent.

Simply record the result. Not just a grade, but a kind of mistake. An invented resource. Right source, wrong conclusion. Too long an answer. Rejection without reason. Good answer, but too expensive to repeat. Only this entry tells if the first location from the public map leads to your destination.

16. Self-reported scores are not independent runs

A supplier's self-assessment is not automatically false. However, it is a different type of proof than an independent run. A model card can be fair and yet not comparable to a competitor's card. The difference may be in the number of attempts, agent scaffold, date, allowed tools, sample or variant selection policy.

A stronger conclusion would require an independent, pre-registered run of the same SWE-bench Verified harness for ChatGPT, Gemini, Claude, and Grok in the same week, with the same scaffold, same number of attempts, same commit, and public logs. This would still not decide the Czech language or the law, but it would solve at least one source of confusion.

For the Czech language, a public Czech-language preference leaderboard with described sampling and without private cherry-picking, where closed and open models would stand side by side, would help. For MMLU, a reannotated successor without the estimated key error rate of 6.49 percent, on which the order of the frontier models would stably vary more than the label noise, would help [10].

It would also help to show that the style-controlled Arena score predicts your company metrics better than the custom set: time to merge, number of complaints, number of hallucinated paragraphs, or proportion of answers returned by the editor. Until this happens, the public leaderboard remains a job map. It is not a judgment.

17. A hidden load is the first place that measures another task

The first place is comfortable. It saves decision-making, fits into the presentation and gives the purchase a simple sentence. But that is precisely its hidden cost. When the first place measures another role, it brings certainty to the company where the question should have been.

Rankings are not the enemy. Without public benchmarks, we would only have advertising and personal impressions. Arena helps to see preferences. SWE-bench helps to see programmer patches in a particular harness. MMLU and GPQA help track academic knowledge. BenCzechMark and CzechBench remind that the Czech language must have its own measurement. Every map is useful when we know where it leads.

Therefore, it is not accurate to say that benchmarks should not be trusted. It is more accurate to say that we should believe them to the extent that they describe themselves. If Arena says that people in an anonymous pairwise comparison preferred one answer, let's not read the truth of the paragraph from it. If SWE-bench says that the patch has passed the tests, don't read it as security of the design. If BenCzechMark says that the model survived the Czech duel, let's not read it as the performance of a closed model that was not included in the set.

The question is not which AI is the best. It's: which task, which prompt population, and which harness scale does it measure — and which error will you see in time on your own set?

Evidence record

How this article was made

Method, the role of AI, corrections and source details in one place.

Sources and further reading22 sources
  1. Peer-reviewed studyChatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML/PMLR 235; arXiv:240304132. https://proceedings.mlr.press/v235/chiang24b.html
    Studies: Chiang, W.-L. et al. · 2024
  2. Other sourceDoes style matter? Disentangling style and substance in Chatbot Arena. LMSYS Org, 28/08/2024. https://www.lmsys.org/blog/2024-08-28-style-control/
    Methodology: Li, T., Angelopoulos, A., Chiang, W.-L. · 2024
  3. Official statisticsCode / data: LMArena, Contextual Bradley-Terry (style control) and dataset arena-human-preference-140k . https://github.com/lmarena/arena-rank
  4. Official statisticsThe Leaderboard Illusion. arXiv:250420879; NeurIPS 2025 Datasets & Benchmarks. https://arxiv.org/abs/250420879
    Studies: Singh, S. et al. · 2025
  5. Other sourcePlatform: Arena / LMArena (formerly LMSYS Chatbot Arena). Public leaderboard. https://lmarena.ai / https://arena.ai
  6. Peer-reviewed studySWE-bench: Can Language Models Resolve Real-world GitHub Issues? ICLR 2024. https://www.swebench.com/
    Studies: Jimenez, C.E. et al. · 2024
  7. Other sourceupdate 2025). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/
    Institution: OpenAI · 2024
  8. Other sourceLeaderboard: SWE-bench Verified / Full / Lite / Multilingual. https://www.swebench.com/
  9. Peer-reviewed studyMeasuring Massive Multitask Language Understanding. ICLR; arXiv:200903300. https://arxiv.org/abs/200903300
    Studies: Hendrycks, D. et al. · 2021
  10. Peer-reviewed studyAre We Done with MMLU? NAACL 2025. https://aclanthology.org/2025.naacl-long.262/
    Studies: Gema, A.P. et al. · 2025
  11. Other sourceClaude Fable 5 & Claude Mythos 5 System Card. https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c342ee809620.pdf
    Model card: Anthropic · 2026
  12. Other sourceClaude Fable 5 and Claude Mythos 5. https://www.anthropic.com/news/claude-fable-5-mythos-5
    Notification: Anthropic (6/9/ · 2026
  13. Other sourceGPT-5,6. https://openai.com/index/gpt-5-6/
    Model card: OpenAI (07/09/ · 2026
  14. Other sourceGemini 3,1 Pro — Model Card. https://deepmind.google/models/model-cards/gemini-3-1-pro/
    Model card: Google DeepMind (2/19/ · 2026
  15. Other sourceGrok 4 Model Card. https://data.x.ai/2025-08-20-grok-4-model-card.pdf
    Model card: xAI (20/08/ · 2025
  16. Other sourceGrok 4,20 System Card. https://data.x.ai/2026-04-07-grok-4-20-model-card.pdf
    Model card: xAI (4/7/ · 2026
  17. Peer-reviewed studyBenCzechMark: A Czech-Centric Multitask and Multimetric Benchmark… TACL. https://doi.org/10,1162/tacl,a,32
    Studies: Fajcik, M. et al. · 2025
  18. Other sourceLeaderboard: CZLC, BenCzechMark. Hugging Face Space. https://huggingface.co/spaces/CZLC/BenCzechMark
  19. Peer-reviewed studyCzechBench leaderboard / CTU diploma thesis. https://huggingface.co/spaces/CIIRC-NLP/czechbench_leaderboard
    Benchmark: Jirkovsky, A. et al. · 2024
  20. Systematic reviewsystematic review. https://doi.org/10,18653/v1/2026,gem-main,50
    Overview of contamination: Are LLM Benchmarks Already Contaminated? GEM · 2026
  21. Other sourceStyle plugin: LMSYS / Arena, Does Sentiment Matter in AI? https://arena.ai/blog/sentiment-control/
  22. Peer-reviewed studyGPQA: Rein, D. et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark. https://arxiv.org/abs/231112022