Technology, decision-making & responsibility
ChatGPT, Claude, Gemini, Copilot, Perplexity, Grok, Mistral or your own model? The question sounds simple only until we notice that we are not comparing a single thing. A model may be excellent while the product is unusable. A citation may be genuine while the conclusion is wrong. A long context window may hold a document without understanding it. And the cheapest subscription may become the most expensive once verification, errors and dependency are included.
1. A meeting in which everyone is partly right
Imagine the management team of a medium-sized Czech company. The director wants ChatGPT because they regard it as the most versatile. IT proposes Microsoft Copilot because the company lives in Outlook, Teams, Excel and SharePoint. A developer argues for Claude or a specialised coding agent. Marketing wants Gemini because of Google services. The analyst asks for Perplexity because its answers show sources. A colleague who follows events on X trusts Grok. The security manager would prefer to run an open-weight model inside the company’s own infrastructure.
At first glance, this is a dispute between brands. In reality, each person is describing a different problem. The director is comparing the breadth of a finished application. IT is comparing integration with company data and identity. The developer is comparing a workbench that can read a repository and run tests. Marketing is comparing available modalities. The analyst is comparing a research process. The security manager is comparing data transfer, key management and auditability.
They can therefore all be partly right while the shared decision is still wrong. If the company buys one licence based on the winner of an online ranking, it answers a question no one formulated precisely. “The best AI” is not a product property like weight or dimensions. It is shorthand for a collection of different requirements: writing quality, source handling, speed, price, repeatability, connectors, permissions, export, privacy and the consequences of error.
The first important distinction is therefore not which vendor is ahead. It is what exactly is being compared. Otherwise we place an engine, a car, a navigation system and a taxi service side by side and wonder why each wins a different discipline.
A model is not a product. A product is not an agent. An agent is not a company system. And a benchmark is not real work.
— Jiný Kontext
2. A single answer has at least six layers
When a paragraph, table or corrected block of code appears on screen, it is easy to think that “the model wrote it”. Yet the result usually emerges through a chain of several layers. Each can improve quality, reduce it or change the type of error. Two products based on the same underlying model therefore need not give the same answer, and the same product may behave differently depending on the mode, account or enabled tools.
- Foundation model: language and multimodal capabilities, knowledge and typical limitations.
- Compute mode: a fast answer, deeper reasoning, length limits and allocated time.
- Work apparatus: system instructions, planning, memory, verification loops and file handling.
- Sources and tools: web search, databases, code, calculators, connectors and company documents.
- Permissions: what the system may read, create, change, send or delete.
- Governance and responsibility: data retention, audit, approval, contractual terms and reversibility.
The foundation model determines an important part of the capability, but it is not the only actor. An application may insert its own instructions before a query, attach history, retrieve documents, shorten an answer or pass it to another tool. A coding agent may verify the same proposal with a test while an ordinary chat merely prints code. A search product may ground an answer in the current web, but it may also select the wrong source. An enterprise assistant may find an internal memo an isolated model could never see, while also inheriting chaos in company permissions.
This is where the first error in many public comparisons appears. The headline announces a contest between two models, but the test actually compares two applications with different tools, limits and instructions. Or it measures a bare API without search and presents the result as a verdict on a finished product. An honest comparison must state not only the name, but also the application version, mode, available tools, attached context, time limit and verification method.
This layering is not an academic detail. It tells us where to look when something fails. If the answer did not know about yesterday’s change in law, a web source may have been missing. If it overlooked a paragraph in the middle of a contract, context selection may have failed. If an agent overwrote the correct file, the problem may have been the permission model or approval step. “The AI made a mistake” begins the diagnosis; it does not end it.
3. The best system is the one whose error you can afford
Users usually look for the system with the highest probability of a correct answer. That is reasonable but incomplete. It matters just as much what an error looks like, whether it can be recognised, how much correction costs and whether the result is reversible. A typo in a slogan draft is a cheap error. A fabricated citation in a legal filing is expensive. A badly named file can be restored. A message already sent to a client, or a change to a production system, sometimes cannot.
NIST uses the term confabulation for confidently presented false content and notes that it may include false facts, fabricated citations and internally inconsistent explanations. Risk does not rise only with the frequency of error. It also rises when a professional style encourages disproportionate trust.[17]
| Failure type | What it looks like | How to test it | What limits the harm |
|---|---|---|---|
| Confabulation | An invented fact, number, source or degree of certainty | Find primary support for every material claim | Citations, independent review, no autonomous action |
| Omission | The answer is correct but a condition or exception is missing | A checklist of required points and boundary-case tests | Structured output and a mandatory “what is missing” field |
| Context error | The system uses an old, foreign or irrelevant document | Show the source, date and permission used | Versioning, a bounded corpus and a document owner |
| Action error | Correct reasoning leads to the wrong intervention | Test in isolation and record tool actions | Change preview, human approval and rollback |
This changes the selection question. We do not ask only which system produces the best average answer. We ask which system fails in a way our process will catch. A creative team can tolerate ten unusable ideas because it can easily recognise one good one. An accounting department cannot tolerate one hidden change of a decimal point. A developer can accept a bad proposal if a test fails reliably. A doctor, judge or public official must not mistake fluent prose for a reviewable expert conclusion.
Capability is therefore only one axis. A second is observability, a third reversibility and a fourth impact. For an organisation, a product with a slightly weaker model but a clear source trail, constrained permissions and approvals may be safer and economically better than a benchmark winner without those supports.
4. A benchmark is a map. It is not a verdict
Benchmarks are necessary because they replace some impressions with a repeatable test. Trouble starts when one score is made to represent the system’s entire capability. Stanford’s HELM project bases evaluation on broad coverage, multiple metrics and standardised conditions. At the same time, it explicitly acknowledges incompleteness: no single ranking can measure every domain, language, interaction pattern and risk.[18]
Every test contains hidden decisions. Who selected the tasks? In what language? Does the model receive tools? How much time and how many attempts? Is the exact string assessed or the functional result? Is cost included? May the system search the web? Is the test set older than the training data? The answer to any of these questions can change the order.
A telling example came from programming evaluations. In July 2026, OpenAI published an audit of the 731-task public split of SWE-Bench Pro. Its automated procedure labelled 200 tasks, or 27.4 per cent, as likely problematic; human annotation labelled 249, or 34.1 per cent. Problems included overly narrow tests, incomplete prompts and checks that allowed unfinished solutions to pass. This is an audit of one dataset by one organisation, not proof that benchmarks are worthless. It does show that an exact percentage can precisely measure a defective condition.[20]
- The exact version of the model, application and work apparatus.
- Reasoning mode, tools, time limit and number of attempts.
- Task population, language, period and scoring method.
- Uncertainty interval, failed runs and cost per attempt.
- Known dataset limitations, possible contamination and the amount of human intervention.
For a purchase, a small internal benchmark is therefore more important than somebody else’s overall ranking. Thirty real company tasks — with real documents, Czech language, approved data and actual controls — may say more than a thousand questions from a domain the organisation will never use. A public benchmark can help select candidates. The work itself must make the decision.
5. Compare ecosystems, not mascots
The market can be described without a false podium. The following map does not say who is first overall. It shows the kind of work around which each ecosystem builds its product, the kind of context it provides and the failure a buyer should deliberately seek during a trial. Features and plans change; the provider’s current page and the terms of the specific account are therefore always decisive.
| Ecosystem | Primary working position | What to verify in a pilot | Typical reason for choosing it |
|---|---|---|---|
| OpenAI / ChatGPT / Codex | General work environment, research, creation, analysis and code | Which mode and tool actually answered; limits and export | One environment for varied work |
| Anthropic / Claude | Long documents, focused knowledge work and programming | Repeatability, file handling and the shared usage budget | Strong emphasis on text and agentic workflows |
| Google / Gemini | Multimodality, search and Google applications | Consumer versus Workspace mode; regional availability | Work already happens in Gmail, Docs, Drive or NotebookLM |
| Microsoft / GitHub Copilot | Microsoft 365, Graph, company permissions, IDEs and repositories | Overshared data, document age, action scope and audit | The context is already inside work systems |
| Perplexity | Web research, source synthesis and a choice of multiple models | Whether citations actually support the sentence | A fast start to research with visible links |
| xAI / Grok | The current web and social stream, text and media | Source balance, volatility and separating signal from popularity | Immediate access to a live information environment |
| Mistral / open weights | API, self-hosting, private cloud or local operation | Licence, infrastructure management, security and real TCO | Greater control over location and model replacement |
The rows are not a verdict on model quality. The same vendor offers several classes, modes and products, each of which may change within months. The map is meant to prevent a different error: a company buying an excellent model in a product that cannot see its work, cannot complete it safely or locks it into an environment that later becomes difficult to leave.
6. OpenAI: the advantage of a broad workbench
OpenAI’s ecosystem today is not limited to chat. Its business offering combines conversation, analysis, creation, programming, work agents and connectors to other services. The official business overview also distinguishes a standard business plan from contract-based enterprise options with broader governance, retention rules and access control.[1]
The strength of this breadth is reduced friction when moving between tasks. In one environment, a user can analyse a spreadsheet, find sources, prepare copy, work with an image and hand part of the task to a coding apparatus. That is practical value a bare model test does not capture. Versatility also creates a new requirement: the user must know whether they are using a fast answer, deeper research, a connected service or an agent able to change files.
What needs to be tested
The first test concerns consistency across modes. Asking the same question of ordinary chat, a research mode and an agent is not the same experiment. The second concerns costs and limits: the marketing word “unlimited” may still carry guardrails, different model limits or additional credits. The third concerns data. OpenAI states that data from ChatGPT Business, Enterprise and the API is not used to train models by default, and that qualifying organisations can obtain retention and residency options.[2] That is a material property of a specific contractual regime, not an automatic property of every personal account and every connected tool.
For an individual who wants one broad service, this integration may be decisive. For an organisation, however, saying “we have ChatGPT” is not enough. It needs to know the workspace type, data owner, permitted connectors, sharing rules, export, action logs and the boundary beyond which a human must approve the result.
7. Anthropic: focused work is not the same as flawless work
Claude is positioned as an environment for documents, research, code and longer workflows. The official plan overview links paid accounts with projects, research, connectors and tools for programming and other creation. It also notes usage limits and that capacity depends on the length and complexity of conversations, the selected model and the features used.[3]
This matters more than a simple message count. A long document, several files and an agent run may consume a very different share of capacity from a short question. A user therefore buys more than access to a model name. They buy the throughput of a working day: how often a task can be repeated, how long an output can be, whether work can continue after interruption and what additional compute costs.
A pleasant style can hide the same old problem
Well-structured prose looks like evidence of understanding. It is not automatically so. In editorial or analytical work, reviewers must specifically check whether the system replaced the author’s thesis with its own addition, omitted an inconvenient exception or merged two different sources into a smooth but nonexistent consensus. Better wording lowers editing costs; it can also raise the cost of detecting a factual error because the answer does not look suspicious.
Consumer and work modes must be separated here as well. In its explanation of consumer terms, Anthropic describes a choice over whether new or resumed conversations may be used to improve models and distinguishes these accounts from services under commercial terms, the API and enterprise plans.[4] A company decision must therefore not be based only on what an employee sees in personal settings.
8. Google: sometimes the model wins; sometimes the location of the data does
Google connects Gemini with a broader service bundle: search, cloud storage, Gmail, Docs, media creation and NotebookLM. Official consumer-plan pages describe this combination as part of one subscription, not merely access to one model version.[5] For someone whose work already lives in this environment, integration may save more time than a small difference in a general-reasoning test.
Multimodality also changes the unit of work itself. A task need not be “write a text” but “review the PDF, images, recording and spreadsheet, then prepare a brief”. The ability to accept more formats is only the entrance. The system must still recognise what matters, separate observation from interpretation and state what could not be read reliably.
Two similar accounts may have different data regimes
With Google, it is especially important not to blur the boundary between the consumer application and Workspace. The Privacy Hub states that when activity retention is enabled, shared content may be used to improve services with the help of human reviewers.[6] Documentation for qualifying Workspace editions, by contrast, describes enterprise protections under which content is not used without permission to train generative models outside the domain or reviewed by humans.[7]
The difference is not cosmetic. The same employee may have a personal account, a work account and different permissions to the same file. A pilot must record which profile is being tested, what is enabled and who owns the configuration. Otherwise the company tests the convenience of a consumer application and accidentally presents it as a security property of an enterprise deployment.
9. Microsoft and GitHub: the intelligence of permissions
Microsoft 365 Copilot derives value from grounding answers in Word, Excel, Outlook, Teams and data available through Microsoft Graph. The offering presents it directly as an assistant inside work applications and company context.[9] In practice, a somewhat less impressive general answer can create more value if it correctly finds the latest minutes, spreadsheet and email on which the decision actually rests.
The same property is a security mirror. Microsoft states that Copilot respects existing user permissions and that Graph data may appear in an answer when the user already has access. The system does not repair historic oversharing by itself. If an employee could open a sensitive folder for years without knowing it existed, AI retrieval may expose that forgotten access. The documentation also states that prompts, responses and Graph data are not used to train foundation models in the enterprise regime.[10]
GitHub Copilot demonstrates a similar principle in development. The result is not produced by a model in a vacuum. Context includes the open file, selected code, other parts of the workspace, frameworks and dependencies. Business plans add licence administration, policy management and other organisational controls.[11]
- Find public links, “everyone” groups and historically overshared folders.
- Assign an owner and validity date to material documents.
- Separate retrieval from actions that change or send data.
- Introduce audit, approval and a trial group with limited scope.
- Verify that users can see the file and version from which an answer was drawn.
The greatest competitive advantage of enterprise AI may therefore not be the model. It may be orderly data. And the greatest weakness may not be a hallucination. It may be a truthful answer based on a document that was true two years ago.
10. Perplexity and Grok: currency is not the same as truth
Perplexity is useful to understand as a research architecture that combines search, sources and multiple models. The service itself notes that the available model may change and that a model inside Perplexity is not the same as the same model in its native application: search, citations, tools, system instructions and service limits all affect the answer.[12]
That is both an advantage and a warning. A user can compare several sources quickly and begin research with a visible trail. They must not confuse a link with evidence. A citation may lead to an existing page that supports only part of a sentence, discusses another population or repeats a claim from another unverified article. The quality of a research product is therefore not measured by the number of blue links, but by the extent to which every material conclusion follows from a primary source.
Grok builds much of its differentiation around the live web, X, voice and multimedia; its business offering adds administration, data isolation and security controls.[15] Access to a fresh stream does not automatically produce a better picture of reality. The event people write about fastest may also be the least verified. Post popularity measures distribution, not truth.
Three tests for a current answer
The first is temporal: when did the event happen and when was the source updated? The second concerns provenance: is it a primary document, eyewitness account, press release or repetition? The third seeks contradiction: did the system find a credible source that challenges the conclusion? A product that answers one minute earlier is not necessarily better if no one can determine an hour later why it answered as it did.
Consumer and enterprise regimes differ here too. Perplexity states that paid consumer accounts can opt out of data use for training, while enterprise data is not used for training.[13] When sensitive material is involved, the entire chain of providers, search and retention rules must therefore be verified, not just the selected model.
11. Mistral and open weights: control does not come with one download
Alongside finished services and an API, Mistral offers private-cloud, local and self-hosted deployment. Its official overview emphasises that data in such deployments can remain within the customer’s environment.[14] For a European organisation, a regulated industry or a company with sensitive know-how, system location may matter more than the final percentage point in a public test.
“The weights can be downloaded” does not mean “it is free, private and open”. Operation requires hardware or cloud capacity, updates, monitoring, access control, backups, testing, security response and someone who understands the licence. Data stays local only if logging, embeddings, retrieval, telemetry and connected tools also remain within the designated environment. A single external connector can turn a local model into a non-local process.
The Open Source Initiative also distinguishes open weights from a genuinely open AI system. Its definition requires not only parameters, but sufficient information about the data and code needed to study, modify and reproduce the system.[16] Some widely available models are therefore more accurately described as open-weight than fully open source.
| Option | What you gain | What you take on | Hidden question |
|---|---|---|---|
| Finished consumer application | Fast start, UI, updates and tools | Dependency on service rules and limits | How is data used and retained? |
| Enterprise SaaS | Identity management, contractual protections and audit | Integration, role configuration and governance | What happens with overshared data? |
| API | Your own product and choice of work apparatus | Development, tests, monitoring and tool costs | Who is responsible for final behaviour? |
| Self-hosted or local model | Control over execution and location | The entire infrastructure, updates and security | Is the complete chain genuinely local? |
An open weight is an opportunity to build. It is not a finished enterprise system. It can substantially reduce dependency on one API, enable specialisation and provide control over data. It also transfers responsibility from the vendor to the operator. That is a strategic choice, not a free lunch.
12. Citations do not remove hallucination. They move the point of verification
Before web search, a user had to check whether the model had invented a fact. With citations, the user checks something more complex: whether the source exists, whether it is credible, whether it addresses the same question and whether it actually supports the sentence to which it is attached. A citation is navigation to evidence. It is not evidence automatically.
One answer can combine three truthful sources into a false conclusion. The first describes a global population, the second Czech households and the third another period. Every link works, but the synthesis is methodologically wrong. At other times the system uses a high-quality secondary article even though a primary document is available. Or it cites a page that has since changed and no longer contains the original claim.
- Existence: the link works and leads to the described document.
- Authority: the source is primary for the claim or clearly identified as secondary.
- Support: the specific passage actually supports the specific sentence.
- Scope: population, period, definition, unit and limitations match.
The most reliable workflow therefore separates searching from writing. First comes a table of claims, primary sources and limitations. Only then comes the prose. For every number, the population, period, measured indicator and uncertainty are recorded. If a source cannot be opened or the relevant passage cannot be found, a claim is not promoted to fact merely because the model worded it well.
A research system is good when it shortens the route to verification, not when it replaces verification. For a low-risk question, a quick answer with several links may be enough. For a health, legal, financial or publicly consequential conclusion, the last step must be taken by a person who understands the source and accepts responsibility for its use.
13. Long context is not memory, and memory is not understanding
A large context window is often translated into “the model can read the whole book” or “it knows the entire repository”. Technically, the system may accept a large number of tokens. It does not follow that it uses every part equally well. Gemini documentation compares the context window to short-term memory: it is information supplied for a particular answer, not permanent and infallible understanding.[8]
The “Lost in the Middle” study found a marked dependence on the position of relevant information in the models it examined. Performance tended to be better when decisive material appeared at the beginning or end and fell when it was placed in the middle of a long context. The authors also note that adding documents may provide useful information while increasing the amount of material over which the model must decide.[19] The study dates from 2023 and does not claim that every current model behaves identically. It does refute the simple equation “it fits = it will be used reliably”.
Product memory is another layer. It may retain preferences, summaries, project instructions or previous conversations. That supports continuity but creates the risk of a stale assumption. The system may “remember” that a project uses an old technology, that a client prefers a rule since withdrawn or that a proposal was approved even though it later changed.
How to test a long document
Place several known decisive facts in different parts of the document and observe whether the system finds them. Ask about a contradiction between the beginning and an appendix. Require a page or section citation. Repeat the task with the attachments in a different order. And do not score only the correct answer: record whether the system admitted that material was missing or filled the gap itself.
For a large corpus, well-designed retrieval with a smaller relevant context is often better than mechanically inserting everything. “The whole company in the prompt” is not a knowledge-management strategy. It is often merely a more expensive way to hide the decisive sentence in the middle of noise.
14. Privacy is not a switch. It is the entire data flow
The question “is my data used for training?” is important but too narrow. An organisation must also know where prompts and answers are stored, for how long, who can access them, whether audit records are kept, which third parties process search and whether a connected agent sends data to another tool. “We do not use it for training” does not mean “we retain nothing” or “no subprocessor handles anything”.
Official materials from the providers show material differences between personal and work accounts. OpenAI, Google, Microsoft, Anthropic and Perplexity describe enterprise regimes with different commitments from ordinary consumer use.[2][4][7][10][13] It does not follow that every enterprise service is suitable for every sensitive datum. It follows that security cannot be judged by a logo without the plan name, contract and configuration.
The European Union adds a legal framework. European Commission guidelines state that the transparency obligations under Article 50 of the AI Act apply from 2 August 2026.[21] For high-risk systems, Article 26 requires deployers, among other duties, to provide competent human oversight, monitor operation and, under specified conditions, keep logs.[22] An ordinary chat used to draft a slogan is not the same as a system affecting employment or access to a service. The specific classification requires legal assessment.
The practical minimum is a data-flow map. It begins with the user, continues through the application, model, search, connectors, logs and exports, and ends at the point where data can be deleted or restored. Without that map, a privacy claim describes only one section of the pipeline.
15. Token price is the price of raw material, not of the result
The market uses several pricing languages at once. A consumer pays a monthly subscription with limits. A company pays per seat, for governance and sometimes additional credits. A developer pays for input and output tokens, context caching, search, code execution or other tools. An agent may make dozens of model calls even though the user entered one sentence. Directly comparing price lists without understanding the workflow is therefore misleading.
A cheap model may be economically excellent for classifying a million short items. The same model may be expensive on a task it repeats three times, consumes extensive context for and ultimately requires half an hour of human correction. An expensive model may be unnecessary for predictable extraction. It may still be cheaper for a one-off decision with a high cost of error if it reduces failed attempts and verification work.
Licence or API + tools + human instruction + verification + corrections + cost of errors + interrupted work + future migration.
A result is cheap only after it has passed review and can actually be used.
Variation belongs in the cost as well. If the same task succeeds once in five attempts, the average cost of a correct result is not the price of one run. If the system slows at peak time or exhausts a limit in the middle of a working day, waiting is a cost. If projects, memory and automations are tightly bound to one environment, future departure is also a cost.
It is therefore useful to measure cost per accepted output, not cost per message. Accepted means that the result met criteria defined in advance, passed review and did not cause a later correction. This metric can reverse the order that looked obvious in the price list.
16. What to choose for which kind of work
A sensible choice begins not with a brand but with the type of work and its control mechanism. The following recommendations are not a product ranking. They are a signpost towards a first pilot. For each category, it makes sense to choose two candidates and compare them on the same data, in the same language and with the same time budget.
| Need | First candidate group | Decisive test | Unacceptable shortcut |
|---|---|---|---|
| One general personal AI | Broad work applications from OpenAI, Anthropic or Google | Your own mix of text, data, files and research | Choosing from one impressive demonstration |
| Research and current sources | Perplexity, web modes of large platforms, Grok | Citation support, primary sources and contradictions | Counting links instead of checking them |
| Long-form writing and editing | Claude, ChatGPT or Gemini with project context | Facts, tone, exceptions, consistency and revision | Confusing fluency with correctness |
| Programming | Codex, Claude Code, GitHub Copilot and other agentic tools | A real repository, tests, diff and safe rollback | Measuring only a generated snippet |
| Microsoft 365 or Google Workspace | The native enterprise assistant of that ecosystem | Permissions, source age, audit and oversharing | Ignoring the condition of company data |
| Sensitive or regulated data | An enterprise contract, private cloud or self-hosting | The entire data flow, legal regime and incident process | Relying on a consumer account |
| High volume of repeated tasks | APIs and smaller or cheaper models with controls | Cost per accepted output and boundary cases | Deploying the most expensive model for everything |
Czech use requires a separate test. An English benchmark does not measure whether the system preserves the meaning of a legal term, inflects a name correctly, distinguishes a decimal comma, understands a Czech institution or translates a local abbreviation into a nonexistent English equivalent. The test set should include diacritics, long sentences, tables in Czech formats, local realities and text for which the correct answer is “the source is insufficient”.
Some teams ultimately need several systems: one for research, a second for writing, a third inside the IDE and a fourth in a secure enterprise environment. That is not a failure of standardisation. It is the same decision as using one tool for accounting and another for graphic design. The condition is that the division remains understandable, data does not move without control and people know where final authority resides.
17. A two-week test that says more than a hundred rankings
A choice can be made without a six-month tender and without an impulsive purchase. A short, disciplined pilot is enough. Its purpose is not to prove that a candidate can do something impressive. It is to find the situations in which it fails and determine whether the organisation can catch them.
- Days 1–2: List 20 to 40 real tasks, their frequency, cost of error and current time cost.
- Days 3–4: Prepare identical inputs, mandatory points and an anonymous scoring sheet.
- Days 5–7: Repeat every important task at least three times and keep an error register.
- Days 8–10: Verify connectors, permissions, export, logs, deletion and safe rollback.
- Days 11–12: Calculate the cost of an accepted result, including review and interrupted work.
- Day 13: Run stress tests: conflicting sources, an incomplete prompt, an old document and a prohibited action.
- Day 14: Decide using the matrix, document limitations and set the next review date.
Where possible, the evaluator should not know which system produced the output. For text, they assess factual accuracy, completeness, tone and the number of corrections. For research, source support. For code, test results, diff size and regressions. For agentic work, the number of human interventions, irreversible steps and the quality of the record. Every failure is classified: confabulation, omission, context, tool, permission or interruption.
Conditions must be comparable: the same time, the same tools and the same number of attempts. If one product may search the web and another may not, we are testing finished workflows, which may be entirely appropriate — but we must call it that. If we want to test models, we must standardise the work apparatus as far as possible.
A pilot also needs disqualification conditions. Examples include a single unreported change to production data, use of a prohibited source, a nonexistent citation in a decisive document or an inability to export the record. A high average score must not outweigh an error whose impact the organisation cannot accept.
The final output is not merely the winner’s name. It is an operating rule: what the system is and is not used for, which data it may see, what a person must approve, how output is verified and when the decision will be reopened. Models, plans and terms change. Without a review date, even a good pilot gradually turns into company folklore.
18. Intelligence is not the only hidden cost. Dependency is one too
The first generation of AI subscriptions sold more messages. The next sold a better model, images, voice and search. Today’s products sell something larger: the place where a person searches, writes, analyses, programs, keeps working memory and increasingly takes action.
Providers therefore compete for more than the best answer. They compete for work context: documents, connectors, history, permissions, agents and habits. The more convenient an ecosystem becomes, the higher the cost of leaving may be. Projects can be hard to export, automations rely on proprietary tools and people stop distinguishing their own process from the product in which they built it.
This does not mean integration is bad. Integration often creates the greatest benefit. It means it should be purchased deliberately. An organisation needs portable prompts and data, documented interfaces, the ability to change models, an emergency process without AI and a regular test of whether the original reasons for its choice still hold.
The best system is therefore not one that never makes mistakes. We do not have such a system today. Nor is it necessarily the system with the highest number, longest context or greatest citation count. It is the system whose capabilities match the actual work and whose errors are visible, reversible and economically and humanly tolerable.
When AI merely proposes a sentence, we are choosing an assistant. Once it reads internal data, selects sources, writes code, changes files and prepares decisions, we are also choosing a new structure of responsibility. The question is no longer: “Which AI is the smartest?”
The right question is: Which error will we see in time, who can reverse it and who will bear its cost?
— Jiný Kontext
