Five steel laboratory weights in a row, largest to smallest, the second smallest carrying a teal ring
Part 2 · Performance, licences, the limits of size Figures pulled from the benchmark interface on 8 August 2026

Which AI models are good enough. And why the strongest are not among them

For summarising and answering questions, open models are now close to the best in the world. That holds only for the mid-sized ones. What sits at the very top you cannot run, and one Swiss model does worse than many expect.

By Mike Jenke Reading time 10 minutes As at 10 August 2026

Using AI in regulated settings, 4 parts

  1. Part 1 What is actually allowed?

    Factsheet, data protection, official secrecy. And what running it yourself solves.

  2. Part 2 · you are here Does it have to be the strongest AI model?

    Measured performance on administrative texts, licences, the limits of size.

  3. Part 3 The cost driver is rarely the AI model

    Cloud versus running it yourself, hardware, power, and when it tips.

  4. Part 4 Swiss providers, and what they leave to you

    Ten providers, three operating models, and what to watch for.

In short

Once the legal question is settled, the second one follows: can an open model do the job at all? For public administration there has been a solid measurement since July 2026. It comes from Germany, though, and that needs keeping in mind as you read.

Four terms, in case you started here

  • Model. A program that continues a piece of text. You give it an instruction and a text, it gives you a text back.
  • Weights. The files a model consists of, often several dozen gigabytes. 'Open weights' means these files can be downloaded and kept.
  • Parameters. The number of adjusted values inside the model, given in billions. '31B' means 31 billion. The figure says how much room the model needs, not how well it does your job. Do not rely on the number in the name: 'Gemma 4 E2B' has 5.1 billion, not 2, and 'Mistral Small 4 119B' is larger than a 31B despite the 'Small'. The column in the table is more reliable than the product name.
  • Quantisation. A procedure that shrinks the model files so they fit on cheaper hardware. It costs a little accuracy.

These terms are covered in more detail in Part 1, the arithmetic behind them in our primer 4B, 70B, 397B-A17B.

What was measured on administrative texts

MÖVE is the model comparison for public administration, run by the Innovation Hub of the Bundesdruckerei, the German federal printing office, together with the Federal Office for Information Security and the Fraunhofer Institute AISEC. The tests use German-language data from real administrative settings. The test data stays unpublished, the operators say, so that models cannot be trained on them. The price of that: nobody can verify the measurement from outside.

German administrative language is not Swiss administrative language. Different terms, different enactments, plus Helvetisms and the three other official languages. For the question of which class of model can handle an administrative task, MÖVE is still the best source available in the German-speaking world. As a ranking for your procurement it will not do. A comparable Swiss benchmark does not exist, as far as we could find.

We pulled the values straight from the benchmark's web interface. The measurement dates from 7 July 2026 and covers 50 models. All values in this table run in the same direction: 0 is the worst result, 100 the best. That applies to the 'Rarely invents' column too. A high value there means the model rarely makes something up.

ModelOriginParams in bnSummarisingAnswering questionsRarely inventsLicence
GPT-5.4USA45080.7972.0591.46closed
Claude Sonnet 4.6USA44079.8769.4592.88closed
Gemma 4 31BUSA, Google31.378.6270.6388.66Apache 2.0
Qwen3.6 27BChina, Alibaba27.078.3369.4389.51Apache 2.0
Gemma 4 E2BUSA, Google5.178.3966.8262.10Apache 2.0
Phi-4USA, Microsoft14.776.9768.8683.49MIT
GPT-OSS 120BUSA, OpenAI11775.8170.3382.93Apache 2.0
Apertus 70BSwitzerland70.675.4869.7468.49Apache 2.0
Apertus 8BSwitzerland8.066.2864.4652.72Apache 2.0
Teuken 7B v0.4Germany7.551.6955.8535.09Apache 2.0, but not the successor

MÖVE, the German model comparison for public administration. Measured on German administrative texts, not on Swiss ones. Values pulled from the interface on 8 August 2026, measurement dated 7 July 2026. Extract from 50 models. Scale 0 to 100, higher is better throughout. Parameter counts for the closed models are the benchmark's estimates. We did not take the licence and origin columns from MÖVE; we checked them against the models themselves.

Why GPT-5.4, GPT-5.5 and GPT-5.6 all turn up in this series, and three numbers for Qwen as well: the benchmarks tested different model versions at different times. Values from different rankings are therefore not directly comparable, even where the maker's name is the same.

Three values in this table deserve a second look.

For the core tasks the gap has all but closed. On summarising, Gemma 4 with 31 billion parameters sits 2.17 points behind GPT-5.4, which the benchmark puts at 450 billion. On answering questions it is 1.42 points. At values around 78 on a hundred-point scale, that is a gap a caseworker will barely notice in daily work. This model runs on a single graphics card.

Size says little about reliability. Gemma 4 E2B with 5.1 billion parameters summarises practically as well as the 31B, which is six times its size, 78.39 against 78.62. In the 'Rarely invents' column it drops from 88.66 to 62.10. If you are looking for a model to answer questions from the public, that column matters more than any other.

What 'inventing' means here

  • The model is given a text and asked to summarise it. What gets counted is whether the summary contains anything the source text does not.
  • Typical cases: a deadline that does not appear in the document at all. An amount that is slightly off. An article of law that sounds plausible and does not exist.
  • For an administration this is not a question of quality but a legal problem. An invented detail in a case file is inaccurate processing of personal data.
  • From that follows a rule that holds whatever the model: whatever goes into a formal ruling has to have been checked against the original by a person.

A second, independent measurement shows the same pattern from the other side. In the hallucination ranking of the American vendor Vectara, as at 11 May 2026, a Qwen3 with 8 billion parameters makes something up in 4.8 per cent of summaries. GPT-5.5 is at 9.3 per cent, Gemini 3.1 Pro at 10.4 per cent. Here too: bigger does not mean more reliable. Do not carry these values over to the models in the table above, though. What was measured is an older and smaller Qwen version, not the one we recommend further down.

Careful, this ranking runs the other way from the table above: with Vectara a low value is good. And it comes with a second column that we do not reproduce here, the answer rate. A model that often refuses to answer invents less and is not better for it. Anyone drawing on the ranking for a procurement should read both columns side by side. Vectara, incidentally, sells tools to tackle this problem itself, but the method of measurement is disclosed.

Apertus, two truths in one paragraph

For Swiss readers this is the uncomfortable part. Apertus, developed at ETH Zurich, at EPFL and at the Swiss National Supercomputing Centre CSCS, is assessed by two independent sources, and they say different things. Three of the four rows below come from the same measurement, the German administrative benchmark.

RatingApertus 70BScaleWhat that means
Training data transparency
FAccT study, Ireland
Grade A and A+A+ to F Best of all 39 summaries assessed, ahead of OpenAI, Google, Anthropic, Meta and Mistral.
Transparency
MÖVE, Germany
90.480 to 100, high is good Rank 1 in the field, a clear margin over the next model. Confirms the study, though on the same body of data.
Summarising
MÖVE, Germany
75.480 to 100, high is good Mid-field. Behind Phi-4, which is almost five times smaller.
Rarely invents
MÖVE, Germany
68.490 to 100, high is good The weakest figure among the leading models. Gemma 4 31B reaches 88.66 at less than half the size.

The Intelligence Index by Artificial Analysis additionally carries an estimated index value of 2 points for Apertus 70B. We do not reproduce it in the table: the provider expressly marks it as an estimate without a measurement of its own, and it flatly contradicts the measured MÖVE values. A figure like that does not belong in a procurement note.

Both are true at once. Where you have to document what a model learned from, Apertus has no competition. Where precise answers matter, there are better ones. For a procurement that means: Apertus belongs in the assessment when you have documentation duties, and it has to be measured on the same test cases as the alternatives.

There is one question these figures never ask. MÖVE measures German administrative text. Multilingualism, the very argument Apertus makes for itself, does not appear in it. How well the model works in Italian or French is something none of the scores above will tell you. The Canton of Ticino runs a translation tool on a fine-tuned Apertus 8B covering six languages, described in Part 1. If your work is multilingual, the answer for your own case is not in this table; it is only in your own texts.

Version 1.0 was measured. Version 1.5 came out on 24 July 2026, trained on considerably more data. Published benchmark figures for it do not exist to date, and the technical report still describes only 1.0. Anyone wanting to use Apertus therefore has to test it themselves. How to do that is further down.

The licence has a say too, and catalogues go stale

For procurement, the licence column above matters as much as the performance values. The difference in practice:

Apache 2.0 and MIT are established open-source licences. They permit use, modification and redistribution, commercially too, without asking and without a limit on user numbers. The rights holder cannot withdraw permission once granted for an existing version.

Vendor licences such as the earlier Gemma terms contain a usage policy the vendor can change, and in some cases reporting duties or territorial exclusions. For an administration that wants to run an application for years, that is a different kind of risk.

What we found in our own check

  • MÖVE lists Gemma 4 under the licence 'Gemma License'. The model card itself has carried license: apache-2.0 since March 2026. Google switched with Gemma 4; the benchmark still shows the old one.
  • Where to look this up: every open model has a public page on the Hugging Face platform, comparable to a product data sheet. The licence sits right at the top of that page. Search for the exact model name, not the model family.
  • Note the date you checked in the procurement file. A licence detail taken at second hand is always older than it looks, and vendors change their terms between versions.

Why the open frontier does not exist for you

The trade press says open models have caught up. That is true, and it does not help you. The Intelligence Index by Artificial Analysis, an international benchmarking service with no Swiss connection, put Claude Opus 5 at the top on 6 August 2026 with 63 points, followed by GPT-5.6 with 61. The best open model, Kimi K3, reaches 60.

These points are not percentages. The index bundles several hard tests, and 63 is the best value reached in the whole field, not a grade. For public administration it is the wrong measure anyway: a third of it measures systems working on their own and a quarter measures programming. For summarising and answering questions, the MÖVE table above is what counts.

The index is still worth a look, but for a different column. We calculated from the model files how much room these models need:

ModelFile sizeIndexCards neededCost of the cards alone
Kimi K3
China, Moonshot
2.8 TB60around 30around CHF 350,000
DeepSeek V4 Pro
China
1.6 TBn/aaround 17around CHF 200,000
GLM-5.2
China, Z.ai
1.5 TB53around 16around CHF 187,000
Apertus 1.5 70B
Switzerland
144 GBn/atwo, or one if shrunkaround CHF 23,000
Gemma 4 31B
USA, Google
63 GB30onearound CHF 11,700
Apertus 1.5 8B
Switzerland
18 GBn/aruns on a laptopnone

File sizes calculated from the model files on Hugging Face, 8 August 2026, each at the precision in which the model ships. Card count calculated at 96 GB of memory per card, with no headroom for live requests. Price calculated at the day's price of CHF 11,660 for an RTX PRO 6000 Blackwell at digitec.ch, without server, power supply, cooling and rack space. Index: Artificial Analysis, as at 6 August 2026; where no value exists, n/a is shown. 'Shrunk' means quantisation, a procedure that makes the file smaller and costs a little accuracy.

What you can run yourself does not play in the top league. What fits on a single card sits at 30 in the index, not at 60. Anyone who reads 'open models have caught up' and concludes that this holds for running it yourself is confusing two statements. For administrative tasks that is no disadvantage, because there the MÖVE values count and there the gap is small.

What 'runs on one card' does not mean

The figures above are lower bounds. They say how much room the model files take up when nothing else is running. In operation the context memory has to be added on top, where the model holds the conversations under way. It grows with the length of the inputs and with the number of people using it at once.

'Runs on one card' is therefore not a complete statement. It holds for one person with short inputs. It does not necessarily hold for thirty caseworkers each feeding in a case file at the same time. Do not ask your provider whether a model fits on the card, but how many simultaneous requests it can handle there at your document length.

To put that answer in context: in a department of 60 staff, each putting two case files through a day, the requests arrive spread across the day. On average very few are under way at once, at peak times on a Monday morning considerably more. How many that is at the peak we did not measure for this series, so we give no figure. It comes out of a load test with your own documents.

An old brass balance scale with two empty pans, the beam level
A calibrated scale with empty pans. No outside ranking measures what your own cases weigh.

What follows from this

What we recommend

Start with Gemma 4 31B from Google or Qwen3.6 27B from Alibaba. Both are under Apache 2.0, both sit close to the world's best on administrative tasks, both run on a single card. In practice that means: you open an account with a Swiss provider that carries one of these models and select it there from a list. Which providers those are is in Part 4.

Add Apertus if you have documentation duties. On the question of what a model was trained on, it is the most transparent model in the field. Measure it on your own cases against the other two, rather than relying on the values in this article.

And do not let origin replace the test. That a model comes from Google or from Alibaba makes no difference to data protection as long as the files sit with you or with your Swiss provider; the makers never see it. In a procurement you still have to be able to justify the origin, and depending on the committee that is a political question.

How to test this yourself, in two days

  • Collect cases. 30 to 50 real files from your own organisation, spread across the range of difficulty if you can. Without personal data, as long as the legal basis is unsettled.
  • Run them through. The same cases through two or three models, with the same instruction. With a provider that bills by usage, this step costs less than a franc.
  • Judge blind. Two specialists from the subject area assess the results without knowing which model produced them. Three criteria are enough: factually correct, complete, usable tone.
  • Set the threshold beforehand. For example: if more than one summary in twenty contains a factual error, the model is unfit for this purpose. Without a threshold set in advance, the test turns into a matter of taste.
  • Document the outcome. It later becomes part of your procurement file and your impact assessment.
In the comparison matrix We list the open models with licence, operating model, hardware requirement and DACH risk, Apertus 70B among them. It shows which licence applies, what running it yourself demands in hardware and where a model is a problem for the DACH region. To the matrix →

Disclosure

  • digitario advises administrations and companies on selecting, procuring and introducing AI systems. This article ends with an offer to talk.
  • It was written without a commission, without payment and without prior sight by any of the providers named. None of them supplied material that is not publicly available.
  • The performance values come from a third-party benchmark that we neither run nor influence. Where we did checks of our own, the text says so.

Open questions

Which model suits your cases?

Rankings do not answer that, your own files do. In an initial conversation we settle which cases are suitable for a test, which models make the shortlist and what you measure the result against.

30 minutes · free of charge · directly with Mike Jenke
This article is current as at 10 August 2026. Part 2 of a four-part series. The performance values come from the German MÖVE benchmark, measurement dated 7 July 2026, pulled by us from the interface on 8 August. We checked licences and file sizes against the model cards on the same day. Models and values change quickly in this field; the test method at the end stays valid.

Sources

  1. Germany. MÖVE, model comparison for public administration. Innovation Hub of the Bundesdruckerei with the Federal Office for Information Security and the Fraunhofer Institute AISEC. Values pulled from the interface on 8 August 2026, measurement dated 7 July 2026 · moeve.bundesdruckerei.de
  2. United States. Vectara, Hallucination Leaderboard, HHEM-2.3, as at 11 May 2026 · github.com/vectara
  3. International. Artificial Analysis, Intelligence Index v4.1.1, as at 6 August 2026 · artificialanalysis.ai
  4. Ireland. Blankvoort, Pandit and Gahntz, assessing the training data transparency of GPAI models, FAccT 2026, AI Accountability Lab at Trinity College Dublin · doi.org/10.1145/3805689.3806755
  5. Our own research. Licences, parameter counts and file sizes of every model named, taken from the model cards on Hugging Face on 8 August 2026 · huggingface.co/swiss-ai
  6. Our own research. Card price for the worked example, day price at digitec.ch on 8 August 2026
  7. Our own work. digitario, comparison matrix of AI coding tools · digitario.ch/ki-coding-tools-vergleich
Mike Jenke
Mike Jenke digitario GmbH · Zurich

Advises public administrations and companies on procuring and introducing AI systems. Maintains the comparison tables on this site.