The AI model costs almost nothing. Everything else costs
A department of 60 people pays around 36 francs a year for the model calls. For a ready-made interface the same department pays between 25,000 and 58,000. Anyone discussing AI costs who knows only the first figure is budgeting for the wrong thing.
Using AI in regulated settings, 4 parts
-
Part 1
What is actually allowed?
Factsheet, data protection, official secrecy. And what running it yourself solves.
-
Part 2
Does it have to be the strongest AI model?
Measured performance on administrative texts, licences, the limits of size.
-
Part 3 · you are here
The cost driver is rarely the AI model
Cloud versus running it yourself, hardware, power, and when it tips.
-
Part 4
Swiss providers, and what they leave to you
Ten providers, three operating models, and what to watch for.
In short
- The model calls are the smallest item, by a distance. Summarising a ten-page case file costs 0.12 Rappen. What actually hits the budget is the interface your people work with, and the upkeep.
- On token costs alone, running it yourself practically never pays. For a rented card to come out cheaper, some 1,400 staff would have to summarise 20 case files every working day. With upkeep it is around 3,850.
- The strongest objection to running it yourself comes out of your own sums. A Swiss provider on a clean contract usually solves the data question just as well and costs a fraction.
In almost every conversation about open models, somebody says that running it yourself works out cheaper. Worked through, that is not true for most of public administration, and not by a small margin.
Every figure here is calculated with provider prices as at 8 August 2026.
Three terms, in case you started here
- Model. A program that continues a piece of text. ChatGPT is an interface to such a model, not the model itself.
- Token. The unit models are billed in. A fragment of text, sometimes a word, sometimes a syllable. As a rule of thumb, one A4 page of administrative German comes to roughly 500 tokens. That lets you convert any price list into pages.
- Graphics card. The component a model runs on. The name is misleading: this has nothing to do with what appears on a screen. Whether the card sits with you or with a provider decides the whole cost comparison below.
These terms are covered in more detail in Part 1, the arithmetic behind them in our primer 4B, 70B, 397B-A17B.
Four routes, four price tags
Before a single figure is named, it has to be clear what you are actually buying. Between a model and a member of staff working with it sit several layers, and depending on the route you pay for a different number of them.
| Route | What you get | What is still missing | 60 people cost |
|---|---|---|---|
| Ready-made interface Safe Swiss Cloud, PeakPrivacy, AlpineAI |
Your people open a page in the browser, log in and type their instruction into a field, much as with ChatGPT. They drag documents in with the mouse. | Roll-out, training and sign-off on the categories of data. Plus the link to your specialist system, if you want one. | CHF 2,100 to 4,800 a month |
| Model access Infomaniak, Safe Swiss Cloud, PHOENIQS |
An access key. With it a program can address the model. No interface, no screen. | The interface. Somebody has to build one or set up a ready-made one and look after it. | CHF 3 a month |
| Rented computing power Exoscale, cloudscale, Nine |
Remote access to a server with a graphics card that sits in the provider's data centre. Nobody sends you hardware. | Load the model, install the serving software, set up access and logging, provide the interface, look after all of it continuously. | CHF 407 to 1,572 a month, without upkeep |
| Your own hardware purchase |
A card in your own server room. | The same as above, plus procurement, a spare unit and rack space. | CHF 461 a month, without upkeep |
Calculated for 60 staff, each having two ten-page case files summarised on each of the 21 working days, so 2,520 requests a month. Prices are mostly quoted excluding VAT. Model costs with Gemma 4 31B at Infomaniak. Licence prices: CHF 35 per person at Safe Swiss Cloud, CHF 74 to 80 at AlpineAI. Rental prices for a card with 96 GB, the lower value for office hours, the upper for round-the-clock operation. Purchase written off over four years plus electricity, without a server. Upkeep is not included in the bottom two rows, more on that shortly. The biggest lever is not in the table: licences you pay per person, tokens per request. If only 20 of your 60 people need access, the first row drops to a third; the second stays much the same.
The jump between rows one and two is the most important finding in this series. The same amount of work costs 3 francs a month as a bare model call and 2,100 to 4,800 as a finished product. What you are paying for is not the model but the interface, the operation, the support and the liability. Anyone wanting to provide those layers themselves saves the licence fee and takes on the work.
What one request costs at a provider
Several Swiss providers run open models and bill by tokens. The prices, all retrieved on 8 August 2026, in francs per million tokens:
| Model | Provider | Input | Output |
|---|---|---|---|
| Nemotron 3 Nano 30B | Infomaniak | 0.05 | 0.20 |
| Gemma 4 31B | Safe Swiss Cloud | 0.136 | 0.374 |
| gpt-oss-120b | PHOENIQS | 0.1154 | 0.4617 |
| Gemma 4 31B | Infomaniak | 0.20 | 0.40 |
| Mistral Small 4 119B | Infomaniak | 0.20 | 0.75 |
| Apertus 1.5 70B | PHOENIQS | 0.6195 | 2.2201 |
| Apertus 1.5 70B | Infomaniak | 0.70 | 2.50 |
| Qwen3.5 397B-A17B | Infomaniak | 0.80 | 3.60 |
List prices in CHF per million tokens, retrieved on 8 August 2026. Careful at small volumes: Safe Swiss Cloud also applies a minimum spend of CHF 95 a month, PHOENIQS a subscription of CHF 50 a month. Anyone using only a few francs' worth pays the minimum there, and the cheapest token price becomes the most expensive offer. Swisscom runs Apertus as well but publishes no prices for it.
At 500 tokens per A4 page, these prices convert into work. A ten-page case file in, one page of summary out, comes to 5,000 tokens of input and 500 of output:
Ten pages summarised. A thousand of them cost CHF 1.20.
Around four times as much. A thousand of them cost CHF 4.75.
30 pages of transcript down to 3 pages of minutes, with Apertus 70B.
This order of magnitude is also why the question 'cloud or run it ourselves' is rarely a question of cost.
The model call does not replace working time, it shortens it. A summary that goes into a formal ruling has to be read and answered for. Half an hour becomes ten minutes, not nothing.
What the hardware costs
Anyone wanting to run it themselves can rent or buy. Worth being clear about: with every rental offer the card sits in the provider's data centre. You get remote access, not a delivery. Only with a purchase does hardware stand in your own building.
Three rental tiers
AIVITY runs cards with 32 to 96 GB and supplies the models pre-installed. The smallest tier is enough for models up to around 10 GB, the middle one up to around 20 GB. Shared means other customers are running on the same card. Anyone citing data sovereignty as the reason has to check that.
Exoscale and cloudscale.ch rent out an RTX PRO 6000 with 96 GB billed by the second, Exoscale at CHF 2.1528 an hour, cloudscale at less than CHF 2.20. The range above is calculated with the Exoscale price: the lower value for 189 office hours a month, the upper for 730 hours of round-the-clock operation.
Nine.ch rents GPU servers from CHF 1,104 a month, two locations in Zurich. For a 70B model this is the realistic entry tier. Which of the cards on offer you need for it is something to settle in the quote.
When you buy, it is the memory that counts, not the speed
A model has to fit into the card's memory in full, otherwise it does not run at all. That is why the last column gives the price per gigabyte.
| Device | Price | Memory | Speed | Per month over 4 years | CHF per GB |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell Max-Q graphics card, needs a server |
11,660 | 96 GB | 1,792 GB/s | 243 | 121 |
| RTX PRO 4500 Blackwell graphics card, needs a server |
3,144 | 32 GB | n/a | 66 | 98 |
| GeForce RTX 5090 graphics card for games |
4,999 | 32 GB | 1,792 GB/s | 104 | 156 |
| Apple Mac Studio M3 Ultra complete computer |
4,739 | 96 GB | 819 GB/s | 99 | 49 |
Day prices in CHF at digitec.ch, retrieved on 8 August 2026. 'Speed' is the memory bandwidth; it determines how fast the model produces text. On the Mac the memory is shared between the computer and the model, and only part of it is available to the model, which is why the price per gigabyte comes out higher in practice than shown. Straight-line depreciation over 48 months, with no residual value.
The card built for games is the most expensive on memory. The RTX 5090 costs 1,855 francs more than the professional RTX PRO 4500 and has the same memory. Fast it certainly is. For use in a data centre the professional cards are still the better choice, among other reasons because their memory carries error correction. That guards against a bit flipping at random in memory and quietly producing a wrong result.
The Mac is considerably cheaper on memory and comes as a complete computer. For a first attempt in your own building it is often the quickest route: one device, one power cable, no server procurement. For multi-user operation it is ruled out, because the widely used server software for that relies on NVIDIA technology, which the Mac does not have. On top of that, there is no error correction, no remote management and no form factor for the server rack.
What the electricity costs when the card runs continuously
An RTX PRO 6000 Blackwell draws at most 600 watts. At 730 operating hours a month that is 438 kilowatt hours. At the median price of 27.7 Rappen that the Federal Electricity Commission gives for 2026, that comes to 121 francs for the card alone. With power supply losses, the rest of the server and the cooling we allow a factor of 1.8, so around 218 francs. If the card runs only during office hours, it is 57 francs.
Round-the-clock operation, written off over four years, plus electricity. What is needed around it comes on top of that.
Round the clock, in the provider's data centre, with network and replacement.
This comparison is tilted in favour of buying, and heavily so. On the left is a bare card; on the right, a ready-to-run service with network, storage, spare parts and somebody who responds when it fails. The server around the card, a second unit for outages and the rack space are not included in the 461 francs. Anyone doing the sums honestly lands at roughly double for the purchase, and thus close to the rental.
On the assumptions: the electricity price is the median for households in 2026. Commercial tariffs are usually below that, so our electricity bill is more likely too high than too low. The factor of 1.8 for server and cooling is an assumption from practice, not a measurement. The 600 watts apply to the workstation version. For the more frugal Max-Q variant, the one in the purchase table, we have no figure of our own for power draw. Our electricity bill is therefore an upper bound.
When running it yourself starts to pay
The break-even point is simple to work out: fixed monthly costs divided by the cost of one request at a provider. Below it the provider is cheaper, above it running it yourself.
| Operating model | Model and task | Break-even | Which means |
|---|---|---|---|
| Managed card, middle tier CHF 690 a month |
Gemma 4 31B summarising case files |
575,000 requests a month | around 1,400 staff with 20 case files each a day |
| The same card, but with upkeep CHF 690 + 1,250 a month |
Gemma 4 31B summarising case files |
1,617,000 requests a month | around 3,850 staff with 20 case files each a day |
| Your own card, office hours CHF 407 a month |
Apertus 70B, shrunk minutes of meetings |
28,600 requests a month | around 270 staff with 5 sets of minutes each a day |
Our own calculation with Infomaniak's token prices. With the cheaper prices from Safe Swiss Cloud the first break-even point would sit at around 796,000 requests instead of 575,000. One case file: 10 pages in, 1 page out. One set of minutes: 30 pages of transcript in, 3 pages out. Converted at 500 tokens per A4 page, 21 working days a month, office hours at 9 hours a day. The number of requests per person per day is deliberately generous. Two caveats: Apertus 70B takes up 144 GB unquantised and fits on a single 96 GB card only when shrunk; what that costs in accuracy we did not measure. And we did not check whether a single card can process the volumes named at all. For the top two rows that is unlikely; they show where the arithmetic lands, not an achievable scenario.
For municipalities and mid-sized administrations the break-even point is out of reach. Large cantons and the federal government reach it as soon as large volumes of text run through an expensive model, as in the third row. An organisation with 6,000 staff therefore calculates differently from a municipality with 40.
The item that appears on no price list
A self-hosted model needs attention: setting up, keeping current, watching throughput, fixing outages, setting up access and logging. Reckoning a post at a fully loaded cost of 150,000 francs a year gives:
More than the purchased card including electricity.
More realistic as soon as several teams draw on it.
For comparison, round-the-clock operation including cooling.
Ten per cent of a post is convenient arithmetic and awkward in practice. You cannot advertise a ten per cent post for a specialist who knows how to run models. In practice this item means either a share of an existing post that has to be given the time, or a contract with a service provider at hourly rates above the ones calculated here. The figure of 150,000 francs in fully loaded cost is on the low side for this qualification.
The Swiss Federal Audit Office supplied a reference point from the other direction in 2025. In its report on federal chatbot projects it found 21 of them. For the category with self-hosted open models it gives a range of 400,000 to 2 million francs.
Those millions are not running costs. They are project costs for projects that build a whole application: links to specialist systems, interface, legal review, security concept, training, project management. The infrastructure this article is about is a small part of it. The Audit Office also noted that all of these projects were still at the research or development stage.
The strongest objection to running it yourself
It goes like this: if data sovereignty is the decisive reason, then a Swiss provider with a clean data processing agreement and a data centre in the country solves the same problem, at a fraction of the cost and without anyone at your end having to be reachable.
For most administrative work that is the right answer. That leaves three cases for running it yourself. Data that for legal reasons cannot leave the building and for which even a domestic processor will not do. Sustained load on a scale where token billing genuinely tips the balance. And a model that has to be adapted so deeply that no provider carries it in that form.
What we recommend
Do not justify running it yourself on cost. Below the break-even point the sums do not hold. Two other justifications do: data that cannot leave the building, and a model that has to be tuned to your language. For the second you need open weights, not necessarily your own hardware. An adapted model can also be housed with a provider that runs models exclusively for a single customer.
Budget for the layers, not for the model. The model calls are a rounding error. What you have to budget for is the interface, the roll-out and the upkeep. For a department of 60 people the realistic order of magnitude lies between 20,000 and 58,000 francs a year. At the lower end sits a lean self-hosted setup: 407 francs of rent plus 1,250 francs of upkeep, times twelve. At the upper end sit ready-made licences at 25,000 to 58,000, depending on the provider and on how many of your people really need access.
And budget the roll-out separately. Legal review, impact assessment, integration, training and acceptance testing fall due once and appear in none of the tables above. In federal projects it is exactly these items that dominate.
Before you compare costs
- Only compare what contains the same layers. A licence price per person includes the interface, a token price does not. The difference is not a saving but an open task.
- When running it yourself, count the upkeep in before you compare hardware. Ten per cent of a post alone costs more than the purchased card including electricity, and ten per cent is not enough for a service running round the clock.
- With offers that carry a minimum spend, check whether your usage comes anywhere near it. At small volumes the minimum decides, not the token price.
- With a rental offer, ask where the card is physically located, and have the location guaranteed in the contract. A selection in the order form is not enough.
- Plan for the model change. When a new version appears, that is not an installation but acceptance testing all over again against your own test cases.
Disclosure
- digitario advises administrations and companies on selecting, procuring and introducing AI systems. This article ends with an offer to talk.
- It was written without a commission, without payment and without prior sight by any of the providers named. None of them supplied material that is not publicly available.
- All prices come from public provider pages, and the arithmetic is set out in the text. Where we could not measure something, it says so.
Open questions
Do you need a figure for the budget request?
In an initial conversation we go through your expected volumes, map them onto the four routes and work the order of magnitude through, roll-out and upkeep included. You can take the result away and work with it yourself.
Sources
- Our own research. Provider prices from Infomaniak, Safe Swiss Cloud, PHOENIQS, Exoscale, cloudscale.ch, Nine.ch, AIVITY, AlpineAI and PeakPrivacy, each from the provider's own pages on 8 August 2026
- Our own research. Purchase prices and technical data of the cards and devices, digitec.ch on 8 August 2026. Day prices, not procurement terms.
- Switzerland. Federal Electricity Commission ElCom, press release on electricity prices for 2026 of 9 September 2025 · elcom.admin.ch
- Switzerland. Swiss Federal Audit Office, report 24181 of 25 February 2025 on federal chatbot projects · efk.admin.ch
- Our own work. The full calculation including every step exists as a script in the project and can be repeated with your own figures.