Inference Ecosystem

Inference Ecosystem
Getting started with Apertus is easier than ever.
September 17, 2026

Apertus is being deployed across a vibrant ecosystem from on‑premise deployments to high‑throughput APIs. Find out more in this update.

Apertus 1.5 adds image understanding competitive with leading open-weight models at scale, experimental audio understanding, stronger tool use, and thinking mode from our second-generation post-training pipeline. You can read more about this in our July announcement.

Some of these capabilities need to be enabled through configuration changes and upgrades on the inference side. Pioneering organizations have gone above-and-beyond to support our work this summer, in some cases contributing their engineer’s expertise and code commits to key open source components to benefit the entire community.

§   All logos shown above are trademarks of their respective owners.

Over the summer, we became thanks to our enthusiastic community the highest downloaded model from Europe on Hugging Face, with over 4 million downloads in total across official Apertus releases @swiss-ai

A graph of Apertus Model downloads over time

Model name and versionRelease yearDownloads/month (thousands)
Apertus 8B Instruct2025472
Mistral Small 3.1 24B2025462
Ministral 3 14B Instruct2025417
Devstral Small 2 24B2025275
Apertus v1.5 8B2026201
Mistral Medium 3.5 128B202689
Mistral Small 4202654
Apertus 70B Instruct202522
Apertus v1.5 70B202617

Monthly top model downloads, main official repositories, as of 21.9.2026 from Hugging Face

Community Evaluation

One month before the Apertus 1.5 release, 15 organizations evaluated the 8B and 70B model weights, represented by roughly 50 engineers and experts. This pre-release exercise covered multiple areas of deployment readiness: quantization, alignment, multimodal input, tokenizer distribution, and general compatibility. It encouraged the production of quantizations and community builds for several inference frameworks.

We are grateful to evaluators and contributors at Artificialy, Begasoft, Exoscale, Federal Court (BGER), Infomaniak, Liip, OnPrem.ai, Phoeniqs, Public AI, Puzzle AG, stepping stone, Swisscom, Switch and VSHN who participated in the pre-release feedback round.

A limited-scale performance evaluation was run with several providers of the 70B model. To give an idea of the difference in the real-world speed of the various APIs on offer, the table below shows anonymized scores on simulated workload performance (Latency, Throughput) using the open source tool GuideLLM, in order of ascending relative latency:

ProviderLatency (ms)Input (tokens/s)Output (tokens/s)TTFT (ms)
CSCS-17967103
P31.091576187
P51.1815058171
P11.3114356469
P42.63923581
P23.297929223

CSCS is our own research data center in Lugano, from where the tests were run on 16.9.2026.
TTFT denotes time to first token. Lower latency and TTFT, and higher throughput, are better.


General Availability

The fully open model is available for download in two configurations - an 8B version that runs on many laptops and workstations, and the full-scale 70B parameter model for server-class hardware.

We are glad to announce general availability of third-party services that provide access and support of Apertus 1.5:

ProviderModelsLocation¹Per-Token²Docs
Swisscom70B🇨🇭cloud.swisscom.ch
Infomaniak70B🇨🇭infomaniak.com
PHOENIQS70B🇨🇭kvant.cloud
Safe Swiss Cloud70B🇨🇭safeswisscloud.com
OnPrem.ai70B🇨🇭onprem.ai
stepping stone8B🇨🇭stepping-stone.ch
Public AI8B, 70B🌍publicai.co
Featherless70B🌍featherless.ai
AWS Sagemaker8B, 70B🌍aws.amazon.com
Microsoft Azure8B, 70B🌍github.com
Google Vertex AI8B, 70B🌍cloud.google.com
Information is current as of September 2026.
1. Physical location of the datacenter where user data is processed.
2. Per-token rates are available for metered Apertus 1.5 model usage.

Providers

You can find the links to current general-access providers who we are working with on our Get Started page. More technical instructions for deploying our models to several additional cloud providers are also available in the technical deployment guides. Please contact us if there are others that you are using.

Swisscom

“Switzerland shouldn’t only consume AI, it should be able to build it. Apertus 1.5 shows the Swiss research community can deliver a genuinely open, genuinely capable model — and Swisscom’s job is to make it usable: hosted in Switzerland, hardened for regulated industries, available from day one.” — Sarah Levy, Head of Swiss AI Platform

Swisscom is offering 1000 keys for Apertus during the Swiss {ai} Weeks hackathons.

Documentation: cloud.swisscom.ch

⚙️ 70B  ✅ High availability  ✅ Security certification  ✅ Strategic Partner

Infomaniak

“Robust and sovereign European cloud offering LLM access with full data protection, 100 % renewable energy, and a clear no‑logs policy—letting you integrate any open‑source model into your applications while keeping your data securely hosted in Swiss data centers.”

The Euria app supports Apertus 1.5 on mobile phones and in the kSuite collaboration suite.

To get started: Discover Euria, or connect to AI services for API access to Apertus.

⚙️ 70B  ✅ Swiss data center  ✅ Per-token rates ✅ Code contributor

PHOENIQS

“With Apertus 1.5’s support for agents, tools and out-of-the-box EU AI Act compliance, thanks to Switzerland, the Continent now joins the race hereby dominated by the Americans and Chinese.”

The PHOENIQS AI platform brings Apertus to regulated enterprise and public sector workloads.

See model documentation

⚙️ 70B  ✅ Swiss data center ✅ Enterprise scalability ✅ Consulting support

OnPrem AI

“Replace any cloud AI with local enterprise AI servers. It is literally plug&play, thanks to compatible APIs and the latest LLM models, managed through a user-friendly interface.”

Offers quantized, full-featured Apertus builds at roughly half the rated GPU power of the reference deployment, approximately 35% less energy per generated token than the FP8 alternative, and 99% of FP8 MMLU quality in a 48 GiB checkpoint.

Details: Blog post, Models

⚙️ 70B  ✅ On premise deployment ✅ Consulting support ✅ Code contributor

stepping stone

“Use AI models directly on Swiss infrastructure — without sending data abroad. stepping stone offers leading open-source models as a managed service: ready to use, with data sovereignty and personalised support.”

The model deployment is openly documented and operated on Swiss infrastructure.

Consult the solution offer

⚙️ 8B  ✅ Swiss data center ✅ ISO certification ✅ Consulting support

Safe Swiss Cloud

“Choose from a rich catalog of sovereign LLMs – all with the same strict privacy and compliance guarantees. Safe Swiss Cloud’s Private AI (PAI) services combine a broad selection of open-source LLMs with a consistent security, privacy and compliance foundation. You keep full control over data, infrastructure and model choice, while we provide the sovereign hosting and operational excellence”

Apertus 1.5 is available, optimized for multilingual dialogue use cases.

Get started with fully private AI

⚙️ 70B  ✅ Swiss data center ✅ ISO certification ✅ Per-token rates

Public AI

“The Public AI Inference Utility is a nonprofit, open-source project. Our team builds products and organizes advocacy to support the work of public AI model builders like the Swiss National AI Initiative, AI Singapore, AI Sweden, and the Barcelona Supercomputing Center.”

An Apertus demo is freely accessible on Hugging Face and on publicai.co, while a not-for-profit cooperative brings local focus to the vision.

Join the active & growing community

⚙️ 8B & 70B  ✅ Free demo ✅ Large community ✅ Social enterprise


What This Means for You

From academic research groups that demand strict data sovereignty to startups that require seamless API access, the breadth of providers ensures that Apertus can be plugged into your workflow in the way that best matches your operational and ethical requirements.

  • On‑Premise deployments leverage the model securely behind your own firewalls, ideal for handling sensitive datasets or regulated environments.  
  • Cloud and SaaS partners deliver ready‑to‑use endpoints, automatic scaling, and simplified billing, letting you focus on building rather than infrastructure.  
  • Community initiatives like Public AI keep a free or low-cost gateway running on shared compute, preserving inclusivity with an easy way to demo new services.

Please check back on our Get Started website for links, contact us if you are aware of other providers, or have a question about getting Apertus running on another platform.

Looking Ahead

Just as Apertus 1.5 expanded its knowledge base and instruction‑following prowess, this growing partnership network expands its reach without diluting the open‑source principle that powers it. 

We thank every provider and contributor who has turned the vision of a truly universal, multilingual LLM into a practical reality, fully in the sense of the Open project that is one of our primary research goals. Here is what’s next on the Apertus Roadmap:

Apertus 2.0 - 2027 - a magnitude larger model, developed to the same standard of transparency and trustworthiness: deeper understanding, better agentic capabilities, a more globally fair vocabulary, reliability in long-running processes, and more.

Further development of Apertus 2.5 will be based on feedback from our users. Stay tuned for the latest updates via our Inside Apertus newsletter. Join us at upcoming events to learn what innovative solutions emerge as developers worldwide integrate these models into new research, products, and services.

If you build with or use Apertus, reach out for a Showcase listing or through our Contact page. Don’t hesitate to send us a link or screenshot of a cool prompt or use case! We will promote them in presentations, give coverage in the News area, as well as make them visible to our network at the ETH and EPFL AI Centers and CSCS.


🌍 Tag #Apertus and connect with us on social media: Mastodon / LinkedIn / X / Bluesky

⭕️ Join the discussion in the community forums on Hugging Face / Discord