The Need for Swiss Sovereign AI

SOKKAN Inference, a Swiss company specializing in AI, has launched a new inference tier named "Swiss." This tier addresses a critical demand from clients who require their data to remain within Switzerland's borders, a requirement often stemming from strict data privacy regulations or corporate policy. While SOKKAN already offered EU-sovereign tiers from French datacenters, the "Swiss" tier is designed for customers who cannot tolerate their data leaving the country, or even, in the longer term, leaving their own physical premises.

This move is particularly significant given the current landscape of AI hardware. SOKKAN needed to procure GPUs in 2026, a period marked by a severe memory shortage impacting the availability and pricing of high-end AI accelerators. The company's search for second-hand NVIDIA RTX 3090 cards proved fruitless, with auctions consistently exceeding market value. This scarcity created an opening for alternative hardware solutions.

Why Intel Arc Pro B60?

The Intel Arc Pro B60 emerged as a viable alternative. Each card boasts 24 GB of VRAM, a 2-slot blower-style cooler suitable for dense server environments, and a native x8 PCIe interface. Crucially, these cards were available near their Manufacturer's Suggested Retail Price (MSRP) at CHF 614 per unit. SOKKAN acquired four of these cards, totaling 96 GB of VRAM, for approximately CHF 2,450. This provided a cost-effective solution for building a sovereign inference tier, especially when compared to the inflated prices of sought-after NVIDIA cards during the shortage.

Four Intel Arc Pro B60 GPUs installed in a server chassis.

Performance Benchmarks and Real-World Use

The "Swiss" tier is powered by a single machine located in Meyrin, Geneva, utilizing this cluster of four Intel Arc Pro B60 GPUs. The hardware is entirely owned by SOKKAN, ensuring complete control over the infrastructure and data flow. This sovereign setup is primarily used for specific inference tasks, particularly for models that are sensitive to latency and require predictable performance, or for clients with stringent data residency requirements.

SOKKAN's initial benchmarks reveal that the system can handle a range of inference workloads. For instance, a medium-sized LLM inference test showed an average latency of 400ms with a throughput of 10 tokens per second. This performance is adequate for many interactive AI applications where real-time responses are not absolutely critical. However, the company is transparent about the limitations. The performance does not scale linearly with additional GPUs in all scenarios, and certain highly demanding, compute-intensive tasks might still push the limits of this setup.

The choice of Intel Arc Pro B60 was also influenced by its specific features. The 24 GB of VRAM per card is substantial, allowing for larger models or batch sizes compared to GPUs with less memory. The blower-style cooler is designed to exhaust heat directly out of the chassis, which is beneficial in densely packed server racks where airflow can be a challenge. The native x8 PCIe interface, while not as fast as x16, is sufficient for inference workloads where bandwidth is not typically the primary bottleneck.

What It Doesn't Do

It is important to be clear about the capabilities and limitations of this sovereign tier. SOKKAN explicitly states that this setup is not designed for training large-scale AI models. Training requires significantly more computational power, memory bandwidth, and specialized interconnects than what is offered by four Arc Pro B60 cards. The focus is strictly on inference – taking a pre-trained model and using it to make predictions or generate outputs.

Furthermore, while the performance is acceptable for many use cases, it does not compete with the top-tier NVIDIA offerings in raw throughput or speed for highly optimized, large-scale inference tasks. Customers requiring the absolute fastest inference speeds for massive deployments might find this tier insufficient. The benchmarks provided by SOKKAN are honest about the scaling limitations and the specific performance characteristics of the hardware for different model sizes and query types.

Broader Implications for Sovereign AI

SOKKAN's approach highlights a growing trend in the AI industry: the demand for specialized, sovereign AI infrastructure. As data privacy concerns and geopolitical considerations intensify, companies are increasingly looking for solutions that offer control, transparency, and data residency guarantees. Building custom inference tiers using readily available, albeit sometimes less mainstream, hardware like Intel Arc Pro GPUs demonstrates a pragmatic approach to meeting these demands.

This strategy allows smaller companies to compete with larger cloud providers by offering tailored solutions that meet specific niche requirements. It also diversifies the hardware ecosystem for AI, reducing reliance on a single vendor and potentially mitigating supply chain risks. The success of SOKKAN's "Swiss" tier could encourage other organizations to explore similar custom hardware solutions for their sovereign AI needs, especially as specialized AI hardware becomes more accessible and competitive.

The choice of Intel hardware is noteworthy. While NVIDIA has dominated the AI GPU market for years, companies like Intel are making inroads with professional-grade offerings. The Arc Pro B60, with its significant VRAM and competitive pricing, presents a compelling option for specific inference workloads, particularly when data sovereignty and cost are primary drivers. This development signals a potential shift in the enterprise AI hardware landscape, where a more diverse range of vendors may find traction in specialized market segments.