I am now genuinely angry and, quite honestly, appalled by how this problem has been handled.
I rented an RS 8000 G12 with 16 dedicated CPU cores based on the AMD EPYC™ 9645. netcup explicitly advertises its Root Servers as providing CPU and RAM resources exclusively for the customer’s own applications, and states that the performance of the booked CPU cores is guaranteed.
I run compute-intensive applications on this server. For these workloads, average throughput is not the only thing that matters; stable and predictable execution time is equally important. Even brief interruptions or a small amount of CPU steal can significantly delay critical computation phases, so the impact on my workload is disproportionately large.
Despite this, under normal load I repeatedly observe substantial CPU steal across all 16 assigned CPU cores. Even more troubling is the extreme fluctuation: at times everything appears normal, only for the steal value to rise sharply shortly afterwards. At peak moments, I have observed values of approximately 10%.
My repeated longer measurements include:
• 2.11% average CPU steal over five minutes
• 4.32% average CPU steal in another measurement
• 3.80% average CPU steal in a further five-minute measurement
• short-term peaks of approximately 10%
• all 16 assigned CPU cores are affected almost uniformly
• in one measurement, the individual assigned cores showed between 4.18% and 4.49% steal
• in the latest measurement, they showed between 3.69% and 3.95%
• approximately 40% of the total CPU capacity was still idle at the same time
With 16 promised dedicated CPU cores, 4.32% steal is equivalent to approximately 0.69 CPU cores being unavailable on average even though the assigned cores have runnable work. For a compute-intensive and latency-sensitive application, this is not insignificant. A short but concentrated interruption during a critical computation phase can cause considerably more damage than the average value initially suggests.
I want to make it clear that I am not claiming CPU steal has already been proven to be the sole cause of the poor overall performance. Steal is simply the clearly measurable symptom visible from within the guest operating system. Other possible causes include limitations or problems involving CPU frequency and boost, SMT sharing, memory latency or memory bandwidth, NUMA placement, storage latency, or other bottlenecks on the physical host. This is precisely why I expect a serious technical investigation of the complete system, rather than a superficial standard response.
For a short period, the problem appeared to have disappeared. Shortly afterwards, CPU steal returned to almost 4%, with further significantly higher peaks. This cannot reasonably be described as a stable or permanent resolution.
The direct comparison is particularly frustrating: according to my latest real-world tests, this RS 8000 G12 is not even faster for my comparable workload and is sometimes slower than the VPS 8000 ARM G11 I previously rented. I pay approximately twice as much for the RS 8000 G12.
Did I really pay twice the price only to receive a worse and more erratic machine? That is exactly what it feels like at the moment. This is utterly disappointing for a product explicitly advertised with dedicated CPU cores and guaranteed CPU performance.
The way support has handled this is also unacceptable. I submitted my support request five days ago. Of course, I understand that part of this period fell on a weekend. However, several business days have now passed without a technical assessment, a concrete action plan, or a realistic timeframe. This is clearly inconsistent with the communicated expectation of receiving a response on the same day.
I provided support with reproducible measurements and a detailed technical description. Despite that, there has been silence. For a paid production server with a demonstrable and reproducible performance problem, leaving the customer without an answer for days is unacceptable.
I do not want a generic acknowledgement. I expect:
1. A concrete investigation of the physical host.
2. Verification of CPU pinning, SMT sharing, host quotas, CPU frequency, and boost availability.
3. Verification of memory latency, memory bandwidth, NUMA placement, and storage performance.
4. A clear explanation for the nearly uniform steal across all 16 assigned CPU cores and the severe fluctuations.
5. If the problem cannot be resolved permanently and immediately, migration to a different physical host.
Has anyone else had similar experiences with an RS 8000 G12? I am particularly interested in heavily fluctuating steal values, unexpectedly poor performance in compute-intensive applications, and differences between physical hosts.
To netcup: Five days of silence regarding such a clearly documented problem is enough. I am paying approximately twice as much as before and currently receiving worse, less reliable performance. Please finally treat this case seriously and provide a permanent solution.