RS 8000 G12: Up to 10% CPU Steal, Worse Than My ARM VPS — Support Remains Silent

  • I am now genuinely angry and, quite honestly, appalled by how this problem has been handled.

    I rented an RS 8000 G12 with 16 dedicated CPU cores based on the AMD EPYC™ 9645. netcup explicitly advertises its Root Servers as providing CPU and RAM resources exclusively for the customer’s own applications, and states that the performance of the booked CPU cores is guaranteed.

    I run compute-intensive applications on this server. For these workloads, average throughput is not the only thing that matters; stable and predictable execution time is equally important. Even brief interruptions or a small amount of CPU steal can significantly delay critical computation phases, so the impact on my workload is disproportionately large.

    Despite this, under normal load I repeatedly observe substantial CPU steal across all 16 assigned CPU cores. Even more troubling is the extreme fluctuation: at times everything appears normal, only for the steal value to rise sharply shortly afterwards. At peak moments, I have observed values of approximately 10%.

    My repeated longer measurements include:

    • 2.11% average CPU steal over five minutes
    • 4.32% average CPU steal in another measurement
    • 3.80% average CPU steal in a further five-minute measurement
    • short-term peaks of approximately 10%
    • all 16 assigned CPU cores are affected almost uniformly
    • in one measurement, the individual assigned cores showed between 4.18% and 4.49% steal
    • in the latest measurement, they showed between 3.69% and 3.95%
    • approximately 40% of the total CPU capacity was still idle at the same time

    With 16 promised dedicated CPU cores, 4.32% steal is equivalent to approximately 0.69 CPU cores being unavailable on average even though the assigned cores have runnable work. For a compute-intensive and latency-sensitive application, this is not insignificant. A short but concentrated interruption during a critical computation phase can cause considerably more damage than the average value initially suggests.

    I want to make it clear that I am not claiming CPU steal has already been proven to be the sole cause of the poor overall performance. Steal is simply the clearly measurable symptom visible from within the guest operating system. Other possible causes include limitations or problems involving CPU frequency and boost, SMT sharing, memory latency or memory bandwidth, NUMA placement, storage latency, or other bottlenecks on the physical host. This is precisely why I expect a serious technical investigation of the complete system, rather than a superficial standard response.

    For a short period, the problem appeared to have disappeared. Shortly afterwards, CPU steal returned to almost 4%, with further significantly higher peaks. This cannot reasonably be described as a stable or permanent resolution.

    The direct comparison is particularly frustrating: according to my latest real-world tests, this RS 8000 G12 is not even faster for my comparable workload and is sometimes slower than the VPS 8000 ARM G11 I previously rented. I pay approximately twice as much for the RS 8000 G12.

    Did I really pay twice the price only to receive a worse and more erratic machine? That is exactly what it feels like at the moment. This is utterly disappointing for a product explicitly advertised with dedicated CPU cores and guaranteed CPU performance.

    The way support has handled this is also unacceptable. I submitted my support request five days ago. Of course, I understand that part of this period fell on a weekend. However, several business days have now passed without a technical assessment, a concrete action plan, or a realistic timeframe. This is clearly inconsistent with the communicated expectation of receiving a response on the same day.

    I provided support with reproducible measurements and a detailed technical description. Despite that, there has been silence. For a paid production server with a demonstrable and reproducible performance problem, leaving the customer without an answer for days is unacceptable.

    I do not want a generic acknowledgement. I expect:

    1. A concrete investigation of the physical host.
    2. Verification of CPU pinning, SMT sharing, host quotas, CPU frequency, and boost availability.
    3. Verification of memory latency, memory bandwidth, NUMA placement, and storage performance.
    4. A clear explanation for the nearly uniform steal across all 16 assigned CPU cores and the severe fluctuations.
    5. If the problem cannot be resolved permanently and immediately, migration to a different physical host.

    Has anyone else had similar experiences with an RS 8000 G12? I am particularly interested in heavily fluctuating steal values, unexpectedly poor performance in compute-intensive applications, and differences between physical hosts.

    To netcup: Five days of silence regarding such a clearly documented problem is enough. I am paying approximately twice as much as before and currently receiving worse, less reliable performance. Please finally treat this case seriously and provide a permanent solution.

    Edited once, last by oldman: Translated the post into English (July 30, 2026 at 4:08 PM).

  • oldman July 30, 2026 at 4:05 PM

    Changed the title of the thread from “RS 8000 G12: Bis zu 10 % CPU-Steal, schlechter als mein ARM-VPS – Support schweigt” to “RS 8000 G12: Up to 10% CPU Steal, Worse Than My ARM VPS — Support Remains Silent”.
  • I share your frustration—my experience with the RS 8000 G12 has been very similar.

    Hitting a ceiling of 20k IOPS on an RS 16000 G12 with 32 dedicated cores clearly points to a host-level bottleneck, whether it's storage limits, CPU pinning, or NUMA configuration.

    The main issue is the wasted investment. Hiring an external team to configure the server and purchasing 3rd-party software licenses becomes a direct financial loss when the underlying hardware cannot deliver basic expected performance.

    The lack of support response only makes things worse. Leaving tickets unanswered for days on production-tier servers leaves customers with no clear path forward.

  • Update: I want to provide a fair and accurate update on the situation.

    The server has now been restored to normal. CPU steal is currently consistently at 0%, and after further real-world testing, the performance now meets my expectations. The technical performance issue therefore appears to have been resolved, and I appreciate that corrective action was taken.

    However, I still have not received any official email response regarding my support requests—not even a confirmation of what was changed, what the root cause was, or whether the fix is considered permanent.

    I am pleased that the server is now performing as expected, but the complete lack of official communication remains disappointing. I would still appreciate a written response and a brief technical explanation from netcup.

  • Since few people here on forum got similar problems in this particular location I want to ask:

    Did you checked all other factors? I've seen ST too, but not as huge as yours. Now its permanent 0 also. But unfortunatelly disk I/O is still random and CPU is not performing as it should 1258/6000 in geekbench 6 (original ~2200/11900). It's not as bad as it was earlier today or yestarday, but in my case, i still see random drops and spikes. Serioulsy i don't believe clocks are promised 3.7 Ghz per core. They seems to be default 2.3 for some reason or sometimes throttle.

  • Well, apparently I celebrated a little too early. 😅

    I had just posted that the server was finally back to normal: CPU steal was consistently at 0%, CPU performance had improved significantly, and the machine finally met my expectations.

    Then the disk apparently said, “Hold my beer.”

    After a reboot, storage performance looked completely healthy: about 1.8 GiB/s direct write, 2.1 GiB/s direct read, and roughly 2 ms latency. But after restarting a large parallel download, the storage problem became reproducible again. I/O pressure rose to about 98–99%, several processes became stuck in uninterruptible D state (including jbd2/ext4 writeback), roughly 4.2 GiB of dirty data accumulated, write latency climbed into the 9–35 second range, and actual write throughput sometimes collapsed to around 1 MiB/s.

    The network still performs normally and CPU steal remains at 0%, so this new bottleneck is in the storage/writeback path under this workload. A reboot temporarily clears it, but restarting the same workload brings it back.

    The timing is almost funny: I had barely announced that everything was resolved before the server found an entirely new way to misbehave. At this point, I am almost afraid to post another “resolved” update—apparently the server treats it as a challenge.

    Joking aside, this is very frustrating. I am now also testing whether more conservative parallelism and file-allocation settings can avoid the problem, but I would still appreciate a technical explanation of this storage behaviour.

  • any reply from support?

    No, I have not received a single email or any other official reply from support.

    Judging by what has happened so far, they seem to prefer fixing things silently in the background without telling the customer what was changed, what the cause was, or whether the fix is permanent.

    And since it is the weekend now, I do not expect them to take any action today either. So for the moment, all I can do is keep testing and try to guess whether something has been changed behind the scenes.

  • Looks like You are right - they like to fix things silently. My machine is now performing ~8/10 and im fine with it comparing to my other machines in other company.
    Dedicated cpu and ram BUT shared storage. Funny thing, because it's root of all problems with them. Some fio tests shows write latency tails exceeding much much above 100ms. Phenomenal was p99.99 = 6377ms latency spike.

    Isn't this like playing video game with 200 FPS but with random screen stutters every few seconds? :D

    Edited once, last by cysiek (August 1, 2026 at 11:54 AM).

  • My latest FIO test still shows very poor storage performance:

    Metric Result
    Read IOPS 14,716
    Write IOPS 4,922
    Combined IOPS ≈19,638
    Read bandwidth 57.4 MiB/s
    Write bandwidth 19.2 MiB/s
    Average read latency 2.99 ms
    Average write latency 4.06 ms
    99th percentile read latency 10.99 ms
    99th percentile write latency 12.63 ms
    Disk utilization 64.8%

    Should netcup be concerned about potential legal consequences in Germany if its servers are advertised as using high-performance NVMe storage while customers receive substantially lower performance, especially when support tickets concerning the issue remain unanswered for several days?

  • Well, I'm getting angry.

    Did I rated my machine 8/10? Yes, based on this values (Time shifted -2h from Berlin time):
    [23:32:16] Processed: 55000 | Δt: 0.69s | Total: 8.96s | AVG EPS: ~6140, LAP EPS: ~7215
    [23:32:17] Processed: 60000 | Δt: 0.61s | Total: 9.57s | AVG EPS: ~6270, LAP EPS: ~8166
    [23:32:17] Processed: 65000 | Δt: 0.62s | Total: 10.19s | AVG EPS: ~6382, LAP EPS: ~8119
    [23:32:18] Processed: 70000 | Δt: 0.61s | Total: 10.80s | AVG EPS: ~6482, LAP EPS: ~8138
    Total 100 000 message processed in 15.03s

    currently its drowning again:
    [11:12:19] Processed: 55000 | Δt: 1.38s | Total: 14.82s | AVG EPS: ~3712, LAP EPS: ~3623
    [11:12:21] Processed: 60000 | Δt: 1.33s | Total: 16.15s | AVG EPS: ~3716, LAP EPS: ~3762
    [11:12:22] Processed: 65000 | Δt: 1.27s | Total: 17.41s | AVG EPS: ~3733, LAP EPS: ~3952
    [11:12:23] Processed: 70000 | Δt: 1.23s | Total: 18.64s | AVG EPS: ~3756, LAP EPS: ~4081
    Total 100 000 message processed in 26.57s

    These tests include cli php script talking to mariadb (read then write). Its like iddling for this kind of setup. I think storage is simply broken.
    2026-08-01_13-39.png

    I think netcup team should have a closer look at this, because such unstability is not acceptable.I bet it's faulty backend driver for storage.