Note on identifiers. Institution, hosts, addresses, and credentials have been replaced with generic placeholders. Nothing here identifies the client or the servers involved. Malware hashes and technique names are kept — they are public threat intel and the useful part.

Summary

An on-premise Proxmox host was found mining Monero at the expense of the applications running on it. What looked at first like a capacity problem — "why is CPU and RAM always maxed out?" — turned out to be a userland rootkit hiding a miner behind a fake Proxmox system service.

The host was rebuilt from scratch. Monitoring and SIEM were then deployed on a separate VPS so a repeat of the same class of incident would be visible instead of silent.

This document covers the detection path, the forensic evidence, the rootkit's techniques, and the rebuild and hardening that followed.

Timeline

WhenEvent
—CPU/RAM on the Proxmox host pegged high, sustained. Attributed to workload.
—The owner's contact asks why resources are always exhausted. Investigation starts.
—Fake systemd units and an LD_PRELOAD rootkit discovered. Mining confirmed.
—Forensic evidence captured before any remediation touching the host.
—Decision: rebuild, not clean. Too much was uncertain (ld.so.preload locked, self-updating miner).
—Rebuild: Proxmox install, VM split, OS install, hardening (SSH key-only, fail2ban).
—Proxmox Backup Server deployed on a second node. Backup job scheduled.
—Monitoring + Wazuh SIEM deployed on a separate VPS. Failover runbook written.

The evidence

Forensic triage was captured with a purpose-built script and shipped off-host as a tarball. The artefact set:

FileWhat it is
ringkasan-forensik.txtTriage summary: process, network, units, hashes
PVE-1.serviceThe fake unit — miner launched under a Proxmox-sounding name
PVE-1.timer, PVE-1-update.timerPersistence + self-update schedules
PVE-1-maintain.service, PVE-1-maintainWatchdog that restarts the miner
PVE-1-update.service, PVE-1-updateDownloader that re-fetches the payload daily
ld-svcWrapper that installs the LD_PRELOAD library then execs the miner
PVE-1The XMRig binary itself, renamed

Hashes (MD5, as captured at triage):

69087359c827ba0aabe46b0abd432ae2  ld-svc
e1b3e738928012a07dfce8659b3ff31d  PVE-1
9f1fdc6df811a3176eabdf8ade5dda17  PVE-1-maintain
9616da25f24861035d0967d5604ba416  PVE-1-update

What the rootkit did

1. Disguised as a Proxmox service

The systemd unit was named PVE-1, described as "Proxmox VE node health monitor", and marked Managed-by: install_and_mine.sh. To systemctl and to a casual top, it looked like part of Proxmox. The process name was even paired with kworker/3:0 so it resembled a kernel worker thread.

2. Userland rootkit via LD_PRELOAD

The wrapper ld-svc installs /var/lib/systemd/.hide/libhide.so and sets:

HIDE_PROCS=PVE-1,kworker/3:0,PVE-1-maintain
HIDE_PATHS=/var/lib/systemd/PVE-1,/var/lib/systemd/ld-svc,...
HIDE_PORTS=20004
HIDE_SCOPE=all

Result: the miner's process, files, and listening port are filtered out of ps, ls, and netstat for normal tools. HIDE_NETLINK=1 and HIDE_DIRECT_PROC/FILES=1 close the easier detection paths.

3. Global preload, deliberately NOT used

The unit file contains a revealing comment from the attacker:

# Do NOT write /etc/ld.so.preload globally -- that contaminates sshd and
# the next ECDH KEX hangs forever.

They had tested it in the field and learned the failure mode. Instead they scoped LD_PRELOAD per-unit and polluted login sessions separately through /etc/environment and /etc/profile.d/sysenv.sh. /etc/ld.so.preload was left present but locked (lsattr → Permission denied, i.e. immutable).

4. Persistence with self-update and anti-lockstep jitter

Two timers:

  • PVE-1-maintain.timer — OnUnitActiveSec=180s, RandomizedDelaySec=180
  • PVE-1-update.timer — daily (86400s), RandomizedDelaySec=1800

The randomized delays are the interesting part: they exist to avoid lockstep across an infected fleet, so mass remediation (kill and restart on many hosts at once) does not reveal the campaign. RestartSec=42 on the service plus a jitter mechanism inside PVE-1-maintain serve the same goal.

5. Mining configuration

--url=gulf.moneroocean.stream:20004
--user=<monero wallet address>          # redacted
--pass=<worker tag>                      # redacted
--donate-level=0
--tls --no-color --log-file=/dev/null

MoneroOcean is a pool that pays in multiple coins; donate-level=0 means the operator kept every cent. Logging was nulled to avoid local traces.

6. Live C2 / staging connection

ss at triage time showed a process named log in SYN-SENT to an external address on port 80 — a beacon or secondary downloader separate from the miner. This is why "kill the miner" was not the remediation: there was a second channel that could re-establish the foothold.

Detection signals (blue team takeaways)

These are the things that would have caught it sooner, and now drive the detection approach in the companion monitoring stack:

  • A systemd unit whose name and description do not match anything in the distro package set. Diff systemctl list-units --all against known-good.
  • LD_PRELOAD set in any unit's Environment= outside the distribution's own units.
  • A process whose exe (from /proc/<pid>/exe) is under /var/lib/systemd/ or any other improbable path.
  • An outbound connection to a mining-pool port (e.g. :20004, :3333, :4444) from a hypervisor that has no business mining.
  • /etc/ld.so.preload existing at all, and lsattr failing on it.
  • High sustained CPU on a hypervisor host with no matching VM workload.

Rebuild and hardening

Cleaning was rejected because the rootkit was self-updating, the preload file was immutable, and a second network channel existed. Rebuild path:

  1. Fresh Proxmox VE install on both physical nodes.
  2. VMs re-created: application database host and application host.
  3. Operating systems installed clean.
  4. SSH hardened — key-only authentication, password auth disabled.
  5. fail2ban deployed for brute-force protection.
  6. Proxmox Backup Server installed as a VM on the second node.
  7. Backup job scheduled (snapshot mode, zstd compression, retention 7/4/3, least-privilege token).
  8. Monitoring and Wazuh SIEM deployed on a separate VPS — off the compromised network.
  9. Failover runbook written (node-1 death → production moves to node-2).

Lessons learned

  • Resource exhaustion is a security symptom, not only a capacity one. The first question was "is it underpowered?"; the right question was "what is running that I did not start?"
  • Rootkits hide from the same userland tools you are diagnosing with. A miner that hides ps output will not show up in top. Compare against package manifests and inspect /proc directly.
  • Evidence first. Capture before touching anything. The tarball was taken while the miner was still running, which is why the unit files and hashes survived.
  • An off-network vantage point matters. Monitoring on a separate VPS means a compromised production host cannot tamper with its own alarm.
  • Least privilege everywhere. The backup token is scoped, not root. SSH is key-only. This limits the next incident's blast radius.