Cloud VM Workload Forensics

A research project investigating host-level forensic evidence collection, chain-of-custody preservation, and hybrid anomaly detection for cloud IaaS Linux workloads.

What this is

Viewing

CVAF | the framework

A seven-phase forensic pipeline for Linux VM workloads covering auditd-based evidence acquisition, cryptographic chain-of-custody, semantic normalization toward an OCSF-aligned representation, hybrid statistical and heuristic anomaly detection, knowledge-graph construction, visualization, and forensic reporting.

Viewing

CVAF-DS | the dataset

A controlled dataset collected from full-boot KVM virtual machines using native Linux auditd. The collection includes host-level activity such as kernel module operations, credential-file changes, persistence mechanisms, and other events relevant to VM forensics.

Architecture

The CVAF framework is a seven-phase pipeline. Evidence flows from kernel-level acquisition through cryptographic integrity, OCSF-aligned normalization, hybrid detection, per-feature attribution, knowledge-graph construction, and finally reporting. Each phase is detailed below with what it produces and how it feeds the next.

CVAF FRAMEWORK Seven-phase forensic pipeline for Linux VM workloads in cloud IaaS INPUT Live Linux VM workloads running inside cloud IaaS Linux guests PHASE 1 Acquisition Kernel-level evidence capture KERNEL SPACE audit subsystem syscall hooks · filter rules audit_rule dispatch → netlink USER SPACE auditd daemon rule dispatcher · log writer audit.log stream Native Linux auditd Per-scenario rule set Full-boot capture window KERNEL-LEVEL EVIDENCE PHASE 2 Integrity Cryptographic chain-of-custody HASH CHAIN r₁ H₁ → r₂ H₂ → r₃ H₃ → r₄ H₄ Hₙ = SHA-256( recₙ ‖ Hₙ₋₁ ) append-only ledger SHA-256 per record Immutable append-only Tamper-detection CHAIN OF CUSTODY PHASE 3 Normalization Semantic representation RAW AUDITD OCSF type:1300 → class_uid pid:4521 → process.pid uid:0 → user.uid exe:bash → file.path Field mapping Schema alignment Type coercion Missing-value handling OCSF-aligned Cross-source compatible Structured output OCSF ALIGNED PHASE 4 Detection Hybrid anomaly scoring Heuristic 25+ signature rules Extended IF isolation forest RRCF random cut forest → MERGE + SCORE Consensus scoring Anomaly ranking Three independent detection methods Statistical + heuristic HYBRID DETECTION PHASE 5 Attribution Per-feature explanation FEATURE IMPORTANCE exe cmdline syscall uid Shapley value Per-feature contribution Explains each flag Analyst-readable Non-black-box output Per-anomaly reasoning EXPLAINABLE PHASE 6 Graph Entity and relationship model proc file user net 5 entity types process · file · user network · kernel module Relation extraction Cross-event linkage Timeline joining KNOWLEDGE BASE PHASE 7 Reporting Analyst-facing deliverables 3D Visualization Interactive entity graph view WebGL · Three.js Alert Stream Severity-ranked anomaly feed JSON · structured Forensic Report With attribution and timeline Markdown · PDF Evidence Export OCSF records + raw audit data NDJSON Analyst-readable Machine-readable Court-admissible DELIVERABLE OUTPUT OUTPUT Attributed anomalies, knowledge graph, and forensic report

Seven-phase CVAF framework, end to end.

Research objectives

The research examines a forensic pipeline for Linux workloads running inside IaaS virtual machines. The work focuses on evidence acquisition, integrity preservation, semantic representation, anomaly detection, explanation, and forensic analysis, with evaluation based on controlled scenarios and ground truth recorded independently of the detector.

Methodology

Evidence is acquired through native Linux auditd on full-boot KVM virtual machines and hash-chained during acquisition. The normalized records are processed by heuristic rules and statistical anomaly scoring components, including an Extended Isolation Forest variant and a buffer-based Robust Random Cut Forest approximation. Detection output includes Shapley-style feature attribution. Attack scenarios are mapped to MITRE ATT&CK techniques, while ground truth is recorded independently from the detection process.

auditd KVM / libvirt MITRE ATT&CK SHA-256 chain-of-custody Extended Isolation Forest variant Robust Random Cut Forest approximation Shapley-style attribution OCSF-aligned schema

Architecture

CVAF-DS is collected on an isolated two-role virtual machine lab. An attack scenario is executed, its evidence is captured through the same native auditd path CVAF uses in live operation, and a ground-truth record is written independently before the cycle repeats with natural variation.

CVAF-DS DATASET Collection methodology for a controlled Linux VM forensic dataset INPUT Full-boot KVM virtual machines, native Linux auditd, isolated network ROLE 01 Victim VM full-boot Linux guest · auditd ON EVIDENCE SOURCE attack traffic MITRE ATT&CK scenario ROLE 02 Attacker VM isolated role · scenario driver GROUND TRUTH SOURCE evidence ground truth COLLECTION LOOP each scenario runs through this cycle repeatedly with natural variation STAGE 01 Execute attack scenario STAGE 02 Capture auditd stream STAGE 03 Record ground truth STAGE 04 Iterate re-run with variation REPEATS N TIMES CVAF-DS public Linux VM forensic dataset DELIVERABLE OUTPUT Reproducible dataset with independent ground truth

Collection loop for each CVAF-DS attack scenario.

Research objectives

Build a dataset that lets the CVAF framework, along with other Linux host-based detection approaches, be evaluated against real kernel-level evidence rather than synthetic or container-only traces, with ground truth that is reproducible and recorded independently of any detector, and that supports direct comparison against existing public intrusion-detection datasets.

Methodology

Each scenario is mapped to a named MITRE ATT&CK technique and executed repeatedly with natural variation in timing and parameters, interleaved with continuous low-privilege benign background activity so that attack evidence remains a minority class. A cloud instance-metadata scenario is collected via a link-local address and a minimal responder mimicking a real metadata service, since no such endpoint exists on a standalone virtual machine. This is the dataset's primary differentiator from existing public Linux intrusion-detection benchmarks.

SSH Brute-Force · T1110.001 Privilege Escalation · T1548.003 Passwd/Shadow Tampering · T1098 Credential Dumping · T1003.008 New User / Backdoor · T1136.001 SSH Config Tampering · T1098.004 Kernel Module Load/Unload · T1547.006 Recon / Enumeration · T1082 / T1046 Cloud Metadata Access · T1552.005 Log Tampering · T1070.002

Collecting on real virtual machines with real auditd, rather than a synthetic or pre-cleaned log source, surfaced genuine infrastructure findings along the way. They are documented here because they affect reproducibility for anyone working with a similar environment.

Audit-log parsing

The auditd build used in this environment concatenated some interpreted fields without a separator, breaking standard query tooling for certain record types. Resolved by querying the raw log directly for those cases.

Snapshot-safe automation

Automation that must survive a VM snapshot revert cannot live in temporary or ephemeral storage inside the guest, since that is wiped on revert.

Snapshot / audit-state sync

A previously-working audit rule intermittently stopped producing evidence after a snapshot was merely created. Resolved with a full audit-subsystem restart, reload, and a live-fire probe before trusting the environment for an attack action.

Network evidence structure

For the network-based scenario, the record carrying the detection tag does not itself contain the destination address. That lives in a separate, related record, which evidence-matching has to account for.

Remote process lifecycle

Background services started over SSH need proper session detachment to survive the session closing. Killing them reliably by process name was unreliable, whereas identifying them by the port they were bound to was not.

Dataset-scale statistics, run counts, and validation results are reported with the dataset research publication.