A research project investigating host-level forensic evidence collection, chain-of-custody preservation, and hybrid anomaly detection for cloud IaaS Linux workloads.
A seven-phase forensic pipeline for Linux VM workloads covering auditd-based evidence acquisition, cryptographic chain-of-custody, semantic normalization toward an OCSF-aligned representation, hybrid statistical and heuristic anomaly detection, knowledge-graph construction, visualization, and forensic reporting.
A controlled dataset collected from full-boot KVM virtual machines using native Linux auditd. The collection includes host-level activity such as kernel module operations, credential-file changes, persistence mechanisms, and other events relevant to VM forensics.
The CVAF framework is a seven-phase pipeline. Evidence flows from kernel-level acquisition through cryptographic integrity, OCSF-aligned normalization, hybrid detection, per-feature attribution, knowledge-graph construction, and finally reporting. Each phase is detailed below with what it produces and how it feeds the next.
Seven-phase CVAF framework, end to end.
The research examines a forensic pipeline for Linux workloads running inside IaaS virtual machines. The work focuses on evidence acquisition, integrity preservation, semantic representation, anomaly detection, explanation, and forensic analysis, with evaluation based on controlled scenarios and ground truth recorded independently of the detector.
Evidence is acquired through native Linux auditd on full-boot KVM virtual machines and hash-chained during acquisition. The normalized records are processed by heuristic rules and statistical anomaly scoring components, including an Extended Isolation Forest variant and a buffer-based Robust Random Cut Forest approximation. Detection output includes Shapley-style feature attribution. Attack scenarios are mapped to MITRE ATT&CK techniques, while ground truth is recorded independently from the detection process.
CVAF-DS is collected on an isolated two-role virtual machine lab. An attack scenario is executed, its evidence is captured through the same native auditd path CVAF uses in live operation, and a ground-truth record is written independently before the cycle repeats with natural variation.
Collection loop for each CVAF-DS attack scenario.
Build a dataset that lets the CVAF framework, along with other Linux host-based detection approaches, be evaluated against real kernel-level evidence rather than synthetic or container-only traces, with ground truth that is reproducible and recorded independently of any detector, and that supports direct comparison against existing public intrusion-detection datasets.
Each scenario is mapped to a named MITRE ATT&CK technique and executed repeatedly with natural variation in timing and parameters, interleaved with continuous low-privilege benign background activity so that attack evidence remains a minority class. A cloud instance-metadata scenario is collected via a link-local address and a minimal responder mimicking a real metadata service, since no such endpoint exists on a standalone virtual machine. This is the dataset's primary differentiator from existing public Linux intrusion-detection benchmarks.
Collecting on real virtual machines with real auditd, rather than a synthetic or pre-cleaned log source, surfaced genuine infrastructure findings along the way. They are documented here because they affect reproducibility for anyone working with a similar environment.
The auditd build used in this environment concatenated some interpreted fields without a separator, breaking standard query tooling for certain record types. Resolved by querying the raw log directly for those cases.
Automation that must survive a VM snapshot revert cannot live in temporary or ephemeral storage inside the guest, since that is wiped on revert.
A previously-working audit rule intermittently stopped producing evidence after a snapshot was merely created. Resolved with a full audit-subsystem restart, reload, and a live-fire probe before trusting the environment for an attack action.
For the network-based scenario, the record carrying the detection tag does not itself contain the destination address. That lives in a separate, related record, which evidence-matching has to account for.
Background services started over SSH need proper session detachment to survive the session closing. Killing them reliably by process name was unreliable, whereas identifying them by the port they were bound to was not.
Dataset-scale statistics, run counts, and validation results are reported with the dataset research publication.