Big DataCycle II · 2026-1
Distributed detection of transactional anomalies in SAP Business One for mining audit
A Big Data architecture for finding financial and operational anomalies across roughly 30 GB of real ERP history from a Peruvian mining company. Project proposal.
Manual audit of enterprise ERP systems does not scale with transaction volume. Detecting duplicate payments, split purchase orders and anomalous behaviour through traditional rules is expensive and brittle. This project proposes a distributed architecture that does it with machine learning instead, over real historical data extracted from the SAP Business One database of a Peruvian mining company — around 30 GB covering 2017 to 2026.
The pipeline runs SQL Server (SAP B1) → PySpark on EC2 → Amazon S3 in raw and curated zones → ML models → DynamoDB → a Tableau audit dashboard. Detection layers four techniques: Benford’s law for statistical analysis, Isolation Forest for unsupervised detection, a deep-learning autoencoder, and a hybrid ensemble that ranks the anomalies it finds. The targets are duplicate payments, split purchase orders, backdated documents, atypical financial movements and ghost assets.
Sensitive data is protected before anything leaves the ERP: critical identifiers hashed with SHA-256, and amounts obfuscated with distribution-preserving techniques so the statistical properties survive. This document is the project proposal — the implementation lives in Spark notebooks and EMR jobs.